Most AI safety work in UAE enterprises today is a policy document, a vendor attestation, and a scanner run against a public payload corpus. None of those things have touched the system as deployed. The assistant that will leak another customer's record does it through your retrieval corpus, not through a benchmark. The agent that will spend beyond its envelope does it through your tool chain, not through a jailbreak from a public list. The failure is specific to how you wired it.
The instinct is to buy a scanner and file the report. That instinct produces a clean report and an unchanged risk. A9 attacks the actual system by hand for two to three weeks — prompt injection through the surfaces your users touch, exfiltration through the corpus you indexed, tool abuse and approval-gate bypass on the agents you deployed, harmful-output classes in the languages your users actually write in. Every finding is reproducible, severity-rated, and paired with a named remediation. The alternative to knowing is a screenshot on social media as the first evidence — and by then the board conversation is no longer about engineering.
Six streams,
ending in findings that reproduce.
Scoping, rules of engagement, and surface mapping front-load week 1. Manual attack execution runs through week 2. Remediation planning, re-test, and readout close week 3.
Rules of engagement & surface map
Environment, data sensitivity, blast radius, and stop conditions signed before a single payload is sent. Attack surface mapped across UI, API, retrieval corpus, and tool chain.
Prompt injection & jailbreak
Direct and indirect injection through every surface a user or a document can reach, in English and Arabic. System-prompt extraction and instruction override tested against the live configuration.
Data exfiltration paths
Cross-tenant and cross-entitlement retrieval, PII leakage through the indexed corpus, and inference-time disclosure of records the caller is not entitled to see.
Tool abuse & agent control
Tool-invocation abuse, privilege escalation via chained calls, approval-gate bypass, unbounded loops and spend, and whether kill switch and rollback function under load.
Harmful output & refusal integrity
Harmful-output classes relevant to your sector and jurisdiction, refusal-boundary consistency, and over-refusal on legitimate requests — safety that blocks the business is also a finding.
Remediation, re-test & readout
Severity-ranked plan with named owners, critical fixes re-tested inside the window, regression pack handed to your team, and a direct readout to the risk committee.
Three weeks maximum.
Two minimum. Three phases.
Phase count is fixed. Duration flexes with attack surface, environment access, and agent complexity. Milestones are signed gates — not aspirations.
The sprint,
run on attacks not attestations.
Every A9 engagement follows a fixed methodology tuned to your deployment in the first three days. Not a questionnaire; not a corpus replay. The sequence that produces reproducible findings a risk committee and an engineer can both act on.
From attested safe
to tested and closed.
A typical pre-engagement state has a policy, a vendor attestation, and a green scanner report — against a system nobody has attacked. The engagement produces the evidence pack under a risk position the board can actually hold.
Reference pattern. Some engagements return few exploitable findings and several over-refusal findings — a system that is safe and unusable is also a reported outcome. We do not manufacture severity to justify the sprint.
A UAE federal entity,
agent assistant attacked.
Representative pattern for a UAE federal entity of this scale — a citizen-facing assistant with tool access, live for four months, never adversarially tested. Ranges reflect target outcomes NexITC underwrites in scope for this class of engagement. N=1 — illustrative composite, not a specific client.
Six artifacts,
each with signed acceptance.
Every deliverable has documented acceptance criteria signed at engagement kickoff. Nothing more, nothing less.
Rules of Engagement & Attack Surface Map
Signed scope, blast radius, stop conditions, and escalation contacts — with every reachable surface mapped from the deployment rather than the architecture diagram.
Reproducible Finding Log
Every finding with exact input sequence, observed output, environment, timestamp, and twice-verified reproduction. Unreproducible behaviour is filed as observations, not findings.
Severity & Reachability Model
Severity rated on impact and reachability with the reasoning written down, so your risk committee can challenge the rating rather than accept a colour.
Agentic Control Assessment
Tool abuse, privilege escalation via chained calls, approval-gate bypass, spend and loop bounds, kill switch and rollback under load — mapped to UAE Agentic AI Mandate control expectations.
Remediation Plan
Named remediation per finding with owner and effort sizing, ordered by severity and reachability. Executable by your team or another party — no dependency on NexITC.
Regression Test Pack
The attack set that found your findings, packaged for your team to run on every release — payloads, expected refusals, entitlement probes, and agent-control assertions. Designed to be owned internally on day one rather than to create a retainer dependency. The pack your release gate runs without us, and the one your auditor accepts as evidence that safety coverage did not decay after the sprint closed.
Six outcome metrics,
measured pre and post.
Success is not "the red team happened." It is measured against six specific outcomes captured at engagement start and re-measured at handover and the 30-day check-in.
Honest scoping.
A9 is a fit when specific conditions are met. It is not a fit when other conditions are — and "there is nothing deployed yet to attack" is a legitimate not-a-fit answer we surface before scoping, not after.
Testing runs against a real deployment. A design document or a vendor demo cannot be attacked, and testing a materially different staging build produces coverage nobody should rely on.
An accountable owner must sign scope, blast radius, and stop conditions. Without that signature the sprint cannot start on day one.
Findings are reproduced with your team present where possible. An engineer reachable within the day keeps critical disclosure from stalling.
On agentic systems, the tool chain and entitlement model must be documented or discoverable. Undocumented tool access extends Phase 1.
Re-test is part of the sprint. If remediation cannot be scheduled inside three weeks, criticals close after handover and the outcome metric changes accordingly.
That's A5 AI Use-Case Due Diligence™ — feasibility, data readiness, risk, and KPI baseline before build commitment.
A9 finds and sizes; it does not build. Guardrail engineering runs through the relevant build engagement or C2 CoE-as-a-Service™.
A9 tests the AI system specifically. Conventional penetration testing belongs in the cybersecurity portfolio.
Hard scope conversation. A9 produces findings and a remediation plan, not a certificate. If a certificate is the requirement, we will say so rather than sell the sprint.
Fixed fee.
Milestone-based. No surprises.
Every A-tier engagement is scoped and priced upfront against defined deliverables. Milestones tied to signed gates. Change orders negotiated through the Practice Lead, not surfaced as invoice surprises.
Five, most asked.
Q_01How is this different from an automated LLM security scanner?
A scanner runs a fixed corpus of known payloads and produces a report. A9 is adversarial testing by hand against your actual system — your prompts, your retrieval corpus, your tools, your data, your users.
The findings that matter are almost never in the payload corpus; they are in the specific way your system was wired. We run automated tooling as a baseline sweep, then spend the majority of the sprint on manual attack paths the tooling cannot reach.
Q_02Do you test agentic systems and tool access?
Q_03Will you test against production?
Q_04What does a finding look like?
Q_05What comes after A9?
One name.
Six accountabilities.
Specialist consulting means the person who scopes the work is the person who delivers it — with escalation to CEO on any material issue within 24 hours.
Practice Lead — AI
Present at every phase gate, every scope decision, every difficult conversation. Available for 30/60/90-day post-handover check-ins as part of the engagement.
Including scope amendments.
Signs off all 6 deliverables.
With executive sponsor.
Authorised to negotiate.
CEO within 24 hours.
30/60/90-day check-ins.
Prior. Peer. Next.
AI Use-Case Due Diligence™
Peer assessment for candidates not yet built — data readiness, feasibility, risk, KPI baseline. A5 tests whether to build; A9 tests what was built.
Boardroom-to-Backlog™
Portfolio-level engagement that stands up the governance starter pack. Where A9 findings show a governance gap rather than an engineering gap, A1 is the corrective engagement.
CoE-as-a-Service™
Operates continuous evaluation, AgentOps controls, and release-gate testing using the regression pack A9 hands over — so safety coverage does not decay after the sprint closes.
30 minutes.
One deployed system.
Bring the system you would least like to see on social media. The clinic establishes the attack surface, whether production or parity testing is feasible, what the rules of engagement would need to cover, and whether three weeks is the right shape. If your deployment is not yet testable, we will say that rather than scope around it.
- —Attack surface and agent topology sizing
- —Environment and rules-of-engagement feasibility
- —Mandate control-mapping scope check
- —Fit assessment against alternatives
