Skip to main content
NexITC
B18 · AI · 8–14 WEEKS · BUILD

AI compute.
Inside your boundary.

B18 · Sovereign AI Infrastructure Build™ deploys AI inference infrastructure — GPU clusters, model serving pipelines, security controls — inside your sovereign boundary. Data does not leave. Inference happens on named hardware. Audit trails satisfy PDPL, ADHICS v2, and sector-specific supervisor expectations.

DURATION
8–14 wks
DELIVERABLES
7 named
COMMERCIAL
Fixed fee
B18·PROJECTION / INFERENCE LATENCY
B18
BEFORE
N/A
NO SOVEREIGN AI · BLOCKED
B18
AFTER
1.6s
P95 INFERENCE · IN-BOUNDARY
WK 00
WK 04
WK 08
WK 12
STEADY
GPU UTILISATION
75%
P95 LATENCY
1.6s
USE CASES UNBLOCKED
5+
SCENARIO · UAE GOVT SOVEREIGN AI · N=1
ILLUSTRATIVE
§ 00 · THESIS
01
WHY SOVEREIGN
COMPUTE MATTERS.

Every regulated UAE enterprise reaches the same point in every AI conversation. The use case is clear. The commercial case is clear. The data classification is clear. And then someone in the room says the sentence that stops the meeting: this data cannot leave the country. Or this data cannot leave this building. Or this data cannot touch a foreign cloud provider under any interpretation of the contract we already signed with our regulator.

The instinct at that moment is to postpone. Wait for the vendor to open a UAE region. Wait for the regulator to update the guidance. Wait for someone else to solve it. The instinct is wrong. Sovereign AI infrastructure — GPU compute inside a named boundary, model serving on hardware you point to, audit trails a supervisor can inspect — is not novel work. It is deployment work. B18 does that deployment on a fixed scope in 8–14 weeks so the use case stops waiting.

STATE · BLOCKED
AI use case defined. Data classification prohibits foreign cloud. Deployment stalled.
STATE · OPERATIONAL
GPU compute inside sovereign boundary. Inference under supervisor-grade audit. Use cases unblocked.
§ 01 · WORK STREAMS

Six streams,
ending in operational compute.

Requirements and architecture front-load weeks 1–3. Provisioning and cluster configuration overlap through weeks 3–9. Model serving, monitoring, and security hardening close weeks 9–14.

STREAM 01
WK 01–02

Requirements & workload profile

AI workloads catalogued: LLM inference, vision, embeddings, agentic. Utilisation and latency targets committed. Data classification confirmed against sovereign residency requirements.

STREAM 02
WK 02–04

Architecture design

Compute, storage, network, security architecture designed to the workload profile. Sovereign hardware or in-country IaaS provider selected on evidence, not vendor preference.

OUTCOME
1
SOVEREIGN AI CLUSTER LIVE
+ SUPERVISOR EVIDENCE PACK
STREAM 03
WK 03–09

Provisioning & cluster configuration

GPU cluster procured or provisioned. Cluster configured for the workload mix. Network segmentation and encryption at rest wired from day one.

STREAM 04
WK 06–11

Model serving pipeline

Model deployment, versioning, and A/B testing pipeline built. Autoscaling policies configured per workload. Quantisation tuned against gold-standard accuracy sets.

STREAM 05
WK 09–13

Security hardening & audit

Access governance, audit logging, key management. Every inference call attributable. Supervisor-grade evidence pack drafted with your compliance team.

STREAM 06
WK 13–14

Performance validation & handover

Load testing to committed SLA. GPU utilisation baselined. Operator training. 30/60/90-day post-handover check-ins.

EXPLICITLY NOT COVERED
Model training pipelines and MLOps release governance
beyond initial deployment. That's B3 MLOps Factory™.
The GenAI application or agent layer built on top
of the sovereign infrastructure. That's B21 Sovereign AI Platform Build or B14 Agentic Workflow Agent Build™.
§ 02 · TIMELINE

Fourteen weeks maximum.
Eight minimum. Four phases.

Phase count is fixed. Duration flexes with hardware procurement lead-time and environment complexity. Milestones are signed gates — not aspirations.

WK 01020304050607080910111213 · 14Phase 1 · RequirementsPhase 2 · Provisioning & clusterPhase 3 · Model serving & securityPhase 4 · ValidationArchitecture signedEND WK 04 · GATE 01Cluster provisionedEND WK 09 · GATE 02Security hardenedEND WK 13 · GATE 03Handover completeEND WK 14 · GATE 04OPERATING RHYTHMDaily standup · Weekly sponsor check-in · Bi-weekly Practice Lead reviewNAMED ACCOUNTABILITYPractice Lead — AI (CEO escalation available)
§ 03 · APPROACH

Compute selected on evidence,
not on brand recognition.

Every hardware or sovereign IaaS decision runs a six-criteria scorecard in weeks 1–2. Each criterion scored 1–5 against evidence — vendor spec sheets are read carefully but never the final word. Signed by your infrastructure lead before Phase 2 begins.

SOVEREIGN COMPUTE SELECTION SCORECARD · TEMPLATE
CRITERIA · 06 · WEIGHTED 1–5
ILLUSTRATIVE SAMPLE RENDERING — actual scores are engagement-specific and derived from evidence gathered during discovery.
01
Workload fit
Measured throughput on your workload mix, not vendor benchmarks. LLM inference, vision, embeddings each behave differently on the same silicon.
4/5
02
UAE data residency
In-country hardware and support presence. For IaaS, contractual guarantees on where inference and logs are processed.
5/5
03
Three-year TCO
Capex plus opex plus support plus depreciation, modelled against realistic utilisation growth. Not sticker price.
3/5
04
Security & audit posture
Native support for logging, key management, and access governance to the standard your supervisor expects.
4/5
05
Serviceability in-country
Response SLA for hardware faults, spare parts availability, engineer certification. Sovereign compute that can't be serviced fast is not sovereign for long.
4/5
06
Ecosystem & framework support
Model framework compatibility (PyTorch, TensorFlow, ONNX), MLOps tool integration, existing skills on your team.
3/5
!
DISCLOSURE · VENDOR-NEUTRALITY
NexITC maintains commercial arrangements with several sovereign compute and GPU vendors — these are how specialist consultancies build sustainable practices. We do not disclose which arrangements exist publicly because we do not want them to influence hardware or provider choice by anyone reading this page. The scorecard exists precisely so selection happens on evidence, not on economics. In practice, we have recommended vendors with which we have no partnership when the scorecard result favoured them.
§ 04 · ARCHITECTURE

From blocked deployments
to in-boundary compute.

A typical pre-engagement state has AI use cases stalled at the data-classification gate — the workload works technically but cannot be deployed on foreign cloud. The engagement stands up compute inside your sovereign boundary and moves the stalled workloads into deployment path.

BEFORE · T=0
TYPICAL STATE
USE_CASE_01
Doc analysis (blocked)
DATA CLASS · RESTRICTED
USE_CASE_02
Citizen services LLM
DATA CLASS · SENSITIVE
USE_CASE_03
Internal search
SUPERVISOR RESTRICT
USE_CASE_04
Ops assistant
PDPL RESIDENCY
COMPUTE · FOREIGN
Foreign cloud AI compute available in principle. Blocked by policy for these workloads.
OPERATIONAL REALITY
  • Multiple AI use cases stalled at deployment gate
  • Board pressure to ship; compliance pressure to wait
  • Foreign cloud AI declined by procurement or legal
  • In-house GPU capability limited or unmapped
B18 · DEPLOY
AFTER · STEADY STATE
TARGET-STATE
PLATFORM_01
Sovereign GPU Cluster
Compute · Storage · Network · Encryption
PLATFORM_02
Model Serving Pipeline
Deployment · Versioning · Autoscaling · Monitoring
↓ IN-BOUNDARY · METERED · AUDITED · ATTRIBUTABLE ↓
USE-CASE APPLICATIONS · RETAINED
Point at sovereign endpoints
STEADY-STATE OUTCOME
  • Sovereign AI compute operational, use cases unblocked
  • P95 inference latency within SLA on real workload
  • GPU utilisation baselined at 70–80% steady state
  • Every inference call attributable to a logged user or agent

Reference pattern. Some engagements retain a foreign-cloud path for workloads whose data classification permits it, alongside the sovereign path. What always changes is that the workloads that couldn't deploy now can.

§ 05 · REPRESENTATIVE SCENARIO

A government
sovereign platform, live.

Representative pattern for a UAE federal or emirate-level government entity requiring in-boundary AI compute. Ranges reflect target outcomes NexITC underwrites in scope for this class of engagement. N=1 — illustrative composite, not a specific client.

SCENARIO / B18 / UAE GOVT · 3 LLM DEPLOYMENTS
DURATION · 12 WKS
P95 INFERENCE LATENCY
1.6s
on 3 LLM deployments · in-boundary
GPU UTILISATION
75%
steady-state, workload-tuned
USE CASES UNBLOCKED
5+
downstream AI initiatives enabled
SITUATION

UAE government entity. Three AI use cases identified — document analysis, citizen-services assistant, internal knowledge search. All blocked by data-classification requirements that preclude foreign cloud AI. Board approval given contingent on sovereign deployment path.

ENGAGEMENT

12-week B18. Weeks 1–4 requirements, workload profile, architecture. Weeks 4–9 hardware provisioning inside the entity's sovereign data centre, cluster configuration. Weeks 9–12 model serving for three LLMs, security hardening, load testing, handover to the entity's infrastructure team.

OUTCOME

P95 inference latency under 2 seconds across all three LLMs. GPU utilisation optimised to 75% at steady state. Full audit trail for every inference call, accepted by the entity's internal audit. Foundation enabled five downstream AI use cases that could not otherwise have shipped.

§ 06 · DELIVERABLES

Seven artifacts,
each with signed acceptance.

Every deliverable has documented acceptance criteria signed at engagement kickoff. Nothing more, nothing less.

D_01

AI Infrastructure Architecture

Compute, storage, network, and security architecture designed to your workload profile — signed off by your infrastructure lead.

D_02

GPU Cluster Configuration

Cluster provisioned, tuned for the workload mix, and baselined for utilisation and throughput.

D_03 · CORE

Model Serving Pipeline

Deployment, versioning, A/B testing, and autoscaling for the workloads in scope. Quantisation tuned against accuracy tolerances you set.

D_04

Inference Endpoints & Autoscaling

Endpoint management, autoscaling policies, and SLA-tuned routing across the cluster.

D_05 · CORE

Monitoring Dashboards

Latency (P50/P95/P99), GPU utilisation, cost per inference, availability — one dashboard your infrastructure lead checks daily.

D_06

Security & Audit Controls

Access governance, audit logging, encryption, key management — every inference call attributable to a logged user or agent.

D_07 · BOARD-READY

Cost Model & Capacity Plan

Three-year cost model against your workload growth curve, capacity plan for the next expansion, and a supervisor-grade evidence pack for your compliance team. The document your CFO signs before the next AI use case begins.

HANDOVER
WK 14
§ 07 · OUTCOMES

Six outcome metrics,
measured pre and post.

Success is not "the cluster is racked." It is measured against six specific outcomes captured in a baseline report at engagement start and re-measured at cluster steady state.

THE INFERENCE LATENCY JOURNEY · REPRESENTATIVE
Seconds to under two, across the four phases.
1.6sP95 STEADY
8s6s4s2s0BlockedBaselinePRE-ENGAGEMENT4.2sFirst deploymentEND WK 092.3sTunedEND WK 131.6sSteady state30 DAYS POST
01 · LATENCY
<2s
P95 inference latency at steady state on the workload profile.
02 · UTILISATION
70–80%
GPU utilisation steady-state, workload-tuned.
03 · DEPLOYMENT
<48hrs
Model deployment time from registry commit to endpoint.
04 · COST
Meas.
Cost per inference, baselined at go-live.
05 · AVAILABILITY
99.5+%
Uptime SLA adherence, measured monthly.
06 · AUDIT
100%
Inference calls attributable to a logged user or agent.
§ 08 · FIT

Honest scoping.

B18 is a fit when specific conditions are met. It is not a fit when other conditions are. We say so before the scope conversation, not after the commercial commitment.

PREREQUISITES
Move fast when these five conditions are in place at kickoff.
01
One or more AI use cases actually blocked by data-residency requirements

Not "we want sovereign for future flexibility." A named workload that cannot deploy today because of a residency rule.

02
Infrastructure lead as day-to-day counterpart

Signs off architecture, cluster configuration, and validation. Typically 30% time commitment through the engagement.

03
Physical space or IaaS commitment already made

Data centre floor space allocated, or a sovereign IaaS provider selected. B18 does not include site selection or IaaS procurement negotiation.

04
Compliance and supervisor context defined

Which residency framework applies (PDPL, ADHICS v2, CBUAE) and what evidence your compliance team needs to sign off.

05
Named executive sponsor with capex authority

GPU procurement decisions land at the executive level. Sponsor needs to be able to approve within engagement timeline, not next quarter.

NOT SUITABLE IF
Four patterns indicate a different engagement is a better fit.
You want a proof-of-concept, not a production cluster

B18 is production scope. For POC-scale work, start with B1 Pilot Factory™ on foreign cloud where residency allows.

You need the GenAI application layer, not the compute

B18 delivers infrastructure. For the platform on top, B21 Sovereign AI Platform Build is the sequenced next engagement.

You need MLOps pipeline for supervised ML models

That's B3 MLOps Factory™ — often run on B18 infrastructure once B18 is complete.

Hardware procurement lead-time exceeds 12 weeks

We can accelerate architecture and provisioning-plan work into weeks 1–4, then pause. B18 clock restarts when hardware arrives.

§ 09 · COMMERCIAL

Fixed fee.
Milestone-based.

Total engagement fee agreed in the scope statement. Not time-and-materials. Not day rate. Every engagement is preceded by a scope conversation to ensure fit before commitment.

STANDARD MODEL
ENGAGEMENT MODEL
Fixed fee
PAYMENT CADENCE
Milestone-based

Payment schedule aligned to engagement phases and defined delivery milestones agreed upfront.


INCLUDED IN SCOPE
  • All 7 named deliverables with acceptance criteria
  • Named Practice Lead throughout the engagement
  • Bi-weekly executive sponsor reviews
  • 30/60/90-day post-handover check-ins
  • Written scope amendment process for any changes
01

Signed scope statement

Every engagement begins with a signed scope statement fixing deliverables, timeline, milestones, and commercial terms. No verbal agreements. No moving targets.

02

No scope creep

Scope changes require a signed scope amendment. If scope changes, so does the commercial arrangement — always in writing, always signed by both parties.

03

Named accountability

The Practice Lead is accountable for commercial and delivery outcomes throughout the engagement, with escalation to the CEO within 24 hours if needed.

§ 10 · QUESTIONS

Five, most asked.

Q_01Why is sovereign AI infrastructure actually necessary?

Three concrete reasons, not one abstract one. First, PDPL data-residency requirements for personal data create a legal boundary that foreign cloud AI services cannot always cross.

Second, government classification policies restrict which workloads may leave sovereign infrastructure at all.

Third, contractual obligations — CBUAE supervisory expectations, sector-specific supervisor access rights, procurement clauses — often close the door on foreign AI compute before technical questions get asked.

B18 exists because the question has already been settled for these organisations. What remains is execution.

Q_02Does this only support LLMs, or the full AI workload mix?
The full mix. B18 deploys GPU compute plus a model-serving pipeline sized for the workload mix you actually run — LLM inference, vision models, embeddings, and agentic AI systems that call multiple models. Model quantisation and autoscaling are configured per workload profile, not per model, so cost stays proportional to real utilisation.
Q_03Can this run on our existing data-centre hardware, or do we need to procure GPUs?
Either. If GPU-capable hardware is already on the floor, we deploy against it. If it needs procuring, hardware selection runs the same evidence-based scorecard as any other platform choice — Gartner class, three-year TCO at your workload profile, service and support in-country. Sovereign cloud deployment (UAE-resident IaaS providers) is also supported when you prefer OpEx to CapEx.
Q_04How do you keep GPU costs from escalating?
Three levers, tuned to what you actually run. Utilisation monitoring surfaces underused GPUs weekly, not quarterly. Model quantisation reduces compute per inference where accuracy tolerances allow — measured on a gold-standard set before deployment. Autoscaling caps peak burn while meeting SLA. The KPI dashboard makes cost-per-inference visible from day one — the number that actually matters, rather than headline compute spend.
Q_05What comes after infrastructure is deployed?
Either B21 Sovereign AI Platform Build to add the model and agent platform layers on top, or B3 MLOps Factory™ to productionise the model release pipeline. Governance runs through C2 CoE-as-a-Service™ once multiple AI workloads share the infrastructure. B18 is the foundation; what runs on it is a separate engagement.
§ 11 · NAMED ACCOUNTABILITY

One name
on the engagement letter.

A named Practice Lead is accountable for delivery, commercial outcomes, and the client relationship throughout the engagement. Not a project manager who disappears after kickoff. Not a partner who nods at the SOW and vanishes.

THE ROLE

Practice Lead — AI

Present at every phase gate, every scope decision, every difficult conversation. Available for 30/60/90-day post-handover check-ins as part of the engagement.

SIX ACCOUNTABILITIES
01
Commercial arrangement

Including scope amendments.

02
Deliverables acceptance

Signs off all 7 deliverables.

03
Bi-weekly reviews

With executive sponsor.

04
Change orders

Authorised to negotiate.

05
Escalation path

CEO within 24 hours.

06
Post-handover

30/60/90-day check-ins.

§ 13 · BOOK A CLINIC

Thirty minutes.
No slide deck.

A structured 30-minute scope conversation with the Practice Lead. You describe the workloads blocked, the residency requirement, and the infrastructure available. We describe whether B18 is the right engagement — and if not, what is.

Book a clinic →Email directly
DURATION
30 minutes
PREPARATION
None required
FOLLOW-UP
Written scope, 5 business days