Skip to content
EU AI Stack

The real TCO of a sovereign AI deployment over three years

Fryderyk Pryjma5 min read
Editorial infographic on a dark navy background comparing per token metered inference with per rack owned capacity in gold type

Over three years the cost gap between metered API inference and an owned sovereign deployment is decided by two numbers: sustained tokens per day and how much of your accelerator capacity sits idle. Below roughly a million tokens a day, per token pricing usually wins on every line. Above steady, predictable volume with utilisation over about 40 percent, owned or dedicated capacity is frequently cheaper by year two, and the sovereignty properties arrive as a side effect instead of a premium. The lines most budgets forget are egress, staff time, model refresh and exit, not the hardware.

Why do sovereign AI budgets get this wrong?

Almost every comparison I am handed puts GPU list price next to a per million token rate and calls it a total cost of ownership. It is not. It is one line out of nine, and it is the line that moves least over a three-year horizon. The costs that decide the outcome are the ones nobody has to sign a purchase order for: engineer hours, network egress, the model refresh you did not plan for in month eighteen, and the exit you will eventually need to price.

The second mistake is comparing a peak-load private cluster against average-load API spend. Metered inference is elastic by construction, so a fair comparison sizes owned capacity for the sustained load and treats bursts as either queued or overflowed to an API. Sizing for the worst hour of the quarter is what makes on-premise look absurd on paper.

What belongs in a three-year model?

Nine lines cover it. The table below is the one I use in cost reviews, kept deliberately short so it fits on a single page next to the architecture diagram.

Cost lineMetered APIEuropean managed hostingOwned on-premise
ComputePer token, scales to zero, no idle costMonthly reserved instance or dedicated GPU hourCapital cost amortised over 36 to 48 months
Utilisation riskNone, you pay for what you sendPartial, reservations are committedHigh, idle accelerators are pure loss
Egress and transferLow for text, material for documents and audioProvider dependent, check switching termsZero inside the building, relevant at replication
Operations staffLow, integration onlyModerate, shared responsibility model0.3 to 1 FTE for platform and lifecycle
Model refreshIncluded, and forced on your scheduleNegotiated per contract termYour project every 9 to 18 months
Compliance evidenceVendor documents plus your own DPIA and logsShared, provider supplies part of the packYours end to end, and reusable across audits
Exit costLow technically, high in prompt and tooling lock-inData Act switching terms apply from 12 January 2027Residual hardware value and decommissioning
Three-year cost lines for the three deployment models, with the item that usually surprises the budget owner.

Two lines in that table are worth arguing about internally. Operations staff is the one teams underestimate, because platform work does not stop after go-live. Model refresh is the one they forget entirely, and it is the line that turns a favourable owned-capacity model into a stalled one when nobody owns the upgrade.

Where is the break-even in practice?

Break-even is a utilisation question, not a headcount question. With one dedicated accelerator node running a mid-sized open-weight model at steady load, owned capacity typically overtakes metered pricing somewhere between month 14 and month 22, and the range is that wide because it depends almost entirely on whether the node runs at 20 percent or 60 percent average utilisation. Below 20 percent it never overtakes within the horizon.

The practical consequence is a hybrid answer more often than a pure one. Route the predictable, high-volume, sensitive workload onto owned or dedicated European capacity, and keep a metered endpoint for spikes, experiments and the models you are not ready to host. That shape wins the cost argument and the sovereignty argument at the same time, provided the routing rule is written down and enforced rather than left to whoever writes the next integration.

Sovereignty stops being a premium the moment your utilisation is high enough to make owned capacity the cheaper option anyway.

Which regulatory costs change the arithmetic?

Two instruments touch the money directly. The Data Act applies from 12 September 2025, and its switching provisions remove charges for cloud switching from 12 January 2027, with reduced egress charges in the transition. That matters because exit cost has historically been the strongest argument against moving workloads at all, and it is being priced down by law rather than by negotiation.

The Cyber Resilience Act adds the other side. Reporting duties for actively exploited vulnerabilities start on 11 September 2026, with the full product requirements from 11 December 2027. If you host your own appliance, that lifecycle work is yours to staff. If you buy managed capacity, it is a contract clause you must verify. Either way it is a real line in the model, and leaving it out is how a favourable comparison quietly stops being true.

Evidence work is the one cost that pays itself back. The residency diagram, access model and retention table you build for a sovereign deployment are the same artefacts a buyer's auditor requests under NIS2, so a single pack serves the cost model, the security questionnaire and the audit.

Frequently asked questions

Is on-premise AI cheaper than an API?
Only at sustained volume. With steady load and average accelerator utilisation above roughly 40 percent, owned capacity usually overtakes metered pricing during the second year. With bursty or low-volume usage, per token pricing stays cheaper across the whole three-year horizon because you never pay for idle hardware.
What is the most commonly missed cost line?
Operations staff and model refresh. Platform work continues after go-live, and hosted models need an upgrade project every 9 to 18 months. Budgets that include hardware and electricity but no named owner for those two lines understate owned capacity by a wide margin.
Does the Data Act reduce our exit cost?
Yes. The Data Act applies from 12 September 2025 and its switching provisions remove charges for cloud switching from 12 January 2027, with reduced egress in the transition period. Put the dated position in your model rather than a historical egress quote.
Should we run a hybrid setup instead of choosing one model?
In most regulated environments, yes. Predictable sensitive workloads go to owned or dedicated European capacity, spikes and experiments stay on a metered endpoint. The condition is a written routing rule, otherwise sensitive traffic drifts to whichever endpoint is easiest to call.
ShareLinkedInXEmail

Related articles

Next step

Need this as an outcome, not an article? Sovereign deployment.

Sovereignty is a buyer's problem stated as a hosting question. We answer it as an architecture decision: which layer must stay in your jurisdiction, what that costs over three years, and what an auditor accepts as proof.

Explore Sovereign deployment