Ask a vendor whether their AI runs in the EU and you will get a yes. Ask where a single agent request goes, hop by hop, and the yes starts to need footnotes. The gap between those two answers is where most sovereignty claims quietly fail.
This is not usually dishonesty. It is scope. “Our inference runs in Frankfurt” is true and verifiable and describes one hop of a chain that has five or six. Everybody involved is answering the question they were asked.
So here is the longer question: for one agent request, what leaves your perimeter, where does it land, and who could be compelled to hand it over?
The six hops
A single agent turn that feels like one operation is a distributed transaction across several vendors. Each one is a jurisdiction.
Hop 1 — the model call
The one everyone checks, and usually the one that is fine. If you have picked an EU region or an EU-headquartered provider, the prompt and completion stay where you think.
Two caveats that survive a regional selection. Fallback routing: many gateways fail over to another region on capacity or timeout, and the failover region is frequently not the one in the contract. Control plane versus data plane: your inference may run in Frankfurt while the account, keys and usage metadata live in a US control plane. The second is often out of scope in the compliance answer and in scope for a regulator.
Hop 2 — embeddings
Retrieval means embedding, and embedding is a separate model call to a possibly separate vendor. Teams that carefully pinned their chat model to an EU region routinely leave embeddings on whatever the SDK defaulted to.
The content is the same content. If the document was confidential when the chat model saw it, it was confidential when the embedding model saw it. Vector databases deserve the same question and rarely get it.
Hop 3 — the tools
This is the hop people forget entirely, and it is the one with the widest spread.
An agent that searches the web has sent your query to a search provider. An agent that reads a document has touched a storage vendor. An agent that files a ticket has written customer data into a SaaS product with its own region, its own subprocessors and its own retention policy.
The governance layer can be perfectly EU-hosted and the agent can still push data to eight services in one turn, because calling tools is what the agent is for.
Hop 4 — observability
Traces, spans and prompt logs. LLM observability platforms exist to capture inputs and outputs, which means the observability vendor holds a copy of exactly the data you were careful about, often in plaintext, usually in a US region, frequently on a longer retention than the primary system.
This is the most common single failure I see. The pipeline is clean, the audit is clean, and the debugging tool that someone added in week two has a full transcript archive.
Hop 5 — evaluation
Sampled production traffic scored by a judge model, or exported to an annotation vendor, or fed back as training data. Every one of those is a copy leaving the perimeter, and evals are usually set up by a different team from the one that answered the DPIA.
Hop 6 — the error path
When something fails, the payload frequently goes somewhere it was never meant to: a crash reporter, an APM tool, a Slack alert with the request body attached, a support ticket with a transcript pasted in. The happy path was reviewed. The failure path was written at 2am.
One request, traced
A claims assistant at an insurer. A handler asks it to summarise a claim and draft a response. One sentence of work. Here is the trace.
- The claim reference is embedded to search internal policy documents. Claim text leaves for the embedding provider.
- The vector store returns passages. Wherever that store runs, it now holds vectors derived from claim data.
- The model call goes to the EU region, correctly, with the claim and the retrieved passages in the prompt.
- The agent calls a document tool to pull the original PDF from a storage vendor.
- It calls the claims system to check status, writing the query into that vendor's logs.
- Every step emits a trace containing the full prompt, including the policyholder's name and the claim detail.
- The turn is sampled for evaluation and scored by a judge model.
Seven external touches. One of them — hop three — is the one the compliance answer described. The policyholder's data is now, legitimately and contractually, at rest in several places, at least one of which is a debugging tool that nobody listed on the DPIA because it was added by an engineer solving a different problem.
Nothing here was negligent. Each piece was a reasonable engineering decision. The failure is that no one was counting.
What to ask, and what a good answer sounds like
These are the questions that separate a traced system from an assumed one. They are also the questions worth answering about your own stack before a customer asks.
“Where does the embedding call go?” A good answer names a provider and a region without checking. A weak answer is “same as the model”, which is true surprisingly rarely.
“What is in your traces, and for how long?” A good answer distinguishes metadata from content and states a retention period. If prompts are captured in full, that vendor is processing whatever your prompts contain.
“What happens on failover?” A good answer says failover stays in region and the request fails if it cannot. A weak answer says it has not come up.
“Which tools can the agent reach, and can you enumerate them?” If the answer is a list, the system is constrained. If the answer describes what the agent is supposed to do, it is not.
“Where is the audit log?” Often the most revealing one, because the record is the most complete artefact in the system and it is routinely hosted somewhere nobody checked.
Why the paperwork does not catch this
Most of these hops are contracted, covered by a DPA, and invisible to the person answering the questionnaire. The architecture diagram shows a box marked “LLM” and the agent's actual behaviour is emergent: which tools it calls and in what order depends on what it decides, turn by turn.
That is a genuinely new problem. You can review a static integration. You cannot review, in advance, the set of tools an autonomous system will choose to chain, because the whole point is that it chooses.
The honest conclusion is that a design-time answer cannot settle a run-time question. You can either constrain what the agent may reach, or you can find out afterwards where it went.
What actually holds
Enumerate at the boundary, not on the diagram. The list of destinations an agent can reach should be a configured set, not an emergent one. If a tool is not on the list, the call does not happen. This is the only version of the question that has a checkable answer.
Make region a policy, not a default. “This agent may only call models and tools in these regions” is a rule that can be enforced per call and logged. A dropdown selected once during setup is not.
Redact before the hop, not after. Stripping PII at the boundary means the downstream vendor never held it, which is a materially different statement to a regulator than a deletion request after the fact.
Include observability and evals in scope. If your trace vendor holds prompts, it is a processor of whatever was in those prompts. Treat it like one.
Keep the record where you keep the data. An EU-hosted pipeline with a US-hosted audit log has moved the evidence, and the evidence is frequently the most sensitive artefact in the system, because it is the part that is complete.
The question worth asking
Not “is your AI EU-hosted”. It will be answered yes and the answer will be true and it will not tell you much.
Ask instead: for one request, list every third party that receives any part of it, including telemetry, evaluation and the error path, with the region for each.
A vendor who can answer that has done the work. A vendor who cannot has told you something useful too: not that they are careless, but that nobody has asked them to trace it, which means nobody has traced it.
Sovereignty is not a hosting region. It is knowing every hop and being able to stop the ones you did not authorise, in the moment, rather than discovering them in an incident review.