How the VIDAI Control Plane routes your request: a customer-facing view¶
This page explains what the control plane is doing between your SDK and the real provider, so you can read response headers, write sane monitoring, and understand the canonical errors you'll see.
It's written for SDK users (app developers). The admin's view is a separate concern; ops configures the routes, you just call the models.
The two-name model¶
Every model you can call has two names inside the control plane:
| Name | Who sees it | Example |
|---|---|---|
| Public name | You, in your SDK code. Also appears in logs, dashboards, and billing. | gpt-4o |
| Upstream name | Only the real provider. You never write it. | gpt-4o-2024-11-20 |
Always call the model by its public name. Ops can change the upstream name (to pin a specific snapshot, to swap vendors, to rebalance cost) without touching your code.
Why this matters¶
Two real scenarios where aliasing saves you a deploy:
Version pinning. Ops points the public name gpt-4o at the
upstream gpt-4o-2024-11-20. Next quarter, they re-pin to
gpt-4o-2024-12-15. Your code keeps calling "gpt-4o".
Vendor switching. Ops creates a public model called
fast-model pointing at OpenAI's gpt-4o-mini. Six months later
Anthropic is cheaper for the same workload, so ops repoints
fast-model to Anthropic's claude-haiku-4-5. Your code still
calls "fast-model". The control plane translates the wire format
internally.
What the control plane does between your SDK and the provider¶
For every request, in this order:
- Auth: checks your API key.
- Access control: checks you're allowed to call this model
(
allowed_modelson your key, if configured). - Routing rules: ops can have active rules that rewrite the model or provider (e.g. "traffic from key X goes to the cheaper upstream"). You don't see this happen; you just get a response.
- Model resolution: the public name you asked for gets resolved to a provider + upstream name.
- Protocol translation: if the upstream speaks a different wire format than your SDK (OpenAI SDK hitting an Anthropic upstream, say), the control plane translates the request. After the response comes back, it translates that too.
- Circuit-breaker check: if recent requests to this provider have been failing, the control plane may route to a fallback instead.
- Fallback: if the primary is unhealthy and ops configured a fallback chain, the control plane tries the chain in order.
- Upstream call: the real provider gets hit.
- Response translation (if needed) and return to you.
Only steps 1, 2, 8, 9 are things you can directly observe. The rest happen on the server side.
Response headers: reading what actually happened¶
Every control plane response carries headers prefixed x-vidai-*.
These tell you what routed where. Useful for monitoring dashboards,
cost attribution, and debugging.
| Header | Meaning |
|---|---|
x-vidai-requested-model |
The public name your SDK asked for |
x-vidai-model |
The public name that actually served the request (differs from requested when a routing rule or fallback fired) |
x-vidai-fallback |
"true" if fallback fired for this request |
x-vidai-fallback-from |
If fallback fired, the name of the provider that failed |
x-vidai-routing-rule |
If a routing rule fired, its ID |
x-vidai-provider-key |
Which provider key served the request (multi-key setups) |
Note: the public name in x-vidai-model is still the name you
can refer to in ops discussions. If you want to see what the real
provider got on the wire (the upstream name after aliasing), that
lives in the control plane's usage logs; ask ops for access to those
if you need to build cost dashboards.
Reading headers from each SDK¶
OpenAI SDK: use with_raw_response:
raw = client.chat.completions.with_raw_response.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "hi"}],
)
print(raw.headers.get("x-vidai-fallback"))
parsed = raw.parse() # then work with the typed response
Anthropic SDK: use with_raw_response:
raw = client.messages.with_raw_response.create(
model="claude-haiku-4-5",
max_tokens=20,
messages=[{"role": "user", "content": "hi"}],
)
print(raw.headers.get("x-vidai-fallback"))
parsed = raw.parse()
google-genai SDK: headers aren't exposed on the high-level
generate_content call. For now, if you need the routing trail,
either:
- Inspect resp.model_version / usage info, which often reflects
the served upstream model, or
- Drop to httpx for debug requests (not a production pattern).
LangChain: the SDK response headers aren't surfaced through
LangChain's AIMessage type. For debugging, either call the
vendor SDK directly (bypass LangChain briefly), or read the
response_metadata on the returned message. Some wrappers
populate it with provider response metadata, which may or may not
include the vidai headers depending on version.
LangGraph: same as LangChain; state updates don't propagate raw headers. Use the underlying ChatModel's debug hooks if you need this level of detail.
Canonical errors: what you'll see when something goes wrong¶
| Status | Cause | What to check |
|---|---|---|
401 AuthenticationError |
Missing / invalid API key | Authorization: Bearer <key> is present and the key hasn't been revoked |
403 ModelNotAllowed |
Your key isn't authorised to call this model | Ask ops for the list of allowed_models on your key |
404 ModelNotFound |
The model name isn't registered in the control plane (or your key isn't allowed to see it) | Typo in the model name, or ops hasn't registered this alias for you |
429 RateLimitError |
RPM / TPM / concurrent-request limit tripped | Wait (check Retry-After header) or ask ops to raise your limit |
| 500 Upstream error surfacing | The real provider returned an error and no fallback is configured | If you expected fallback to fire, check with ops whether a fallback chain is configured for this model + your key |
All of these surface as the native exception type of your SDK
(openai.AuthenticationError, anthropic.NotFoundError, etc.),
so your existing error-handling code keeps working.
Cross-provider: what's supported today¶
See the matrix in Client integrations. Short version:
- From the OpenAI SDK, you can hit any upstream (OpenAI, Anthropic, Gemini). The control plane translates the response back to OpenAI shape.
- From the google-genai SDK, same: any upstream, Gemini-shape response back.
- From the Anthropic SDK, only Anthropic upstreams work today. Cross-provider from Anthropic SDK is on the roadmap.
- From Google ADK: the
Geminiclass (orLiteLlm("openai/…")) can reach any upstream.AnthropicLlm/Claudeonly reaches Anthropic.
If fallback fires during your request¶
From your SDK's point of view, a fallback is invisible: you
get a normal successful response. The only signal is in the
response headers (x-vidai-fallback: true). Your SDK code doesn't
need any retry logic for this.
If the primary AND every fallback entry are unhealthy, the control plane falls back to trying the primary anyway (the "optimistic primary" rule: better to try than hard-reject). If everything really is down, you'll get the primary's error surfaced to your SDK as if no fallback was configured.
Further reading¶
- Per-SDK guides: Client integrations links to each one.
- Cross-provider matrix: also in Client integrations.
- Provider catalogue: Provider catalogue for the upstream side (where to grab API keys, what each preset configures).