Skip to content

How the VIDAI Control Plane routes your request: a customer-facing view

This page explains what the control plane is doing between your SDK and the real provider, so you can read response headers, write sane monitoring, and understand the canonical errors you'll see.

It's written for SDK users (app developers). The admin's view is a separate concern; ops configures the routes, you just call the models.

The two-name model

Every model you can call has two names inside the control plane:

Name Who sees it Example
Public name You, in your SDK code. Also appears in logs, dashboards, and billing. gpt-4o
Upstream name Only the real provider. You never write it. gpt-4o-2024-11-20

Always call the model by its public name. Ops can change the upstream name (to pin a specific snapshot, to swap vendors, to rebalance cost) without touching your code.

Why this matters

Two real scenarios where aliasing saves you a deploy:

Version pinning. Ops points the public name gpt-4o at the upstream gpt-4o-2024-11-20. Next quarter, they re-pin to gpt-4o-2024-12-15. Your code keeps calling "gpt-4o".

Vendor switching. Ops creates a public model called fast-model pointing at OpenAI's gpt-4o-mini. Six months later Anthropic is cheaper for the same workload, so ops repoints fast-model to Anthropic's claude-haiku-4-5. Your code still calls "fast-model". The control plane translates the wire format internally.

What the control plane does between your SDK and the provider

For every request, in this order:

  1. Auth: checks your API key.
  2. Access control: checks you're allowed to call this model (allowed_models on your key, if configured).
  3. Routing rules: ops can have active rules that rewrite the model or provider (e.g. "traffic from key X goes to the cheaper upstream"). You don't see this happen; you just get a response.
  4. Model resolution: the public name you asked for gets resolved to a provider + upstream name.
  5. Protocol translation: if the upstream speaks a different wire format than your SDK (OpenAI SDK hitting an Anthropic upstream, say), the control plane translates the request. After the response comes back, it translates that too.
  6. Circuit-breaker check: if recent requests to this provider have been failing, the control plane may route to a fallback instead.
  7. Fallback: if the primary is unhealthy and ops configured a fallback chain, the control plane tries the chain in order.
  8. Upstream call: the real provider gets hit.
  9. Response translation (if needed) and return to you.

Only steps 1, 2, 8, 9 are things you can directly observe. The rest happen on the server side.

Response headers: reading what actually happened

Every control plane response carries headers prefixed x-vidai-*. These tell you what routed where. Useful for monitoring dashboards, cost attribution, and debugging.

Header Meaning
x-vidai-requested-model The public name your SDK asked for
x-vidai-model The public name that actually served the request (differs from requested when a routing rule or fallback fired)
x-vidai-fallback "true" if fallback fired for this request
x-vidai-fallback-from If fallback fired, the name of the provider that failed
x-vidai-routing-rule If a routing rule fired, its ID
x-vidai-provider-key Which provider key served the request (multi-key setups)

Note: the public name in x-vidai-model is still the name you can refer to in ops discussions. If you want to see what the real provider got on the wire (the upstream name after aliasing), that lives in the control plane's usage logs; ask ops for access to those if you need to build cost dashboards.

Reading headers from each SDK

OpenAI SDK: use with_raw_response:

raw = client.chat.completions.with_raw_response.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "hi"}],
)
print(raw.headers.get("x-vidai-fallback"))
parsed = raw.parse()   # then work with the typed response

Anthropic SDK: use with_raw_response:

raw = client.messages.with_raw_response.create(
    model="claude-haiku-4-5",
    max_tokens=20,
    messages=[{"role": "user", "content": "hi"}],
)
print(raw.headers.get("x-vidai-fallback"))
parsed = raw.parse()

google-genai SDK: headers aren't exposed on the high-level generate_content call. For now, if you need the routing trail, either: - Inspect resp.model_version / usage info, which often reflects the served upstream model, or - Drop to httpx for debug requests (not a production pattern).

LangChain: the SDK response headers aren't surfaced through LangChain's AIMessage type. For debugging, either call the vendor SDK directly (bypass LangChain briefly), or read the response_metadata on the returned message. Some wrappers populate it with provider response metadata, which may or may not include the vidai headers depending on version.

LangGraph: same as LangChain; state updates don't propagate raw headers. Use the underlying ChatModel's debug hooks if you need this level of detail.

Canonical errors: what you'll see when something goes wrong

Status Cause What to check
401 AuthenticationError Missing / invalid API key Authorization: Bearer <key> is present and the key hasn't been revoked
403 ModelNotAllowed Your key isn't authorised to call this model Ask ops for the list of allowed_models on your key
404 ModelNotFound The model name isn't registered in the control plane (or your key isn't allowed to see it) Typo in the model name, or ops hasn't registered this alias for you
429 RateLimitError RPM / TPM / concurrent-request limit tripped Wait (check Retry-After header) or ask ops to raise your limit
500 Upstream error surfacing The real provider returned an error and no fallback is configured If you expected fallback to fire, check with ops whether a fallback chain is configured for this model + your key

All of these surface as the native exception type of your SDK (openai.AuthenticationError, anthropic.NotFoundError, etc.), so your existing error-handling code keeps working.

Cross-provider: what's supported today

See the matrix in Client integrations. Short version:

  • From the OpenAI SDK, you can hit any upstream (OpenAI, Anthropic, Gemini). The control plane translates the response back to OpenAI shape.
  • From the google-genai SDK, same: any upstream, Gemini-shape response back.
  • From the Anthropic SDK, only Anthropic upstreams work today. Cross-provider from Anthropic SDK is on the roadmap.
  • From Google ADK: the Gemini class (or LiteLlm("openai/…")) can reach any upstream. AnthropicLlm / Claude only reaches Anthropic.

If fallback fires during your request

From your SDK's point of view, a fallback is invisible: you get a normal successful response. The only signal is in the response headers (x-vidai-fallback: true). Your SDK code doesn't need any retry logic for this.

If the primary AND every fallback entry are unhealthy, the control plane falls back to trying the primary anyway (the "optimistic primary" rule: better to try than hard-reject). If everything really is down, you'll get the primary's error surfaced to your SDK as if no fallback was configured.

Further reading