Skip to content

OpenAI Python SDK: using the VIDAI Control Plane as the backend

The OpenAI SDK is the easiest path. Two env vars and your existing code works unchanged.

TL;DR

export OPENAI_API_BASE=https://your-vidaiserver.example.com/v1
export OPENAI_API_KEY=sk-your-vidai-key
from openai import OpenAI
client = OpenAI()   # picks up the env vars automatically
resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "hello"}],
)

That's it. The SDK talks to the control plane instead of api.openai.com.

Prerequisites

  • openai Python SDK >=1.0 (any modern version works; current tests run on 2.x)
  • The control plane URL and an API key (get from your ops team)

New project: OpenAI SDK

  1. Install the SDK:
    pip install openai
    
  2. Set env vars:
    export OPENAI_API_BASE=https://your-vidaiserver.example.com/v1
    export OPENAI_API_KEY=sk-your-vidai-key
    
  3. Write code as if you were talking to OpenAI:
    from openai import OpenAI
    client = OpenAI()
    resp = client.chat.completions.create(
        model="gpt-4o-mini",
        messages=[{"role": "user", "content": "summarize LangGraph in one sentence"}],
        max_tokens=120,
    )
    print(resp.choices[0].message.content)
    

Passing config explicitly instead of env vars

If you can't set env vars (shared process, tests, etc.), pass the same values as constructor kwargs:

client = OpenAI(
    base_url="https://your-vidaiserver.example.com/v1",
    api_key="sk-your-vidai-key",
)

Existing project: OpenAI SDK

If your code already uses the OpenAI SDK, you usually don't need to change a line of Python. Just change the env vars:

# Before
export OPENAI_API_KEY=sk-openai-...        # your real OpenAI key

# After
export OPENAI_API_BASE=https://your-vidaiserver.example.com/v1
export OPENAI_API_KEY=sk-your-vidai-key     # your VIDAI Server-issued key

If your code constructs the client with explicit api_key= / base_url= (e.g. because you have multiple clients in one process), update those two kwargs:

# Before
client = OpenAI(api_key=os.environ["OPENAI_API_KEY"])

# After
client = OpenAI(
    base_url="https://your-vidaiserver.example.com/v1",
    api_key=os.environ["OPENAI_API_KEY"],  # now holds a VIDAI Server key
)

Everything else (streaming, tools, JSON mode, vision, embeddings, multi-turn, retry behaviour) works exactly as before.

Cross-provider routing

Ops can register a model in the control plane that routes to a non-OpenAI upstream (Anthropic, Gemini, etc.). From the OpenAI SDK, you don't see this; you just call the model by its registered name:

# Ops registered "claude-haiku-4-5" on an Anthropic upstream.
# Your OpenAI SDK code calls it the same way.
resp = client.chat.completions.create(
    model="claude-haiku-4-5",
    messages=[{"role": "user", "content": "hello"}],
    max_tokens=100,
)

The control plane translates your OpenAI-shape request into Anthropic's wire format, calls Anthropic, translates the response back, and returns a standard OpenAI ChatCompletion to your SDK. Your code sees a normal OpenAI response.

The same pattern works for Gemini upstreams.

Streaming

Streaming works unchanged:

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "count to five"}],
    stream=True,
    stream_options={"include_usage": True},
)
for chunk in stream:
    delta = chunk.choices[0].delta if chunk.choices else None
    if delta and delta.content:
        print(delta.content, end="", flush=True)

Reading the control plane's response headers

The control plane adds x-vidai-* headers to every response that tell you what routed where (which model actually served, whether fallback fired, etc.). Use with_raw_response to read them:

raw = client.chat.completions.with_raw_response.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "hi"}],
)
print("served by:", raw.headers.get("x-vidai-model"))
print("fallback?", raw.headers.get("x-vidai-fallback"))
parsed = raw.parse()   # typed ChatCompletion

See routing-and-headers.md for the full header reference and what each one means.

Verify it works: 5-line probe

from openai import OpenAI
client = OpenAI(
    base_url="https://your-vidaiserver.example.com/v1",
    api_key="sk-your-vidai-key",
)
resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "say: ok"}],
    max_tokens=5,
)
print("status:", resp.choices[0].finish_reason)
print("body:", resp.choices[0].message.content)

If this returns a response, your setup is correct. If it fails with AuthenticationError (401), the key is wrong. If it fails with NotFoundError (404), the model isn't registered in the control plane; ask ops which models are available for your key.