Skip to content

google-genai SDK: using the VIDAI Control Plane as the backend

The google-genai SDK speaks Gemini's native wire protocol. Point it at the control plane via http_options.

TL;DR

from google import genai
from google.genai import types

client = genai.Client(
    api_key="sk-your-vidai-key",
    http_options=types.HttpOptions(base_url="https://your-vidaiserver.example.com"),
)
resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="hello",
)
print(resp.text)

Prerequisites

  • google-genai SDK >=1.0 (current tests run on 1.73.1)
  • The control plane URL and an API key
  • A Gemini model (or an alias) registered at the control plane

New project: google-genai

  1. Install the SDK:
    pip install google-genai
    
  2. Create a client pointed at the control plane:
    from google import genai
    from google.genai import types
    
    client = genai.Client(
        api_key="sk-your-vidai-key",
        http_options=types.HttpOptions(
            base_url="https://your-vidaiserver.example.com",
        ),
    )
    
    resp = client.models.generate_content(
        model="gemini-2.5-flash",
        contents="summarize LangGraph in one sentence",
    )
    print(resp.text)
    

Env var option

The SDK reads GOOGLE_API_KEY / GEMINI_API_KEY from the env for the key, but has no env-var shortcut for base_url. Put the URL in your own env var and read it at client creation:

export GOOGLE_API_KEY=sk-your-vidai-key
export VIDAI_SERVER_URL=https://your-vidaiserver.example.com
import os
from google import genai
from google.genai import types

client = genai.Client(
    http_options=types.HttpOptions(base_url=os.environ["VIDAI_SERVER_URL"]),
)

Existing project: google-genai

Add http_options=... to your Client(...) and swap the API key:

# Before
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])

# After
client = genai.Client(
    api_key=os.environ["GEMINI_API_KEY"],   # now holds a VIDAI Server key
    http_options=types.HttpOptions(
        base_url="https://your-vidaiserver.example.com",
    ),
)

Everything else (generate_content, generate_content_stream, chats, function-calling, multi-turn) works unchanged.

Cross-provider routing

Ops can register a Gemini-name model that routes to an OpenAI or Anthropic upstream. Your code doesn't change:

# Ops registered "gemini-2.5-flash" to route to an OpenAI upstream
# under the hood. You call it via the Gemini SDK normally:
resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="hello",
)

The control plane translates your Gemini-shape request into OpenAI's wire format, calls OpenAI, translates the response back to Gemini shape, and your SDK sees a normal GenerateContentResponse.

Streaming

stream = client.models.generate_content_stream(
    model="gemini-2.5-flash",
    contents="count to five",
)
for chunk in stream:
    if chunk.text:
        print(chunk.text, end="", flush=True)

Function calling

from google.genai import types

tools = [
    types.Tool(
        function_declarations=[
            types.FunctionDeclaration(
                name="get_weather",
                description="Get the current weather",
                parameters=types.Schema(
                    type="OBJECT",
                    properties={"city": types.Schema(type="STRING")},
                    required=["city"],
                ),
            )
        ]
    )
]

resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="Weather in Paris?",
    config=types.GenerateContentConfig(tools=tools),
)
# Inspect resp.candidates[0].content.parts for function_call parts.

Reading the control plane's response headers

The control plane adds x-vidai-* headers to every response that indicate what actually routed. Unlike the OpenAI and Anthropic SDKs, the google-genai SDK doesn't expose a with_raw_response-style hook for the high-level generate_content call, so the headers aren't easily reachable from the typed API.

If you need the routing trail for debug dashboards, either: - Use resp.model_version / usage metadata on the response, which often reflects the actually-served model. - Issue the probe requests via httpx directly (not a production pattern) to inspect the raw headers.

See routing-and-headers.md for what the headers mean and which ones exist.

Verify it works: probe

from google import genai
from google.genai import types

client = genai.Client(
    api_key="sk-your-vidai-key",
    http_options=types.HttpOptions(base_url="https://your-vidaiserver.example.com"),
)
resp = client.models.generate_content(
    model="gemini-2.5-flash",
    contents="say: ok",
)
print("text:", resp.text)
print("usage:", resp.usage_metadata)

If this returns a GenerateContentResponse with text and usage, your setup is correct. If it fails with an auth error, the key is wrong. If the model is unknown, check with ops for the list of registered models on your key.