google-genai SDK: using the VIDAI Control Plane as the backend¶
The google-genai SDK speaks Gemini's native wire protocol. Point
it at the control plane via http_options.
TL;DR¶
from google import genai
from google.genai import types
client = genai.Client(
api_key="sk-your-vidai-key",
http_options=types.HttpOptions(base_url="https://your-vidaiserver.example.com"),
)
resp = client.models.generate_content(
model="gemini-2.5-flash",
contents="hello",
)
print(resp.text)
Prerequisites¶
- google-genai SDK
>=1.0(current tests run on1.73.1) - The control plane URL and an API key
- A Gemini model (or an alias) registered at the control plane
New project: google-genai¶
- Install the SDK:
- Create a client pointed at the control plane:
from google import genai from google.genai import types client = genai.Client( api_key="sk-your-vidai-key", http_options=types.HttpOptions( base_url="https://your-vidaiserver.example.com", ), ) resp = client.models.generate_content( model="gemini-2.5-flash", contents="summarize LangGraph in one sentence", ) print(resp.text)
Env var option¶
The SDK reads GOOGLE_API_KEY / GEMINI_API_KEY from the env for
the key, but has no env-var shortcut for base_url. Put the URL in
your own env var and read it at client creation:
export GOOGLE_API_KEY=sk-your-vidai-key
export VIDAI_SERVER_URL=https://your-vidaiserver.example.com
import os
from google import genai
from google.genai import types
client = genai.Client(
http_options=types.HttpOptions(base_url=os.environ["VIDAI_SERVER_URL"]),
)
Existing project: google-genai¶
Add http_options=... to your Client(...) and swap the API key:
# Before
client = genai.Client(api_key=os.environ["GEMINI_API_KEY"])
# After
client = genai.Client(
api_key=os.environ["GEMINI_API_KEY"], # now holds a VIDAI Server key
http_options=types.HttpOptions(
base_url="https://your-vidaiserver.example.com",
),
)
Everything else (generate_content, generate_content_stream,
chats, function-calling, multi-turn) works unchanged.
Cross-provider routing¶
Ops can register a Gemini-name model that routes to an OpenAI or Anthropic upstream. Your code doesn't change:
# Ops registered "gemini-2.5-flash" to route to an OpenAI upstream
# under the hood. You call it via the Gemini SDK normally:
resp = client.models.generate_content(
model="gemini-2.5-flash",
contents="hello",
)
The control plane translates your Gemini-shape request into OpenAI's wire
format, calls OpenAI, translates the response back to Gemini shape,
and your SDK sees a normal GenerateContentResponse.
Streaming¶
stream = client.models.generate_content_stream(
model="gemini-2.5-flash",
contents="count to five",
)
for chunk in stream:
if chunk.text:
print(chunk.text, end="", flush=True)
Function calling¶
from google.genai import types
tools = [
types.Tool(
function_declarations=[
types.FunctionDeclaration(
name="get_weather",
description="Get the current weather",
parameters=types.Schema(
type="OBJECT",
properties={"city": types.Schema(type="STRING")},
required=["city"],
),
)
]
)
]
resp = client.models.generate_content(
model="gemini-2.5-flash",
contents="Weather in Paris?",
config=types.GenerateContentConfig(tools=tools),
)
# Inspect resp.candidates[0].content.parts for function_call parts.
Reading the control plane's response headers¶
The control plane adds x-vidai-* headers to every response that indicate
what actually routed. Unlike the OpenAI and Anthropic SDKs, the
google-genai SDK doesn't expose a with_raw_response-style hook
for the high-level generate_content call, so the headers aren't
easily reachable from the typed API.
If you need the routing trail for debug dashboards, either:
- Use resp.model_version / usage metadata on the response, which
often reflects the actually-served model.
- Issue the probe requests via httpx directly (not a production
pattern) to inspect the raw headers.
See routing-and-headers.md for what the headers mean and which ones exist.
Verify it works: probe¶
from google import genai
from google.genai import types
client = genai.Client(
api_key="sk-your-vidai-key",
http_options=types.HttpOptions(base_url="https://your-vidaiserver.example.com"),
)
resp = client.models.generate_content(
model="gemini-2.5-flash",
contents="say: ok",
)
print("text:", resp.text)
print("usage:", resp.usage_metadata)
If this returns a GenerateContentResponse with text and usage,
your setup is correct. If it fails with an auth error, the key is
wrong. If the model is unknown, check with ops for the list of
registered models on your key.