Skip to content

Anthropic Python SDK: using the VIDAI Control Plane as the backend

The Anthropic SDK supports a base_url kwarg that points it at the control plane. One-line change from existing code.

TL;DR

import anthropic
client = anthropic.Anthropic(
    base_url="https://your-vidaiserver.example.com",
    api_key="sk-your-vidai-key",
)
msg = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=100,
    messages=[{"role": "user", "content": "hello"}],
)
print(msg.content[0].text)

Prerequisites

  • anthropic Python SDK >=0.40 (current tests run on 0.96.0)
  • The control plane URL and an API key
  • A Claude-family model registered at the control plane (claude-haiku-4-5, etc.)

New project: Anthropic SDK

  1. Install the SDK:
    pip install anthropic
    
  2. Write code pointing at the control plane:
    import anthropic
    
    client = anthropic.Anthropic(
        base_url="https://your-vidaiserver.example.com",
        api_key="sk-your-vidai-key",
    )
    
    msg = client.messages.create(
        model="claude-haiku-4-5",
        max_tokens=200,
        messages=[{"role": "user", "content": "summarize LangGraph in one sentence"}],
    )
    print(msg.content[0].text)
    

Env var option

The Anthropic SDK reads ANTHROPIC_API_KEY from the environment, but does NOT have a standard env var for base_url. If you want to keep base_url out of your code:

export ANTHROPIC_API_KEY=sk-your-vidai-key
client = anthropic.Anthropic(
    base_url=os.environ["VIDAI_SERVER_URL"],
)

Existing project: Anthropic SDK

If your code already uses anthropic.Anthropic(...), change two things:

# Before
client = anthropic.Anthropic(api_key=os.environ["ANTHROPIC_API_KEY"])

# After
client = anthropic.Anthropic(
    base_url="https://your-vidaiserver.example.com",
    api_key=os.environ["ANTHROPIC_API_KEY"],   # now holds a VIDAI Server key
)

Nothing else changes; streaming, tool use, vision, multi-turn, prompt caching, and extended thinking all behave as they did against api.anthropic.com.

Cross-provider routing

The Anthropic SDK works against any upstream the control plane has configured: Claude on Anthropic / Bedrock / Vertex natively, and OpenAI / Gemini upstreams via wire-format translation. Your code is the same; ops sets the alias and target upstream when registering the model. See the cross-provider matrix in Client integrations for the combinations.

Streaming

Works unchanged:

with client.messages.stream(
    model="claude-haiku-4-5",
    max_tokens=200,
    messages=[{"role": "user", "content": "count to five"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)
    final = stream.get_final_message()
    print("\nusage:", final.usage)

Tool use

Works unchanged:

tools = [
    {
        "name": "get_weather",
        "description": "Get current weather for a city",
        "input_schema": {
            "type": "object",
            "properties": {"city": {"type": "string"}},
            "required": ["city"],
        },
    }
]

msg = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=200,
    tools=tools,
    messages=[{"role": "user", "content": "What's the weather in Paris?"}],
)
# Inspect msg.content for `tool_use` blocks, execute the tool, and
# send the result back as a follow-up user message with `tool_result`.

Reading the control plane's response headers

The control plane adds x-vidai-* headers that tell you what actually routed (model served, fallback fired, etc.). Read them via with_raw_response:

raw = client.messages.with_raw_response.create(
    model="claude-haiku-4-5",
    max_tokens=20,
    messages=[{"role": "user", "content": "hi"}],
)
print("served by:", raw.headers.get("x-vidai-model"))
print("fallback?", raw.headers.get("x-vidai-fallback"))
parsed = raw.parse()   # typed Message

See routing-and-headers.md for the full header list.

Verify it works: 5-line probe

import anthropic
client = anthropic.Anthropic(
    base_url="https://your-vidaiserver.example.com",
    api_key="sk-your-vidai-key",
)
msg = client.messages.create(
    model="claude-haiku-4-5",
    max_tokens=10,
    messages=[{"role": "user", "content": "say: ok"}],
)
print("stop_reason:", msg.stop_reason)
print("content:", msg.content[0].text)

If this returns a Message with content, your setup is correct. If it fails with AuthenticationError, the key is wrong. If NotFoundError, the model isn't registered in the control plane; check with ops.