Skip to content

OpenAI Responses API: using the VIDAI Control Plane as the backend

The Responses API is OpenAI's newer stateful call shape (single request → structured output + tool calls + reasoning). The control plane accepts it on the same base URL as the Chat Completions endpoint — set your OpenAI client's base_url and key and use client.responses.create(...) unchanged.

TL;DR

from openai import OpenAI

client = OpenAI(
    base_url="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)

resp = client.responses.create(
    model="gpt-4o-mini",
    input="reply with exactly: ok",
)
print(resp.output_text)

Prerequisites

  • pip install "openai>=1.60" (Responses API needs the newer client).
  • Control plane base URL and an API key from API Keys.
  • A model registered on the Models page that supports the Responses API on its upstream.

Tools

Attach tools the Responses way (tools=[{...}]) and the model uses them across turns. Every tool round-trip is a control-plane request:

resp = client.responses.create(
    model="gpt-4o-mini",
    input=[{"role": "user", "content": "What's the weather in Kyoto?"}],
    tools=[{
        "type": "function",
        "name": "get_weather",
        "parameters": {"type": "object", "properties": {"city": {"type": "string"}}},
    }],
)

Structured output

response_format={"type": "json_schema", ...} works the same way as on Chat Completions. The control plane forwards unchanged.

Streaming

with client.responses.stream(
    model="gpt-4o-mini",
    input="count to five",
) as stream:
    for event in stream:
        if event.type == "response.output_text.delta":
            print(event.delta, end="", flush=True)

Continuation via previous_response_id

The Responses API supports server-side continuation with previous_response_id. The control plane forwards this transparently as long as your model + upstream combination supports it — see the OpenAI docs for the applicable model list.

Verify it works

from openai import OpenAI
client = OpenAI(
    base_url="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)
print(client.responses.create(model="gpt-4o-mini", input="reply with exactly: ok").output_text)

If something's off

Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with the request body and the model name. We'll get it sorted.

Where to go next