Skip to content

Guidance: using the VIDAI Control Plane as the backend

Guidance's model constructors accept base_url and api_key. Set them to the control plane's values and every guidance program you evaluate (including constrained grammars, tool calls, and multi-turn state) runs through it.

TL;DR

import guidance
from guidance import models, gen

lm = models.OpenAI(
    "gpt-4o-mini",
    base_url="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)

lm += "Reply with exactly: " + gen("answer", max_tokens=5)
print(lm["answer"])

Prerequisites

  • pip install guidance (0.1+ works; older versions use a different constructor shape).
  • Control plane base URL and an API key from API Keys.
  • A model registered on the Models page.

Constrained generation

Guidance's grammar constraints (select, gen(regex=...), json) run on the model's tokens directly. When the model is served through a remote endpoint like the control plane, some constraints work via the underlying model's structured-output mode and some are skipped — Guidance falls back gracefully. The behaviour matches what happens when you point Guidance at api.openai.com.

from guidance import select

lm += "Sentiment: " + select(["positive", "negative", "neutral"], name="s")

JSON output

The JSON constraint works transparently through the model's response_format support:

from guidance import json as g_json

lm += "Extract fields: " + g_json(
    name="row",
    schema={
        "type": "object",
        "properties": {"city": {"type": "string"}, "count": {"type": "integer"}},
        "required": ["city", "count"],
    },
)
print(lm["row"])

Anthropic models

For Claude-family models registered on the control plane, use Guidance's Anthropic constructor. Note the Anthropic base URL is the root (no /v1):

from guidance import models

lm = models.Anthropic(
    "claude-haiku-4-5",
    base_url="https://your-vidai-server.example.com",
    api_key="your-vidai-key",
)

Verify it works

from guidance import models, gen

lm = models.OpenAI(
    "gpt-4o-mini",
    base_url="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)
lm += "Reply with exactly: " + gen("answer", max_tokens=5)
print(lm["answer"])

A row appears on Request Logs.

If something's off

Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with the guidance program that reproduces the issue. We'll get it sorted.

Where to go next