Databricks Genie: using the VIDAI Control Plane as the backend¶
You reach the control plane from Databricks by registering it as
an External Model endpoint in Model Serving. Once that
endpoint exists, Genie, ai_query(), the AI Playground, and any
notebook that speaks to a Databricks-served endpoint all route
through the control plane.
TL;DR¶
Serving → Create serving endpoint → External model:
- Provider:
OpenAI - Task:
Chat - openai_api_base:
https://your-vidai-server.example.com/v1 - openai_api_key: your VIDAI API key
- Served entity name: the model you want to expose (e.g.
gpt-4o-mini)
Then in a notebook:
Prerequisites¶
- A Databricks workspace with permission to create Model Serving endpoints.
- The control plane's base URL and an API key from the API Keys page.
- A model registered on the Models page —
whatever name you give it here is what your
ai_query()/ Playground / Mosaic AI calls will use.
Register the endpoint¶
In your Databricks workspace go to Serving → Create serving endpoint and pick External model.
Name — anything memorable, e.g. vidai-control-plane. This
is the string you'll pass to ai_query() and to the OpenAI
SDK's base_url when writing agents.
Provider — OpenAI. The control plane's /v1/chat/completions
endpoint speaks the OpenAI wire shape, which is what this
provider type expects.
Task — Chat.
Served entities — add one entity per model you want callable from Databricks. The name of the entity is what Databricks callers will pass as the model name; it should match a model name registered on the control plane's Models page (or an alias ops has set up).
Provider config:
openai_api_base: https://your-vidai-server.example.com/v1
openai_api_key: "{{secrets/your-scope/vidai_api_key}}"
Use a Databricks secret scope for the key in production
(openai_api_key_plaintext is fine for a first-time smoke test
but Databricks warns against it).
Leave openai_api_type and openai_api_version blank — the
control plane serves the OpenAI wire shape on this path, not
Azure OpenAI.
Save. Databricks provisions the endpoint in about 30 seconds.
Alternative: the Custom provider type¶
If your Databricks reporting expects the endpoint to be labelled
as something other than "OpenAI", use provider type Custom
instead:
custom_provider_url: https://your-vidai-server.example.com/v1/chat/completions
authentication:
bearer_token:
token: "{{secrets/your-scope/vidai_api_key}}"
The wire shape and behaviour are identical; only the Databricks- side label changes.
Call it from ai_query()¶
The SQL function is the shortest path from a notebook to your new endpoint:
SELECT ai_query(
'vidai-control-plane',
'Classify this ticket as URGENT, NORMAL, or LOW: ' || ticket_body
) AS priority
FROM support_tickets
WHERE created_at >= current_date() - 7;
The first argument is the endpoint name you registered. The
second is the prompt. See the Databricks ai_query docs for
advanced shapes (structured output, temperature, max tokens).
Call it from an OpenAI-SDK notebook (Mosaic AI Agent Framework)¶
Mosaic AI agents and any Python notebook that uses the OpenAI SDK reach the endpoint via its Databricks-side invocation URL:
from databricks.sdk.runtime import spark
from openai import OpenAI
workspace_host = spark.conf.get("spark.databricks.workspaceUrl")
databricks_token = dbutils.notebook.entry_point.getDbutils() \
.notebook().getContext().apiToken().get()
client = OpenAI(
base_url=f"https://{workspace_host}/serving-endpoints/vidai-control-plane/invocations",
api_key=databricks_token,
)
resp = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "hello from a notebook"}],
)
print(resp.choices[0].message.content)
The model= you pass is what the control plane resolves — so
if ops has registered gpt-4o-mini as an alias that routes to
Claude Haiku, that's what will serve.
Call it from Genie¶
Genie itself has no LLM configuration screen. It uses whichever Model Serving endpoint your workspace admin has configured for Genie usage. Once your VIDAI endpoint is registered, the workspace admin points Genie at it via the standard Databricks Genie configuration flow — no VIDAI-side setup required.
Every question a user asks in Genie now runs through the control plane. You'll see the calls arrive in Request Logs attributed to the API key you set on the endpoint.
Streaming¶
Streaming works end-to-end for callers that request it. Mosaic AI agents that use the OpenAI SDK path support it directly:
stream = client.chat.completions.create(
model="gpt-4o-mini",
messages=[{"role": "user", "content": "count to five"}],
stream=True,
)
for chunk in stream:
delta = chunk.choices[0].delta if chunk.choices else None
if delta and delta.content:
print(delta.content, end="", flush=True)
ai_query() does not stream — it returns the full result as a
SQL value. Genie's UX does not stream either.
Cross-provider routing¶
Databricks callers can request any model registered on the control plane by name, including models on non-OpenAI upstreams:
-- Ops registered 'claude-haiku-4-5' on an Anthropic upstream.
-- Databricks calls it the same way as any other model.
SELECT ai_query('vidai-control-plane', 'summarise: ' || body)
FROM my_docs;
The control plane translates the OpenAI-shape request into the upstream's wire format, calls the upstream, translates the response back, and returns a standard OpenAI response to Databricks. Same pattern for Gemini, Vertex, Bedrock.
Attributing Databricks traffic¶
Every call from Databricks arrives at the control plane authenticated with the single API key you set on the External Model endpoint. That key's owner is the attribution unit on Request Logs and Chargeback.
A tidy pattern for a real workspace:
- Create an application on the Applications
page named after the workspace (e.g.
databricks-prod-us-east). - Mint an agent inside that application.
- Use that agent's API key as the
openai_api_keyon the Databricks endpoint.
Now the "Chargeback by application" section on the Chargeback tab reports the Databricks workspace's spend as its own line item. If you have multiple workspaces, one application per workspace keeps them separable.
Verify it works¶
From a Databricks notebook:
The cell returns ok. A new row appears within a few seconds on
the control plane's Request Logs page,
attributed to the key you set on the endpoint.
If something's off¶
Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with the Databricks endpoint config (masking the API key) and the notebook cell that reproduces the issue. We'll get it sorted.
Where to go next¶
- Client integrations overview
- API Keys — mint the workspace key
- Applications — one application per Databricks workspace
- Cost Engine → Chargeback
- Routing — cost-saver, A/B and budget-breaker rules apply to Databricks traffic transparently
- Request Logs