Skip to content

Databricks Genie: using the VIDAI Control Plane as the backend

You reach the control plane from Databricks by registering it as an External Model endpoint in Model Serving. Once that endpoint exists, Genie, ai_query(), the AI Playground, and any notebook that speaks to a Databricks-served endpoint all route through the control plane.

TL;DR

Serving → Create serving endpoint → External model:

  • Provider: OpenAI
  • Task: Chat
  • openai_api_base: https://your-vidai-server.example.com/v1
  • openai_api_key: your VIDAI API key
  • Served entity name: the model you want to expose (e.g. gpt-4o-mini)

Then in a notebook:

SELECT ai_query('vidai-control-plane', 'reply with: ok') AS answer;

Prerequisites

  • A Databricks workspace with permission to create Model Serving endpoints.
  • The control plane's base URL and an API key from the API Keys page.
  • A model registered on the Models page — whatever name you give it here is what your ai_query() / Playground / Mosaic AI calls will use.

Register the endpoint

In your Databricks workspace go to Serving → Create serving endpoint and pick External model.

Name — anything memorable, e.g. vidai-control-plane. This is the string you'll pass to ai_query() and to the OpenAI SDK's base_url when writing agents.

Provider — OpenAI. The control plane's /v1/chat/completions endpoint speaks the OpenAI wire shape, which is what this provider type expects.

Task — Chat.

Served entities — add one entity per model you want callable from Databricks. The name of the entity is what Databricks callers will pass as the model name; it should match a model name registered on the control plane's Models page (or an alias ops has set up).

Provider config:

openai_api_base: https://your-vidai-server.example.com/v1
openai_api_key: "{{secrets/your-scope/vidai_api_key}}"

Use a Databricks secret scope for the key in production (openai_api_key_plaintext is fine for a first-time smoke test but Databricks warns against it).

Leave openai_api_type and openai_api_version blank — the control plane serves the OpenAI wire shape on this path, not Azure OpenAI.

Save. Databricks provisions the endpoint in about 30 seconds.

Alternative: the Custom provider type

If your Databricks reporting expects the endpoint to be labelled as something other than "OpenAI", use provider type Custom instead:

custom_provider_url: https://your-vidai-server.example.com/v1/chat/completions
authentication:
  bearer_token:
    token: "{{secrets/your-scope/vidai_api_key}}"

The wire shape and behaviour are identical; only the Databricks- side label changes.

Call it from ai_query()

The SQL function is the shortest path from a notebook to your new endpoint:

SELECT ai_query(
  'vidai-control-plane',
  'Classify this ticket as URGENT, NORMAL, or LOW: ' || ticket_body
) AS priority
FROM support_tickets
WHERE created_at >= current_date() - 7;

The first argument is the endpoint name you registered. The second is the prompt. See the Databricks ai_query docs for advanced shapes (structured output, temperature, max tokens).

Call it from an OpenAI-SDK notebook (Mosaic AI Agent Framework)

Mosaic AI agents and any Python notebook that uses the OpenAI SDK reach the endpoint via its Databricks-side invocation URL:

from databricks.sdk.runtime import spark
from openai import OpenAI

workspace_host = spark.conf.get("spark.databricks.workspaceUrl")
databricks_token = dbutils.notebook.entry_point.getDbutils() \
    .notebook().getContext().apiToken().get()

client = OpenAI(
    base_url=f"https://{workspace_host}/serving-endpoints/vidai-control-plane/invocations",
    api_key=databricks_token,
)

resp = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "hello from a notebook"}],
)
print(resp.choices[0].message.content)

The model= you pass is what the control plane resolves — so if ops has registered gpt-4o-mini as an alias that routes to Claude Haiku, that's what will serve.

Call it from Genie

Genie itself has no LLM configuration screen. It uses whichever Model Serving endpoint your workspace admin has configured for Genie usage. Once your VIDAI endpoint is registered, the workspace admin points Genie at it via the standard Databricks Genie configuration flow — no VIDAI-side setup required.

Every question a user asks in Genie now runs through the control plane. You'll see the calls arrive in Request Logs attributed to the API key you set on the endpoint.

Streaming

Streaming works end-to-end for callers that request it. Mosaic AI agents that use the OpenAI SDK path support it directly:

stream = client.chat.completions.create(
    model="gpt-4o-mini",
    messages=[{"role": "user", "content": "count to five"}],
    stream=True,
)
for chunk in stream:
    delta = chunk.choices[0].delta if chunk.choices else None
    if delta and delta.content:
        print(delta.content, end="", flush=True)

ai_query() does not stream — it returns the full result as a SQL value. Genie's UX does not stream either.

Cross-provider routing

Databricks callers can request any model registered on the control plane by name, including models on non-OpenAI upstreams:

-- Ops registered 'claude-haiku-4-5' on an Anthropic upstream.
-- Databricks calls it the same way as any other model.
SELECT ai_query('vidai-control-plane', 'summarise: ' || body)
FROM my_docs;

The control plane translates the OpenAI-shape request into the upstream's wire format, calls the upstream, translates the response back, and returns a standard OpenAI response to Databricks. Same pattern for Gemini, Vertex, Bedrock.

Attributing Databricks traffic

Every call from Databricks arrives at the control plane authenticated with the single API key you set on the External Model endpoint. That key's owner is the attribution unit on Request Logs and Chargeback.

A tidy pattern for a real workspace:

  1. Create an application on the Applications page named after the workspace (e.g. databricks-prod-us-east).
  2. Mint an agent inside that application.
  3. Use that agent's API key as the openai_api_key on the Databricks endpoint.

Now the "Chargeback by application" section on the Chargeback tab reports the Databricks workspace's spend as its own line item. If you have multiple workspaces, one application per workspace keeps them separable.

Verify it works

From a Databricks notebook:

SELECT ai_query('vidai-control-plane', 'reply with exactly: ok') AS answer;

The cell returns ok. A new row appears within a few seconds on the control plane's Request Logs page, attributed to the key you set on the endpoint.

If something's off

Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with the Databricks endpoint config (masking the API key) and the notebook cell that reproduces the issue. We'll get it sorted.

Where to go next