Skip to content

LlamaIndex: using the VIDAI Control Plane as the backend

LlamaIndex's OpenAI (and Anthropic) LLM classes accept api_base and api_key directly. Configure the LLM once, set it as the global default (or pass it to your index / query engine), and every RAG query, agent step, and reranker call routes through the control plane.

TL;DR

from llama_index.core import Settings
from llama_index.llms.openai import OpenAI

Settings.llm = OpenAI(
    model="gpt-4o-mini",
    api_base="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)

Any downstream LlamaIndex call — VectorStoreIndex, RetrieverQueryEngine, ChatEngine, ReActAgent — inherits this LLM automatically.

Prerequisites

  • pip install llama-index-llms-openai (or llama-index-llms-anthropic).
  • Control plane base URL and an API key from API Keys.
  • A model registered on the Models page.

Query engine over a directory

from llama_index.core import SimpleDirectoryReader, VectorStoreIndex

docs = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(docs)
qe = index.as_query_engine()
print(qe.query("What does this repo do?"))

The completion call inside query() routes through the control plane. If you also want embeddings served through the control plane, set an embedding model the same way (see the LlamaIndex embeddings docs for provider-specific classes).

Anthropic LLM

For Claude-family models registered on the control plane. Note the Anthropic base URL is the root (no /v1):

from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings

Settings.llm = Anthropic(
    model="claude-haiku-4-5",
    base_url="https://your-vidai-server.example.com",
    api_key="your-vidai-key",
)

Agent frameworks

ReActAgent, FunctionCallingAgent, and OpenAIAgent all use whichever LLM is set on Settings.llm (or passed directly). Every step of the agent's reasoning loop and every tool call is a control-plane request:

from llama_index.core.agent import ReActAgent
from llama_index.core.tools import FunctionTool

def get_stock(symbol: str) -> float:
    return 137.42

agent = ReActAgent.from_tools([FunctionTool.from_defaults(get_stock)])
print(agent.chat("What's the price of AAPL?"))

Streaming

query() supports streaming via .query_stream() and the streaming=True flag on chat engines:

resp = qe.query("Summarise chapter 1", streaming=True)
resp.print_response_stream()

LlamaIndex.js

The TypeScript version follows the same shape via the OpenAI provider:

import { OpenAI } from "@llamaindex/openai";
import { Settings } from "llamaindex";

Settings.llm = new OpenAI({
  model: "gpt-4o-mini",
  apiKey: "your-vidai-key",
  baseURL: "https://your-vidai-server.example.com/v1",
});

Verify it works

from llama_index.llms.openai import OpenAI

llm = OpenAI(
    model="gpt-4o-mini",
    api_base="https://your-vidai-server.example.com/v1",
    api_key="your-vidai-key",
)
print(llm.complete("reply with exactly: ok").text)

If something's off

Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with your Settings.llm config and the query that reproduces the issue. We'll get it sorted.

Where to go next