LlamaIndex: using the VIDAI Control Plane as the backend¶
LlamaIndex's OpenAI (and Anthropic) LLM classes accept
api_base and api_key directly. Configure the LLM once, set
it as the global default (or pass it to your index / query
engine), and every RAG query, agent step, and reranker call
routes through the control plane.
TL;DR¶
from llama_index.core import Settings
from llama_index.llms.openai import OpenAI
Settings.llm = OpenAI(
model="gpt-4o-mini",
api_base="https://your-vidai-server.example.com/v1",
api_key="your-vidai-key",
)
Any downstream LlamaIndex call — VectorStoreIndex,
RetrieverQueryEngine, ChatEngine, ReActAgent — inherits
this LLM automatically.
Prerequisites¶
pip install llama-index-llms-openai(orllama-index-llms-anthropic).- Control plane base URL and an API key from API Keys.
- A model registered on the Models page.
Query engine over a directory¶
from llama_index.core import SimpleDirectoryReader, VectorStoreIndex
docs = SimpleDirectoryReader("./docs").load_data()
index = VectorStoreIndex.from_documents(docs)
qe = index.as_query_engine()
print(qe.query("What does this repo do?"))
The completion call inside query() routes through the control
plane. If you also want embeddings served through the control
plane, set an embedding model the same way (see the LlamaIndex
embeddings docs for provider-specific classes).
Anthropic LLM¶
For Claude-family models registered on the control plane. Note
the Anthropic base URL is the root (no /v1):
from llama_index.llms.anthropic import Anthropic
from llama_index.core import Settings
Settings.llm = Anthropic(
model="claude-haiku-4-5",
base_url="https://your-vidai-server.example.com",
api_key="your-vidai-key",
)
Agent frameworks¶
ReActAgent, FunctionCallingAgent, and OpenAIAgent all use
whichever LLM is set on Settings.llm (or passed directly).
Every step of the agent's reasoning loop and every tool call is
a control-plane request:
from llama_index.core.agent import ReActAgent
from llama_index.core.tools import FunctionTool
def get_stock(symbol: str) -> float:
return 137.42
agent = ReActAgent.from_tools([FunctionTool.from_defaults(get_stock)])
print(agent.chat("What's the price of AAPL?"))
Streaming¶
query() supports streaming via .query_stream() and the
streaming=True flag on chat engines:
LlamaIndex.js¶
The TypeScript version follows the same shape via the OpenAI provider:
import { OpenAI } from "@llamaindex/openai";
import { Settings } from "llamaindex";
Settings.llm = new OpenAI({
model: "gpt-4o-mini",
apiKey: "your-vidai-key",
baseURL: "https://your-vidai-server.example.com/v1",
});
Verify it works¶
from llama_index.llms.openai import OpenAI
llm = OpenAI(
model="gpt-4o-mini",
api_base="https://your-vidai-server.example.com/v1",
api_key="your-vidai-key",
)
print(llm.complete("reply with exactly: ok").text)
If something's off¶
Raise an issue at github.com/vidaiUK/vidai-quickstart/issues
with your Settings.llm config and the query that reproduces
the issue. We'll get it sorted.
Where to go next¶
- Client integrations overview
- Request Logs — inspect agent + retriever calls
- Routing