Haystack (deepset): using the VIDAI Control Plane as the backend¶
Haystack's OpenAIGenerator and OpenAIChatGenerator accept an
api_base_url and api_key. Configure them once, add the
generator to your pipeline, and every generation step routes
through the control plane.
TL;DR¶
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.utils import Secret
from haystack.dataclasses import ChatMessage
llm = OpenAIChatGenerator(
model="gpt-4o-mini",
api_base_url="https://your-vidai-server.example.com/v1",
api_key=Secret.from_token("your-vidai-key"),
)
result = llm.run(messages=[ChatMessage.from_user("Reply with: ok")])
print(result["replies"][0].text)
Prerequisites¶
pip install haystack-ai(2.x — the component-based API).- Control plane base URL and an API key from API Keys.
- A model registered on the Models page.
RAG pipeline¶
Drop the generator into a standard Haystack RAG pipeline:
from haystack import Pipeline
from haystack.components.builders import PromptBuilder
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever
pipe = Pipeline()
pipe.add_component("retriever", InMemoryBM25Retriever(document_store=store))
pipe.add_component("prompt", PromptBuilder(template="Given {{documents}}, answer: {{query}}"))
pipe.add_component("llm", OpenAIChatGenerator(
model="gpt-4o-mini",
api_base_url="https://your-vidai-server.example.com/v1",
api_key=Secret.from_token("your-vidai-key"),
))
pipe.connect("retriever", "prompt.documents")
pipe.connect("prompt", "llm.messages")
pipe.run({"retriever": {"query": "..."}, "prompt": {"query": "..."}})
Every generation step is a control-plane request that lands on Request Logs.
Function calling and agents¶
Haystack's OpenAIChatGenerator supports tools= for function
calling and integrates with the Agent component:
from haystack.components.agents import Agent
from haystack.tools import tool
@tool
def get_weather(city: str) -> str:
return f"Sunny in {city}."
agent = Agent(
chat_generator=llm,
tools=[get_weather],
)
result = agent.run(messages=[ChatMessage.from_user("Weather in Kyoto?")])
Streaming¶
Pass a streaming_callback to the generator:
def on_chunk(chunk):
print(chunk.content, end="", flush=True)
llm = OpenAIChatGenerator(
model="gpt-4o-mini",
api_base_url="https://your-vidai-server.example.com/v1",
api_key=Secret.from_token("your-vidai-key"),
streaming_callback=on_chunk,
)
Environment-based config¶
Set once, reuse across components:
llm = OpenAIChatGenerator(
model="gpt-4o-mini",
api_base_url="https://your-vidai-server.example.com/v1",
# api_key picked up from OPENAI_API_KEY
)
Verify it works¶
from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.utils import Secret
from haystack.dataclasses import ChatMessage
llm = OpenAIChatGenerator(
model="gpt-4o-mini",
api_base_url="https://your-vidai-server.example.com/v1",
api_key=Secret.from_token("your-vidai-key"),
)
print(llm.run(messages=[ChatMessage.from_user("reply with exactly: ok")])["replies"][0].text)
If something's off¶
Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with your pipeline / generator config and the query that reproduces the issue. We'll get it sorted.