Skip to content

Haystack (deepset): using the VIDAI Control Plane as the backend

Haystack's OpenAIGenerator and OpenAIChatGenerator accept an api_base_url and api_key. Configure them once, add the generator to your pipeline, and every generation step routes through the control plane.

TL;DR

from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.utils import Secret
from haystack.dataclasses import ChatMessage

llm = OpenAIChatGenerator(
    model="gpt-4o-mini",
    api_base_url="https://your-vidai-server.example.com/v1",
    api_key=Secret.from_token("your-vidai-key"),
)

result = llm.run(messages=[ChatMessage.from_user("Reply with: ok")])
print(result["replies"][0].text)

Prerequisites

  • pip install haystack-ai (2.x — the component-based API).
  • Control plane base URL and an API key from API Keys.
  • A model registered on the Models page.

RAG pipeline

Drop the generator into a standard Haystack RAG pipeline:

from haystack import Pipeline
from haystack.components.builders import PromptBuilder
from haystack.components.retrievers.in_memory import InMemoryBM25Retriever

pipe = Pipeline()
pipe.add_component("retriever", InMemoryBM25Retriever(document_store=store))
pipe.add_component("prompt", PromptBuilder(template="Given {{documents}}, answer: {{query}}"))
pipe.add_component("llm", OpenAIChatGenerator(
    model="gpt-4o-mini",
    api_base_url="https://your-vidai-server.example.com/v1",
    api_key=Secret.from_token("your-vidai-key"),
))

pipe.connect("retriever", "prompt.documents")
pipe.connect("prompt", "llm.messages")

pipe.run({"retriever": {"query": "..."}, "prompt": {"query": "..."}})

Every generation step is a control-plane request that lands on Request Logs.

Function calling and agents

Haystack's OpenAIChatGenerator supports tools= for function calling and integrates with the Agent component:

from haystack.components.agents import Agent
from haystack.tools import tool

@tool
def get_weather(city: str) -> str:
    return f"Sunny in {city}."

agent = Agent(
    chat_generator=llm,
    tools=[get_weather],
)

result = agent.run(messages=[ChatMessage.from_user("Weather in Kyoto?")])

Streaming

Pass a streaming_callback to the generator:

def on_chunk(chunk):
    print(chunk.content, end="", flush=True)

llm = OpenAIChatGenerator(
    model="gpt-4o-mini",
    api_base_url="https://your-vidai-server.example.com/v1",
    api_key=Secret.from_token("your-vidai-key"),
    streaming_callback=on_chunk,
)

Environment-based config

Set once, reuse across components:

export OPENAI_API_KEY=your-vidai-key
llm = OpenAIChatGenerator(
    model="gpt-4o-mini",
    api_base_url="https://your-vidai-server.example.com/v1",
    # api_key picked up from OPENAI_API_KEY
)

Verify it works

from haystack.components.generators.chat import OpenAIChatGenerator
from haystack.utils import Secret
from haystack.dataclasses import ChatMessage

llm = OpenAIChatGenerator(
    model="gpt-4o-mini",
    api_base_url="https://your-vidai-server.example.com/v1",
    api_key=Secret.from_token("your-vidai-key"),
)
print(llm.run(messages=[ChatMessage.from_user("reply with exactly: ok")])["replies"][0].text)

If something's off

Raise an issue at github.com/vidaiUK/vidai-quickstart/issues with your pipeline / generator config and the query that reproduces the issue. We'll get it sorted.

Where to go next