Observability: Prometheus metrics¶
The control plane exposes a Prometheus scrape endpoint so your existing monitoring stack can ingest control plane metrics (request volume, latency, error rates, token counts, in-flight load, provider health) without going through the dashboard or the database.
This page is the operator's reference: what the control plane exposes, on which port, in what shape, and how to scrape it safely.
📌 Worth knowing. Prometheus metrics are an enterprise feature. The endpoint is compiled into the enterprise build only; community / unlicensed deployments won't expose it.
When you'd reach for this¶
- You run Prometheus / VictoriaMetrics / Mimir elsewhere and want the control plane to land in your existing dashboards alongside the rest of your platform.
- You need alerting (on error-rate spikes, latency budgets, provider degradation) and your alerting stack is already wired up.
- You're capacity-planning and want historical request volume at second-level granularity (the dashboard's Cost Insights panel rolls up by day).
- You're debugging upstream incidents and want to see per-provider in-flight load in real time.
For the in-console story, see Dashboard and Request Logs. The dashboard is the admin's daily surface; this endpoint is for your monitoring team's stack.
What's exposed¶
A separate HTTP listener serves two endpoints:
| Endpoint | Format | Auth | Purpose |
|---|---|---|---|
GET /metrics |
Prometheus text format v0.0.4 | None | Scrape target. |
GET /health |
{"status":"healthy"} |
None | Liveness probe. Minimal: does not include database latency. |
Both endpoints are HTTP-only and unauthenticated. The control plane's posture is "network ACL is the only access control": see Security posture below before you expose this off-host.
⚠️ Watch out. The
/healthhere is a liveness probe, not a readiness check. It returnshealthyas long as the metrics server is running, regardless of whether the control plane can reach Postgres or its upstreams. For deeper health, use the control plane's admin health endpoint athttp://<gateway-host>:<gateway-port>/admin/health.
The metric catalogue¶
Fifteen metrics across four shapes. Every metric carries the
vidai_ prefix.
| Metric | Type | Labels | What it counts |
|---|---|---|---|
vidai_requests_total |
counter | provider, model, status_code |
Every request the control plane routed. |
vidai_request_duration_seconds |
histogram | provider, model |
End-to-end request latency. Buckets span 0.1s → 60s. |
vidai_queue_duration_seconds |
histogram | provider |
Time spent waiting in the queue before dispatch. Buckets 1ms → 5s. |
vidai_active_requests |
gauge | (none) | Total in-flight across the control plane. |
vidai_active_requests_by_provider |
gauge | provider |
In-flight, broken out per provider. |
vidai_prompt_tokens_total |
counter | provider, model |
Tiktoken-counted prompt tokens (control-plane-side). |
vidai_completion_tokens_total |
counter | provider, model |
Tiktoken-counted completion tokens (control-plane-side). |
vidai_vendor_input_tokens_total |
counter | provider, model |
Vendor-reported input tokens (from the upstream's response). |
vidai_vendor_output_tokens_total |
counter | provider, model |
Vendor-reported output tokens. |
vidai_cached_tokens_total |
counter | provider, model |
Cache-hit tokens: Anthropic cache_read_input_tokens, OpenAI cached_tokens, Gemini cachedContentTokenCount. |
vidai_reasoning_tokens_total |
counter | provider, model |
OpenAI o-series reasoning tokens. Not populated for Anthropic or Gemini. |
vidai_cost_dollars_total |
counter | provider, model, currency |
Aggregator-reported cost (e.g. OpenRouter's per-request charge). Most providers don't populate this; use the control plane's own cost engine via the dashboard for full coverage. |
vidai_errors_total |
counter | error_type, provider |
error_type ∈ {unauthorized, forbidden, rate_limited, server_error, client_error}. |
vidai_rate_limit_hits_total |
counter | provider |
429s seen per provider. |
vidai_provider_up |
gauge | provider |
1 = up, 0 = down. Driven by the circuit-breaker state. |
📌 Worth knowing. All token counters appear twice: once control-plane-counted (
prompt_tokens_total,completion_tokens_total, via tiktoken) and once vendor-reported (vendor_input_tokens_total,vendor_output_tokens_total). They usually agree within a few percent. Material disagreement points at a tokeniser mismatch with the upstream: useful when you're auditing billing reconciliation against vendor invoices.
Configuration¶
The metrics endpoint is configured in vidai-server.yaml,
which ships in the deployment bundle. The metrics: block
controls everything:
metrics:
# Master switch. Off by default in fresh installs;
# the deployment bundle ships with this set to true.
enabled: true
# Port for the metrics HTTP listener.
# Bundle default: 9099 (avoids collisions with common
# Prometheus exporter ports). Source default: 9091.
port: 9099
# Bind address. Source default: 127.0.0.1 (loopback only,
# per SEC-001 F-08). The deployment bundle overrides to
# 0.0.0.0 so a sibling Prometheus container in the same
# Compose network can reach it.
bind: "0.0.0.0"
After editing the YAML, restart the control plane container so the new metrics listener picks up the change.
⚠️ Watch out. When you bind to
0.0.0.0, the metrics endpoint is reachable by anything that can reach the host on that port. Combined with the no-auth posture, this means a network-level rule (firewall, security group, Compose-internal-only port) is the only thing keeping your metrics private. Don't bind 0.0.0.0 on a host that's directly internet-exposed.
Scraping: the minimal recipe¶
A working prometheus.yml job, assuming the control plane is
reachable at gateway.internal:9099:
scrape_configs:
- job_name: vidai-server
metrics_path: /metrics
static_configs:
- targets: ['gateway.internal:9099']
labels:
deployment: production # or whatever you label by
scrape_interval: 15s
Confirm the scrape works:
You should see output starting with # HELP vidai_requests_total …
and a series of metrics with their current values.
Security posture¶
Three things to know:
- No authentication. Anyone who can reach the port can read every metric. Token counts and request volumes are not directly sensitive but they're not nothing; they reveal traffic patterns and provider mix.
- HTTP only. TLS isn't terminated by the metrics listener. If you want HTTPS, front it with a reverse proxy or your cluster's mesh.
- Loopback by default in source; bundle overrides. A
fresh source build sticks to
127.0.0.1. The deployment bundle's default0.0.0.0is intentional for the single-host Compose case where a sibling Prometheus container is the only client. Tighten this when you move to a multi-host or shared-network deployment.
The recommended posture for production:
- Bind to a private interface or to
127.0.0.1plus a reverse proxy. - Network ACL: only your Prometheus / monitoring source can reach the metrics port.
- Reverse proxy with basic auth or mTLS in front of the endpoint when scraping across a public network.
Wiring into a Grafana dashboard (a starting point)¶
You'll typically want at least these four panels in a control plane overview dashboard:
| Panel | Query |
|---|---|
| Request rate by provider | sum(rate(vidai_requests_total[5m])) by (provider) |
| Error rate % | sum(rate(vidai_errors_total[5m])) / sum(rate(vidai_requests_total[5m])) * 100 |
| p95 latency by provider | histogram_quantile(0.95, sum by (le, provider) (rate(vidai_request_duration_seconds_bucket[5m]))) |
| In-flight per provider | vidai_active_requests_by_provider |
💡 Pro tip. Start with the four panels above, then add token-volume + cost-per-provider panels once you've calibrated thresholds. The 15-metric surface gives you a lot of room without needing to instrument the control plane further.
There is no prebuilt Grafana dashboard to import; the four queries above plus the request-logs detail in the console cover most operational questions.
Troubleshooting¶
curl /metrics returns "connection refused"¶
The metrics listener didn't start. Three usual suspects:
metrics.enabled: falseinvidai-server.yaml. Flip it totrueand restart.- Community build. Check the control plane logs at startup
for a line like
metrics: enterprise feature not available. Prometheus metrics are gated to the enterprise build. - Port already in use. Another process bound
9099(or9091) before the control plane came up. Pick a free port and updatemetrics.port.
curl /metrics returns "connection refused" only from outside the host¶
Bind address is 127.0.0.1 and you're scraping across hosts.
Either:
- Set
metrics.bind: "0.0.0.0"(read Security posture first), or - Front the endpoint with a reverse proxy that binds the external interface and forwards to loopback.
Metric values look stuck¶
Metrics are recorded inside the batch-flush worker, not
inline on the request hot path. If telemetry.batch_size is
high and traffic is low, batches take longer to fill and
counters appear to lag by up to batch_timeout_ms. This is
expected: metrics will catch up on the next flush.
Limitations¶
What's deliberately not in this metrics surface:
- No per-key, per-user, or per-application labels. The cardinality cost outweighs the value at the metrics layer; for that level of breakdown, use Request Logs or BI Tables.
- No business / compliance metrics. The dashboard's Compliance and VIDAI Impact panels live in the database, not the metrics endpoint. Their semantic fits the daily reporting cadence rather than the second-level scrape cadence.
- No webhook / circuit-breaker event stream. Webhooks are the right surface for event-driven workflows; see Webhooks.
Where to go next¶
- Dashboard: the in-console rollups of the same telemetry, on a daily cadence.
- Request Logs: the per-request forensic surface.
- Webhooks: push-based event delivery for circuit trips, guardrail blocks, budget breaches, deny-rule fires.
- Settings: the worker-flags toggle that
controls which pipelines feed the in-console rollups
(independent of the metrics endpoint, which is always on
when
metrics.enabled: true).