Skip to content

Observability: Prometheus metrics

The control plane exposes a Prometheus scrape endpoint so your existing monitoring stack can ingest control plane metrics (request volume, latency, error rates, token counts, in-flight load, provider health) without going through the dashboard or the database.

This page is the operator's reference: what the control plane exposes, on which port, in what shape, and how to scrape it safely.

📌 Worth knowing. Prometheus metrics are an enterprise feature. The endpoint is compiled into the enterprise build only; community / unlicensed deployments won't expose it.


When you'd reach for this

  • You run Prometheus / VictoriaMetrics / Mimir elsewhere and want the control plane to land in your existing dashboards alongside the rest of your platform.
  • You need alerting (on error-rate spikes, latency budgets, provider degradation) and your alerting stack is already wired up.
  • You're capacity-planning and want historical request volume at second-level granularity (the dashboard's Cost Insights panel rolls up by day).
  • You're debugging upstream incidents and want to see per-provider in-flight load in real time.

For the in-console story, see Dashboard and Request Logs. The dashboard is the admin's daily surface; this endpoint is for your monitoring team's stack.


What's exposed

A separate HTTP listener serves two endpoints:

Endpoint Format Auth Purpose
GET /metrics Prometheus text format v0.0.4 None Scrape target.
GET /health {"status":"healthy"} None Liveness probe. Minimal: does not include database latency.

Both endpoints are HTTP-only and unauthenticated. The control plane's posture is "network ACL is the only access control": see Security posture below before you expose this off-host.

⚠️ Watch out. The /health here is a liveness probe, not a readiness check. It returns healthy as long as the metrics server is running, regardless of whether the control plane can reach Postgres or its upstreams. For deeper health, use the control plane's admin health endpoint at http://<gateway-host>:<gateway-port>/admin/health.


The metric catalogue

Fifteen metrics across four shapes. Every metric carries the vidai_ prefix.

Metric Type Labels What it counts
vidai_requests_total counter provider, model, status_code Every request the control plane routed.
vidai_request_duration_seconds histogram provider, model End-to-end request latency. Buckets span 0.1s → 60s.
vidai_queue_duration_seconds histogram provider Time spent waiting in the queue before dispatch. Buckets 1ms → 5s.
vidai_active_requests gauge (none) Total in-flight across the control plane.
vidai_active_requests_by_provider gauge provider In-flight, broken out per provider.
vidai_prompt_tokens_total counter provider, model Tiktoken-counted prompt tokens (control-plane-side).
vidai_completion_tokens_total counter provider, model Tiktoken-counted completion tokens (control-plane-side).
vidai_vendor_input_tokens_total counter provider, model Vendor-reported input tokens (from the upstream's response).
vidai_vendor_output_tokens_total counter provider, model Vendor-reported output tokens.
vidai_cached_tokens_total counter provider, model Cache-hit tokens: Anthropic cache_read_input_tokens, OpenAI cached_tokens, Gemini cachedContentTokenCount.
vidai_reasoning_tokens_total counter provider, model OpenAI o-series reasoning tokens. Not populated for Anthropic or Gemini.
vidai_cost_dollars_total counter provider, model, currency Aggregator-reported cost (e.g. OpenRouter's per-request charge). Most providers don't populate this; use the control plane's own cost engine via the dashboard for full coverage.
vidai_errors_total counter error_type, provider error_type ∈ {unauthorized, forbidden, rate_limited, server_error, client_error}.
vidai_rate_limit_hits_total counter provider 429s seen per provider.
vidai_provider_up gauge provider 1 = up, 0 = down. Driven by the circuit-breaker state.

📌 Worth knowing. All token counters appear twice: once control-plane-counted (prompt_tokens_total, completion_tokens_total, via tiktoken) and once vendor-reported (vendor_input_tokens_total, vendor_output_tokens_total). They usually agree within a few percent. Material disagreement points at a tokeniser mismatch with the upstream: useful when you're auditing billing reconciliation against vendor invoices.


Configuration

The metrics endpoint is configured in vidai-server.yaml, which ships in the deployment bundle. The metrics: block controls everything:

metrics:
  # Master switch. Off by default in fresh installs;
  # the deployment bundle ships with this set to true.
  enabled: true

  # Port for the metrics HTTP listener.
  # Bundle default: 9099 (avoids collisions with common
  # Prometheus exporter ports). Source default: 9091.
  port: 9099

  # Bind address. Source default: 127.0.0.1 (loopback only,
  # per SEC-001 F-08). The deployment bundle overrides to
  # 0.0.0.0 so a sibling Prometheus container in the same
  # Compose network can reach it.
  bind: "0.0.0.0"

After editing the YAML, restart the control plane container so the new metrics listener picks up the change.

⚠️ Watch out. When you bind to 0.0.0.0, the metrics endpoint is reachable by anything that can reach the host on that port. Combined with the no-auth posture, this means a network-level rule (firewall, security group, Compose-internal-only port) is the only thing keeping your metrics private. Don't bind 0.0.0.0 on a host that's directly internet-exposed.


Scraping: the minimal recipe

A working prometheus.yml job, assuming the control plane is reachable at gateway.internal:9099:

scrape_configs:
  - job_name: vidai-server
    metrics_path: /metrics
    static_configs:
      - targets: ['gateway.internal:9099']
        labels:
          deployment: production    # or whatever you label by
    scrape_interval: 15s

Confirm the scrape works:

curl http://gateway.internal:9099/metrics | head -50

You should see output starting with # HELP vidai_requests_total … and a series of metrics with their current values.


Security posture

Three things to know:

  1. No authentication. Anyone who can reach the port can read every metric. Token counts and request volumes are not directly sensitive but they're not nothing; they reveal traffic patterns and provider mix.
  2. HTTP only. TLS isn't terminated by the metrics listener. If you want HTTPS, front it with a reverse proxy or your cluster's mesh.
  3. Loopback by default in source; bundle overrides. A fresh source build sticks to 127.0.0.1. The deployment bundle's default 0.0.0.0 is intentional for the single-host Compose case where a sibling Prometheus container is the only client. Tighten this when you move to a multi-host or shared-network deployment.

The recommended posture for production:

  • Bind to a private interface or to 127.0.0.1 plus a reverse proxy.
  • Network ACL: only your Prometheus / monitoring source can reach the metrics port.
  • Reverse proxy with basic auth or mTLS in front of the endpoint when scraping across a public network.

Wiring into a Grafana dashboard (a starting point)

You'll typically want at least these four panels in a control plane overview dashboard:

Panel Query
Request rate by provider sum(rate(vidai_requests_total[5m])) by (provider)
Error rate % sum(rate(vidai_errors_total[5m])) / sum(rate(vidai_requests_total[5m])) * 100
p95 latency by provider histogram_quantile(0.95, sum by (le, provider) (rate(vidai_request_duration_seconds_bucket[5m])))
In-flight per provider vidai_active_requests_by_provider

💡 Pro tip. Start with the four panels above, then add token-volume + cost-per-provider panels once you've calibrated thresholds. The 15-metric surface gives you a lot of room without needing to instrument the control plane further.

There is no prebuilt Grafana dashboard to import; the four queries above plus the request-logs detail in the console cover most operational questions.


Troubleshooting

curl /metrics returns "connection refused"

The metrics listener didn't start. Three usual suspects:

  1. metrics.enabled: false in vidai-server.yaml. Flip it to true and restart.
  2. Community build. Check the control plane logs at startup for a line like metrics: enterprise feature not available. Prometheus metrics are gated to the enterprise build.
  3. Port already in use. Another process bound 9099 (or 9091) before the control plane came up. Pick a free port and update metrics.port.

curl /metrics returns "connection refused" only from outside the host

Bind address is 127.0.0.1 and you're scraping across hosts. Either:

  • Set metrics.bind: "0.0.0.0" (read Security posture first), or
  • Front the endpoint with a reverse proxy that binds the external interface and forwards to loopback.

Metric values look stuck

Metrics are recorded inside the batch-flush worker, not inline on the request hot path. If telemetry.batch_size is high and traffic is low, batches take longer to fill and counters appear to lag by up to batch_timeout_ms. This is expected: metrics will catch up on the next flush.


Limitations

What's deliberately not in this metrics surface:

  • No per-key, per-user, or per-application labels. The cardinality cost outweighs the value at the metrics layer; for that level of breakdown, use Request Logs or BI Tables.
  • No business / compliance metrics. The dashboard's Compliance and VIDAI Impact panels live in the database, not the metrics endpoint. Their semantic fits the daily reporting cadence rather than the second-level scrape cadence.
  • No webhook / circuit-breaker event stream. Webhooks are the right surface for event-driven workflows; see Webhooks.

Where to go next

  • Dashboard: the in-console rollups of the same telemetry, on a daily cadence.
  • Request Logs: the per-request forensic surface.
  • Webhooks: push-based event delivery for circuit trips, guardrail blocks, budget breaches, deny-rule fires.
  • Settings: the worker-flags toggle that controls which pipelines feed the in-console rollups (independent of the metrics endpoint, which is always on when metrics.enabled: true).