Skip to content

Guardrails

Guardrails inspect every request that flows through the VIDAI Control Plane and decide what to do with it: let it through, mask sensitive parts, quietly log a flag, or block / redirect it entirely. They run in two layers: a fast regex tier that matches against the request body in microseconds, and an optional ML tier that runs trained models for prompt-injection, toxicity, PII, and intellectual- property detection. Together they give you content-policy enforcement that's fast on the happy path, sophisticated on the edge cases, and entirely server-side.

This page is where you author rules, manage which models the ML tier runs, and set system defaults. The composition with users / groups / keys (who gets which rules) lives on those pages; this page is the rule-shop.


When you'd open this page

  • A new compliance rule comes in (e.g. "block credit-card numbers from being sent to upstreams"). Author the rule here, then attach it to the right teams / applications via their guardrail policy.
  • A leaked-secrets scan should fire on every key. The preset secrets category covers most of this; this page is where you add custom secret patterns specific to your org.
  • An admin reports a false positive. Find the firing rule, adjust its pattern or move it from block to log_only while you investigate.
  • A new model is being added to the ML tier. Enable / disable it on the VidaiGuard tab and verify it's reachable.
  • An audit asks "what content rules does the control plane enforce?". The rules list is the answer.

How guardrails fit with everything else

Guardrails run first in the request pipeline:

request → guardrails → routing rule → alias → fallback (when unhealthy)

Two things to internalise:

  1. Guardrails see the original request body, before any routing rewrites. If a routing rule would have rewritten the call's model, the guardrail evaluation has already happened on the original model name.
  2. A guardrail outcome can short-circuit the rest of the pipeline. A block returns 403 immediately; a redirect rewrites the target model AND skips routing-rule evaluation; a mask rewrites the body and continues; a log_only records and continues unchanged.

📌 Worth knowing. Guardrails apply to the request body only. Response-side guardrails (inspecting the upstream's reply for sensitive content before returning it to the client) aren't supported today; if you need to filter what the control plane returns to clients, that has to happen at the application layer.


The page at a glance

The page has four tabs:

  • Rules: the regex tier. Available on both Community and Enterprise. Author and manage individual pattern-based rules.
  • VidaiGuard: Enterprise. The ML tier (injection / toxicity / PII detection via the VidaiGuard sidecar). On Community, this tab renders a locked-card stub.
  • Key Overrides: per-key guardrail overrides: which keys set their own policy instead of inheriting from user/group/default. (You can also reach this from the API Keys page; it's surfaced here for review.)
  • Settings: guardrail-wide toggles for the deployment.

See Licensing & tiers for the full feature breakdown.

Guardrails: Rules tab

The Rules list columns:

Column What it shows
Name Friendly name.
Category One of pii / secrets / prompt_injection / profanity / internal_data / custom. Coloured badge.
Pattern The regex (truncated; hover for full).
Severity low / medium / high / critical. Surfaced in audit + dashboard panels.
Action What happens on a match: Block / Mask / Log only / Redirect. Coloured chip.
Status Enabled or Disabled.

Above the list:

  • Search box: narrows by name, category, or pattern. GitHub-style multi-word search (each word AND-matches across any of the three fields).
  • Category + Severity dropdowns: strict-equality filters layered on top of search. Useful when search alone returns too many matches.
  • Column sort: click a column header to sort. Newest-first by default.
  • Bulk select + delete: checkbox per row + Select All on the visible page. Clears on tab/page change.
  • Pagination: pick 10 / 25 / 50 / 100 entries. Default 25.

💡 Pro tip. The page summary at the top shows Total Rules and Active Rules at a glance. Active means enabled-and-included-in-the-default-policy. A high total with a low active count usually means an admin has been authoring drafts or experimenting; clean those up when they go stale.


What's authored here vs what fires for whom

This page is the rule library. The rules you create here exist as candidates; whether they actually fire on a given key's traffic depends on the guardrail policy attached to the calling key, its owner, the owner's group / application, and the system default. That composition story lives on the Groups, Applications, and API Keys pages.

The short version:

  • A rule with action block is in the global default until somebody overrides downstream.
  • A team's policy on Groups can restrict which rules fire for that team's members (whitelist, blacklist, category filter).
  • A key's policy on API Keys can override the team / user policy entirely for that key.

Groups: Guardrail policy walks through the override-toggle pattern and the "empty-whitelist = all rules fire" gotcha.


What you do on this page

Author a custom regex rule

Goal: block US Social Security Numbers from leaving the control plane.

(There's already a preset for this, pii_ssn_us, so this walkthrough mostly mirrors what the preset does.)

  1. Click Add Rule. The editor opens.

Add rule modal

  1. Fill in:
  2. Name: SSN block (or whatever you'll recognise in the audit log).
  3. Pattern: the regex. For US SSN: \b(?!000|9\d\d|6{3})\d{3}-(?!00)\d{2}-(?!0000)\d{4}\b
  4. Category: pii.
  5. Severity: high.
  6. Action: block.
  7. Save.

The new rule appears on the list, enabled by default. Traffic that includes a string matching the pattern starts returning 403 content_policy_violation immediately.

⚠️ Watch out. Patterns are evaluated against the raw request body, including system prompts, conversation history, and tool definitions. A pattern that's too permissive can fire on unexpected fragments. Test with realistic traffic before locking down on block.

💡 Pro tip. Start with action log_only to observe what your rule catches without affecting traffic. Promote to mask or block once you've confirmed no false positives over a representative sample.


Mask a pattern instead of blocking

Goal: redact emails in request bodies before forwarding to the upstream, but don't reject the call.

  1. Add Rule.
  2. Pattern: \b[A-Za-z0-9._%+-]+@[A-Za-z0-9.-]+\.[A-Z|a-z]{2,}\b
  3. Category: pii. Severity: medium. Action: mask.
  4. Save.

Now any email address in a request body gets replaced with [REDACTED-PII] before the control plane forwards to the upstream. The client gets a normal response; the upstream never saw the email.

📌 Worth knowing. The mask string is [REDACTED-{CATEGORY}]: [REDACTED-PII] for pii, [REDACTED-SECRETS] for secrets, etc. The format is consistent so downstream automated tools can parse for redacted spans.


Redirect a request to a safer model when content trips

Goal: when an inbound request contains anything that looks like internal infrastructure data, route it to an on-prem model rather than blocking or sending it to a cloud upstream.

  1. Add Rule.
  2. Pattern: matches your internal hostnames, RFC1918 IPs, etc. (the preset internal_data category covers most of this).
  3. Category: internal_data. Severity: high. Action: redirect + target model: our-onprem-llm.
  4. Save.

Requests matching the rule have their target model rewritten to our-onprem-llm ahead of routing-rule evaluation. Routing rules don't then fire on the redirected request. Guardrail-redirect is the higher-priority decision.

⚠️ Watch out. The redirected target must be a registered model the calling key is allowed to use. If the key's allowed-models doesn't include the target, the control plane returns 403 model_not_allowed (the permissions re-check fires after the redirect, same as for routing rules). Verify the target is reachable before wiring up critical redirect rules.

📌 Worth knowing. Guardrail-redirect bypasses routing rules. If you have routing logic that's supposed to apply to the same source models, the redirect-target call doesn't re-evaluate against routing. Plan accordingly: either the redirect target handles its own onward routing, or the source-side routing rule needs to be replicated as a guardrail- redirect-or-routing-rule decision elsewhere.


Log a pattern silently

Goal: track how often colleagues paste internal Jira ticket references into prompts, but don't change anything.

  1. Add Rule.
  2. Pattern: matches your Jira project codes (PROJ-\d{4,}).
  3. Category: internal_data. Severity: low. Action: log_only.
  4. Save.

Matching requests are logged with the firing rule id and flow to the upstream unchanged. The dashboard's Compliance Insights surface counts these as observations, not enforcements. See Compliance.

💡 Pro tip. log_only is the safest action when you're authoring a new rule. Log for a week, review the dashboard's guardrail_logonly source bucket and the matching request-logs rows, then promote to mask or block once you trust the pattern.


Edit / disable / delete a rule

  • Edit: click the row. Same form as create. Saving propagates immediately.
  • Disable: inline toggle. Disabled rules are kept on the list but never fire. Reversible.
  • Delete: trash icon. Permanent.

📌 Worth knowing. Disabling a preset rule is sometimes the right call (your org genuinely doesn't want phone-number masking applied because it breaks a downstream workflow). The rule stays on the list with "Disabled" so a future admin can see the deliberate decision.


Manage the ML tier (VidaiGuard)

🔒 Enterprise edition. The VidaiGuard ML tier is part of the Enterprise tier. On Community deployments, the VidaiGuard tab still shows in the page navigation but its content collapses to a locked-card stub. The Rules tab (regex) is unaffected: it's a Community baseline.

Click the VidaiGuard tab.

VidaiGuard tab

Four ML models ship by default:

Model What it detects
Injection Prompt-injection / jailbreak attempts ("ignore previous instructions," "DAN" patterns, role-confusion attacks).
Toxicity Toxic / offensive language, including more nuanced patterns than regex profanity rules can catch.
PII Named-entity recognition for people, organisations, locations, plus structured personal data the regex tier might miss.
IP Protection A GLiNER-based detector for entity classes like trade-secrets, code patterns, and IP-relevant keywords.

Each model has an inline Enabled toggle. Disable a model when you don't want it running (e.g. you're using your own ML detector elsewhere and don't need duplication). The control plane caches per-model status, so disabling skips the inference round-trip entirely.

The page also shows the control plane's connection status to the ML sidecar. If the sidecar is down, all ML rules are silently skipped (the request flows through with regex-tier results only). The status indicator surfaces this so you know the ML tier isn't actually running.

💡 Pro tip. The ML tier runs after the regex tier. If a regex rule already masked PII content, the PII ML model is skipped (deduplication: no point running ML over content the regex tier already redacted). Other ML models still run.

⚠️ Watch out. The ML tier inspects only the most recent user message, not the full conversation history. Multi-turn jailbreak attacks that build context across turns may not trigger; consider a regex-tier backstop for long-context attacks if your threat model warrants.


How a request is evaluated

Step-by-step through the guardrail layers for a single request:

  1. Regex layer runs first. Every active regex rule the key's effective policy includes is matched against the request body in parallel. The first matching rule's action wins (block / mask / log_only / redirect).
  2. If the regex layer blocked or redirected, the request short-circuits there. The ML tier doesn't run.
  3. If the regex layer masked, the masked body proceeds to the ML tier. ML categories that overlap with the masked rule's category are skipped (no point running PII ML over content with PII already masked out).
  4. If the regex layer logged or passed cleanly, the ML tier runs over the original (or masked) body. Each enabled ML model is evaluated; the first ML model to declare a violation determines the action (typically block, depending on the model's configured action).
  5. The result is stamped onto the request log with a single (guardrail_status, guardrail_rule_id, guardrail_category) triple. When multiple guardrails would have fired, the last write wins: ML can overwrite regex's earlier decision if it has the stronger action.

This means a request blocked by ML even though regex masked something earlier shows up as blocked in Request Logs, with the ML rule id, not the regex rule id.


When a rule looks like it's "not firing"

Two real cases where an admin authors a guardrail rule, expects it to fire, and sees zero hits in the logs.

Case 1: Within-regex shadow (the control plane tells you)

Regex rules are evaluated first-match-wins in the order they're loaded. If two rules' patterns overlap, the earlier-loaded one catches the traffic and the later one never fires.

The control plane watches for this and surfaces it on the dashboard misconfig rail as Guardrail rule potentially shadowed:

A guardrail rule has been enabled for ≥ 7 days and recorded 0 firings in that window, AND there's at least one earlier-loaded rule whose pattern overlaps it. The earlier rule is catching the traffic first; this one never fires.

The misconfig click-through lands on the rule's detail page, which shows the candidate shadower(s): the earlier rules the analyser thinks are pre-empting it. From there:

  • Re-order: change priority so the shadowed rule fires first (only possible if the re-order is acceptable; first rule wins for all matching content, not just the overlap).
  • Tighten the pattern: narrow the shadowed rule so it matches content the earlier rule doesn't.
  • Change the action: if the earlier rule's action is fine for this content, you might not actually need the shadowed rule. Delete it.

Case 2: ML preempts regex (you have to spot this yourself)

When an ML detector and a regex rule cover the same content category (PII, prompt injection, secrets, etc.), the ML layer evaluates first. If ML matches and blocks, the regex rule never sees the request.

This isn't a bug; it's how the layer ordering works. But it can be surprising when:

  • You authored a regex log_only rule expecting to observe injection attempts while still allowing them through to audit, but the ML INJECTION classifier is on and blocks every match before your logonly rule can stamp.
  • You authored a regex mask rule expecting to redact PII, but the ML PII classifier blocks (or masks) first, so your rule's per-pattern stamping doesn't show in request logs.

The control plane doesn't surface this case as a misconfig today. The check: open VidaiGuard status above the rule list. If an ML detector is enabled in the same category as your regex rule, and your regex rule is logonly while the ML is enforcing, the ML is preempting.

To see your regex rule fire instead: either disable the ML detector in that category (likely undesirable, since ML usually catches more than your regex), or accept that the regex rule is a fallback if the ML is later disabled.

📌 Worth knowing. This layer-ordering case explains why the dashboard's Compliance Insights Runway tile can read zero even when you have logonly rules configured. Every observed request is already enforced by the ML layer; there's no observation gap to promote. Honest answer, but worth understanding when you see it.


Reference

Permissions

Role Sees Guardrails page Can do
admin Yes Create / edit / delete / enable rules. Manage VidaiGuard ML toggles.
bi_read_only No Page hidden.
user No Page hidden.

Categories (preset)

Category What it covers
pii Personal identifiable information: SSN, IBAN, Aadhaar, phone, credit card, email, date-of-birth patterns.
secrets API keys and credentials: OpenAI/Anthropic keys, AWS keys, GitHub tokens, private keys, OAuth tokens.
prompt_injection Jailbreak / prompt-injection patterns: "ignore previous instructions," role-confusion, DAN-style.
profanity Offensive language: profanity, slurs.
internal_data Internal infrastructure leakage: RFC1918 IPs, *.internal / *.corp domains, DB connection strings.
custom Free-form bucket for user-authored rules that don't fit the presets.

Severities

Severity Use for
critical Things that should never reach an upstream under any circumstances (live credentials, production secrets, customer-PII clearly out of scope).
high Real policy violations that warrant blocking by default (most secrets, most PII).
medium Policy concerns that warrant logging or masking but not blocking unless context dictates.
low Observations rather than violations: patterns you want to track over time.

Actions

Action What happens Telemetry stamp
block Request returns 403 content_policy_violation with the rule's id + category + severity in the response message. Upstream is never called. guardrail_status=blocked
mask The matched substring is replaced with [REDACTED-{CATEGORY}] in the request body before forwarding. The upstream sees the masked body; the client gets a normal response based on it. guardrail_status=masked
log_only Match is recorded but the request flows unchanged. Useful for observing patterns before promoting them to mask / block. guardrail_status=logged
redirect The target model is rewritten to a guardrail-specified safer model. Routing-rule evaluation is bypassed for the redirected request. The upstream-side call goes to the redirect target. guardrail_status=redirect

📌 Worth knowing. block returns a structured error with type: "content_policy_violation" and code: "content_guardrail_blocked" so client SDKs can distinguish guardrail-blocks from rate-limits or upstream errors.

Composition with user / group / key policies

The rules authored on this page are the library. Whether a rule fires for a given key's traffic depends on the key's effective guardrail policy, which is composed across:

  1. The key's own override (if set on API Keys).
  2. The user / agent's override (if set).
  3. The group / application's policy (set on Groups / Applications).
  4. System default (rules with action block and the default categories enabled).

The merge is wholesale per top-level field: a key overriding rules: [X] replaces the team's rules: [Y, Z] entirely (it doesn't union). See Groups: Guardrail policy for the worked examples.

Audit log records

Every rule action is logged: create, edit (with before/after), enable / disable, delete. The audit also records when the ML sidecar's reachability changes state.

Audit Log shows these.


Limitations

A few things guardrails don't do today:

  • Request-body only. Response-side guardrails (filtering what the upstream returns to the client) aren't supported. If you need to redact something the upstream might say, do it at the application layer.
  • ML inspects last user message only. The ML tier looks at the most recent user turn in a chat conversation, not the full multi-turn history. Multi-turn jailbreak attacks that build context across turns may not trigger; layer regex rules as a backstop if your threat model warrants.
  • Mask redactions are static strings. The replacement is always [REDACTED-{CATEGORY}]; there's no per-rule custom replacement. If you need different redaction behaviour, model it via separate categories.
  • First-match-wins on regex. When several regex rules could match the same content, only one fires (the first in the priority order). The ML tier runs separately. Per-rule firing attribution (knowing every rule that would have matched) isn't surfaced today.
  • No response-time trend telemetry on individual ML models. Aggregate counts are on the dashboard's Compliance panel; per-model latency / throughput surfaces aren't broken out.

Common questions

My custom mask rule isn't firing on test inputs that I know match the pattern. Why?

Two usual suspects: - The ML INJECTION-detection model is firing first on your test content and short-circuiting the request before the regex tier evaluates. The ML model has been observed false-positiving on inputs that contain trigger-shaped tokens (especially data: … or "process this …" style content). Try the same test with the Injection ML model temporarily disabled on the VidaiGuard tab. If your mask rule fires then, the ML model was eating the request first. - The rule isn't in the calling key's effective policy. Check API Keys → effective guardrail config for the key.

A block rule fired and I want to know what content tripped it.

Request Logs → click the row → detail. The guardrail rule id + category are stamped on the row. If you need to inspect the actual content that matched, the request body is preserved in the log detail unless body-redaction is enabled.

Why does log_only still cost me ML inference time?

Because log_only is "observe + continue": the regex tier records the match and lets the request through, so the ML tier still gets to evaluate (and incurs its latency). If the rule was mask or block, the matching ML category would be skipped (deduplication). log_only is therefore the most "pure observability" action.

A redirect rule is firing but the redirect target isn't getting the call.

Probably the calling key isn't allowed to call the redirect target. The control plane re-checks allowed-models after rewriting; if the target isn't in the key's list, the request returns 403 model_not_allowed. Check the key's allowed-models on API Keys.

Can I write a rule that fires on something in the system prompt only, not the user message?

Not directly: patterns match against the whole request body. The body has structure (system / user / assistant turns), but rules don't selectively scope. Workaround: include surrounding context in your pattern to anchor it (e.g. require the match be near a known system-prompt marker).

Why are there 42 preset rules across 5 categories?

The preset library covers the most common compliance / security patterns out of the box: 12 PII, 8 secrets, 7 prompt-injection, 9 profanity, 6 internal-data. The exact list is enabled by default and can be selectively disabled via the team / key policy. Custom rules sit on top.

Where do I see how often each rule has fired?

The dashboard's Compliance Insights panel shows firing counts by category and source (regex / ML). Per-rule firing-count granularity is available via Request Logs filtered by guardrail_rule_id.

The Testing tab from older releases: where did it go?

Removed before general release. To experiment with a new rule, set its action to log_only, send test traffic from a real test key, and review the matching request-logs rows.

Can I integrate my own ML model alongside VidaiGuard?

Today, no: the ML tier slot expects the control-plane-managed sidecar. If you have a specific ML detector you'd like the control plane to integrate with, raise it with your VIDAI support contact.

A guardrail blocked a legitimate customer query. What's the fastest fix?

Two options: - Soft fix: change the rule's action from block to log_only (or mask). Saves immediately; the rule keeps observing without rejecting. - Hard fix: tighten the pattern so it stops matching the legitimate query. Faster than rewriting the policy; just edit the rule.

Request Logs shows the exact text that matched, which usually makes it clear which option applies.


Where to go next

  • Compliance & governance: guardrails are one of the two enforcement engines in the leading compliance story (they enforce the sensitive-data obligation). Read this for how guardrails fit the end-to-end arc.
  • Groups → Guardrail policy: how rules combine across team / user / key. The composition mechanics aren't on this page; they're on Groups.
  • API Keys → Guardrail policy: per-key overrides for "this specific key needs different rules."
  • Compliance: how guardrail outcomes surface in the dashboard's compliance attribution.
  • Routing: composes downstream of guardrails. Guardrail-redirect bypasses routing rules; routing rules fire only on requests that passed guardrails cleanly.
  • Request Logs: every guardrail outcome is recorded per request.
  • Audit Log: every rule edit is logged.