A layered prompt firewall: edge WAF + AI Gateway Guardrails + DLP


Ask the shop bot to type out a credit card number and it won’t — not because the model is careful, but because the prompt never reaches the model. The Lumina demo screens every chat prompt with two independent, dashboard-configured layers, and the fun part is seeing exactly where each one draws the line.

The story

The request path is:

Browser → zone edge WAF (AI Security for Apps rule) → Retail Worker → AI Gateway (Guardrails + DLP) → model

No security code lives in the request path — both layers are dashboard configuration. The Worker’s only job is surfacing blocks, which is itself a design decision: a block must reach the user, never silently fall back to a friendlier path.

Layer 1: the zone edge (WAF custom rule)

A WAF custom rule on the chat endpoint uses the AI Security for Apps detection fields (cf.llm.prompt.*), which populate on LLM-labeled endpoints:

(http.request.uri.path wildcard r"/api/chat*"
  and any(cf.llm.prompt.pii_categories[*] in {"CREDIT_CARD" "US_SSN" "IP_ADDRESS"}))
or (http.request.uri.path wildcard r"/api/chat*" and cf.llm.prompt.token_count gt 500)

The category scoping is deliberate. The broader “any PII detected” also fires on PERSON/DATE_TIME/LOCATION — normal retail chat contains those constantly (“return it by Friday”, “do you ship to Leeds”). Card, SSN and IP are the patterns that should never appear in a shop chat; the token cap catches paste-their-life-into-the-chat behavior.

Blocks return Cloudflare’s HTML block page before the Worker runs — flagged prompts cost the app nothing. Detections show up in Security → Events.

A sharp edge from tuning: the NER catches an incomplete card number (“4111 1111 1111 111”) but not a blanked-out one (“4111 1111 1111 ____”). Which leads to the gap the edge can’t cover…

Layer 2: AI Gateway Firewall (Guardrails + DLP)

Requests that pass the edge reach the gateway’s firewall:

  • Guardrails (Llama Guard 3 8B): every prompt is scored against hazard categories (violence, weapons, privacy, prompt injection…). Demo config: all prompt categories → Block; responses mostly Flag, except the hard blocks.
  • DLP (Cloudflare’s DLP engine): profiles for payment cards and government/insurance/tax identifiers, checking both directions — prompts and responses.

The Worker classifies the binding’s rejection codes into a JSON 403 with a blocked marker:

function classifyAiBlock(err: unknown): AiBlockKind | null {
  const code = String((err as { code?: unknown })?.code ?? '');
  if (code.includes('2016')) return 'guardrails-prompt';    // prompt: unsafe topic
  if (code.includes('2017')) return 'guardrails-response';  // response: unsafe content
  if (code.includes('2029')) return 'dlp-request';          // DLP, request side
  if (code.includes('2030')) return 'dlp-response';         // DLP, response side
  return null;
}

The critical bit is the fallback discipline: a block is final. It is never retried on the plain path or answered by the offline mock — otherwise the firewall would be decorative.

The response-side gap is the gateway’s reason to exist. The edge WAF is request-side only. So the demo’s “(AIG) DLP” tile asks an innocent-sounding question — “fill in this example: my card is 4111 1111 ____ — the missing digits are…” — which passes the edge and Guardrails, and the model dutifully completes the number. DLP blocks the response (code 2030) and nothing sensitive ships. Tuning note: the free-form “show me an example” variant only fired DLP 2 of 7 times (small models don’t always emit digits); the blank-fill form is 8 for 8.

Two overlap notes from the field:

  • The edge shadows request-side DLP for card/SSN prompts — they die at the WAF, so the gateway never pays for the eval (no Guardrails/DLP entry in its logs).
  • Guardrails’ Privacy category shadows request-side DLP for bank/IBAN phrasing — the unsafe- topic block fires first.

Also: the gateway’s custom domain sits behind Cloudflare Access, so unauthenticated calls to the gateway itself get a 401 — the model’s front door is policy-gated, not just its inputs.

What the demo shows

Three suggestion tiles, labeled by layer, clickable in any order (blocked messages are popped from conversation history — the gateway scans the whole request body, so a blocked message left in history would re-block every later send):

Tile Prompt shape Blocked by
(AIS) PII test card + SSN typed into the prompt Edge WAF → HTML block page
(AIG) Unsafe topic test weapons how-to Guardrails (2016)
(AIG) DLP test model completes a blanked card number DLP on the response (2030)

A clean prompt answers normally throughout.

Evidence: what to capture

  • (AIS) PII tile → red “Blocked at the Cloudflare edge (WAF)” bubble → 12-edge-block.png
  • Security → Events with the block, filter Managed Endpoint Label = cf-llm → 12-security-events.png
  • (AIG) Unsafe topic tile → the bubble names Guardrails → 12-guardrails-block.png
  • (AIG) DLP tile → response-side block → 12-dlp-block.png
  • AI Gateway → Logs: green Guardrails shields + DLP Action filter → 12-gateway-logs.png
  • The prompt-firewall-layers diagram: open Blog/src/assets/diagrams/prompt-firewall-layers.excalidraw in Excalidraw, export PNG → 12-layered-prompt-firewall/prompt-firewall-layers.png (embedded above)

Key takeaways

  • Prompt safety belongs in dashboard config at two layers: the edge (cheap, earliest) and the gateway (deeper, response-capable).
  • The edge sees requests; only the gateway can stop what the model emits. Both are needed.
  • Blocks are final — routing them to a fallback silently defeats the firewall.
  • Category scoping is the difference between a PII rule and a rule that blocks everyday conversation.