A layered prompt firewall: edge WAF + AI Gateway Guardrails + DLP
Ask the shop bot to type out a credit card number and it won’t — not because the model is careful, but because the prompt never reaches the model. The Lumina demo screens every chat prompt with two independent, dashboard-configured layers, and the fun part is seeing exactly where each one draws the line.
The story
The request path is:
Browser → zone edge WAF (AI Security for Apps rule) → Retail Worker → AI Gateway (Guardrails + DLP) → model
No security code lives in the request path — both layers are dashboard configuration. The Worker’s only job is surfacing blocks, which is itself a design decision: a block must reach the user, never silently fall back to a friendlier path.
Layer 1: the zone edge (WAF custom rule)
A WAF custom rule on the chat endpoint uses the AI Security for Apps detection fields
(cf.llm.prompt.*), which populate on LLM-labeled endpoints:
(http.request.uri.path wildcard r"/api/chat*"
and any(cf.llm.prompt.pii_categories[*] in {"CREDIT_CARD" "US_SSN" "IP_ADDRESS"}))
or (http.request.uri.path wildcard r"/api/chat*" and cf.llm.prompt.token_count gt 500)
The category scoping is deliberate. The broader “any PII detected” also fires on
PERSON/DATE_TIME/LOCATION — normal retail chat contains those constantly (“return it by
Friday”, “do you ship to Leeds”). Card, SSN and IP are the patterns that should never appear
in a shop chat; the token cap catches paste-their-life-into-the-chat behavior.
Blocks return Cloudflare’s HTML block page before the Worker runs — flagged prompts cost the app nothing. Detections show up in Security → Events.
A sharp edge from tuning: the NER catches an incomplete card number (“4111 1111 1111 111”) but not a blanked-out one (“4111 1111 1111 ____”). Which leads to the gap the edge can’t cover…
Layer 2: AI Gateway Firewall (Guardrails + DLP)
Requests that pass the edge reach the gateway’s firewall:
- Guardrails (Llama Guard 3 8B): every prompt is scored against hazard categories (violence, weapons, privacy, prompt injection…). Demo config: all prompt categories → Block; responses mostly Flag, except the hard blocks.
- DLP (Cloudflare’s DLP engine): profiles for payment cards and government/insurance/tax identifiers, checking both directions — prompts and responses.
The Worker classifies the binding’s rejection codes into a JSON 403 with a blocked marker:
function classifyAiBlock(err: unknown): AiBlockKind | null {
const code = String((err as { code?: unknown })?.code ?? '');
if (code.includes('2016')) return 'guardrails-prompt'; // prompt: unsafe topic
if (code.includes('2017')) return 'guardrails-response'; // response: unsafe content
if (code.includes('2029')) return 'dlp-request'; // DLP, request side
if (code.includes('2030')) return 'dlp-response'; // DLP, response side
return null;
}
The critical bit is the fallback discipline: a block is final. It is never retried on the plain path or answered by the offline mock — otherwise the firewall would be decorative.
The response-side gap is the gateway’s reason to exist. The edge WAF is request-side only. So the demo’s “(AIG) DLP” tile asks an innocent-sounding question — “fill in this example: my card is 4111 1111 ____ — the missing digits are…” — which passes the edge and Guardrails, and the model dutifully completes the number. DLP blocks the response (code 2030) and nothing sensitive ships. Tuning note: the free-form “show me an example” variant only fired DLP 2 of 7 times (small models don’t always emit digits); the blank-fill form is 8 for 8.
Two overlap notes from the field:
- The edge shadows request-side DLP for card/SSN prompts — they die at the WAF, so the gateway never pays for the eval (no Guardrails/DLP entry in its logs).
- Guardrails’ Privacy category shadows request-side DLP for bank/IBAN phrasing — the unsafe- topic block fires first.
Also: the gateway’s custom domain sits behind Cloudflare Access, so unauthenticated calls to the gateway itself get a 401 — the model’s front door is policy-gated, not just its inputs.
What the demo shows
Three suggestion tiles, labeled by layer, clickable in any order (blocked messages are popped from conversation history — the gateway scans the whole request body, so a blocked message left in history would re-block every later send):
| Tile | Prompt shape | Blocked by |
|---|---|---|
| (AIS) PII test | card + SSN typed into the prompt | Edge WAF → HTML block page |
| (AIG) Unsafe topic test | weapons how-to | Guardrails (2016) |
| (AIG) DLP test | model completes a blanked card number | DLP on the response (2030) |
A clean prompt answers normally throughout.
Evidence: what to capture
- (AIS) PII tile → red “Blocked at the Cloudflare edge (WAF)” bubble →
12-edge-block.png - Security → Events with the block, filter
Managed Endpoint Label = cf-llm→12-security-events.png - (AIG) Unsafe topic tile → the bubble names Guardrails →
12-guardrails-block.png - (AIG) DLP tile → response-side block →
12-dlp-block.png - AI Gateway → Logs: green Guardrails shields +
DLP Actionfilter →12-gateway-logs.png - The prompt-firewall-layers diagram: open
Blog/src/assets/diagrams/prompt-firewall-layers.excalidrawin Excalidraw, export PNG →12-layered-prompt-firewall/prompt-firewall-layers.png(embedded above)
Key takeaways
- Prompt safety belongs in dashboard config at two layers: the edge (cheap, earliest) and the gateway (deeper, response-capable).
- The edge sees requests; only the gateway can stop what the model emits. Both are needed.
- Blocks are final — routing them to a fallback silently defeats the firewall.
- Category scoping is the difference between a PII rule and a rule that blocks everyday conversation.