The agent-ready storefront: one Worker, two audiences


A storefront SPA is invisible to machines: one JavaScript app at one URL, no pages to crawl, no markup to parse. But your customers’ agents — ChatGPT, Claude, Perplexity — increasingly shop, compare and recommend on their users’ behalf. The Lumina demo serves both audiences from the same Worker: humans keep the SPA; agents get real URLs, semantic markup and machine-readable data.

The story

The insight that shaped the whole surface: an agent that can’t read your store invents its own version of it. Machine-readable surfaces fix that, and each one has a distinct job:

Surface What it does
/robots.txt Content-Signal: search=yes, ai-input=yes, ai-train=no, use=reference — agents may read and cite the content, not train on it. Allow: /api/ai/ sits above the Disallow: /api/
/sitemap.xml Generated from D1 on every request (19 URLs) — it can never drift from the catalog because it is the catalog
/llms.txt + /llms-full.txt An agent-facing index plus a complete Markdown store description generated from D1 + the KB summaries
/openapi.json An OpenAPI 3.1 contract of the agent surface — an agent reads the contract before touching the store
/api/ai/* Public JSON mirrors: catalog, products, collections, shipping, returns
/products/:slug Server-rendered pages with canonical links, meta descriptions and schema.org Product + BreadcrumbList JSON-LD
Accept: text/markdown Every content page negotiates a Markdown rendering for agent consumers

How it works

The sitemap is a query, not a file:

const paths = ['/', '/collections', '/featured', ...COLLECTIONS.map((c) => `/collections/${c.slug}`),
               ...products.map((p) => productPath(p))];
const xml = paths.map((p) => `<url><loc>${origin}${p}</loc></url>`).join('\n');

Markdown negotiation is one header check per page — the same handler serves HTML or Markdown:

function wantsMarkdown(request: Request): boolean {
  return (request.headers.get('accept') ?? '').includes('text/markdown');
}

Product pages render Product + BreadcrumbList JSON-LD in one script tag (USD offers with InStock/OutOfStock, material/color/size), and the published shipping/returns pages are distilled from the same internal KB the chat bot cites — so the bot’s answers and the public pages share one source of truth. Verify it from a terminal:

curl -H "Accept: text/markdown" https://retail.mattwynne.solutions/products/merino-coat
curl https://retail.mattwynne.solutions/robots.txt
curl https://retail.mattwynne.solutions/openapi.json

The validation war story: logs beat assumptions

An external review kept reporting the discovery endpoints “inaccessible” while crawling the HTML fine. The instinct is to fiddle with robots.txt — the right move was to ask the zone’s own logs. GraphQL over Security Events answered both halves:

  1. Every crawler request that reached the edge returned 200 — including a full GPTBot crawl session (robots, HTML pages, product pages, media).
  2. The “inaccessible” URLs never appear in the logs from that crawler’s ASN at all — the browsing tool declined those fetches before any HTTP request was made. The verdicts were retrieval-layer limits, not site controls.

The investigation did surface two real finds: a zone-wide HEAD/OPTIONS block (a managed WAF rule — meaning every curl -I verification was doomed from the start; an exception was scoped, and the app’s read-method guard widened to accept GET ∪ HEAD), and an intermittent www.mattwynne.solutions/robots.txt 525 (a proxied www record with no Worker route — tracked for cleanup). Both were found by logs, not by guessing.

The dashboard piece that closes the loop: AI Crawl Control (Security → AI Crawl Control) left in “allow” mode for the major crawlers — robots.txt + Content-Signal govern them — and its telemetry shows which crawlers arrived, which paths they hit, and their robots.txt compliance.

What the demo shows

Browse the site as a machine: /robots.txt (Content-Signal + the /api/ai/ carve-out), /sitemap.xml (19 URLs from D1), /llms.txt and /llms-full.txt, /openapi.json, view-source on a product page for the JSON-LD, and curl -H "Accept: text/markdown" for the same product as clean Markdown. Then AI Crawl Control for the “did the bots show up” dashboard.

Evidence: what to capture

  • https://retail.mattwynne.solutions/robots.txt showing Content-Signal → 14-robots.png
  • /sitemap.xml showing D1-derived product URLs → 14-sitemap.png
  • /llms-full.txt excerpt → 14-llms-full.png
  • /openapi.json in a viewer (or DevTools JSON view) → 14-openapi.png
  • View-source on /products/merino-coat showing Product + BreadcrumbList JSON-LD → 14-json-ld.png
  • Terminal: curl -H "Accept: text/markdown" .../products/merino-coat → 14-markdown-curl.png
  • Security → AI Crawl Control dashboard → 14-crawl-control.png
  • The agent-surface-map diagram: open Blog/src/assets/diagrams/agent-surface-map.excalidraw in Excalidraw, export PNG → 14-agent-ready-storefront/agent-surface-map.png (embedded above)

Key takeaways

  • Same origin, two audiences: the SPA for humans, a crawlable information architecture for agents.
  • Generated-from-D1 beats hand-maintained: sitemap, llms-full and JSON mirrors can’t drift from the catalog.
  • Content-Signal lets you say “read me, don’t train on me” in the one file every crawler starts with.
  • When a validator says “inaccessible”, query Security Events before touching config — logs told the whole story here.