The AI stylist: an agent loop with its own MCP tool server


Most “AI agent” demos are a chat box with a system prompt. The interesting parts of an agent are everything around the model: when it decides to use tools, how those tools are built and reached, what happens when a tool fails, and how you watch it run. This post is the full loop. The next post covers who’s allowed to use which tool (the governance half).

The story

Every chat turn starts on the plain path: a small model, a system prompt grounded in a live catalog snapshot, fast and cheap. A keyword heuristic decides when the turn deserves more:

export function looksLikeInternalQuery(message: string): boolean {
  const text = message.toLowerCase();
  return (
    /\b(return|returns|refund|exchanges)\b/.test(text) ||
    /\b(ship|shipped|shipping|delivery|tracking)\b/.test(text) ||
    /\bwhere\b.*\bmy order\b/.test(text) ||
    /\b(loyalty|member|premium|tier|rewards?)\b/.test(text) ||
    /\bcloudflare\b|\bwrangler\b|\bworkers?\b|\b(kv|r2|d1|vectorize|mcp)\b/.test(text) ||
    // ... care questions, account questions, docs/deploy intents
  );
}

“What’s your return policy for sale items?” matches → agent mode. “What styles well with a camel coat?” stays plain. The expensive path only runs when a tool could genuinely matter.

The agent loop

Agent turns run on llama-3.3-70b-instruct-fp8-fast with a bounded loop — three rounds, no more:

const client = new PortalMcpClient(env.MCP_PORTAL_URL, env.CF_ACCESS_CLIENT_ID, env.CF_ACCESS_CLIENT_SECRET);
await client.initialize();
const tools = (await client.listTools()).filter((t) => !t.name.startsWith('portal_'));

for (let round = 0; round < MAX_TOOL_ROUNDS; round++) {
  const res = await runModel(env, chatMessages, tools, conversationId);
  const calls = extractToolCalls(res);
  if (!calls.length) return { text: extractContent(res), toolsUsed };

  for (const call of calls) {
    const output = await client.callTool(call.name, call.args);   // → portal → MCP server
    chatMessages.push({ role: 'tool', name: call.name, content: output });
  }
}

Three details make this feel production-shaped rather than toy-shaped:

The tool roster is discovered, not hardcoded. tools/list returns whatever the portal currently exposes — Lumina’s internal tools and tools from other registered MCP servers (the demo includes Cloudflare’s documentation server, so the bot can answer “how do I deploy a Worker?” from real docs). Add a server in the portal, and the bot gains capabilities with zero code changes. Portal administration tools (portal_*) are filtered out — the bot consumes servers, it doesn’t administer the portal.

Failure is honest. A failed tool call produces a tool message that says the capability is unavailable and forbids retrying — so the bot’s answer is “the order-status lookup tool is currently unavailable” rather than an invented policy:

chatMessages.push({
  role: 'tool', name: call.name,
  content: `Tool "${call.name}" is unavailable right now. Do not retry it; answer without it
            and be transparent that this capability is currently disabled.`,
});

The system prompt forbids invention. “Never invent policy, prices, or order data. If a tool fails or returns nothing useful, say so plainly.” Combined with tool results that carry source keys (source: returns-and-exchanges.md), answers come back grounded and citable.

The tool server (a second Worker)

lumina-internal-mcp is a small Worker exposing three tools over the MCP protocol (Streamable HTTP), each returning text the model can cite:

Tool Backend Notes
search_internal_kb(query) AI Search (AutoRAG) Scored chunks + source keys; “no results” instructs the model to say so
check_order_status(order_id) D1 orders Translates confirmed/fulfilled/cancelled into customer language
get_customer_profile(email) D1 users Name, tier, synthesized CRM notes

The server fails closed — no bearer credential configured means 401 for everything — and its governance model is deliberately thin: it registers all tools statically and lets the portal decide what clients may see (that’s the next post).

Watching the agent: GenAI traces

The loop instruments itself with Workers’ custom spans API using OpenTelemetry GenAI attributes — the structure that makes an agent surface in the dashboard Agents tab:

return tracing.enterSpan('invoke_agent', async (span) => {
  span.setAttributes({
    'gen_ai.operation.name': 'invoke_agent',
    'gen_ai.agent.name': 'lumina-stylist',
    'gen_ai.agent.id': 'lumina-stylist-production',
    'gen_ai.conversation.id': conversationId,
  });
  return runAgentLoop(env, conversation, plan, conversationId);
});

Each model call becomes a chat child span; each portal tool call an execute_tool span with gen_ai.tool.name, arguments and result. One conversationId (generated in the browser) ties a multi-turn conversation into one replayable session — sessions, span waterfall and token usage in the dashboard, no third-party APM. Plain-chat turns get an invoke_agent + chat pair too, so all chat traffic is visible in the same place.

What the demo shows

Ask “What is your return policy for sale items?” → the chat meta line grows a tool badge (🔧 lumina_search_internal_kb) and the answer cites the KB doc. Ask “Where is my order ord_…?” (after placing one) → real order data. Ask a Cloudflare question → the docs server’s tool fires. Toggle a tool off in the portal (next post) → the bot explains it can’t verify.

Evidence: what to capture

  • Agent chat turn with the tool badge in the meta line + grounded answer → 06-agent-tool-badge.png
  • Order-status turn with a real ord_… id → 06-order-status.png
  • Cloudflare-docs turn → 06-docs-tool.png
  • Dashboard → Workers & Pages → the retail Worker → Agents tab: session list → 06-agents-tab.png
  • One session opened: span waterfall (invoke_agent → chat / execute_tool) → 06-span-waterfall.png
  • The agent-mcp-path diagram: open Blog/src/assets/diagrams/agent-mcp-path.excalidraw in Excalidraw, export PNG → 06-ai-stylist-agent/agent-mcp-path.png (embedded above)

Key takeaways

  • Gate agent mode with a cheap heuristic — the 70B model only runs when tools can matter.
  • Discover the tool roster at runtime from the portal; agents gain/lose capabilities without deploys.
  • Honesty is a feature: failed tools and empty results are fed back as messages, never papered over.
  • invoke_agent → chat/execute_tool GenAI spans make any custom agent loop show up in the Agents tab.