The AI stylist: an agent loop with its own MCP tool server
Most “AI agent” demos are a chat box with a system prompt. The interesting parts of an agent are everything around the model: when it decides to use tools, how those tools are built and reached, what happens when a tool fails, and how you watch it run. This post is the full loop. The next post covers who’s allowed to use which tool (the governance half).
The story
Every chat turn starts on the plain path: a small model, a system prompt grounded in a live catalog snapshot, fast and cheap. A keyword heuristic decides when the turn deserves more:
export function looksLikeInternalQuery(message: string): boolean {
const text = message.toLowerCase();
return (
/\b(return|returns|refund|exchanges)\b/.test(text) ||
/\b(ship|shipped|shipping|delivery|tracking)\b/.test(text) ||
/\bwhere\b.*\bmy order\b/.test(text) ||
/\b(loyalty|member|premium|tier|rewards?)\b/.test(text) ||
/\bcloudflare\b|\bwrangler\b|\bworkers?\b|\b(kv|r2|d1|vectorize|mcp)\b/.test(text) ||
// ... care questions, account questions, docs/deploy intents
);
}
“What’s your return policy for sale items?” matches → agent mode. “What styles well with a camel coat?” stays plain. The expensive path only runs when a tool could genuinely matter.
The agent loop
Agent turns run on llama-3.3-70b-instruct-fp8-fast with a bounded loop — three rounds, no
more:
const client = new PortalMcpClient(env.MCP_PORTAL_URL, env.CF_ACCESS_CLIENT_ID, env.CF_ACCESS_CLIENT_SECRET);
await client.initialize();
const tools = (await client.listTools()).filter((t) => !t.name.startsWith('portal_'));
for (let round = 0; round < MAX_TOOL_ROUNDS; round++) {
const res = await runModel(env, chatMessages, tools, conversationId);
const calls = extractToolCalls(res);
if (!calls.length) return { text: extractContent(res), toolsUsed };
for (const call of calls) {
const output = await client.callTool(call.name, call.args); // → portal → MCP server
chatMessages.push({ role: 'tool', name: call.name, content: output });
}
}
Three details make this feel production-shaped rather than toy-shaped:
The tool roster is discovered, not hardcoded. tools/list returns whatever the portal
currently exposes — Lumina’s internal tools and tools from other registered MCP servers (the
demo includes Cloudflare’s documentation server, so the bot can answer “how do I deploy a
Worker?” from real docs). Add a server in the portal, and the bot gains capabilities with zero
code changes. Portal administration tools (portal_*) are filtered out — the bot consumes
servers, it doesn’t administer the portal.
Failure is honest. A failed tool call produces a tool message that says the capability is unavailable and forbids retrying — so the bot’s answer is “the order-status lookup tool is currently unavailable” rather than an invented policy:
chatMessages.push({
role: 'tool', name: call.name,
content: `Tool "${call.name}" is unavailable right now. Do not retry it; answer without it
and be transparent that this capability is currently disabled.`,
});
The system prompt forbids invention. “Never invent policy, prices, or order data. If a
tool fails or returns nothing useful, say so plainly.” Combined with tool results that carry
source keys (source: returns-and-exchanges.md), answers come back grounded and citable.
The tool server (a second Worker)
lumina-internal-mcp is a small Worker exposing three tools over the MCP protocol
(Streamable HTTP), each returning text the model can cite:
| Tool | Backend | Notes |
|---|---|---|
search_internal_kb(query) |
AI Search (AutoRAG) | Scored chunks + source keys; “no results” instructs the model to say so |
check_order_status(order_id) |
D1 orders |
Translates confirmed/fulfilled/cancelled into customer language |
get_customer_profile(email) |
D1 users |
Name, tier, synthesized CRM notes |
The server fails closed — no bearer credential configured means 401 for everything — and its governance model is deliberately thin: it registers all tools statically and lets the portal decide what clients may see (that’s the next post).
Watching the agent: GenAI traces
The loop instruments itself with Workers’ custom spans API using OpenTelemetry GenAI attributes — the structure that makes an agent surface in the dashboard Agents tab:
return tracing.enterSpan('invoke_agent', async (span) => {
span.setAttributes({
'gen_ai.operation.name': 'invoke_agent',
'gen_ai.agent.name': 'lumina-stylist',
'gen_ai.agent.id': 'lumina-stylist-production',
'gen_ai.conversation.id': conversationId,
});
return runAgentLoop(env, conversation, plan, conversationId);
});
Each model call becomes a chat child span; each portal tool call an execute_tool span with
gen_ai.tool.name, arguments and result. One conversationId (generated in the browser) ties
a multi-turn conversation into one replayable session — sessions, span waterfall and token
usage in the dashboard, no third-party APM. Plain-chat turns get an invoke_agent + chat
pair too, so all chat traffic is visible in the same place.
What the demo shows
Ask “What is your return policy for sale items?” → the chat meta line grows a tool badge
(🔧 lumina_search_internal_kb) and the answer cites the KB doc. Ask “Where is my order
ord_…?” (after placing one) → real order data. Ask a Cloudflare question → the docs server’s
tool fires. Toggle a tool off in the portal (next post) → the bot explains it can’t verify.
Evidence: what to capture
- Agent chat turn with the tool badge in the meta line + grounded answer →
06-agent-tool-badge.png - Order-status turn with a real
ord_…id →06-order-status.png - Cloudflare-docs turn →
06-docs-tool.png - Dashboard → Workers & Pages → the retail Worker → Agents tab: session list →
06-agents-tab.png - One session opened: span waterfall (
invoke_agent→chat/execute_tool) →06-span-waterfall.png - The agent-mcp-path diagram: open
Blog/src/assets/diagrams/agent-mcp-path.excalidrawin Excalidraw, export PNG →06-ai-stylist-agent/agent-mcp-path.png(embedded above)
Key takeaways
- Gate agent mode with a cheap heuristic — the 70B model only runs when tools can matter.
- Discover the tool roster at runtime from the portal; agents gain/lose capabilities without deploys.
- Honesty is a feature: failed tools and empty results are fed back as messages, never papered over.
invoke_agent→chat/execute_toolGenAI spans make any custom agent loop show up in the Agents tab.