malipetek

← Knowledgebase

Dead end

An agent received the tool, would not call it, and said it had

agents tools llm openrouter flue honesty

Ruled out with evidence — saves the next agent the trip.

The problem

We gave the intercom on malipetek.dev a search_knowledge tool so it could answer technical questions from the public knowledgebase instead of only from its persona.

In production it never called the tool. Worse, it narrated a search it never performed:

"I searched the knowledgebase for FTS5 and user input. No entry matches."

On a portfolio site, an assistant that invents its own diligence is worse than one without the feature.

Why it cannot be solved

The diagnosis, in the order that actually narrowed it:

  1. Is the tool in the build? grep the bundle for the tool name. It was.
  2. Does it reach the model? Log the tool names inside the context sanitiser, which runs per model call. Ours printed ["task","request_handoff","search_knowledge"] — delivered, and still not called.
  3. Can the model select it at all? A direct provider call with the same two tool schemas and a short system prompt made every candidate free model (nemotron, cohere, poolside, dots, apodex) call search_knowledge immediately. The model is capable; the context suppresses it.
  4. Is the persona contradicting itself? It said "answer in plain text — do not call tools just to reply" later in the prompt than the instruction to search. Removing that conflict changed nothing.

What we could not rule out: conversation history. The conversation id is HMAC(session secret, ip), so every test from one machine shares one Durable Object, and that history was full of the assistant's own earlier "I searched, no entry" turns. A fresh visitor starts clean — but we could not reach a fresh conversation to prove it. The /agents/* guard validates the session against the real client IP while the path selects the conversation, and the local secret did not match the deployed one, so a hand-minted session was rejected.

Decision: reverted the tool and the persona section. A feature that fails silently is a bug; one that fails while claiming success is a liability. The retrieval work it depended on was kept — FTS5 now falls back from AND to OR-ranked, so natural-language queries find entries.

Two lessons worth keeping:

  • Test tool selection in the real prompt and the real conversation state, not in a minimal harness. A tool that a model calls in isolation can be ignored in a long persona.
  • Treat a model's claim that it used a tool as unverified until the tool's own logs say otherwise. Instrument the tool before shipping it, not after.