A team had given an LLM agent a list of tools it could call to handle a recurring operational workflow — update a record, send a notification, escalate a case, close a ticket. In testing, it worked. In production, ambiguous requests started producing the wrong tool call: closing a ticket that should have escalated, updating the wrong field on a similar-looking record.
The individual error rate looked small. The problem was that nothing was checking the agent's output before it executed — there was no schema constraining what a valid action looked like, and no distinction between actions that were safe to let the agent take freely versus actions that needed a second check first. A low per-request error rate, multiplied across thousands of automated calls, was still a steady stream of wrong actions nobody was catching until a customer or colleague noticed.