Case Study — AI Agent Development

Stopping an agent from taking
the wrong action, at scale.

Illustrative scenario. Composed from patterns we've diagnosed across multiple agent engagements — not a write-up of one named client project. See our case studies note for why.
Industry: Internal operations automation Timeline: ~6 weeks Related: AI Agent Development
The problem

The agent wasn't wrong often. It was wrong automatically.

A team had given an LLM agent a list of tools it could call to handle a recurring operational workflow — update a record, send a notification, escalate a case, close a ticket. In testing, it worked. In production, ambiguous requests started producing the wrong tool call: closing a ticket that should have escalated, updating the wrong field on a similar-looking record.

The individual error rate looked small. The problem was that nothing was checking the agent's output before it executed — there was no schema constraining what a valid action looked like, and no distinction between actions that were safe to let the agent take freely versus actions that needed a second check first. A low per-request error rate, multiplied across thousands of automated calls, was still a steady stream of wrong actions nobody was catching until a customer or colleague noticed.

What it was built on.

A representative setup for this kind of workflow automation — illustrative of the pattern, not a specific project's exact tools:

LLM with function/tool calling Internal REST APIs as tools No output schema validation No action-level permissions No human-in-the-loop checkpoint
The approach

Guardrails before more capability.

Rather than trying to make the model smarter at disambiguating intent, we constrained what it was allowed to do when it wasn't sure:

What changed once actions were gated.

This is the pattern of improvement we typically see once schema enforcement and gating go in — illustrative, not a specific client's measured figures:

Before
Wrong actions executed silently, caught only after the fact
After
Malformed or out-of-scope calls rejected before execution
Ongoing
Irreversible actions confirmed, not assumed
Related

If this looks familiar.

This is a pattern-level breakdown of what our AI agent development work is built to prevent. If your agent is taking occasional wrong actions in production, the fix usually isn't a smarter model — it's tighter constraints on what it's allowed to do.

Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll tell you honestly whether your agent has this gap, and what it would take to close it.

Book a free 30-min call → More case studies →