Case Studies

How we've approached
AI failure modes.

Each scenario below walks through a problem, the stack, the approach, and the outcome pattern we've seen when it's fixed — the same failure modes described in What we've seen break, in more depth.

Three of these are illustrative scenarios, composed from patterns we've diagnosed across multiple engagements — not write-ups of a single named client project. The LLM inference cost case study is different: it's a real, single-client engagement, with the client's name changed and some details generalized under NDA.
RAG · Retrieval
Fixing a RAG pipeline that answered confidently and wrongly

No retrieval eval layer meant bad chunks went straight to the model, which hallucinated on top of them. Here's the fix, step by step.

Read the scenario →
AI Agents · Guardrails
Stopping an LLM agent from taking the wrong action at scale

No output schema enforcement meant ambiguous prompts led to the wrong tool call — automatically, repeatedly. Here's how we scoped it back.

Read the scenario →
Computer Vision · Edge
Closing the gap between lab accuracy and field accuracy

A vision model that hit 94% in testing dropped to 61% on the actual production line. Here's what the gap was made of, and how we closed it.

Read the scenario →
LLM Integration · MLOps
Cutting inference cost on an underwriting pipeline that outgrew its budget

A real client engagement (name withheld under NDA) — model routing, context trimming, and caching cut per-document inference cost roughly 70%.

Read the case study →
Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll tell you honestly whether your situation looks like one of these, and what it would take to fix it.

Book a free 30-min call → Other ways to reach us →