Case Study — LLM Integration & MLOps

Cutting inference cost on an
underwriting pipeline that outgrew its budget.

Client name withheld under NDA. This is a real engagement — the client name "Veridian Capital" is a placeholder, and some identifying details have been generalized, under the confidentiality terms of our agreement with the actual client.
Client: Veridian Capital (name changed) Industry: Fintech underwriting Timeline: ~5 weeks Related: LLM Integration, MLOps
The problem

A $0.002 call doesn't look like a budget line. Until it is one.

Veridian's underwriting team had an LLM reading loan application documents and extracting structured risk signals — a strong use case, and it worked well in staging. Each call cost a fraction of a cent, so nobody modeled what it would look like once every application, every amendment, and every re-submission ran through the same pipeline.

By the time it reached production volume, the pipeline was making hundreds of thousands of calls a day, almost all of them at the same model tier regardless of how straightforward the document was. The monthly inference bill had become a line item finance was asking about — and growing faster than application volume, because retries and long-context re-reads were quietly adding up underneath it.

What it was built on.

The pipeline as we found it:

Single top-tier LLM for every document Full document text on every call, no trimming No response caching Synchronous, uncapped retries on failure No per-call cost tracking
The approach

Spend the expensive model only where it earns its cost.

None of this needed a different LLM provider — it needed the pipeline to stop treating every document identically:

The bill stopped tracking volume 1:1.

Figures are approximate and rounded to protect the client's specific financials, per our NDA.

Before
Per-document inference cost rising in lockstep with application volume
After
Roughly 70% lower inference cost per document, same extraction accuracy
Ongoing
Cost per call tracked and alerted on, not discovered at month-end
Related

If this looks familiar.

This is the same failure mode we write about in why LLM inference costs blow up at scale, and the kind of problem our LLM integration and MLOps work is built to catch before the invoice does.

Get started

Tell us what
you're trying to build.

Book a 30-minute call — we'll tell you honestly whether your inference bill has room to shrink, and by how much.

Book a free 30-min call → More case studies →