Veridian's underwriting team had an LLM reading loan application documents and extracting structured risk signals — a strong use case, and it worked well in staging. Each call cost a fraction of a cent, so nobody modeled what it would look like once every application, every amendment, and every re-submission ran through the same pipeline.
By the time it reached production volume, the pipeline was making hundreds of thousands of calls a day, almost all of them at the same model tier regardless of how straightforward the document was. The monthly inference bill had become a line item finance was asking about — and growing faster than application volume, because retries and long-context re-reads were quietly adding up underneath it.