Token-based pricing is deceptive at small scale. $0.002 per call is a number nobody budgets around — it's effectively free in a demo, in staging, in the first week of soft launch. The problem is that "per call" cost was never the number that mattered. The number that matters is cost at your actual production volume, and that number is usually invisible until the first full month's invoice arrives.
It gets worse in three predictable ways that rarely show up in a quick back-of-envelope estimate:
- Retries and chained calls. An agent or pipeline that makes 2–4 model calls per user request multiplies the per-call cost before you've even accounted for volume.
- Context growth. A chat or RAG system that keeps accumulating conversation history or retrieved context pushes input tokens up over a session — the tenth message in a conversation costs more than the first.
- Silent model upgrades. A provider's "recommended" model sometimes has a higher per-token price than what you benchmarked against. If nothing is pinned or monitored, cost per call can quietly rise.