Tool — Speech-to-Text

Transcription,
hosted or self-run.

Whisper is OpenAI's speech-to-text model — and unlike a pure API service, it's also available as an open-weight model you can self-host. We reach for that self-hosted path when per-minute transcription costs at volume, or data-residency rules, make a hosted API the wrong fit.

Serving UK and EU clients: GDPR, EU AI Act, and data residency are covered on our Trust & Safety page.

What it is

Open weights,
real flexibility.

Whisper is notable among speech-to-text options for being open-weight — it can run as a hosted API call or be deployed entirely inside a client's own infrastructure. For high-volume transcription where cost compounds fast, or for clients where audio can't leave their environment, self-hosting is a real option most proprietary transcription services don't offer.

For real-time, low-latency conversation — like a live voice agent — we more often reach for Vapi's Deepgram integration instead; Whisper tends to be the better fit for batch or privacy-sensitive transcription.

How we build it

Matched to the job,
not a default.

Whisper and Deepgram solve overlapping but distinct problems.

01Self-hosted deployment

Running Whisper inside a client's own infrastructure when audio data can't leave their environment.

02Batch transcription at scale

High-volume, non-real-time transcription where self-hosting controls cost better than per-minute API pricing.

03Real-time alternative

For live conversation, we typically route to Deepgram via Vapi instead — lower latency is the priority there.

Where this fits

Where this fits

One option in our speech-to-text toolkit, chosen by the specific requirement.

Common questions

Before you
book a call.

The questions we get asked most about Whisper and speech-to-text — answered straight, no sales pitch.

Why would you self-host Whisper instead of using a hosted transcription API?

Mainly cost at volume and data residency — self-hosting avoids per-minute API charges at scale and keeps audio data inside a client's own infrastructure when that's a requirement.

Do you use Whisper for live voice agents?

Usually not — for real-time conversation we typically use Deepgram through Vapi for lower latency. Whisper tends to be the better fit for batch transcription or privacy-sensitive, non-real-time use cases.

Is self-hosted Whisper as accurate as a hosted API?

Accuracy is comparable since it's the same underlying model — the difference is in deployment, latency characteristics, and cost structure, not transcription quality.

Get started

Tell us what
you're trying to transcribe.

Book a 30-minute call — we'll tell you honestly whether self-hosted Whisper fits your use case, or whether a hosted option is simpler.