SPEECH TO TEXT / JULY 2026

How many words
fall through the cracks?

Compare transcription errors, throughput, API cost, and the open-weight models that can listen locally on your Mac.

00:04.2

The frontier is not one model.

00:08.7

It is the edge between cost speed and accuracy.

00:13.1

Every dropped word changes the meaning.

1 substitution / 20 words · 5.0% WER
01 / ACCURACYScribe v2

2.2% WER · ElevenLabs

02 / BALANCEMAI-Transcribe-1.5

2.4% WER · 265.9× realtime

03 / SPEEDParakeet 0.6B v3

904.9× realtime · open weights

04 / LOW COSTStepAudio 2.5

$0.37 / 1,000 minutes

THE ERROR MAP

Fewer mistakes, less friction.

Word error rate ↓ lower is better
2%
3%
4%
5%
6%
$0.25
$0.5
$1
$2
$4
$8
$16
API price per 1,000 audio minutes → log scale
HOSTED MODEL

MAI-Transcribe-1.5

Microsoft

AA-WER
2.4%
Speed
265.9×
Price
$6 / 1K min
Frontier
No

PRIVATE TRANSCRIPTION

Let your Mac do the listening.

Change memory and precision to plan a local setup. Availability varies by runtime, so treat these as weight-based estimates—not measured throughput on every Apple chip.

Best ranked fitVoxtral Mini

3.8% provider AA-WER

16 GB
Planning precision

7 of 8 models fit with a 20% system reserve.

BENCHMARK YOUR AUDIO

WER is the start, not the answer.

01

Use your domain

Meetings, calls, lectures, and medical dictation have different failure modes. Keep a human transcript set.

02

Normalize once

Standardize punctuation, casing, numbers, and hesitations before comparing word error rate.

03

Slice the errors

Break results down by language, accent, noise, speaker overlap, proper nouns, and audio duration.

04

Time the whole path

Include upload, endpointing, diarization, and post-processing—not just model inference.