Benchmark · October 2026

Live captions, measured in real time.

By Saamer Mansoor, ConferenceCaptioning · Updated · Method · Data · Changelog

We streamed 2 minutes 53 seconds of a real NASA news conference through ConferenceCaptioning’s on-device engines at real-time speed, the way a microphone would, and compared the captions with a corrected, word-for-word transcript. Several speakers, technical vocabulary and unusual names. Nothing left the device.

Apple M416 GB unified memorymacOS 27 1× real-time playback173 s of continuous speechEnglish

Check it yourself. The audio is NASA’s Artemis II post-splashdown news conference (April 10, 2026). Open it at 6:38 (the link jumps there); the test runs from 6:38 to 9:31, from “They are Amit Kshatriya…” to “…I’ll hand it over to Dr. Glaze.” Press play and watch our captions beside it, or score the engine of your choice on the same clip.

Most accurate
97.2%
Mix · word accuracy (2.8% word error rate)
Fastest captions
0.05 s
Rapid · average delay
Words checked
562
in 173 s of real-time speech
Privacy
On-device
audio never leaves the device; both engines transcribe locally

Accuracy

Word accuracy is the share of the words in a corrected, word-for-word reference that were captioned correctly: 100% minus the word error rate, where word error rate is (wrong + missed + extra words) ÷ words in the reference. The reference has 562 words (not counting filler sounds).

RapidSystem speech service
4.8% word error rate
94.2% accuracy counting every word
MixLarge model + pause detection
2.8% word error rate
97.2% accuracy counting every word
0%25%50%75%100%

Bars run from 0% to 100%. Filler sounds (“uh”, “um”) are ignored in the main score; the line under each engine also gives the accuracy when they count.

EngineWords in referenceWrongMissedExtraTotal errorsWord accuracyWord error rateAccuracy, every word
Rapid56223402795.2%4.8%94.2%
Mix56211411697.2%2.8%97.2%

What the errors were

We sorted every counted error into what actually went wrong, so you can see whether a score reflects hard names and word endings or real mistakes.

Kind of errorWhat it looks likeRapidMix
Names and titlesSpeaker names such as Amit Kshatriya, Howard Hu and Shawn, written differently from the NASA spelling.64
Word endings“touched” vs “touch”, “return” vs “returned”, singular vs plural.33
Similar-sounding wordsA different word that sounds close, such as “saved” vs “shaped” or “rode” vs “wrote”.103
Words dropped, added or swappedMostly short words such as “the”, “that” or “it” that were missed, swapped or extra.86
Total2716

Examples: “Amit Kshatriya” heard as “Amet Shatriya” or “Ahmed Shatriya”, “touched” written as “touch”, “packed” heard as “peck”. Names like these are the hardest part of any live captioning job, which is why we kept them in the score.

What counted as correct

  • Case, punctuation and number formatting (“25,000 ft” or “twenty-five thousand feet”).
  • The same word written differently: “Dr.” or “doctor”, “programs” or “programmes”, “we were” or “we’re”, “it” or “it’s”.
  • Words the speaker said twice by accident (“then then take questions”, “as soon as as soon as”): the engine may caption them once or twice.
  • A few stretches no listener can make out on the recording: two place names (“Michoud”, “Bremen”), a third in “Teams at Stennis”, the surname “Radigan”, one garbled phrase (“…in an incredible way”) and a speaker’s self-correction (“That is… that is a thousand people”). We marked those spots as wildcards in the reference, about 25 words in all.

What still counted

  • Every other wrong, missed or extra word, including all of the people’s names, whether or not an engine could be expected to know them.
  • Filler sounds count only in the stricter score under each bar.
  • The reference started from the video’s automatic captions and was corrected by hand against the recording.

Caption delay

How long after a word was spoken it appeared on screen. The solid bar is the average; the white marker is the 95th percentile, the delay you would notice on a slow moment. The grey bars are published reference points, not products we tested.

RapidSystem speech service
95th percentile 0.07 s
MixLarge model + pause detection
95th percentile 1.68 s
Human captioner (CART)2014 study, mean latency
Automatic captions, 20142014 study, per-word latency
0 s2.5 s5 s7.5 s10 s

Rapid shows words almost immediately and refines them as it hears more. Mix waits for a natural pause before it commits a phrase, which is why it is later but more accurate. Reference points: a professional human captioner (CART) averaged 4.2 s, and automatic captions of the day averaged 7.9 s per word, in a 2014 classroom study by Kushalnagar, Lasecki and Bigham (ACM Transactions on Accessible Computing). Today’s cloud services are faster than that; we have not measured them.

Footprint on the device

Peak memory of the app process while captioning, as a share of this Mac’s 16 GB, and average processor use as a share of one core.

Peak memory (share of 16 GB)

Rapid0.5% of 16 GB
Mix4.9% of 16 GB
04 GB8 GB12 GB16 GB

Average CPU (% of one core)

RapidSystem speech service
MixLarge model + pause detection
0%25%50%75%100%

* Rapid runs in a system service outside the app, so its in-app memory and CPU understate its real cost.

Power while captioning

Average power drawn by the whole Mac while each engine captioned the clip, measured with Apple’s powermetrics tool. The Neural Engine is Apple’s low-power accelerator for machine learning.

EngineNeural Engine averageNeural Engine peakCPU averageGPU averageAll compute, averageEnergy over the clip
Rapid7 mW20 mW112 mW1 mW121 mW21 J
Mix63 mW628 mW164 mW1 mW228 mW40 J

Rapid averaged about 121 mW across CPU, GPU and Neural Engine; Mix about 228 mW, of which 63 mW was the Neural Engine. Figures are absolute averages for the whole machine from a single run, so they include whatever else the Mac was doing in the background, and Rapid’s includes the system speech service that runs outside the app.

All the numbers

Both engines ran on the same audio, in the same session, on the same machine. We ran the clip 2 times: word error rate was Rapid 4.8% and Mix 2.8%, average caption delay 0.05 s and 1.24–1.25 s. The table shows one of the runs.

EngineModelWord accuracyWord error rateWords wrong / missed / extraFirst captionDelay avgDelay p95Delay maxCaption complete after speech endsLoad timePeak memoryCPU avg
RapidSystem speech service (iOS / macOS 26+)95.2%4.8%23 / 4 / 01.14 s0.05 s0.07 s1.2 s0.10 s0.9 s74 MB *7% *
MixLarge on-device model + voice-activity detection97.2%2.8%11 / 4 / 15.33 s1.24 s1.68 s1.7 s0.08 s1.0 s800 MB10%

Methodology

The audio

  • Public NASA footage: the Artemis II post-splashdown news conference, from 6:38 to 9:31 (173 s). Several speakers, technical terms and unusual names; a clean broadcast feed.
  • The clip is decoded to 16 kHz mono and streamed in 100 ms chunks at exactly real-time speed (1×), followed by three seconds of silence so the last words can finish.
  • This is a replay of a clean file through the engine code, not a microphone in a room, so absolute delay and resource figures in the app can differ somewhat.

The engines

  • The same code paths the app runs when listening live: Rapid uses the operating system’s speech service; Mix uses a large on-device model that waits for natural pauses to commit a phrase.
  • Models are downloaded and warmed up before the run, as in the app’s setup screen; load time is reported separately.
  • Caption delay: for every caption update, how far the newest word in it was behind the audio, that is, how long after it was spoken it appeared. The 95th percentile is the slower moments.

The scoring

  • The reference transcript started from the video’s automatic captions and was corrected by hand against the recording.
  • Case, punctuation and numbers are normalised; words are aligned with a word-level edit distance; score = (wrong + missed + extra) ÷ reference words. The exact rules are in the scoring script below.
  • Accepted spellings and about 25 unintelligible words are marked in the reference (see “What counted as correct”).

The machine

  • Apple M4 Mac mini, 16 GB unified memory, macOS 27, Low Power Mode off, mains power.
  • Memory and CPU are of the app process; power is whole-machine, from Apple’s powermetrics tool, in a separate run of the same clip.
  • The clip was run 2 times; accuracy was identical each time.

Reproduce it

  1. Download the audio of the NASA video and cut 6:38–9:31, for example yt-dlp -x --audio-format wav <video-url> then ffmpeg -ss 398 -t 173 -i video.wav clip.wav.
  2. Caption the clip with the engine you want to test, at real-time speed, and save its final text to hypothesis.txt.
  3. Download the reference transcript and the scoring script, then run python3 score_wer.py nasa-reference-transcript.txt hypothesis.txt.
  4. The script prints word accuracy and word error rate with and without filler sounds, and the wrong / missed / extra counts. Our engines score the figures above; yours will score what it scores.

Data and files

License

  • Data and reference transcript: Creative Commons Attribution 4.0 (CC BY 4.0). Share and adapt them, including commercially, as long as you credit ConferenceCaptioning and link to https://conferencecaptioning.com/benchmark/.
  • Scoring script: MIT License, © 2026 ConferenceCaptioning (the licence text is at the top of score_wer.py).
  • The audio: NASA’s footage, which NASA says is generally not subject to copyright (see NASA’s media usage guidelines). We do not host it, we link to the original video. Our reference transcript is a corrected, annotated transcription of the words spoken in it.

Cite this page

Saamer Mansoor, ConferenceCaptioning. “Live captioning accuracy and delay benchmark: on-device engines on a NASA news conference.” October 2026. https://conferencecaptioning.com/benchmark/

How to read this

What it does not tell you

  • One clip and one machine. Repeating the same engine gave the same word error rate, but other speakers and rooms will differ, so treat small gaps as a tie.
  • A clean broadcast feed. Room noise, distant microphones and heavy accents will move every number.
  • English only, and only our own engines: we have not compared other captioning products on this clip.

Choosing an engine

  • Rapid: fastest captions and lightest on memory and power; needs iOS or macOS 26.
  • Mix: top accuracy, works on older systems (iOS 17 / macOS 14), captions about a second later.

Changelog

  • October 5, 2026: added the methodology, reproduction steps, data downloads and this changelog; numbers are now in the page itself.
  • October 5, 2026: replaced the September test (a different, noisier clip with four engines) with this NASA news conference clip and the Rapid and Mix engines; reference accepts spelling variants; added power measurements.
  • September 30, 2026: first published benchmark.

Want to hear it on your own event?

Everything above runs on-device, so captions and translations keep flowing even if the venue internet goes down. Tell us about your room, your audio feed and your languages and we’ll help you plan the setup.

See full AV specs → Start free →