We streamed 2 minutes 53 seconds of a real NASA news conference through ConferenceCaptioning’s on-device engines at real-time speed, the way a microphone would, and compared the captions with a corrected, word-for-word transcript. Several speakers, technical vocabulary and unusual names. Nothing left the device.
Check it yourself. The audio is NASA’s Artemis II post-splashdown news conference (April 10, 2026). Open it at 6:38 (the link jumps there); the test runs from 6:38 to 9:31, from “They are Amit Kshatriya…” to “…I’ll hand it over to Dr. Glaze.” Press play and watch our captions beside it, or score the engine of your choice on the same clip.
Word accuracy is the share of the words in a corrected, word-for-word reference that were captioned correctly: 100% minus the word error rate, where word error rate is (wrong + missed + extra words) ÷ words in the reference. The reference has 562 words (not counting filler sounds).
Bars run from 0% to 100%. Filler sounds (“uh”, “um”) are ignored in the main score; the line under each engine also gives the accuracy when they count.
| Engine | Words in reference | Wrong | Missed | Extra | Total errors | Word accuracy | Word error rate | Accuracy, every word |
|---|---|---|---|---|---|---|---|---|
| Rapid | 562 | 23 | 4 | 0 | 27 | 95.2% | 4.8% | 94.2% |
| Mix | 562 | 11 | 4 | 1 | 16 | 97.2% | 2.8% | 97.2% |
We sorted every counted error into what actually went wrong, so you can see whether a score reflects hard names and word endings or real mistakes.
| Kind of error | What it looks like | Rapid | Mix |
|---|---|---|---|
| Names and titles | Speaker names such as Amit Kshatriya, Howard Hu and Shawn, written differently from the NASA spelling. | 6 | 4 |
| Word endings | “touched” vs “touch”, “return” vs “returned”, singular vs plural. | 3 | 3 |
| Similar-sounding words | A different word that sounds close, such as “saved” vs “shaped” or “rode” vs “wrote”. | 10 | 3 |
| Words dropped, added or swapped | Mostly short words such as “the”, “that” or “it” that were missed, swapped or extra. | 8 | 6 |
| Total | 27 | 16 |
Examples: “Amit Kshatriya” heard as “Amet Shatriya” or “Ahmed Shatriya”, “touched” written as “touch”, “packed” heard as “peck”. Names like these are the hardest part of any live captioning job, which is why we kept them in the score.
How long after a word was spoken it appeared on screen. The solid bar is the average; the white marker is the 95th percentile, the delay you would notice on a slow moment. The grey bars are published reference points, not products we tested.
Rapid shows words almost immediately and refines them as it hears more. Mix waits for a natural pause before it commits a phrase, which is why it is later but more accurate. Reference points: a professional human captioner (CART) averaged 4.2 s, and automatic captions of the day averaged 7.9 s per word, in a 2014 classroom study by Kushalnagar, Lasecki and Bigham (ACM Transactions on Accessible Computing). Today’s cloud services are faster than that; we have not measured them.
Peak memory of the app process while captioning, as a share of this Mac’s 16 GB, and average processor use as a share of one core.
* Rapid runs in a system service outside the app, so its in-app memory and CPU understate its real cost.
Average power drawn by the whole Mac while each engine captioned the clip, measured with Apple’s powermetrics tool. The Neural Engine is Apple’s low-power accelerator for machine learning.
| Engine | Neural Engine average | Neural Engine peak | CPU average | GPU average | All compute, average | Energy over the clip |
|---|---|---|---|---|---|---|
| Rapid | 7 mW | 20 mW | 112 mW | 1 mW | 121 mW | 21 J |
| Mix | 63 mW | 628 mW | 164 mW | 1 mW | 228 mW | 40 J |
Rapid averaged about 121 mW across CPU, GPU and Neural Engine; Mix about 228 mW, of which 63 mW was the Neural Engine. Figures are absolute averages for the whole machine from a single run, so they include whatever else the Mac was doing in the background, and Rapid’s includes the system speech service that runs outside the app.
Both engines ran on the same audio, in the same session, on the same machine. We ran the clip 2 times: word error rate was Rapid 4.8% and Mix 2.8%, average caption delay 0.05 s and 1.24–1.25 s. The table shows one of the runs.
| Engine | Model | Word accuracy | Word error rate | Words wrong / missed / extra | First caption | Delay avg | Delay p95 | Delay max | Caption complete after speech ends | Load time | Peak memory | CPU avg |
|---|---|---|---|---|---|---|---|---|---|---|---|---|
| Rapid | System speech service (iOS / macOS 26+) | 95.2% | 4.8% | 23 / 4 / 0 | 1.14 s | 0.05 s | 0.07 s | 1.2 s | 0.10 s | 0.9 s | 74 MB * | 7% * |
| Mix | Large on-device model + voice-activity detection | 97.2% | 2.8% | 11 / 4 / 1 | 5.33 s | 1.24 s | 1.68 s | 1.7 s | 0.08 s | 1.0 s | 800 MB | 10% |
yt-dlp -x --audio-format wav <video-url> then ffmpeg -ss 398 -t 173 -i video.wav clip.wav.hypothesis.txt.python3 score_wer.py nasa-reference-transcript.txt hypothesis.txt.{a|b} lists accepted spellings, {x|~} may be skipped, and * is a wildcard for one unintelligible word.score_wer.py).Saamer Mansoor, ConferenceCaptioning. “Live captioning accuracy and delay benchmark: on-device engines on a NASA news conference.” October 2026. https://conferencecaptioning.com/benchmark/
Everything above runs on-device, so captions and translations keep flowing even if the venue internet goes down. Tell us about your room, your audio feed and your languages and we’ll help you plan the setup.