Captions are only useful to a multilingual audience if the translations keep up with the speaker. We fed an English transcript at conversational speed (3 words per second) into ConferenceCaptioning’s fastest translation mode, Blazing, with one to five target languages running together, and timed how long each word took to appear in every language on screen. Translation runs on the device; nothing is sent to a translation service.
With five languages translating together, the average word appeared on screen 0.27 s after it was available, 0.54 s for 95 out of 100 words, and the slowest word in any run took 0.61 s. Adding languages from one to five changed the average delay only slightly (see the table), because each translation takes a few tens of milliseconds on the device. That is well under a second, so for this part of the pipeline “five languages in under a second” holds. With the refresh interval turned down to 50 ms, five languages averaged 0.08 s (see faster refresh), at the cost of more processor time. It is not the delay an audience sees end to end: it excludes speech recognition, which comes first (see what is not measured).
Time from the moment a word became available in the transcript to the moment its translation was on screen, averaged over every word, with the line marking the 95th percentile. Lower is better. Blazing is the sky-blue bars; grey bars are our other on-device translation mode (Native Translation) for scale.
Bars run from 0 to 4 seconds. The grey Native Translation bars come from an earlier test on the same Mac with a different harness (translation_benchmark_report.md in our repository), so compare the shape, not the last digit.
Refreshing every 500 ms (the default). Each row pools 6 runs of 40 measured seconds (the first 10 s of each 50 s run are warm-up and discarded), across continuous speech and speech with a 2 s pause after every sentence. “Call time” is how long one translation request took.
| Languages | Mean delay | Median | 95th pct | Slowest | Call time (mean) | Call time (95th) | Avg CPU | Peak memory | Runs |
|---|---|---|---|---|---|---|---|---|---|
| 1 language | 0.25 s | 0.20 s | 0.40 s | 0.54 s | 22 ms | 34 ms | 15% | 1.3 GB | 6 |
| 2 languages | 0.25 s | 0.21 s | 0.47 s | 0.56 s | 28 ms | 50 ms | 14% | 1.2 GB | 6 |
| 3 languages | 0.25 s | 0.22 s | 0.48 s | 0.58 s | 38 ms | 65 ms | 14% | 1.3 GB | 6 |
| 4 languages | 0.25 s | 0.23 s | 0.54 s | 0.59 s | 45 ms | 78 ms | 14% | 1.3 GB | 6 |
| 5 languages | 0.27 s | 0.24 s | 0.54 s | 0.61 s | 51 ms | 90 ms | 14% | 1.4 GB | 6 |
CPU is the translation process as a share of one core. Memory is the whole translation process including its language models. Continuous and paused speech gave the same result at 5 languages (mean 0.27 s and 0.25 s). Errors across all timed runs: 0.
The refresh rate (how often the translation page checks for new words) is a setting. Faster refresh shows words sooner but uses more processor time. Continuous speech, one and five languages:
| Refresh | 1 language: mean | 95th pct | Slowest | 5 languages: mean | 95th pct | Slowest | CPU (1 language) |
|---|---|---|---|---|---|---|---|
| 50 ms | 0.06 s | 0.07 s | 0.09 s | 0.08 s | 0.12 s | 0.15 s | 28% |
| 75 ms | 0.07 s | 0.10 s | 0.11 s | 0.10 s | 0.15 s | 0.17 s | 25% |
| 100 ms | 0.08 s | 0.12 s | 0.12 s | 0.11 s | 0.16 s | 0.19 s | 23% |
| 125 ms | 0.09 s | 0.14 s | 0.16 s | 0.12 s | 0.19 s | 0.22 s | 20% |
| 150 ms | 0.10 s | 0.17 s | 0.18 s | 0.13 s | 0.21 s | 0.25 s | 19% |
| 200 ms | 0.14 s | 0.22 s | 0.24 s | 0.17 s | 0.28 s | 0.36 s | 11% |
| 500 ms | 0.25 s | 0.40 s | 0.54 s | 0.27 s | 0.56 s | 0.60 s | 15% |
67 languages can be translated on the device, in three speed groups. Each language is listed once, under the fastest mode that supports it. Only the five Blazing languages named in the table were timed directly in this test; the other times are estimates from the matching measurements and could differ for individual languages.
| Mode | Languages | Expected translation delay | Tested with | Basis |
|---|---|---|---|---|
| Blazing | 44 | about 0.3–0.5 s | up to 5 languages | Measured for Spanish, French, German, Japanese and Portuguese (Brazil); estimated for the rest of the group |
| Native Translation | 1 | about 1.0–1.5 s for one language, 3–4 s for four | 1–4 languages | Measured end to end on a similar Mac; the only on-device option for these languages |
| Language Pack | 22 | about 1.0–1.5 s for one language, 3.4–3.8 s for four | 1–4 languages | Estimated from our test of the same pipeline on Spanish, French, German and Japanese; one-time 620 MB download |
Regional variants (for example Spanish for Mexico or Chile, or English from Australia or India) are counted under their language. Languages are added through a one-time on-device download. The delay is the translation step only, as described above.
The data is licensed CC BY 4.0: share and adapt it, including commercially, with credit to ConferenceCaptioning and a link to https://conferencecaptioning.com/multilingual-translation-benchmark/.
Saamer Mansoor, ConferenceCaptioning. “Live translation speed benchmark: five languages at once.” October 2026. https://conferencecaptioning.com/multilingual-translation-benchmark/
Tell us your languages, your venue and how attendees will follow along and we’ll help you plan the setup.