Independent benchmark · Measured 2026-09-17
Speech recognition accuracy: 92 languages, measured against leading ASR models
Pikka Speech transcribes 92 languages with a median character error rate of 2.9%. OpenAI gpt-transcribe scores 5.1%, ElevenLabs Scribe v2 4.3%, Google Chirp 3 12.2% and Gemini 3.5 Transcribe 6.4% on the same audio — lower is better.
Full table
Character error rate by language
Character error rate (CER) per language, lower is better. Engines marked n/a returned no usable transcript for that language. Bold marks the best result in each row.
| Language | Pikka Speech | OpenAI gpt-transcribe | ElevenLabs Scribe v2 | Google Chirp 3 | Gemini 3.5 Transcribe |
|---|---|---|---|---|---|
| Afrikaans | 2.8% | 3.0% | 1.5% | 10.7% | 5.2% |
| Amharic | 3.9% | 14.9% | 10.7% | 12.3% | 5.9% |
| Arabic | 4.1% | 5.4% | 5.2% | 11.5% | 4.9% |
| Armenian | 12.3% | 14.2% | 12.5% | 21.1% | 20.4% |
| Assamese | 8.5% | 12.3% | 8.2% | 42.7% | 14.6% |
| Asturian | 7.6% | 9.4% | 4.2% | 10.8% | 16.0% |
| Azerbaijani | 0.9% | 2.5% | 0.9% | 28.0% | 2.5% |
| Belarusian | 2.2% | 2.8% | 2.2% | n/a | 3.2% |
| Bengali | 0.9% | 3.2% | 3.3% | 10.2% | 1.9% |
| Bulgarian | 0.4% | 0.4% | 1.0% | 8.2% | 1.6% |
| Burmese | 9.9% | 41.6% | 8.7% | 7.9% | 13.5% |
| Cantonese | 39.6% | 4.8% | 22.7% | 12.7% | 14.8% |
| Catalan | 2.2% | 3.7% | 2.7% | 3.6% | 2.4% |
| Cebuano | 9.8% | 8.3% | 4.5% | n/a | 9.1% |
| Chinese | 4.9% | 6.8% | 7.6% | 12.8% | 7.9% |
| Croatian | 0.4% | 27.8% | 2.4% | 2.7% | 2.4% |
| Czech | 6.0% | 4.2% | 3.1% | 9.1% | 2.2% |
| Danish | 1.8% | 3.3% | 1.2% | 2.2% | 6.6% |
| Dutch | 1.8% | 0.4% | 0.6% | 14.2% | 2.4% |
| English | 3.1% | 1.5% | 2.2% | 9.7% | 1.2% |
| Estonian | 0.5% | 1.9% | 1.7% | 9.2% | 7.3% |
| Finnish | 0.3% | 0.6% | 0.2% | 8.8% | 1.0% |
| French | 3.9% | 2.5% | 0.7% | 16.8% | 3.4% |
| Galician | 1.1% | 1.5% | 1.6% | 6.1% | 3.9% |
| Georgian | 4.1% | 6.8% | 2.6% | 4.5% | 9.1% |
| German | 5.1% | 0.4% | 0.4% | 5.7% | 1.1% |
| Greek | 2.4% | 2.1% | 1.7% | 9.6% | 7.1% |
| Gujarati | 3.4% | 8.3% | 6.9% | 15.6% | 6.5% |
| Hausa | 5.9% | 17.4% | 6.6% | 7.4% | 9.4% |
| Hebrew | 6.7% | 6.1% | n/a | 27.0% | 14.2% |
| Hindi | 1.7% | 3.6% | 4.4% | 17.3% | 1.7% |
| Hungarian | 4.0% | 2.3% | 3.2% | 12.1% | 11.3% |
| Icelandic | 10.9% | 12.2% | 9.1% | 12.1% | 19.6% |
| Indonesian | 1.1% | 2.6% | 1.1% | 14.5% | 0.3% |
| Irish | 22.8% | 32.4% | 18.6% | n/a | 60.9% |
| Italian | 0.3% | 0.6% | 0.7% | 0.8% | 0.5% |
| Japanese | 1.8% | 3.6% | 1.1% | 0.9% | 2.7% |
| Javanese | 3.1% | 11.4% | 2.4% | 6.0% | 10.0% |
| Kannada | 1.6% | 4.6% | 7.3% | 11.8% | 11.2% |
| Kazakh | 0.9% | 4.5% | 2.9% | 15.1% | 2.4% |
| Khmer | 9.6% | 12.3% | 7.1% | 14.3% | 11.4% |
| Korean | 2.2% | 2.9% | 1.3% | 8.3% | 3.8% |
| Kurdish | n/a | 10.9% | 7.5% | n/a | 16.1% |
| Kyrgyz | 5.6% | 5.7% | 6.5% | 31.0% | 7.9% |
| Lao | 17.6% | 62.0% | 7.4% | 10.1% | 34.1% |
| Latvian | 2.4% | 4.0% | 2.7% | 13.9% | 6.3% |
| Lithuanian | 2.1% | 6.5% | 5.3% | 17.9% | 9.7% |
| Luxembourgish | 35.5% | 16.9% | 6.1% | 36.1% | 26.3% |
| Macedonian | 1.0% | 1.3% | 0.3% | 9.3% | 2.3% |
| Malay | 2.2% | 1.3% | 2.2% | 22.2% | 2.3% |
| Malayalam | 1.2% | 4.3% | 7.4% | 11.8% | 3.0% |
| Maltese | 12.0% | 6.7% | 2.3% | 12.9% | 16.0% |
| Maori | n/a | 16.3% | 6.5% | 21.2% | 18.0% |
| Marathi | 3.7% | 7.0% | 6.5% | 26.3% | 4.8% |
| Mongolian | 5.2% | 12.5% | 7.3% | 5.9% | 7.0% |
| Nepali | 0.0% | 7.0% | 7.0% | 45.3% | 4.7% |
| Northern Sotho | 13.2% | 34.7% | 6.4% | 29.2% | 23.6% |
| Norwegian | 0.7% | 1.0% | 3.4% | 4.2% | 3.8% |
| Odia | 7.8% | 10.6% | 6.8% | 17.0% | 8.2% |
| Oromo | n/a | 63.1% | 31.4% | n/a | 39.0% |
| Pashto | n/a | 24.1% | 22.9% | n/a | 46.1% |
| Persian | 1.9% | 2.3% | 4.4% | 3.2% | 2.8% |
| Polish | 1.6% | 1.0% | 1.9% | 12.1% | 3.3% |
| Portuguese | 1.0% | 2.8% | 3.4% | 5.4% | 3.7% |
| Punjabi | 4.6% | 23.9% | n/a | 22.1% | 4.5% |
| Romanian | 2.9% | 0.8% | 1.1% | 60.1% | 2.5% |
| Russian | 0.2% | 0.2% | 1.2% | 3.2% | 0.4% |
| Serbian | 15.1% | 88.1% | 30.0% | 81.3% | 40.9% |
| Sindhi | 4.2% | 98.5% | 5.3% | n/a | 8.9% |
| Slovak | 1.2% | 1.8% | 1.2% | 5.0% | 3.5% |
| Slovenian | 2.0% | 3.5% | 1.7% | 23.5% | 8.0% |
| Somali | n/a | 24.9% | 12.4% | n/a | 17.6% |
| Spanish | 1.2% | 0.4% | 0.5% | 2.1% | 0.8% |
| Swahili | 1.9% | 5.3% | 1.1% | 7.8% | 3.0% |
| Swedish | 4.5% | 3.8% | 2.4% | 4.0% | 6.0% |
| Tagalog | 2.3% | 4.9% | 3.1% | 8.0% | 5.3% |
| Tajik | 2.9% | 58.3% | 2.4% | n/a | 10.0% |
| Tamil | 16.8% | 18.6% | 18.3% | 28.8% | 16.5% |
| Telugu | 1.4% | 3.4% | 6.3% | 23.0% | 1.4% |
| Thai | 4.0% | 4.9% | 4.6% | 29.5% | 5.0% |
| Turkish | 0.7% | 1.0% | n/a | 44.9% | 0.9% |
| Ukrainian | 1.2% | 3.3% | 2.6% | 14.7% | 1.1% |
| Urdu | 4.7% | 90.5% | 14.7% | 29.2% | 11.4% |
| Uzbek | 7.7% | 9.2% | 7.0% | 14.1% | 12.7% |
| Vietnamese | 1.1% | 3.6% | 1.6% | 12.2% | 4.2% |
| Welsh | 5.7% | 13.2% | 4.8% | 25.8% | 21.2% |
| Wolof | n/a | 19.6% | 13.2% | 25.7% | 29.5% |
| Xhosa | 49.0% | 33.7% | 6.5% | 44.8% | 12.1% |
| Yoruba | n/a | 34.6% | 18.7% | 27.1% | 25.6% |
| Zulu | n/a | 22.9% | 7.2% | 10.3% | 14.2% |
| he | 6.7% | 6.1% | 8.8% | n/a | 14.2% |
| pt-BR | 1.0% | 2.8% | n/a | n/a | 3.7% |
Method
How we measure
A short, reproducible method — no vendor-run benchmarks, no cherry-picked clips.
Every engine transcribes the same FLEURS clips — the open multilingual speech corpus from Google used across speech-recognition research — through its own official API.
- 92 languages, 10 utterances each, native speakers reading real sentences.
- Scored as character error rate after normalising case and punctuation, so differences reflect recognition, not formatting.
- Word error rate (WER) is reported in the downloadable data for languages that use word spacing; character-based scripts (Chinese, Japanese, Thai, Khmer, Burmese, Lao) are scored on characters only.
- Every row is published, including the languages where another model leads.
Read the corpus paper for the recording methodology: FLEURS: Few-shot Learning Evaluation of Universal Representations of Speech.
FAQ
Questions about this benchmark
What is character error rate (CER)?
CER is the share of characters that differ from the reference transcript, as a percentage. A CER of 2% means the engine transcribed 98 of every 100 characters correctly. Lower is better.
What audio was used for this benchmark?
FLEURS, the open multilingual speech corpus from Google, used across speech-recognition research. We ran 10 utterances per language across 92 languages — native speakers reading real sentences, identical clips for every engine.
Did every model get the same audio and settings?
Yes for the audio: every engine transcribed the exact same clips through its own official API. Each engine was measured in its default configuration; results were scored after normalising case and punctuation so differences reflect recognition, not formatting.
Why do some cells show n/a?
n/a means that engine returned no usable transcript for that language on the day of the run — usually because the model does not support it. We publish those gaps instead of hiding them.
How often is this table updated?
It is re-run whenever a model we use or benchmark changes. This version was measured on 17 September 2026.
Keep exploring
Which languages Pikka Speech covers: language coverage. How interpretation and captions work together: platform capabilities and hybrid simultaneous interpretation. Or start a free test room.
Hear the difference on your own event
Create a free test room, speak a few sentences in your language, and compare the live captions against any other engine.