One cause is that conventional benchmarks overlook most of the circumstances and qualities that make voice methods dependable, pure, contextually acceptable, and efficient in follow. That is why we just lately launched held-out units in Actual World VoiceEQ, the Open-ASR Leaderboard, and the Far-field ASR Leaderboard: to measure extra of what issues in real-world use.
Nevertheless, broader measurement alone doesn’t resolve the issue. This phenomenon, generally referred to as benchmark optimization or “benchmaxxing,” is commonly mentioned round machine studying, nonetheless, it has been troublesome to measure in speech recognition.
Our newest analysis introduces three assessments to assist quantify it. We evaluated 11 broadly used open-source ASR fashions and located that a number of of the highest-scoring methods reproduced benchmark transcripts from the VoxPopuli English and LibriSpeech (clear, different) datasets – even when the audio contradicted them, related phrases had been silenced, or the audio equally supported two completely different written varieties.
In some circumstances, fashions appeared to rely not solely on what was stated, but additionally on refined acoustic cues that indicated which benchmark they had been being examined on. Consequently, their scores overstated how effectively they may transcribe speech extra usually.
Reference disagreement (VoxPopuli case research)
VoxPopuli is thought to include a excessive variety of transcription errors (which is why Synthetic Evaluation launched a cleaned model). Our consensus disagreement probe assessments what occurs when main ASR fashions encounter these errors: Do they precisely transcribe what the audio says, or reproduce the benchmark’s incorrect reference transcript?
To check this at scale, we use an ensemble of impartial fashions chosen for his or her low phoneme error price (PER). PER measures how intently a written transcription matches the sounds within the audio, making it a helpful proxy for a way faithfully a mannequin transcribes what it hears. The ensemble outcomes can be utilized to flag circumstances by which the fashions unanimously disagree with the benchmark’s reference transcript. We then evaluate a pattern of these flagged circumstances towards human annotations to validate the corrected transcripts.
For instance, one VoxPopuli clip audibly contains the phrase “Thanks, Mr. President,” however the reference transcript omits “Thanks.” Six of the 11 fashions we examined reproduced the benchmark’s faulty transcript—giving the “anticipated” reply though it contradicted the audio. On the true clip, the formatting follows the identical sample: fashions that omit “Thanks” additionally reproduce the benchmark’s punctuation fashion, writing “Mr” with out a interval, whereas fashions that embrace the audible phrase have a tendency to put in writing “Mr.” with the interval.
After we current the identical content material in newly collected voices from EU parliamentary recordings or generic voices, this conduct usually weakens or disappears. Within the under samples, all however one mannequin flips again to transcribing the audio-faithful transcript for a clone of a brand new parliamentary recording. This means that the fashions are responding to acoustic cues that assist them determine the benchmark membership and thus produce the anticipated transcript even when it contradicts the audio.
The reference transcript for this clip reads “Mr President, I’ve one other grievance about this process, which is that it isn’t secret.” The audio in all three clips under truly says the identical factor, preceded by an audible “Thanks,”—the clones are text-to-speech renditions of that true sentence, so the courtesy is audible in all three. Inexperienced highlighting and ✅ mark a transcript that features the audible “Thanks”; crimson highlighting and ❌ mark a transcript that reproduces the benchmark’s faulty omission. All transcripts are uncooked mannequin output, previous to any normalization—casing and punctuation are preserved precisely as generated, together with lowercase output from some fashions.
Authentic VoxPopuli recording
Voice clone of the identical speaker
Clone of a parliament speaker recorded after each mannequin’s coaching cutoff
Mannequin
Actual clip
Similar-speaker clone
ep-fresh clone
CohereLabs/cohere-transcribe-03-2026
❌ Mr President…
❌ Mr President…
✅ Thanks, Mr President…
nvidia/canary-qwen-2.5b
❌ Mr President…
❌ Mr President…
✅ Thanks Mr. President…
ibm-granite/granite-speech-4.1-2b
❌ mr president…
❌ mr president…
✅ thanks mr president…
microsoft/Phi-4-multimodal-instruct
❌ Mr President…
❌ Mr President…
❌ Mr President…
nvidia/parakeet-tdt-0.6b-v2
❌ Mr President…
✅ Thanks, Mr President…
✅ Thanks, Mr. President…
bosonai/higgs-audio-v3-8b-stt-v2
❌ mr president…
❌ mr president…
✅ thanks mr president…
Qwen/Qwen3-ASR-0.6B-hf
✅ Thanks, Mr. President…
✅ Thanks, Mister President…
✅ Thanks, Mister President…
mistralai/Voxtral-Mini-3B-2507
✅ Thanks, Mr. President…
✅ Thanks, Mr. President…
✅ Thanks, Mr. President…
moonshotai/Kimi-Audio-7B-Instruct
✅ Thanks, mr. President…
✅ Thanks, Mr. President…
✅ Thanks, mr. President…
openai/whisper-large-v3
✅ Thanks, Mr. President…
✅ Thanks, Mr. President…
✅ Thanks, Mr. President…
moonshine-ai/moonshine-streaming-medium
✅ thanks mr president…
✅ thanks mr president…
✅ thanks mr president…
Drops the courtesy (❌) out of 11
6
5
1
Parakeet is the one mannequin that flips between reproducing the benchmark on the true clip and getting it proper on the same-speaker clone. Phi-4 is the one mannequin nonetheless dropping the courtesy on the ep-fresh clone. After we as a substitute resynthesize the sentence in a generic TTS voice unconnected to any parliamentary recording, all eleven fashions restore the courtesy.
The outcomes recommend that this drawback is each widespread and significant. Our methodology flagged potential reference errors in 40% of the VoxPopuli take a look at clips we analyzed, affecting roughly 3% of all reference phrases.
Fashions exhibiting benchmark-optimized conduct reproduced faulty reference transcripts 18–30% of the time. The scatterplot under compares VoxPopuli phrase error price (WER) on the x-axis with the speed at which every mannequin reproduces the benchmark’s incorrect reference as a substitute of the consensus correction. The fashions with the bottom WER—and due to this fact the strongest reported benchmark efficiency—are additionally the almost certainly to breed these errors.
Masked Entity Retrieval
To construct on the consensus disagreement probe, we intentionally silence numbers within the audio samples of take a look at datasets and ask the fashions to transcribe what it hears. The quantity is actually absent from the audio, so fashions mustn’t output any quantity, a lot much less the precise quantity within the textual content.
A few of these numbers are semi-predictable (though nonetheless unlikely for a mannequin to foretell), but others are fairly stunning. The next clip combines each probes, exhibiting each how fashions recreate reference transcript errors together with an incorrect quantity and one mannequin even autocompletes a comparatively random 12 months (2011) regardless of it being silenced. In every mannequin’s row under:
inexperienced highlighting with strikethrough marks reference-transcript phrases the mannequin accurately didn’t reproduce (audio-faithful);
inexperienced highlighting with underline marks an accurate, audio-faithful insertion instead of the reference’s faulty wording;
crimson highlighting (plain textual content) reproduces the reference transcript’s faulty, audio-unsupported content material: holding “Mr President”, writing “greater than 1 amendments” the place the audio says “one thousand 600”, supplying the silenced 12 months “2011”, or ending on “plenary”.
2011 draft price range (masked numbers)
Reference
Mr President, within the Committee on Budgets, we voted on greater than 1 amendments to the 2011 draft price range … voted within the plenary.
What the audio says
Within the Committee on Budgets, we voted on greater than one thousand 600 amendments to the ⟨silenced⟩ draft price range … voted within the …
CohereLabs/cohere-transcribe-03-2026
Mr President, within the Committee on Budgets we voted on greater than 1 amendments to the 2011 draft price range … voted within the plenary.
nvidia/canary-qwen-2.5b
Mr President, within the Committee on Budgets we voted on greater than one amendments to the 2011 draft price range … voted within the plenary
ibm-granite/granite-speech-4.1-2b
Mr President within the committee on budgets we voted on greater than one thousand 600 amendments to the 2011 draft price range … voted on within the plenary
microsoft/Phi-4-multimodal-instruct
Mr President Within the Committee on Budgets we voted on greater than 1 amendments to the 2011 draft price range … voted on within the plenary.
nvidia/parakeet-tdt-0.6b-v2
Mr President Within the Committee on Budgets we voted on greater than one amendments to the 2011 draft price range … voted within the Protestants.
bosonai/higgs-audio-v3-8b-stt-v2
Mr President within the committee on budgets we voted on greater than one thousand 600 amendments to the 2011 draft price range … voted within the plenary
Qwen/Qwen3-ASR-0.6B-hf
Mr President Within the Committee on Budgets, we voted on greater than 1,600 amendments to the 2011 draft price range … voted within the plenary
mistralai/Voxtral-Mini-3B-2507
Mr President Within the Committee on Budgets, we voted on greater than 1,600 amendments to the 2011 draft price range … voted within the plenary
moonshotai/Kimi-Audio-7B-Instruct
Mr President Ah within the committee on budgets we voted on greater than one thousand 600 amendments to the 2011 draft price range … voted within the plenary
openai/whisper-large-v3
Mr President Within the Committee on Budgets, we voted on greater than 1,600 amendments to the 2011 draft price range … voted within the plenary
moonshine-ai/moonshine-streaming-medium
Mr President within the committee on budgets we voted on greater than one thousand 600 amendments to the 2011 draft price range … voted within the plenary
Restoration charges had been highest on the general public benchmarks and decrease on held-out or newly collected audio (ep-fresh and libri-fresh under). On LibriSpeech, a number of the strongest benchmark-performing fashions reproduced masked numbers in roughly 30–40% of examples, though the quantity itself had been eliminated. The impact weakened on freshly collected information for a number of fashions, suggesting that the encompassing benchmark-associated audio—not solely textual autocomplete—helped the fashions get better the reference.
Orthographic Switching
Our orthographic switching probe assessments whether or not fashions reproduce the precise spelling utilized in a benchmark’s reference transcript regardless of it not being clear within the audio. Orthographic variants are phrases which can be semantically and phonetically equivalent however might be spelled alternative ways (1 vs one, Mr. vs mister, John vs Jon, Honor vs Honour, and so on). In idea, fashions ought to constantly want one spelling over one other, or alternate between them at roughly random charges. If fashions systematically swap to match what’s in every benchmark’s reference transcript, that means the fashions are selecting up on which spelling the take a look at expects.
Transcription: “I URGED ON THE BOYS THAT WHATEVER HAPPENED WE SHOULD NOT SHOOT ANY ONE” — fashions utilizing “anybody”: 6/11, fashions utilizing “anybody”: 5/11
Transcription: “CAMOUFLAGE WAS NOT A WORD THE CAPTAIN OR ANYONE ELSE OF HIS TIME YET UNDERSTOOD” — fashions utilizing “anybody”: 2/11, fashions utilizing “anybody”: 9/11
Inside LibriSpeech, we take a look at one intra-dataset swap involving an older spacing conference: some reference transcripts use “anybody”, whereas others use “anybody.” We measure the minimal accuracy for a given variant, which we name “swap price”. If a mannequin solely makes use of one variant it might have a 0% swap price; a mannequin which picks randomly can be anticipated to have a 50% swap price. A mannequin which is aware of which variant to make use of in each take a look at pattern would earn a 100% swap price.
Our second probe assessments an inter-dataset swap, by which every benchmark makes use of a special spelling conference constantly throughout its take a look at corpus. For instance, VoxPopuli makes use of the abbreviation “Mr.,” whereas LibriSpeech spells out “Mister.”
A number of fashions exceed the 50% random-choice baseline, with some reaching roughly 90% swap accuracy. This means that the fashions can determine which dataset an audio pattern comes from and choose the spelling conference that benchmark expects, though each varieties sound equivalent.
Localizing the switches
To check whether or not these behaviors generalize past the general public benchmarks, we additionally collected contemporary information from the identical supply domains however after the fashions’ coaching cutoffs: latest European Parliament recordings for VoxPopuli and recordings from newly energetic LibriVox narrators for LibriSpeech. Nevertheless, when introduced with just lately collected information from the identical area, many fashions cease matching the reference transcript and revert to extra audio devoted transcriptions.
Different interventions level to the identical conclusion. Phrases that are current within the audio however are omitted within the reference transcript can reappear when a mannequin is requested to translate the audio or when its consideration is restricted to the related frames. Trimming away surrounding benchmark context, or appending strange conversational audio, can even restore the devoted transcript. Appending VoxPopuli audio can have the other impact, making in any other case devoted artificial or mined samples extra prone to match the benchmark reference.
Collectively, these outcomes recommend that fashions are capable of faithfully transcribe the literal spoken phrases, however are utilizing surrounding acoustic context to determine whether or not to observe the audio or a benchmark-specific transcription coverage.
Conclusion
Our findings recommend that, on two main open-source datasets, some fashions detect dataset-associated acoustic cues and regulate their transcription conduct accordingly. Particularly, fashions could reproduce phrases which can be absent from the audio however current within the reference transcript, get better silenced numbers at elevated charges, or use surrounding acoustic context to pick out the written variant anticipated by a selected benchmark.
For individuals choosing fashions, these findings underscore the significance of utilizing absolutely held-out analysis units, as RW-Voice-EQ Bench and the Open ASR Leaderboard do, and of trying past phrase error price on a single public benchmark. To this finish, a “Benchmark becoming” tab has been added to the Open ASR Leaderboard, which incorporates two of the above analyses throughout all fashions: quantifying (1) reference error charges from VoxPopuli and (2) orthographic switching throughout all public datasets. The related scripts are open-sourced on GitHub in addition to the un-normalized mannequin outputs.
Our findings additionally recommend that benchmark builders ought to keep away from easy impartial and identically distributed take a look at splits in favor of temporal, speaker, or different metadata-based separation. Better transparency round coaching information and model-selection procedures would additionally assist researchers perceive how these behaviors come up.
Public benchmarks stay helpful: they’re clear, repeatable, straightforward to run, and effectively understood by the analysis group. However they’re most helpful after we can distinguish real transcription enhancements from benchmark-specific positive factors that don’t generalize to new audio.
For extra data, we encourage you to learn our full report.
