VOICE-H: a human evaluation benchmark for text-to-speech, with an interactive leaderboard, head-to-head records, price versus performance, rater demographics, and the preference distribution.
VOICE-H: a human benchmark for text-to-speech
Real people judged nine AI voices against the original human recording, on real conversation, and gave the reason behind every preference.
The VOICE-H leaderboard
Evaluating voice quality is challenging: judgments are inherently subjective, and aggregate preference scores rarely reveal what actually drives them. VOICE-H is grounded in real human speech at both ends, supplying both the prompt and the baseline sample that every model is judged against.
We selected quotes and audio clips from an Askable Labs dataset of over 40,000 real user interviews, spanning topics that naturally elicited emotional expression. Participants compared TTS-generated versions of each quote, rated them for naturalness, accuracy, and emotion/tone, and explained the reasoning behind every rating.
Capturing the qualitative nuances behind preference decisions is what turns benchmark results into actionable insights. VOICE-H goes beyond measuring which voices people prefer, to identify the specific qualities that make speech sound natural, engaging, and human.
The benchmark now runs in three languages, and the filter above switches every figure on this page between the pooled result and a single study.
Head to head
Each cell is how often the row's voice was preferred over the column's, across all comparisons. Amber cells win, indigo cells lose.
Price vs performance
Elo against list-price cost per prompt. Top-left is the sweet spot: strong and cheap.
What makes a voice feel human
Emotion has to match the content
Flat, monotone delivery was the single biggest complaint.
“A sounds like someone living the experience. B sounds like someone reading from a script.”
Voices can be too natural
Overly performative delivery drew complaints of its own.
“It felt like the voice was trying to sell me something rather than just talking.”
Accuracy can only lose you votes
Mispronunciations were forgiven. Wrong or missing words were not.
“Grammar may not be good, but B sounded more natural and real.”
Disfluencies have to be earned
Natural hesitation was rewarded. Scripted ums were punished.
“The ums were bad, like someone was told to say um rather than meaning it.”
Recording quality mattered
Real-world audio hurt the human clip, but not always enough to lose.
“Despite the bad audio, this sounded exactly like a real person. Best of the lot.”
In their own words
Comments raters left on the comparisons that went against the human recording, grouped by what they reveal. Follows the study filter; non-English comments are translated with the original underneath.
Who judged
Identity-verified people, not an anonymous crowd. Here is exactly who they were.
How strongly raters preferred
Distribution of the five preference labels. Sides were randomised per comparison, and rater-quality filtering removes a small number of votes from each study before these counts.
Methodology
100 English prompts were curated from over 40,000 real Askable user interviews: 10 seconds to three minutes of contiguous, emotional speech, manually reviewed for transcription accuracy. French and Spanish each use 50 prompts and human baseline recordings from separate native datasets in each language, not from the original English interviews. Nine TTS models plus the original human recording give 45 unique pairs per prompt; each rater judged 15 randomised comparisons for one prompt, rating each sample for naturalness, accuracy, and emotion/tone, then giving an overall preference on a five-point scale with free-text reasoning. Elo scores use ordinal Bradley-Terry preference with weighting, taking into account the strength of the preference, with bootstrap confidence intervals. The All languages view is a single fit over every vote across the three studies, rather than an average of their separate scores. The full methodology, findings, and rater quotes are in the research post.
What's next
VOICE-H is a living benchmark. English, French and Spanish are live, and German is next. We are also recording purpose-made samples to strengthen the human baseline for the next English round.
Want your voice model in the next run?
VOICE-H is built on real human conversation and judged by the Askable Labs panel. If you build voice AI, put it in front of real people, and see not just how it ranks but why.