VOICE-H: a human evaluation benchmark for text-to-speech, with an interactive leaderboard, head-to-head records, price versus performance, rater demographics, and the preference distribution.

HomeHuman Evals › VOICE-H

VOICE-H: a human benchmark for text-to-speech

Real people judged nine AI voices against the original human recording, on real conversation, and gave the reason behind every preference.

The VOICE-H leaderboard

Evaluating voice quality is challenging: judgments are inherently subjective, and aggregate preference scores rarely reveal what actually drives them. VOICE-H is grounded in real human speech at both ends, supplying both the prompt and the baseline sample that every model is judged against.

We selected quotes and audio clips from an Askable Labs dataset of over 40,000 real user interviews, spanning topics that naturally elicited emotional expression. Participants compared TTS-generated versions of each quote, rated them for naturalness, accuracy, and emotion/tone, and explained the reasoning behind every rating.

Capturing the qualitative nuances behind preference decisions is what turns benchmark results into actionable insights. VOICE-H goes beyond measuring which voices people prefer, to identify the specific qualities that make speech sound natural, engaging, and human.

The benchmark now runs in three languages, and the filter above switches every figure on this page between the pooled result and a single study.

ModelElo
View as table · full results →
Each system has its own colour; the human recording is the red dot. Whiskers are 95% bootstrap confidence intervals. Toggle for per-category scores.
Results

Head to head

Each cell is how often the row's voice was preferred over the column's, across all comparisons. Amber cells win, indigo cells lose.

Loses
Wins
Win rate over decisive (non-tie) votes. Hover any cell for the full record.
Value

Price vs performance

Elo against list-price cost per prompt. Top-left is the sweet spot: strong and cheap.

Gemini is not only the best, it is strong for its cost. The Eleven Labs models are expensive without leading.
Findings

What makes a voice feel human

01

Emotion has to match the content

Flat, monotone delivery was the single biggest complaint.

“A sounds like someone living the experience. B sounds like someone reading from a script.”
02

Voices can be too natural

Overly performative delivery drew complaints of its own.

“It felt like the voice was trying to sell me something rather than just talking.”
03

Accuracy can only lose you votes

Mispronunciations were forgiven. Wrong or missing words were not.

“Grammar may not be good, but B sounded more natural and real.”
04

Disfluencies have to be earned

Natural hesitation was rewarded. Scripted ums were punished.

“The ums were bad, like someone was told to say um rather than meaning it.”
05

Recording quality mattered

Real-world audio hurt the human clip, but not always enough to lose.

“Despite the bad audio, this sounded exactly like a real person. Best of the lot.”

In their own words

Comments raters left on the comparisons that went against the human recording, grouped by what they reveal. Follows the study filter; non-English comments are translated with the original underneath.

The panel

Who judged

Identity-verified people, not an anonymous crowd. Here is exactly who they were.

Every rater sat one study, in one language, and saw no more than 15 comparisons. Language categories overlap where derived from panel language profiles.

How strongly raters preferred

Distribution of the five preference labels. Sides were randomised per comparison, and rater-quality filtering removes a small number of votes from each study before these counts.

Method

Methodology

100 English prompts were curated from over 40,000 real Askable user interviews: 10 seconds to three minutes of contiguous, emotional speech, manually reviewed for transcription accuracy. French and Spanish each use 50 prompts and human baseline recordings from separate native datasets in each language, not from the original English interviews. Nine TTS models plus the original human recording give 45 unique pairs per prompt; each rater judged 15 randomised comparisons for one prompt, rating each sample for naturalness, accuracy, and emotion/tone, then giving an overall preference on a five-point scale with free-text reasoning. Elo scores use ordinal Bradley-Terry preference with weighting, taking into account the strength of the preference, with bootstrap confidence intervals. The All languages view is a single fit over every vote across the three studies, rather than an average of their separate scores. The full methodology, findings, and rater quotes are in the research post.

Evolving

What's next

VOICE-H is a living benchmark. English, French and Spanish are live, and German is next. We are also recording purpose-made samples to strengthen the human baseline for the next English round.

English · live French · live Spanish · live German · soon
Get an email when a new language or benchmark lands. No spam, unsubscribe anytime.

Want your voice model in the next run?

VOICE-H is built on real human conversation and judged by the Askable Labs panel. If you build voice AI, put it in front of real people, and see not just how it ranks but why.