Naturalness score

Latency & Performancealso: naturalness ratingalso: TTS naturalness

In one sentence

A naturalness score rates how human synthesized speech sounds on a roughly 1 to 5 scale averaged from listener ratings, and it has become a weaker way to choose a voice as leading systems cluster near the human ceiling of 4.5 to 4.7.

4.14 MOS Independent mean opinion score testing, 2026Last reviewed 31 July 2026

Definition

A naturalness score rates how human synthesized speech sounds, usually on a five-point scale derived from listener ratings.

It is the headline quality number in text-to-speech marketing, and it is a specific application of mean opinion score methodology.

Naturalness scoring measures whether speech sounds like it was produced by a person, and nothing else. It says nothing about intelligibility, whether domain terms are pronounced correctly, how prosody holds over long content, or whether the voice sounds the same from one session to the next. A voice can score highly on naturalness and mispronounce every product name, and the score will not reflect that.

The ceiling problem

  • Human speech itself typically scores around 4.5 to 4.7 rather than 5.0, because listeners reserve the top of the scale.
  • Leading synthesis is now at or near that band on short samples.
  • Once systems cluster near the ceiling, absolute naturalness scores stop discriminating usefully, and comparison-based protocols become necessary.
  • So naturalness score is becoming less useful as a selection criterion precisely because the field improved.

What now differentiates voices, given the ceiling

  • Long-form prosodic stability, since drift over multiple sentences is where an audible difference remains.
  • Pronunciation control and lexicon support.
  • Latency, specifically streaming time to first byte.
  • Consistency, meaning the same voice sounding the same across sessions.
  • Commercial terms, including whether the voice may be used commercially and whether it derives from a real performer.

Methodology sensitivity

  • Sample length, listener pool, whether a human anchor was included and the wording of the scale all move the result.
  • Predicted naturalness, estimated by a model rather than collected from listeners, is common in vendor material and should be labeled as a prediction.
  • Two labs testing the same system can reasonably produce results a third of a point apart.

Common misconception

That a higher naturalness score means a better choice for a business voice agent. Above roughly 4, other properties matter more, and the score is measured on content unlike what a business agent actually says.

Why it matters commercially

It is the number in the marketing material, so it needs to be readable. The practical guidance is to use it as a filter to exclude the poor options, then choose between whatever survives by listening to your own content.

In voice specifically

A page of text is judged on whether it is right. Audio has to be judged on whether it is right and on how it sounds, and there is no instrument for the second half, which is how an average of listener opinions ended up as the headline number for synthesis quality.

Where AsqVox fits

A voice should be tested on the agent's worst-case content, meaning the prices, product names and reference codes AsqVox draws from your uploaded documents, rather than on the scripted samples a naturalness study uses. That is the content where a high-scoring voice most often falls down.

Visual

The metric that stopped discriminating

The metric that stopped discriminatingNaturalness score, 1 to 5recognizably syntheticolder neural systemscurrent commercial, the crowd zoneHUMAN SPEECH SITS HERElisteners rarely award this1MOS5MOS4.14MOSElevenLabsindependent testing4.7MOSSesame CSMstrongest open-source result cited4.5MOSWhere it stops discriminatingabove this, differences need comparison testing, not absolute scores

Use it to exclude. Do not use it to choose.

What now separates voices sits outside the score: long-form prosodic stability, pronunciation and lexicon support, streaming time to first byte, cross-session consistency, and commercial usage rights. And several things move a naturalness score without the voice changing: sample length, listener pool, whether a human anchor was in the test set, the wording of the scale, and whether the figure was predicted by a model rather than collected from listeners.

Statistics

Every figure carries its source and year. Vendor numbers are labelled as vendor numbers, and where no reliable figure exists this page says so rather than borrowing one.

ElevenLabs measured around 4.14 mean opinion score for naturalness in independent testing.

4.14 MOSindependent

Independent mean opinion score testing, 2026 - Independent in the sense that the testers were not the vendor. It is still an averaged listener judgment, so it carries the methodology dependency every naturalness figure carries.

Sesame CSM scored around 4.7, cited as the strongest open-source naturalness result.

4.7 MOSindependent

Reported results, cited as the leading open-source figure, 2026 - This sits inside the human reference band, which is where the scale stops separating systems. A comparison-based protocol is the only way to rank anything up here.

Human speech itself is generally placed at 4.5 to 4.7 rather than 5.0.

4.5 to 4.7industry range

Standard mean opinion score reference band, 2026 - Listeners reserve the top of the scale, including for real people. A synthetic voice at 4.5 is at the human reference, not 90 percent of the way there.

Because the human reference is not 5.0, a percentage-style claim such as a given proportion of human quality does not mean anything on this scale.

Human reference is not 5.0industry range

Follows from the 4.5 to 4.7 human reference, 2026 - The interpretive fact that matters most on this page. It also holds only for the short sample lengths naturalness studies use, not for a long, number-heavy answer.

Streaming synthesis time to first byte typically lands in the 100 to 200 millisecond band, and it now differentiates commercial voices more than naturalness does.

100 to 200 msindustry range

Common production budgeting convention, 2026 - An engineering convention drawn from production practice, and it applies to streaming synthesis only. It is quoted here because latency separates voices that naturalness scores no longer can.

No single authoritative, continuously updated public naturalness leaderboard covers commercial text-to-speech providers the way speech recognition benchmarks cover theirs.

-no reliable figure

So two naturalness figures from two sources cannot be lined up against each other, and a vendor comparison table of them should be read as marketing rather than as evidence.

No standard independent benchmark exists for intelligibility as distinct from naturalness, nor for pronunciation accuracy on domain vocabulary, nor for prosodic stability over long-form content. All three now matter more than naturalness, and none is measured publicly.

-no reliable figure

Which is the deeper problem: the properties that decide whether a business voice agent works are exactly the ones no public number covers.

Examples

In practice

A team selects a voice on the strength of a naturalness-score difference of 0.15 and finds in production that the lower-scoring alternative holds prosody better across multi-sentence answers. Absolute scoring on short samples had concealed the property that mattered. Re-evaluating with comparison-based testing on long-form content reverses the decision.

The everyday version

A naturalness score is how human a voice sounds, rated out of five by listeners. The good ones are now clustered so closely together that the number no longer tells you much. What separates them is whether the voice holds up over a long answer and whether it says your product names correctly, neither of which the score measures.

Usage

Who says it

  • Text-to-speech vendors, prominently in marketing.
  • Speech researchers, with the methodology attached.
  • Buyers, who meet it in comparison material and rarely interrogate it.

Where it turns up

  • In a spec sheet, next to voice quality, available voices and sample audio.
  • A rigorous claim states the listener protocol. Most do not.

Common misuse

  • Comparing scores from different studies as though they share a scale.
  • Quoting predicted scores without labeling them as predictions.
  • Treating small differences near the ceiling as meaningful.

Questions people ask

What is a naturalness score?

It is a rating of how human synthesized speech sounds, usually on a five-point scale derived from listener ratings. It is a specific application of mean opinion score methodology, and it is the headline quality number in text-to-speech marketing. It measures whether the speech sounds like a person, and nothing about intelligibility or pronunciation.

What is a good naturalness score?

The 4 to 4.5 band is current commercial quality, and 4.5 to 4.7 is where human speech itself sits, because listeners rarely award 5 even to a real person. Above roughly 4.5 the score stops discriminating usefully, so differences at the top of the scale need a comparison-based protocol rather than absolute ratings, and small gaps near the ceiling are not meaningful.

Why is a naturalness score less useful now?

Because leading synthesis clusters near the human reference band on short samples, so absolute scores stop separating systems. The field improved its own metric into uselessness. What differentiates voices now sits outside the score: long-form prosodic stability, pronunciation and lexicon support, streaming latency, cross-session consistency, and commercial usage rights.

What is predicted naturalness?

A naturalness score estimated by a model trained on human ratings, rather than collected from listeners. It is cheap and fast, and it is a prediction of a judgment rather than the judgment itself, so it should always be labeled as predicted. Where a naturalness figure appears with no listener protocol described, treat it as unverified.

Share this definition

Last reviewed 31 July 2026. Written and reviewed by Dhruv Dholakia, founder of AsqVox.