Mean opinion score

Latency & PerformanceMOSalso: naturalness MOSalso: listener rating

In one sentence

Mean opinion score (MOS) is a 1 to 5 quality rating for synthesized or transmitted audio, produced by asking human listeners to score samples and averaging their answers, where 5 is excellent.

4.14 MOS Independent mean opinion score testing, 2026Last reviewed 30 July 2026

Not to be confused with Word error rate.

Definition

Mean opinion score is a quality rating for synthesized or transmitted audio. Human listeners score samples and the scores are averaged.

It runs from 1 to 5, where 5 is excellent. It is an average of opinions, which is what the name says and what most people quoting it forget.

MOS came out of telecommunications, where ITU-T recommendations standardized it for judging voice transmission quality, and speech synthesis later borrowed it to judge naturalness. A panel listens to samples and rates each one, usually on a five-point absolute category rating scale, and the arithmetic mean of those ratings is the MOS.

Variants worth telling apart

  • Naturalness MOS, the common usage in text to speech, which asks how human the speech sounds.
  • Intelligibility scoring, a separate protocol which asks whether the words were correctly understood.
  • Comparison MOS, or CMOS. Listeners rate one sample against another instead of on its own, which picks up smaller differences.
  • Predicted or automatic MOS, produced by a model trained on human ratings. Fast and cheap, and what it gives you is a forecast of a judgment rather than the judgment. Label it as such, always.

Why MOS numbers do not travel between studies

  • Listener pool composition, size and native language all move the result.
  • Sample selection. Short, clean, well-punctuated sentences score higher than realistic content.
  • Anchoring. Slipping a genuine human sample into the test set shifts how listeners spread their scores across the range.
  • Scale interpretation varies between listener populations.

The ceiling problem

  • Human speech itself normally comes in around 4.5 to 4.7 instead of 5.0, because listeners keep the top of the scale in reserve. So a synthetic voice on 4.5 has reached the human reference; it is not 90 percent of the way there.
  • Once the field is inside that band MOS stops telling systems apart, and comparison-based protocols become the only option.

Common misconception

That MOS is objective. What it actually is is an average of subjective ratings, and the methodology moves it a long way. Two labs can test the same model and land a third of a point apart without either doing anything wrong. Good for sorting into rough tiers, bad for ranking finely.

Why it matters commercially

MOS is the figure text to speech marketing leads with, so a buyer has to be able to read it. In practice, use it to knock out the weak options and nothing more, then pick between whatever survives by listening to your own material under your own conditions.

In voice specifically

Written output has no equivalent problem, because a page of text is judged on whether it is right. Audio has to be judged on whether it is right and on how it sounds, and there is no instrument for the second half. That is how an average of listener opinions ended up as the headline number for synthesis quality.

Visual

The MOS scale, and where the ceiling actually is

The MOS scale, and where the ceiling actually isMean opinion score, 1 to 5bad, unusablepoor, recognizably syntheticfair to good, older neural TTSvery good, current commercialHUMAN SPEECH SITS HERElisteners rarely award this1MOS5MOS4.14MOSElevenLabsindependent testing4.7MOSSesame CSMbest open source cited4.5MOSThe ceilingabove this, MOS stops discriminating and you need comparison testing

A subjective average, honestly reported. Not a measurement.

Four things move a MOS score without the model changing at all: listener pool, sample length, whether a human anchor is included in the test set, and the wording of the scale. That is why figures from two studies do not sit on a common scale.

Statistics

Every figure carries its source and year. Vendor numbers are labelled as vendor numbers, and where no reliable figure exists this page says so rather than borrowing one.

ElevenLabs scored around 4.14 mean opinion score in independent testing.

4.14 MOSindependent

Independent mean opinion score testing, 2026 - Independent in the sense that the testers were not the vendor. It is still an averaged listener judgment, so it carries the methodology dependency every MOS figure carries.

Sesame CSM scored around 4.7 mean opinion score, cited as the strongest open-source result.

4.7 MOSindependent

Reported mean opinion score results, cited as the leading open-source figure, 2026 - This sits inside the human reference band, which is where the scale stops separating systems. A comparison-based protocol is the only way to rank anything up here.

Human speech itself is generally placed at 4.5 to 4.7 rather than 5.0.

4.5 to 4.7industry range

Standard mean opinion score reference band, 2026 - Listeners reserve the top of the scale, including for real people. Any reading of a MOS figure that treats 5.0 as the human benchmark is reading it wrong.

Since the human reference is not 5.0, a claim phrased as a percentage of human quality, 95 percent for instance, does not mean anything on this scale. On short samples the leading systems already sit at or beside the human band for naturalness.

At the human bandindustry range

Follows from the 4.5 to 4.7 human reference, 2026 - The interpretive fact that matters most on this page. Short samples are doing work in that sentence: the finding holds for the sample lengths MOS studies use, not for a long, number-heavy answer.

No single authoritative, continuously updated public MOS leaderboard covers commercial text to speech providers the way speech recognition benchmarks cover theirs. The figures are strewn across individual studies, each run to its own methodology.

-no reliable figure

Which is why two MOS numbers from two sources cannot be lined up against each other, and why a vendor comparison table of MOS figures should be read as marketing rather than as evidence.

Vendor material leans on predicted MOS models routinely and rarely says that is what they are.

-no reliable figure

Where a MOS figure appears with no listener protocol described, treat it as unverified. Ask how many listeners, drawn from where, hearing what samples.

Examples

In practice

Two voices go into a MOS study and come out 0.15 apart. The same samples then get run through a comparison-based protocol, and listeners come down clearly and consistently on the side of the voice that scored lower, once the content runs long: its prosody survives a multi-sentence answer where the other one wanders. Absolute scores on short samples had buried the difference that actually mattered.

The everyday version

Mean opinion score is the number you get by playing recordings to a roomful of people, having each of them mark it out of five, and taking the average. Handy for eliminating the poor ones. No substitute for sitting down and hearing it read your own price list.

Usage

Who says it

  • Speech researchers and text to speech engineers, with the methodology attached.
  • TTS vendors in marketing material, frequently without the methodology attached.
  • Buyers rarely, since they tend to judge by ear, which is usually the sounder instinct.

Where it turns up

  • Research papers, model cards, vendor comparison pages, and now and then an RFP response covering audio quality.
  • In telecoms it keeps its original meaning of transmission quality, which causes confusion in any conversation that spans both fields.

Common misuse

  • Setting MOS figures from separate studies against each other as though one scale ran through both.
  • Quoting predicted MOS without labeling it as predicted.
  • Implying that 4.3 against 4.2 means something. It is usually inside methodology noise.

Questions people ask

What is a good mean opinion score?

The 4 to 4.5 band is current commercial quality, and 4.5 to 4.7 is where human speech itself sits. Above roughly 4.5 the score stops discriminating usefully, because listeners rarely award 5 even to a real person, so differences at the top of the scale need a comparison-based protocol rather than absolute ratings.

Is mean opinion score objective?

No. It is an average of subjective judgments, and the methodology moves it. Listener pool, sample length, whether a human reference was included and even the wording of the scale all change the result, and two labs testing the same model can reasonably land a third of a point apart. Treat it as a coarse filter rather than a fine ranking.

Why does human speech not score 5 out of 5?

Because listeners hold back the top of the scale. Human speech typically scores 4.5 to 4.7 on a five-point absolute rating, which means a synthetic voice at 4.5 is at the human reference rather than 90 percent of the way to it. It also means a claim like 95 percent of human quality does not mean anything on this scale.

What is predicted MOS?

A score estimated by a model trained on human ratings, rather than collected from listeners. It is cheap and fast, and it is a prediction of a judgment rather than the judgment itself, so it should always be labeled as predicted. Where a MOS figure appears with no listener protocol described, treat it as unverified.

Share this definition

Last reviewed 30 July 2026. Written and reviewed by Dhruv Dholakia, founder of AsqVox.