|Guides

MOS Scores Explained: The Number Behind Call Quality

MOS is the standard voice quality metric. Learn what it measures, how it is calculated, and what counts as a good score.

What MOS stands for

MOS is Mean Opinion Score. It dates back to the 1970s, rooted in work by the CCITT (now ITU-T) that eventually became the P.800 recommendation, as a way to quantify how humans perceive voice quality. The original method was simple: a group of listeners rated audio samples on a scale of 1 to 5, and the average was the MOS.

  • 5: Perfect. Like talking in the same room.
  • 4: Good. Minor imperfections, but natural sounding.
  • 3: Fair. Noticeable distortion or effort required to understand.
  • 2: Poor. Difficult to communicate.
  • 1: Unusable.

Getting a panel of listeners together every time you want to test a call is obviously impractical. So the industry developed algorithmic models that predict what a human panel would score.

How modern MOS is calculated

Today, MOS is almost always computed by an algorithm rather than by human listeners. The two main approaches are:

PESQ (Perceptual Evaluation of Speech Quality) compares the original audio signal to what came out the other end. It literally measures how much the audio was degraded. This is called an "intrusive" or "full-reference" method because it needs the original signal for comparison. PESQ is accurate but requires access to both sides of the conversation, which makes it impractical for live monitoring.

POLQA (Perceptual Objective Listening Quality Analysis) is the successor to PESQ, designed for HD voice and modern codecs. Same basic idea: compare input to output. Higher accuracy with wideband and super-wideband audio.

E-model (ITU-T G.107) takes a different approach. Instead of analyzing the audio itself, it calculates an expected quality score based on network conditions: latency, jitter, packet loss, and codec type. The E-model outputs an "R-factor" that theoretically ranges from 0 to 100, though the narrowband practical maximum is around 93-94. The R-factor maps to a MOS score. This is the method most monitoring tools use because it works with just network metrics and does not need the original audio.

What the numbers mean in practice

Nobody gets a MOS of 5 on a real VoIP call. The codec alone introduces some degradation, so even a perfect network path tops out around 4.4 with common codecs like G.711. Here is a more realistic interpretation:

  • 4.0 to 4.4: Excellent. Users are satisfied. This is the target for business VoIP.
  • 3.6 to 4.0: Good. Occasional minor artifacts. Most users will not complain.
  • 3.1 to 3.6: Acceptable for some use cases. Customer-facing calls will generate complaints.
  • 2.6 to 3.1: Poor. Communication is strained. Users will actively seek alternatives.
  • Below 2.6: Not viable for business use.

The important thing is that MOS is not linear. The difference between 4.0 and 3.5 is much more noticeable than the difference between 4.4 and 4.0. Small drops in the middle of the scale represent significant degradation in the listener's experience.

The codec ceiling

Every codec has a maximum MOS it can achieve even under perfect conditions. This is because compression always removes some information from the audio signal.

G.711 (the traditional PSTN codec, PCM with mu-law/A-law companding): Up to 4.4. This is the baseline for toll quality. It uses 64 kbps of bandwidth per direction.

G.729 (compressed, low bandwidth): Up to 3.9. Popular on bandwidth-constrained links because it only uses about 8 kbps. The tradeoff is that quality ceiling.

Opus (modern, adaptive): Up to 4.5+ at higher bitrates. Opus can adjust its bitrate and complexity on the fly, which makes it excellent for variable network conditions. Most WebRTC applications use Opus.

If your monitoring shows a MOS of 3.9 and you are using G.729, you are actually at the codec's theoretical best. Switching to G.711 or Opus would raise the ceiling. But if you see 3.9 on G.711, something in the network is pulling quality down.

Network factors that lower MOS

The E-model gives us a clear picture of how each network issue impacts the score:

Packet loss has the largest impact per percentage point. Even 1% random packet loss drops MOS by roughly 0.3 to 0.4 points. At 3%, most calls fall below the 3.6 threshold. Bursty loss (several packets in a row) is worse than random loss at the same percentage because the jitter buffer cannot conceal gaps that long.

Latency has a gradual effect. One-way delay under 150ms is generally fine. Between 150ms and 300ms, conversations start to feel awkward because of the delay in responses. Above 300ms, people start talking over each other. The MOS impact is moderate compared to packet loss, but the conversational impact is real.

Jitter primarily affects MOS through its interaction with the jitter buffer. High jitter that exceeds the buffer creates effective packet loss. The jitter buffer itself adds latency. So high jitter forces a tradeoff: increase the buffer (more latency) or accept more discarded packets (lower quality).

Why a single number is both useful and dangerous

MOS gives you a single number to track, trend, and alert on. That is its strength. You can set a threshold (say, 3.8) and get notified when quality drops below it. You can compare quality across locations, providers, or time periods.

The danger is in treating it as the whole story. A MOS of 3.8 caused by high latency is a very different problem from a MOS of 3.8 caused by packet loss. The user experience differs, the troubleshooting differs, and the fix differs. MOS tells you something is wrong. It does not tell you what.

This is why good VoIP monitoring always presents MOS alongside the underlying metrics. The score is a summary. The details are in the latency, jitter, and packet loss numbers that feed into it.

How we use MOS at VoIP Test

When we test your connection with our VoIP quality test, we measure the raw network conditions (packet loss, jitter, latency) and run them through the E-model to produce a MOS estimate. But we also show you the breakdown so you can see exactly which factor is pulling your score down. A single number is a starting point. The underlying data is where the answers live.


For a detailed look at the three metrics that feed into MOS, see Latency, Jitter, and Packet Loss: The Details in our VoIP From the Ground Up series.

Frequently Asked Questions

What is a good MOS score for VoIP?+

A MOS score of 4.0 or above is considered good quality -- callers will not notice any issues. Scores between 3.5 and 4.0 are acceptable but may have minor artifacts. Below 3.5, callers start complaining. The maximum achievable MOS depends on the codec in use; G.711 can reach 4.41 while G.729 tops out at 3.92 even under perfect conditions. Try the MOS Explorer to see how different codecs and network conditions affect the score.

How is MOS calculated for VoIP calls?+

Modern MOS calculations use the ITU-T G.107 E-model, which takes codec type, packet loss, latency, and jitter into account to produce an R-factor between 0 and 100. The R-factor is then converted to a MOS score using a standard formula. This avoids the need for human listening panels, which were used in the original ITU-T P.800 methodology. You can experiment with the E-model inputs using the MOS Explorer.

Why is my MOS score low even though my internet is fast?+

MOS is not related to bandwidth. It is driven by latency, jitter, and packet loss -- metrics that a speed test cannot measure. A connection with plenty of bandwidth but inconsistent packet delivery will produce low MOS scores. Run a VoIP quality test to measure the metrics that actually affect call quality.

mos-scorevoip-basicscall-quality

Share

Opens your messaging app. We do not collect or store any phone numbers.
Opens your email client. We do not collect or store any email addresses through sharing.

Want to know when we publish new articles? Sign up for updates