---
title: >-
  Measure call quality on Kenyan networks: MOS, jitter, packet loss, and what to
  do about each
description: >-
  MOS, jitter, RTT, and packet loss explained for voice in Kenya — how to
  measure each without a QoS lab, and a remediation ladder from codec to
  carrier.
summary: >-
  What MOS 4.2 vs 3.5 means for a business call, how jitter buffers and packet
  loss interact, and a measurement plan built on test calls, listening panels,
  and the call outcomes you already collect.
date: 2026-09-22T00:00:00.000Z
type: blog
---


## Summary

A call either sounds fine or it does not — and "does not" has exactly three measurable root causes: delay, jitter, and packet loss, which roll up into one perceptual score, MOS. This post explains each metric properly, shows what a MOS of 3.5 versus 4.2 means for a business call in Kenya, and lays out a measurement plan that needs no QoS lab: instrument the endpoints you control, run scheduled test calls with a small listening panel, and mine the call outcomes you already collect. It closes with a remediation ladder ordered from the cheapest fix (codec choice) to the last resort (carrier escalation).

## The four numbers that describe a call

### MOS: the score everything rolls up into

Mean Opinion Score comes from ITU-T P.800: a panel of listeners rates call audio from 1 (bad) to 5 (excellent), and the average is the MOS. Narrowband telephony tops out around 4.4 — that is what a clean G.711 landline call scores — so nobody ever gets a 5.

Because convening listener panels for every call is impractical, the industry mostly *estimates* MOS from network measurements using the ITU-T G.107 E-model. The E-model combines delay, loss, and codec impairments into an R-factor from 0 to 100, which maps onto MOS: R 80 is roughly MOS 4.0, R 70 roughly 3.6, R 60 roughly 3.1. Two things follow from that model. First, MOS is not linear — the perceptual drop from 4.2 to 3.5 is far larger than from 4.4 to 4.2. Second, impairments *stack*: 250 ms of delay you could live with plus 2% loss you could live with can together produce a call nobody can live with.

### One-way delay and RTT

Delay does not distort audio; it destroys conversation. ITU-T G.114 recommends keeping one-way mouth-to-ear delay under 150 ms; between 150 and 400 ms callers start talking over each other, and beyond 400 ms interactive conversation breaks down. One caution: the round-trip time you measure at the IP layer is not mouth-to-ear delay. Add codec framing, 20 ms packetisation, and the jitter buffers at both ends — in practice mouth-to-ear is roughly RTT/2 plus 40–80 ms of buffering and processing.

### Jitter: the quiet one

Voice is transmitted as a metronome: one RTP packet every 20 ms. Jitter is how unevenly those packets actually arrive — RFC 3550 defines the standard interarrival-jitter statistic that RTCP reports carry. The receiver fixes uneven arrival with a **jitter buffer**: it holds packets briefly and re-spaces them. That buys smoothness at the price of delay, and packets that arrive later than the buffer can wait for get *discarded* — so excess jitter shows up as packet loss even when nothing was lost on the wire. Under ~30 ms of jitter a modest buffer absorbs everything; above that you are trading between added delay and discard-driven gaps. Mobile data paths, where jitter swings by the minute, are exactly where adaptive jitter buffers earn their keep.

### Packet loss: distribution matters more than the percentage

Loss is the fraction of packets that never play — including the late ones the jitter buffer dropped. At 20 ms packetisation you send 50 packets a second, so 1% loss is one 20 ms hole every two seconds. Whether that is audible depends on two things:

- **Distribution.** 2% loss scattered randomly is largely concealable; 2% concentrated in bursts deletes whole syllables. Mobile networks produce burst loss — a radio handover or a congested cell drops a run of packets, not one in fifty.
- **Codec resilience.** G.711's receiver-side concealment papers over gaps up to about 30 ms; Opus embeds forward error correction so single lost packets are reconstructed from their neighbours. Our [codec selection post](/blog/codec-selection-kenya) has the measured MOS-versus-loss curves, including the ~3% loss crossover where Opus with FEC overtakes G.711 — read that rather than have it repeated here.

## What 3.5 versus 4.2 means for a business call

| MOS | R-factor (G.107) | What callers experience |
|---|---|---|
| 4.3+ | 90+ | Indistinguishable from a good landline; rare end-to-end over mobile |
| 4.0–4.3 | 80–90 | Toll quality; nobody repeats anything |
| 3.5–4.0 | 70–80 | Audible artefacts; conversation flows with occasional effort |
| 3.1–3.5 | 60–70 | Amounts and names get repeated; IVR menus get replayed |
| Below 3.1 | Below 60 | Callers talk over each other, mishear, or give up |

The commonly cited floor for commercial voice service is 3.5, and the table shows why it is the right line for business calls specifically. At 4.2, a payment-reminder call states a KES amount once and it lands. At 3.5, the caller asks for the amount again — annoying, but recoverable. At 3.2, the reference number in your TTS script is misheard, the caller pays the wrong amount, and you have created a support ticket with a phone call that was supposed to prevent one.

Scripted calls suffer *more* from marginal quality than human conversations do, not less. A human agent repeats and rephrases instinctively; a `say` verb reads "your balance is KES 4,850" exactly once, and a burst of packet loss over "850" cannot be inferred from context. That asymmetry is why quality measurement deserves engineering time even when nobody is complaining loudly yet.

## Why these numbers move on Kenyan mobile networks

The volatile segment is almost always the caller's radio leg, and it moves on schedules you can learn. Congestion is diurnal — the measurements in the codec post found 3–8% packet loss on Nairobi Eastlands 3G during the 19:00–21:00 peak versus under 1% in the CBD off-peak. Coverage is geographic: our [network realities briefing](/blog/kenyan-network-realities) maps the 2G-dominant zones where AMR-NB becomes the codec floor and in-band DTMF degrades, and covers how the three carriers reach Sautikit over different interconnect paths with different latency budgets. And the interconnect sets a quality ceiling: once a PSTN leg has been through a G.711 transcode, no codec choice on your side recovers what was discarded.

The operational takeaway for measurement: any quality number you collect is meaningless without **time of day, carrier, and rough location** attached. A single "average MOS" across your traffic hides exactly the pattern you need to find.

## Measuring without a QoS lab

Sautikit's part of this stack is the outcomes layer — call statuses, durations, and hangup causes via the API and per-number events, plus codec negotiation at the SIP edge. Per-call jitter and loss are properties of the media path, and you measure them at the endpoints you control. A workable programme has three layers.

### 1. Instrument the endpoints you control

If you run a browser softphone, WebRTC hands you the numbers free: `RTCPeerConnection.getStats()` exposes jitter, packets lost, and round-trip time for the leg between the browser and the media gateway. Log a stats snapshot every 10 seconds during calls and ship it with your application telemetry. SIP softphones and PBXs similarly log the RTCP metrics both ends exchange. This only covers your segment — the PSTN leg beyond the gateway is invisible from here — but your segment is the one you can actually fix, so measure it first.

### 2. Scheduled test calls and a listening panel

For the end-to-end path, borrow the method behind MOS itself: structured listening, simplified. Place automated test calls to real handsets on each carrier at fixed hours, play a reference script, and have the listener score each call 1–5. The script should be the hard stuff: digit strings, KES amounts, and both English and Swahili sentences — the same material your production calls carry.

```bash
npm install @sautikit/node@0.2.0
```

> 
> 
> `@sautikit/node` is in preview — pin the version; the API may change before 1.0.
> 
> 

```js
import { SautikitClient } from "@sautikit/node";

const client = new SautikitClient({ apiKey: process.env.SAUTIKIT_API_KEY });

// Cron at 08:00, 13:00 and 19:00 EAT: one scored test call per listener.
// The number's voice_callback_url serves the reference script (say/play actions).
const hourKey = new Date().toISOString().slice(0, 13); // "2026-09-22T16"
const { call_id } = await client.calls.create({
  from: process.env.QA_NUMBER,
  to: listener.msisdn,
  idempotencyKey: `qa-call:${listener.msisdn}:${hourKey}`,
});
// Listener logs: { call_id, hour: hourKey, carrier, location, score: 1-5 }
```

Three listeners across carriers and locations, three time slots, five working days: 45 scored calls a week. That is a small sample, but diurnal congestion and carrier-specific degradation are large effects — they show up within two weeks of disciplined scoring. The idempotency key means a re-run cron job never double-dials a panelist.

### 3. Lagging indicators from outcomes you already have

Quality problems leave fingerprints in ordinary call data before anyone files a complaint:

- **Short-call rate.** Completed calls under 10 seconds are heavily enriched with "answered, unusable, hung up".
- **No-input loops.** Rising `getDigits` retries and menu replays often mean callers could not hear the prompt, not that they were confused — see the [DTMF reliability post](/blog/dtmf-detection-reliability) for the detection side.
- **Complaint rate.** The classic lagging indicator. By the time "the line was bad" reaches support, the panel above would have caught it days earlier — but track it anyway as the ground truth your other metrics predict.

For the coarse view — how many calls completed over the window your panel scored — the SDK exposes the workspace rollup directly, so you are not aggregating it yourself:

```js
const week = await client.calls.stats("7d"); // "today" | "7d" | "month"
console.log(week);
```

That rollup is workspace-wide, not per-hour, so the hourly breakdown still comes from events you accumulate. Set the number's `events_url` and bucket per hour:

```js
app.post("/events", express.json(), (req, res) => {
  // Verify X-Sautikit-Signature before trusting the body.
  const { kind, data } = req.body;
  if (kind === "call.completed") {
    const hour = data.ended_at.slice(0, 13);
    const b = stats.get(hour) ?? { completed: 0, short: 0 };
    b.completed += 1;
    if (data.duration_seconds < 10) b.short += 1;
    stats.set(hour, b);
  }
  res.sendStatus(200);
});
```

Chart short-call rate by hour next to your panel scores; when they move together you have a monitoring system, not two anecdotes. If you would rather ask than build, the [Sautikit MCP server](https://sautikit.com/mcp) exposes `list_calls` and `get_call_stats`, so "how did completed-call durations after 7 p.m. compare with mornings last week" becomes a question you put to Claude instead of a script you maintain.

## The remediation ladder

Work upward; each rung is cheaper than the one after it.

1. **Fix the codec on the segment you control.** If your softphone or WebRTC leg rides mobile data, prefer Opus with FEC enabled and G.711 as fallback. The [codec post](/blog/codec-selection-kenya) has the SDP configuration and the measured case for it.
2. **Right-size the jitter buffer.** Prefer adaptive buffers where the client offers them (WebRTC's is adaptive by default). For fixed buffers, start around 40 ms and tune against your measured jitter — every extra millisecond of buffer is a millisecond of conversational delay, so do not solve a 20 ms jitter problem with a 120 ms buffer.
3. **Slow the script down.** TTS pacing is free quality margin: short sentences, amounts repeated once, digits read in pairs ("four eight, two one"), and a `getDigits` confirmation for anything money-critical. A prompt that survives a 3.4 call is worth more than the infrastructure work to guarantee 3.9.
4. **Move the traffic.** If your panel shows the 19:00–21:00 congestion dip, stop dialling into it. Campaign calls should retry into a different hour, not the same congested cell 60 seconds later — broadcast scheduling windows, `retry_on`, and `backoff_minutes` exist for exactly this.
5. **Escalate with evidence.** Degradation on one carrier only, at all hours, in one region, is an interconnect or radio-network problem — no application change fixes it. Build the evidence pack before you escalate: session IDs (the `HD_…` identifier on every call), timestamps in EAT, both numbers, durations and hangup causes, your endpoint's jitter/loss logs, and the geographic pattern. A dated table of scored test calls gets carrier-side tickets moving in a way "customers say calls sound bad" never will.

## Get started

1. [Create a Sautikit workspace](/) and claim a phone number.
2. Top up over **M-Pesa**: KES billing, no card.
3. Point your number's `events_url` at the collector above, start logging short-call rate per hour, and schedule your first week of scored test calls.

**[Start with Sautikit →](/)** &nbsp;·&nbsp; **[See pricing →](/pricing)** &nbsp;·&nbsp; **[Need SMS, WhatsApp & an agent desk? Helloduty →](https://helloduty.com)**

## Further reading

- [Choosing G.711 vs Opus codecs for Kenyan mobile networks](/blog/codec-selection-kenya)
- [The network realities every Kenyan voice developer must know](/blog/kenyan-network-realities)
- [Reliable DTMF detection on Safaricom IVRs](/blog/dtmf-detection-reliability)
- [Webhooks: events, signatures, and retries](/developers/concepts/webhooks)
- [ITU-T G.107: the E-model](https://www.itu.int/rec/T-REC-G.107) · [ITU-T G.114: one-way transmission time](https://www.itu.int/rec/T-REC-G.114) · [ITU-T P.800: methods for subjective determination of transmission quality](https://www.itu.int/rec/T-REC-P.800)
