Percentile latency

Latency & Performancealso: tail latencyalso: p95 latencyalso: p99 latencyalso: latency percentiles

In one sentence

Percentile latency describes how latency is distributed rather than averaged, so a p95 figure means 95 percent of requests were faster than it and the slowest 5 percent, the turns a visitor actually remembers, were slower.

0.78s to 2.98s Artificial Analysis, via Softcery, April 2026Last reviewed 31 July 2026

Definition

Percentile latency describes how latency is distributed instead of averaging it away. A p95 of 2,400 milliseconds means 95 percent of requests came back faster than that, and the slowest 5 percent did not.

It is the standard way to report latency, because the average hides exactly the behavior that decides whether a voice product is acceptable.

An average tells you neither the typical experience nor the bad one. A handful of very slow requests drag the mean upward without describing anyone in particular, which is why percentiles, not averages, are how latency gets reported.

The common percentiles, and what each one tells you

  • p50, the median. Half of requests are faster than this. It is the typical experience.
  • p90 and p95. The edge of normal, where noticeable slowness begins.
  • p99, the tail. One request in a hundred is this slow or slower.

Why the tail matters more in voice than on the web

  • A slow web page is mildly annoying. A slow voice response is a silence, and during it a person decides the system has failed and starts talking again.
  • Across a ten-turn conversation, a p95 event is something the visitor is likely to run into at least once. The tail is not rare from where they sit; it is the thing they meet.
  • A system with an excellent p50 and a poor p99 feels like an agent that works well and occasionally freezes, which erodes trust faster than something that is uniformly mediocre.

Why chaining stages makes the tail worse

  • A voice pipeline runs several services in sequence, and the latency of a turn depends on all of them.
  • If each stage has an independent chance of being slow, the odds that at least one of them is slow climb with the number of stages.
  • So the end-to-end p99 comes out worse than any single component's p99, sometimes by a lot. Component-level percentiles are necessary but not sufficient.

What to put in an agreement

  • A percentile, a target, a measurement point and a measurement window.
  • p95 is a common contractual choice. A p50 on its own is not a meaningful commitment.
  • Without a measurement window, a target can be met on the monthly aggregate while failing badly during peak hours.

Common misconception

That the average latency is a useful summary of it. It is close to the least useful single number available, because it describes neither the typical experience nor the bad one.

Why it matters commercially

Percentiles are how latency gets written into a contract and how quality actually gets experienced. A vendor who reports only averages is either unsophisticated or quietly hiding a tail, and either way the number to ask for is p95 with a measurement window attached.

In voice specifically

On the web the tail is a slow page and a spinner. In voice the tail is a silence with nothing on screen, and a person fills it by concluding the agent has frozen and talking over it. That is why the worst 5 percent of turns, not the median, is what a voice product gets judged on.

Visual

The average describes nobody

The average describes nobody830ms time to first audio840ms time to first audio850ms time to first audio860ms time to first audio870ms time to first audio880ms time to first audio890ms time to first audio900ms time to first audio910ms time to first audio920ms time to first audio930ms time to first audio940ms time to first audio950ms time to first audio960ms time to first audio970ms time to first audio980ms time to first audio990ms time to first audio1000ms time to first audio1010ms time to first audio1020ms time to first audio1030ms time to first audio1040ms time to first audio1050ms time to first audio1060ms time to first audio1070ms time to first audio1080ms time to first audio1090ms time to first audio1100ms time to first audio1110ms time to first audio1120ms time to first audio1130ms time to first audio1140ms time to first audio1150ms time to first audio1160ms time to first audio1170ms time to first audio1180ms time to first audio1190ms time to first audio1200ms time to first audio1210ms time to first audio1220ms time to first audio1230ms time to first audio1240ms time to first audio1250ms time to first audio1260ms time to first audio1270ms time to first audio1280ms time to first audio1290ms time to first audio1300ms time to first audio1310ms time to first audio1320ms time to first audio1330ms time to first audio1340ms time to first audio1350ms time to first audio1360ms time to first audio1370ms time to first audio1380ms time to first audio1390ms time to first audio1400ms time to first audio1410ms time to first audio1420ms time to first audio1430ms time to first audio1440ms time to first audio1450ms time to first audio1460ms time to first audio1470ms time to first audio1480ms time to first audio1490ms time to first audio1500ms time to first audio1510ms time to first audio1520ms time to first audio1530ms time to first audio1540ms time to first audio1550ms time to first audio1560ms time to first audio1570ms time to first audio1580ms time to first audio1590ms time to first audio1600ms time to first audio1610ms time to first audio1620ms time to first audio1630ms time to first audio1640ms time to first audio1650ms time to first audio1660ms time to first audio1670ms time to first audio1680ms time to first audio1690ms time to first audio1700ms time to first audio1710ms time to first audio1720ms time to first audio1730ms time to first audio1740ms time to first audio1750ms time to first audio1760ms time to first audio1770ms time to first audio1780ms time to first audio1790ms time to first audio1800ms time to first audio1810ms time to first audio1820ms time to first audio1830ms time to first audio1840ms time to first audio1850ms time to first audio1860ms time to first audio1870ms time to first audio1880ms time to first audio1890ms time to first audio1900ms time to first audio1910ms time to first audio1920ms time to first audio1930ms time to first audio1940ms time to first audio1950ms time to first audio1960ms time to first audio1970ms time to first audio1980ms time to first audio1990ms time to first audio2000ms time to first audio2010ms time to first audio2020ms time to first audio2030ms time to first audio2040ms time to first audio2050ms time to first audio2060ms time to first audio2070ms time to first audio2080ms time to first audio2090ms time to first audio2100ms time to first audio2110ms time to first audio2120ms time to first audio2130ms time to first audio2140ms time to first audio2150ms time to first audio2160ms time to first audio2170ms time to first audio2180ms time to first audio2190ms time to first audio2200ms time to first audio2210ms time to first audio2220ms time to first audio2230ms time to first audio2240ms time to first audio2250ms time to first audio2260ms time to first audio2270ms time to first audio2280ms time to first audio2290ms time to first audio2300ms time to first audio2310ms time to first audio2320ms time to first audio2330ms time to first audio2340ms time to first audio2350ms time to first audio2360ms time to first audio2370ms time to first audio2380ms time to first audio2390ms time to first audio2400ms time to first audio2410ms time to first audio2420ms time to first audio2430ms time to first audio2440ms time to first audio2450ms time to first audio2460ms time to first audio2470ms time to first audio2480ms time to first audio2490ms time to first audio2500ms time to first audio2510ms time to first audio2520ms time to first audio2530ms time to first audio2540ms time to first audio2550ms time to first audio2560ms time to first audio2570ms time to first audio2580ms time to first audio2590ms time to first audio2600ms time to first audio2610ms time to first audio2620ms time to first audio2630ms time to first audio2640ms time to first audio2650ms time to first audio2660ms time to first audio2670ms time to first audio2680ms time to first audio2690ms time to first audio2700ms time to first audio2710ms time to first audio2720ms time to first audio2730ms time to first audio2740ms time to first audio2750ms time to first audio2760ms time to first audio2770ms time to first audio2780ms time to first audio2790ms time to first audio2800ms time to first audio2810ms time to first audio2820ms time to first audio2830ms time to first audio2840ms time to first audio2850ms time to first audio2860ms time to first audio2870ms time to first audio2880ms time to first audio2890ms time to first audio2900ms time to first audio2910ms time to first audio2920ms time to first audio2930ms time to first audio2940ms time to first audio2950ms time to first audio2960ms time to first audio2970ms time to first audio2980ms time to first audio2990ms time to first audio3000ms time to first audio3010ms time to first audio3020ms time to first audio3030ms time to first audio3040ms time to first audio3050ms time to first audio3060ms time to first audio3070ms time to first audio3080ms time to first audio3090ms time to first audio3100ms time to first audio3110ms time to first audio3120ms time to first audio3130ms time to first audio3140ms time to first audio3150ms time to first audio3160ms time to first audio3170ms time to first audio3180ms time to first audio3190ms time to first audio3200ms time to first audio3210ms time to first audio3220ms time to first audio3230ms time to first audio3240ms time to first audio3250ms time to first audio3260ms time to first audio3270ms time to first audio3280ms time to first audio3290ms time to first audio3300ms time to first audio3310ms time to first audio3320ms time to first audio3330ms time to first audio3340ms time to first audio3350ms time to first audio3360ms time to first audio3370ms time to first audio3380ms time to first audio3390ms time to first audio3400ms time to first audio3410ms time to first audio3420ms time to first audio3430ms time to first audio3440ms time to first audio3450ms time to first audio3460ms time to first audio3470ms time to first audio3480ms time to first audio3490ms time to first audio3500ms time to first audio3510ms time to first audio3520ms time to first audio3530ms time to first audio3540ms time to first audio3550ms time to first audio3560ms time to first audio3570ms time to first audio3580ms time to first audio3590ms time to first audio3600ms time to first audio3610ms time to first audio3620ms time to first audio3630ms time to first audio3640ms time to first audio3650ms time to first audio3660ms time to first audio3670ms time to first audio3680ms time to first audio3690ms time to first audio3700ms time to first audio3710ms time to first audio3720ms time to first audio3730ms time to first audio3740ms time to first audio3750ms time to first audio3760ms time to first audio3770ms time to first audio3780ms time to first audio3790ms time to first audio3800ms time to first audio3810ms time to first audio3820ms time to first audio3830ms time to first audio3840ms time to first audio3850ms time to first audio3860ms time to first audio3870ms time to first audio3880ms time to first audio3890ms time to first audio3900ms time to first audio3910ms time to first audio3920ms time to first audio3930ms time to first audio3940ms time to first audio3950ms time to first audio3960ms time to first audio3970ms time to first audio3980ms time to first audio3990ms time to first audio4000ms time to first audio4010ms time to first audio4020ms time to first audio4030ms time to first audio4040ms time to first audio4050ms time to first audio4060ms time to first audio4070ms time to first audio4080ms time to first audio4090ms time to first audio4100ms time to first audio4110ms time to first audio850ms time to first audiop50 median1500ms time to first audiop902400ms time to first audiop954100ms time to first audiop99 tailSource: Illustrative distribution, not measured dataThe average sits between the cluster and the tail. It describes nobody.query types absorbed, widest and earliest firstOne turn, one p95 eventThree turnsFive turnsTen turns, a short website conversation

Voice quality is decided by the worst turns, not the average one.

The four points are the illustrative distribution from the worked example below, not a measured benchmark. The bars underneath are the chance a visitor meets at least one p95 turn as a conversation runs longer, worked out as one minus 0.95 to the power of the turn count: roughly 40 percent across ten turns, so the tail is not a rare edge case from the visitor's side. Chain several pipeline stages and the end-to-end p99 comes out worse than any single stage's p99. Specify a percentile, a target, a measurement point and a window, because a monthly aggregate can pass while peak hours fail.

Statistics

Every figure carries its source and year. Vendor numbers are labelled as vendor numbers, and where no reliable figure exists this page says so rather than borrowing one.

Sub-second time to first audio is the working design target for conversational voice, which is the line the p50 sits comfortably under and the tail routinely crosses.

Under 1,000 msindustry range

Industry working threshold, 2026 - A shared working convention rather than a published standard. It is the number the median is supposed to clear.

Past roughly 1,500 milliseconds a conversation is widely reported to feel broken, which is why a p90 or p95 sitting above that line matters even when the median looks healthy.

1,500 msindustry range

Industry working threshold, 2026 - A working line, not a measured cliff. A visitor feels it without being told to look for it.

Human conversational turn gaps average around 200 milliseconds, the baseline every latency target is set against.

About 200 msindustry range

Conversational turn-taking baseline, 2026 - A reference point, not a target you built. It is the pace the median is trying to feel like.

Independent measurements put model time to first token across leading realtime models between roughly 0.78 and 2.98 seconds, and these are typically published as central figures rather than as distributions.

0.78s to 2.98sindependent

Artificial Analysis, via Softcery, April 2026, 2026 - Cite it as a central-figure comparison. Because the numbers are not distributions, they say nothing about the tail, which is the part percentile latency exists to expose.

Voice AI providers rarely publish percentile latency distributions, reporting central figures instead, which makes independent tail assessment impossible from published material.

-no reliable figure

A genuine transparency gap in the category. The absence of a p95 in a vendor sheet is itself information.

There is no standard measurement window convention for voice latency service level agreements, so two providers can both claim to meet a p95 target while measuring it over very different windows.

-no reliable figure

Which is why a percentile target without a stated window is close to unenforceable.

Examples

In practice

A deployment reports a median time to first audio of 850 milliseconds and starts getting complaints that the agent freezes. Percentile analysis shows p95 at 2,400 milliseconds and p99 at 4,100, driven by cold starts on a retrieval service that is invoked infrequently. The median was fine the whole time. Warming the service flattens the tail and the complaints stop, with no change to the median at all.

The everyday version

Percentile latency means looking at the slowest responses rather than the typical one. If nineteen replies out of twenty are quick and one takes four seconds, the average still looks fine and your customer remembers the four-second one. In a short conversation they will almost certainly land on it.

Usage

Who says it

  • Infrastructure and reliability engineers, as a matter of standard practice.
  • Procurement, written into service level agreements.
  • Business buyers rarely say it, though they experience it directly as inconsistency.

Where it turns up

  • Next to latency targets, measurement methodology, the measurement window, reporting frequency and the remedy for a breach.
  • A commitment written as a percentile with a window is enforceable. One written as an average is theater.

Common misuse

  • Reporting averages instead of percentiles.
  • Specifying a percentile target with no measurement point or window.
  • Assuming component percentiles predict the end-to-end percentile, which understates the combined tail.

Questions people ask

What does p95 latency mean?

p95 latency is the value that 95 percent of requests came back faster than, with the slowest 5 percent slower than it. It describes the edge of normal experience rather than the typical one, which is why it is a common choice for a service level commitment. A p50, the median, tells you about the typical turn; a p95 tells you about the ones a visitor actually remembers.

Why is average latency a bad measure?

Because it describes neither the typical experience nor the bad one. A small number of very slow requests drag the mean upward without matching anyone in particular, so an average can look healthy while a real tail is making the product feel broken. A vendor who reports only averages is either unsophisticated or hiding that tail.

Why does the latency tail matter more in voice?

Because a slow voice response is a silence, and during it a person concludes the system has failed and starts talking again. Across a ten-turn conversation there is roughly a 40 percent chance of hitting at least one p95 turn, so the tail is not a rare edge case from the visitor's side. A system with a great median and a poor p99 feels like an agent that works well and then freezes, which costs more trust than being uniformly mediocre.

What should a voice latency SLA specify?

Four things: a percentile, a target, a measurement point and a measurement window. p95 is a common contractual choice, and a median alone is not a meaningful commitment. Without a window, a target can pass on the monthly aggregate while failing badly during peak hours, so the window is as important as the number.

Share this definition

Last reviewed 31 July 2026. Written and reviewed by Dhruv Dholakia, founder of AsqVox.