Which numbers are honest, which are marketing — and the questions to ask a vendor so you do not buy a black box on faith.
The first question every CTO asks about speech analytics is "So how accurate is it?" The honest answer is more complicated than a single number — and a vendor who gives you one number with no caveats should already raise a flag.
The base metric is WER (word error rate) — the share of words transcribed incorrectly. Word-level accuracy = 100% − WER. The key point: WER is always measured on a specific type of audio. The same model will post different numbers on a studio recording, a phone line, and a wearable badge on a noisy sales floor.
| Recording conditions | Word-level accuracy | Comment |
|---|---|---|
| Telephony, agent headset | 97%+ | clean channel, separate speaker tracks |
| Audio badge, sales floor | 94–97% | background music, crowd noise, overlapping speech |
| Peak-hour floor, hushed speech | below 94% | an honest vendor flags these recordings for quality |
If someone promises "99% in any conditions", it is either a lab benchmark on clean speech or plain marketing. On a floor with music and parallel conversations you cannot cheat physics: even a human transcriber loses some of the words.
For the business, what matters is scoring accuracy, not transcript accuracy. A checklist of "stated the price", "offered the promo", "closed politely" is robust to isolated recognition errors: you do not need every word transcribed perfectly to score a line item correctly.
How we handle this ourselves — in the Frontline FAQ and the QA methodology.
We take your real audio and show you the transcripts and accuracy in a live demo.
Your email client opens right away — we reply within one business day.