Accurate to what, measured how, on whose audio? The comparison tables give you a percentage per product, which is the one thing accuracy is not a property of. Here is what the number actually measures, why the same tool scores twenty points apart in two reviews, and what to ask instead.
Transcription accuracy is conventionally the inverse of Word Error Rate. WER counts the edits needed to turn what the machine heard into what was actually said, divided by the length of what was said.
WER = (substitutions + deletions + insertions) / reference word count accuracy = 1 - WER
The standard speech-recognition metric, and the one our own harness computes.
A substitution is a word heard as a different word. A deletion is a word missed. An insertion is a word invented. They are weighted identically, which is why a single percentage hides the thing you care about: a transcript that drops a price is not equivalent to one that mishears a name, and both cost the same in WER.
Because the score belongs to a recording, not a product. Room noise, microphone distance, accents, crosstalk and how many people are speaking move WER far more than the badge on the app does. Most note takers are a workflow wrapped around one of a small number of speech engines, so two products can post different numbers while sending audio to the same place, and one product can post two numbers depending on whose meeting it transcribed.
A transcript can be word-perfect and still useless for a sales call, because it does not say which of you said the sentence that mattered. That is diarization, and it is scored separately from WER. A comparison table that gives one accuracy figure per tool has quietly merged two capabilities that fail independently.
Speaker attribution appears when the transcription engine that resolves for a recording performs diarization. When it does not, the words are still there and the transcript is simply flat. In our own recorder that is explicit: attributed segments exist only when the resolved engine diarizes, and each recording stores which engine attributed it, or records that none did. Ask a vendor which engine ran and whether attribution was on, rather than accepting a tick in a feature grid.
The failure mode worth asking about is what happens when transcription cannot run at all. Ours returns nothing and says so; it does not produce plausible text to fill the gap. A note taker that always returns something is not necessarily more accurate, it may simply be less willing to admit a miss, and you will not find that out from a percentage.
Because they transcribed different audio. Word Error Rate is a property of a recording and the engine that processed it, so room noise, accents, crosstalk and speaker count move it far more than the product does. Two reviews of one tool can differ by twenty points without either being wrong.
It depends entirely on which words the missing 5% were. WER weights every error the same, so dropping a price, a date and a name scores identically to mishearing three filler words. Ask what was measured and on what audio before treating a percentage as quality.
No, and this is the most common misreading. Speaker attribution is diarization, scored separately, and it depends on whether the engine that resolved for that recording performs it. A perfectly accurate transcript can still be a flat wall of text with no idea who said what.
Which engine transcribes my audio, was diarization on for this recording, what audio was your number measured on, and what happens when transcription fails. The last one matters most: a tool that always returns text may simply be filling gaps rather than admitting a miss.
We are not going to print one here, because the argument above applies to us too. What we can tell you is the method: we score the configured transcription provider with Word Error Rate against a committed set of reference recordings, hold it to a floor, and run that as a gate rather than a claim. When transcription cannot run, we return no transcript rather than generated text. Ask us about the method and we will answer; a number without one is not worth either of our time.