The material decides the number
Read speech from a benchmark corpus is far easier than a four-person meeting on a speakerphone. An engine tuned for one can look excellent on paper and unusable in practice.
Accuracy
Headline accuracy numbers for Hebrew are quoted constantly and compared almost never. Here is what the number depends on, why published figures disagree so wildly, and how KolWrite measures its own.
You can find claims of a 1.2% word error rate and a 26% word error rate for the same engine. Both can be honest. They are measuring different things.
Read speech from a benchmark corpus is far easier than a four-person meeting on a speakerphone. An engine tuned for one can look excellent on paper and unusable in practice.
Word error rate counts a missed preposition the same as a missed name or figure. For Hebrew, where a fused prefix changes meaning, that flattens exactly the errors that matter most.
Hebrew packs into single tokens what English spreads over several words. One segmentation mistake can register as several errors, so Hebrew WER is not comparable to English WER.
Most public Hebrew test sets are monolingual. Real Israeli speech is not, so benchmark scores systematically overstate performance on actual meetings and calls.
Measured on the material customers actually bring, reported by condition rather than as one number.
Clean single speaker, multi-speaker meeting, phone-quality call and noisy field recording are tracked separately, because they behave differently.
Code-switched recordings are part of the evaluation set rather than excluded from it.
Names, numbers, dates and terminology are scored separately from function words. A transcript that gets every name right is more useful than one with a marginally better raw WER.
Audio type is detected and the transcription approach is chosen to match, rather than sending every recording through one fixed model.
There is no single figure that is meaningful across conditions. On clean, single-speaker Hebrew, strong systems cluster closely together and the differences between them rarely matter. On multi-speaker, noisy or code-switched Hebrew — which is most real-world material — the spread between systems is far wider, and that is where the choice of engine actually shows.
Hebrew fuses prepositions, articles and possessives onto word stems, so a single segmentation error can be counted as several word errors. The same underlying quality therefore scores worse in Hebrew than in English, which makes cross-language WER comparisons misleading.
Test them on your own audio, not on a benchmark. Use a recording that is representative of your hardest case — the noisiest room, the most speakers, the most English mixed in — and compare the errors that would actually cost you time to fix.
Substantially. Recording conditions, number of speakers, microphone distance and how much English is mixed in all move accuracy more than the choice between leading engines does on clean audio.
Take your hardest Hebrew recording — the noisy one, with four speakers and English mixed in — and compare the output yourself.
Test it on your audio