Hungarian speech recognition over a phone line: what we measured
Before planning a phone line we measured whether our speech recogniser understands Hungarian over an eight-kilohertz line. The question was simple: if the error stays below a certain level, the customer can speak freely. Above it, a keypad menu remains.
The result: 14.5 % word error rate on three hundred sentences passed through a real telephone audio path. On clean audio, 12.7 %. The phone line, then, barely hurts, far less than we expected.
Two things are just as important to say. This is read speech, not spontaneous. A caller who hesitates, restarts a sentence, or says a software name will score worse. So 14.5 % is a floor. And this is a simulated audio path, not yet a real carrier line.
Speech synthesis surprised us the other way. Long Hungarian words are the best area; English words inside a Hungarian sentence are the worst. So the fix is not a bigger model but a text rewriter that turns English terms into a pronounceable form. We shipped it, and it measurably helped.
The lesson: no promise without a measurement. If a claim has no measurement behind it, the claim is left out.
See it working, on your own workflows.
Next: What Kolli does · Roadmap