Glossary

What is ASR (speech recognition)?

Also called: STT · speech-to-text

ASR, or automatic speech recognition, is the step that converts the caller's spoken audio into text the system can act on.

Accuracy varies enormously by language, accent, background noise and line quality. A model strong on studio English can be close to unusable on a Rayalaseema speaker on a moving bus.

Why it matters for Indian calling. Almost every failure blamed on 'the AI not understanding' is an ASR failure. If the words never arrive, no language model can rescue the turn.

Related terms

Call meStart free