Glossary

Speech Recognition

Speech Recognition is the speech-to-text half of the Web Speech API, exposed as SpeechRecognition, which turns microphone audio into transcripts with confidence scores. It is specified by the W3C Audio Working Group and the Web Platform Incubator Community Group. A detail that surprises people is that on some browsers, such as Chrome, the audio is sent to a remote service, so recognition does not work offline.

How it works

You create a SpeechRecognition object, set a few properties, attach event handlers, and call start(). The browser asks the user for microphone permission and must show a clear recording indicator. The main properties and defaults are:

  • lang is a BCP 47 tag, and if unset it falls back to the document's language.
  • continuous defaults to false, so you get at most one final result. Set it to true for dictation.
  • interimResults defaults to false. Set it to true to receive partial transcripts while the user is still speaking.
  • maxAlternatives defaults to 1.

The recognizer fires start, audiostart, speechstart, speechend, audioend, result, nomatch, error and end events. The key one is result. Its event.results is a list of SpeechRecognitionResult objects, each holding up to maxAlternatives alternatives sorted by confidence, from highest to lowest. Every alternative has a transcript string and a confidence between 0 and 1. A result's isFinal flag is false for interim text and true once it is settled.

Loop from event.resultIndex to the end of the list. Entries below that index are unchanged, and interim results can be overwritten by later events. stop() ends listening and still returns a result from audio already captured. abort() stops at once and returns nothing. In both cases an end event fires.

The snippet below uses plain objects shaped like an event, since Node has no microphone. It splits final text from interim text.

const result = (transcript, confidence, isFinal) => Object.assign([{ transcript, confidence }], { isFinal });
const results = [result("hello world", 0.93, true), result(" how are", 0.41, false)];
const event = { resultIndex: 0, results };
let finalText = "", interim = "";
for (let i = event.resultIndex; i < event.results.length; i++) {
  const r = event.results[i];
  if (r.isFinal) finalText += r[0].transcript; else interim += r[0].transcript;
}
console.log(JSON.stringify({ finalText, interim }));

Output:

{"finalText":"hello world","interim":" how are"}

What do the Speech Recognition error codes mean?

The error event reports one of eight codes. no-speech means nothing was heard, and audio-capture means the microphone failed. not-allowed means the browser blocked speech input for security, privacy or user-preference reasons. service-not-allowed means the requested speech service is blocked but another might be allowed. language-not-supported and network are self-explanatory. The list also includes aborted and phrases-not-supported.

Common pitfalls

  • Expecting offline use: some engines, including Chrome's, send audio to a server and fail without a network, giving the network error.
  • Treating isFinal as the whole transcript: in continuous mode, join the final results in order. The spec says whitespace is included where needed so they concatenate.
  • Rebuilding text from index 0 every time: start from event.resultIndex, because earlier entries are frozen and later ones can be replaced.
  • Relying on confidence alone: the spec notes the value is engine specific, so use it for ranking alternatives, not as a precise probability.
  • Starting without a gesture or permission: a denied prompt gives not-allowed. Start from a click and show the user what to do when it fails.
  • Not handling end: recognition stops on its own after silence when continuous is false. Restart it in the end handler if you want it to keep listening.

Related terms

  • Speech Synthesis — the text-to-speech half of the Web Speech API
  • DOM — the event system that delivers result and error events
  • HTTPS — the usual transport for pages that capture microphone audio
  • Permissions-Policy header — controls feature access in embedded frames
  • PWA — installed apps often add voice input

See also

  • Term: Speech Synthesis — the matching API for speaking text aloud