Speech Recognition is the speech-to-text half of the Web Speech API, exposed as SpeechRecognition, which turns microphone audio into transcripts with confidence scores. It is specified by the W3C Audio Working Group and the Web Platform Incubator Community Group. A detail that surprises people is that on some browsers, such as Chrome, the audio is sent to a remote service, so recognition does not work offline.
You create a SpeechRecognition object, set a few properties, attach event handlers, and call start(). The browser asks the user for microphone permission and must show a clear recording indicator. The main properties and defaults are:
The recognizer fires start, audiostart, speechstart, speechend, audioend, result, nomatch, error and end events. The key one is result. Its event.results is a list of SpeechRecognitionResult objects, each holding up to maxAlternatives alternatives sorted by confidence, from highest to lowest. Every alternative has a transcript string and a confidence between 0 and 1. A result's isFinal flag is false for interim text and true once it is settled.
Loop from event.resultIndex to the end of the list. Entries below that index are unchanged, and interim results can be overwritten by later events. stop() ends listening and still returns a result from audio already captured. abort() stops at once and returns nothing. In both cases an end event fires.
The snippet below uses plain objects shaped like an event, since Node has no microphone. It splits final text from interim text.
const result = (transcript, confidence, isFinal) => Object.assign([{ transcript, confidence }], { isFinal });
const results = [result("hello world", 0.93, true), result(" how are", 0.41, false)];
const event = { resultIndex: 0, results };
let finalText = "", interim = "";
for (let i = event.resultIndex; i < event.results.length; i++) {
const r = event.results[i];
if (r.isFinal) finalText += r[0].transcript; else interim += r[0].transcript;
}
console.log(JSON.stringify({ finalText, interim }));
Output:
{"finalText":"hello world","interim":" how are"}
The error event reports one of eight codes. no-speech means nothing was heard, and audio-capture means the microphone failed. not-allowed means the browser blocked speech input for security, privacy or user-preference reasons. service-not-allowed means the requested speech service is blocked but another might be allowed. language-not-supported and network are self-explanatory. The list also includes aborted and phrases-not-supported.