Glossary

Speech Synthesis

Speech Synthesis is the text-to-speech half of the Web Speech API, exposed as window.speechSynthesis, which speaks a string aloud with a voice from the user's device or browser. It is defined in the Web Speech API specification from the W3C Audio Working Group and the Web Platform Incubator Community Group. Developers most often search for the valid ranges: rate runs from 0.1 to 10, pitch from 0 to 2, and volume from 0 to 1.

How it works

You build a SpeechSynthesisUtterance, set its properties, and pass it to speechSynthesis.speak(). The object is created with new SpeechSynthesisUtterance("Hello") and has these properties:

  • text is the string to speak.
  • lang is a BCP 47 tag such as en-US.
  • voice is one of the SpeechSynthesisVoice objects from getVoices().
  • rate is relative to the voice's normal speed, default 1, where 2 is twice as fast and 0.5 is half.
  • pitch runs from 0 to 2, default 1.
  • volume runs from 0 to 1, default 1.

The speechSynthesis object keeps one global queue. speak() appends an utterance, and it plays immediately only if nothing is queued and the instance is not paused. The object also has pause(), resume() and cancel(), plus the read-only flags speaking, pending and paused. cancel() empties the whole queue and stops the current utterance at once.

Voices are user agent dependent. getVoices() can return an empty list the first time you call it, because the list may load asynchronously. The voiceschanged event fires when the list changes, so wait for it before choosing a voice. Each voice has a name, lang, voiceURI, localService flag and default flag.

Utterances fire events: start, end, error, pause, resume, boundary (word or sentence) and mark. The spec bars rates below 0.1 and above 10, and a voice may narrow that further. The snippet applies the spec's ranges to a settings object.

const clamp = (v, lo, hi) => Math.min(hi, Math.max(lo, v));
function settings({ rate = 1, pitch = 1, volume = 1 }) {
  return { rate: clamp(rate, 0.1, 10), pitch: clamp(pitch, 0, 2), volume: clamp(volume, 0, 1) };
}
console.log(settings({}));
console.log(settings({ rate: 12, pitch: -1, volume: 1.5 }));

Output:

{ rate: 1, pitch: 1, volume: 1 }
{ rate: 10, pitch: 0, volume: 1 }

What are the speechSynthesis error codes?

The error event on an utterance carries a code from a fixed list. The spec defines canceled, interrupted, audio-busy, audio-hardware, network, synthesis-unavailable, synthesis-failed, language-unavailable, voice-unavailable, text-too-long, invalid-argument and not-allowed. Calling cancel() produces canceled for queued items and interrupted for the one being spoken, so those two are usually expected, not bugs.

Common pitfalls

  • Reading getVoices() once at startup: the list can be empty at that moment. Listen for voiceschanged and read it again.
  • Treating cancel() as an error: a cancelled queue fires error events with canceled or interrupted. Filter those codes out of your error logging.
  • Changing an utterance after speak(): the spec leaves the effect undefined and says it may cause an error. Create a new utterance instead.
  • Reusing an utterance across frames: the SpeechSynthesis object takes exclusive ownership of it, and passing it to another instance should throw.
  • Out-of-range values: a value the synthesizer rejects raises invalid-argument. Clamp rate, pitch and volume first.
  • Assuming the same voices everywhere: the available voices depend on the device, so a hard-coded voice name can produce voice-unavailable. Match on lang and fall back to the default voice.

Related terms

  • Speech Recognition — the speech-to-text half of the same Web Speech API
  • DOM — the event model that delivers utterance events
  • Web Share API — another browser feature that hands content to the operating system
  • PWA — installed web apps often add read-aloud features
  • HTTPS — recommended for any page that also requests microphone input

See also

  • Term: Speech Recognition — the matching API for converting speech to text