Tutorials Audio

Homer Dudley: Taking Speech Apart and Rebuilding It

Profile ยท ~15 min
Historical photograph of Homer Dudley.Image credit & licence

Overview

Homer Dudley (1896โ€“1980) built the Vocoder at Bell Labs in the 1930s: a system that analyses speech into the energy in a small number of frequency bands plus a pitch signal, transmits only those, and resynthesises intelligible speech at the far end. The Voder, demonstrated at the 1939 World's Fair, was a keyboard-played version. Every speech codec descends from this idea.

What You Need

  • Lived: 1896โ€“1980, United States
  • Anchor year: 1939 โ€” the Voder demonstrated
  • Strand: overlooked โ€” a Bell Labs name behind every phone call

Steps

1

The problem as it stood

Transatlantic telephone capacity was scarce and expensive. Sending the speech waveform itself requires the full audio bandwidth, and it was not obvious that anything less could carry intelligible speech.

2

What he actually did

Dudley separated speech into what the vocal tract is doing โ€” a slowly changing filter shape โ€” and what the vocal cords are doing โ€” a buzz at some pitch, or noise. Transmit a description of the filter plus the excitation, and rebuild the speech at the other end from a locally generated source.

3

How it worked

The analyser measures energy in each of a set of bands and detects pitch. Those control signals vary slowly, so they need far less bandwidth than the waveform. At the receiver, a buzz or hiss source is shaped by filters following those measurements. The insight โ€” model the source and the filter separately โ€” underlies linear predictive coding and every mobile phone codec.

4

What it made possible

Low-bandwidth speech transmission, and SIGSALY, the encrypted voice link used between Roosevelt and Churchill. Musically, the vocoder became a defining sound of electronic pop. The source-filter model remains the foundation of speech technology.

5

What happened to him

He remained at Bell Labs. The Voder required a trained operator playing a keyboard, wrist bar and foot pedal, and the World's Fair demonstrators were young women who had trained for months to make it speak.

6

Where the credit landed

Dudley is credited within speech technology and unknown outside it, where the vocoder is thought of as a music effect. The work is Bell Labs work and the patents were the company's. The intellectual line from the Vocoder through Itakura's PARCOR and Atal's linear predictive coding to every mobile phone is direct and almost never traced in public accounts.

Pro Tips

  • Source-filter separation is the founding idea of speech coding.
  • Control signals change slowly, so they need far less bandwidth than the waveform.
  • The vocoder's musical fame has obscured its role in telephony.

What You'll Learn

Speech coding works because a voice is a buzz through a changing tube.

Why speech compresses so much better than music

Speech is produced by a known mechanism: an excitation source driving a resonant tract that changes shape relatively slowly. Because the physical model is known, a codec can transmit model parameters instead of a waveform, which is why phone-quality speech survives at a few kilobits per second while music needs far more. Music has no equivalent constraining model, so audio codecs must fall back on perceptual masking instead.

Where This Fits

This guide covers one specific part of the history of media technology. The wider picture โ€” how each link in the chain from capture through transmission to display was actually built, who built it, and why the credit so often landed somewhere else โ€” is in A History of Broadcast Technology: The Chain From Capture to Screen, which frames the discipline as a whole and links out to the detailed guides underneath it, including this one. If you are starting from scratch rather than solving a specific problem, read that first and come back here.

FAQ

Q: What is a vocoder?
A: A system that analyses speech into the energy in several frequency bands plus a pitch or noise indication, transmits only those slowly changing parameters, and resynthesises speech from them at the far end.

Q: Is the musical vocoder the same thing?
A: Yes, used differently. The musical effect substitutes a musical signal for the synthesiser's internal source, so an instrument is shaped by the filter pattern of a voice.