Speech Processing Basics

Vocal tract modeling, LPC.

Darshan N
Updated: 19 March 2026
12 min read

Speech processing is one of the most impactful application domains of Digital Signal Processing. It forms the backbone of technologies such as voice assistants, automatic speech recognition, text-to-speech synthesis, and voice compression in mobile communication. Understanding how speech is produced, modeled, and analyzed mathematically is essential for any ECE student working toward GATE or a career in embedded systems, telecom, or AI hardware.

Speech Production and LPC Model OverviewExcitationSource(Glottis / Noise)Vocal Tract FilterH(z)(All-pole AR Model)Lip RadiationFilter(High-pass ~6dB/oct)SpeechOutput s(n)(Sampled 8kHz+)LPC Analysis PipelineFrameBlockingWindowing(Hamming)AutocorrelationR(k) computationLevinson-DurbinLPC coefficients a_kVoiced: Periodic pulse excitation (pitch period T0)Unvoiced: White noise excitationTransfer function: H(z) = G / (1 - sum a_k z^-k) for k=1 to p
Figure 1: Source-filter model of speech production and the LPC analysis pipeline used in speech coding.

Core Concept: Source-Filter Model of Speech

Human speech is produced when air from the lungs excites the vocal tract. The key insight exploited in DSP is that this process can be modeled as a source driving a filter. The source is either a periodic pulse train (for voiced sounds like vowels) or a random noise signal (for unvoiced sounds like fricatives). The filter represents the resonant characteristics of the vocal tract cavity, which shapes the spectrum of the output sound.

This is called the source-filter model. The vocal tract is modeled as a time-varying all-pole (AR) filter. Its transfer function is written as H(z) = G / A(z), where A(z) = 1 - sum of a_k * z^-k for k from 1 to p, and p is the model order. The peaks of |H(e^jw)| correspond to formant frequencies, which are the resonant frequencies of the vocal tract and determine vowel identity.

The fundamental frequency (F0), also called pitch, is the rate at which the vocal cords vibrate. For a typical adult male, F0 is around 85-180 Hz. For adult female voices, it is around 165-255 Hz. Pitch detection and voiced/unvoiced decision are fundamental preprocessing tasks in speech analysis.

Linear Predictive Coding (LPC)

Linear Predictive Coding is the most widely used technique in speech analysis and compression. The core idea is that each speech sample can be predicted as a linear combination of past samples. If the prediction is good, the prediction error (called the residual) carries far less energy than the original signal, enabling efficient compression.

The LPC predictor equation is: s(n) = sum of a_k * s(n-k) for k from 1 to p. The prediction error signal e(n) = s(n) - s_hat(n). In the Z-domain, the error filter is A(z) = 1 - sum a_k z^-k, and the synthesis filter is H(z) = 1/A(z). The predictor coefficients a_k are computed by minimizing the mean squared prediction error, which leads to the Yule-Walker equations, solved efficiently using the Levinson-Durbin algorithm.

The autocorrelation method frames speech into short overlapping segments (typically 20-30 ms) using a windowing function such as the Hamming window. Each frame is assumed quasi-stationary, meaning the statistical properties of speech do not change significantly within the frame. The autocorrelation sequence R(k) = sum of s(n) * s(n+k) is computed per frame, and the Levinson-Durbin recursion solves for the p LPC coefficients with O(p^2) complexity.

Mathematical Expression

The LPC synthesis model has transfer function H(z) = G / (1 - sum_{k=1}^{p} a_k z^{-k}), where G is the gain of the excitation. The spectral envelope of speech is approximated by |H(e^{jwT})|. Formant frequencies F_i correspond to the angles of the poles of H(z) on the z-plane: F_i = (angle of pole i) / (2 * pi * T_s), where T_s is the sampling period. For 8 kHz sampling, T_s = 125 microseconds.

The LPC order p determines how many spectral peaks can be modeled. A common rule of thumb used in speech coding is p = f_s / 1000 + 2, where f_s is the sampling rate in Hz. For telephone-quality speech at 8 kHz, p = 10 is standard. This captures roughly 5 formant pairs (each pole pair models one formant).

Practical Understanding

LPC is the core algorithm behind several widely deployed voice codecs. The Code-Excited Linear Prediction (CELP) codec family, used in GSM and 3G mobile networks, is based on LPC analysis. In CELP, the excitation signal is selected from a codebook rather than being directly transmitted, dramatically reducing the bit rate to 4-13 kbps while maintaining intelligibility.

Another important concept is the cepstrum, which separates the excitation and vocal tract contributions in the log spectrum domain. The mel-frequency cepstral coefficients (MFCCs) derived from cepstral analysis form the standard feature set in automatic speech recognition. Understanding LPC is a prerequisite to understanding MFCCs and speech recognition front ends.

Example
Given:
Speech sampled at fs = 8000 Hz, LPC order p = 10
One frame of 20 ms duration is analyzed.

Why this formula applies:
LPC order rule-of-thumb p = fs/1000 + 2 gives p = 10 for 8 kHz.
Frame length = fs * frame_duration = 8000 * 0.020 = 160 samples.
The LPC model H(z) = G / (1 - sum a_k z^-k) approximates the vocal tract.
Formant frequency from pole angle: F_i = angle(pole_i) / (2 * pi * Ts)

Formula:
F_i = angle(z_i) * fs / (2 * pi)

Substitution:
Suppose a pole is located at angle = 0.785 radians (45 degrees) on unit circle.
F_i = 0.785 * 8000 / (2 * pi)

Calculation:
F_i = 6280 / 6.2832 = 1000 Hz

Final Answer:
Formant frequency F_i = 1000 Hz
This corresponds to the first formant F1 of a typical vowel sound.
Exam Tip: In GATE, LPC order p = fs/1000 + 2 is a standard formula. For 8 kHz speech, p = 10. Poles of H(z) inside the unit circle guarantee a stable all-pole synthesis filter. Voiced speech has a pitch period of roughly 5-20 ms at 8 kHz sampling.

Mechanism: Speech Processing Signal Flow

LPC Encoder and Decoder Block DiagramSpeech s(n)Input SignalPre-emphasisH(z)=1-0.97z^-1Framing &WindowingLPC AnalysisLevinson-DurbinCoefficientsa1..ap + Gain GDecoder SideV/UV Decision+ Pitch PeriodExcitationGenerator e(n)LPC SynthesisH(z)=G/A(z)De-emphasisInverse filters_hat(n)ReconstructedKey Parameters Transmitted- LPC coefficients a1 to ap (p=10 for 8kHz)- Gain G (excitation amplitude)- Voiced/Unvoiced flag (1 bit)- Pitch period T0 (if voiced)Formant Locations in Z-planeUnit circlepole (F1)pole (F2)
Figure 2: Complete LPC encoder-decoder mechanism showing the signal flow from input speech to reconstructed speech via LPC coefficients.

Mechanism Summary

  • Speech is modeled as an excitation signal passed through an all-pole vocal tract filter H(z) = G/A(z). The poles of H(z) represent formant frequencies.
  • LPC analysis computes the predictor coefficients a_k by minimizing mean squared prediction error. The Levinson-Durbin algorithm does this efficiently in O(p^2) operations per frame.
  • For voiced speech, the excitation is a periodic pulse train with pitch period T0. For unvoiced speech, white noise is used as the excitation source.
  • Pre-emphasis (filter 1 - 0.97 z^-1) is applied before analysis to boost high-frequency components, improving the spectral flatness and prediction accuracy.
  • LPC coefficients, gain, voiced/unvoiced flag, and pitch period are the minimal set of parameters required for speech synthesis at the decoder.

Quick Revision

  • Source-filter model: excitation (glottal pulse or noise) convolved with vocal tract all-pole filter H(z) produces speech.
  • LPC synthesis filter: H(z) = G / (1 - a_1 z^-1 - a_2 z^-2 - ... - a_p z^-p). Poles give formant frequencies.
  • LPC order rule: p = fs(kHz) + 2. At 8 kHz, p = 10 is standard for telephone speech.
  • Levinson-Durbin algorithm solves Yule-Walker equations in O(p^2). Key for efficient real-time implementation.
  • Framing: 20-30 ms frames assumed quasi-stationary. Hamming window reduces spectral leakage at frame boundaries.
  • Exam trap: LPC poles must be inside the unit circle for the synthesis filter to be stable. The all-pole model cannot represent spectral nulls (zeros), only peaks.
  • Applications: GSM codec (RPE-LTP), CELP in 3G, MFCCs in speech recognition, formant synthesis in TTS systems.

Speech Processing Quiz

Test your knowledge of vocal tract modeling and linear predictive coding fundamentals.

Question 1 of 3

Q1.In Linear Predictive Coding (LPC), the vocal tract is modeled as an all-pole filter of the form H(z) = G / A(z). The prediction order p for telephone-quality speech is typically: