IP Library Granted Patent US 7,062,434
Granted Patent B2
US 7,062,434 · App. 10/242,465 · Granted Jun 13, 2006

Compressed domain voice activity detector

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,062,434
App. No.
10/242,465
Granted
Jun 13, 2006
Kind
B2
Abstract

The system and method of the present invention comprises a compressed domain voice activity detector that detects the presence or absence of voice activity in a digital input signal. The method includes converting a digital input signal into parametric data. The parametric data is subsequently analyzed, and then compared against a background noise threshold to determine if voice activity is present.

Claims (34)

1. A method for determining if a plurality of parametric model data of a compressed bit stream contain voice data, comprising:

computing normalized signal levels for a plurality of frequency sub-bands of said compressed bit stream using at least one of said parametric model data;

determining a stability level for said compressed bit stream using at least one of said parametric model data;

estimating a background noise level for said frequency sub-bands based on at least one of said stability level and said normalized signal levels; and

identifying the presence of voice data in said compressed bit stream based on said estimation and said normalized signal levels.

2. The method according to claim 1 , wherein said parametric model data comprise at least one of:

short term filter coefficients;

overall frame gain;

voice cutoff level; and

pitch.

3. The method according to claim 1 , further comprising:

identifying periods of inactivity between identified voice data; and

removing said periods of inactivity from said compressed bit stream.

4. The method according to claim 2 , wherein said short term filter coefficients comprise Line Spectral Frequency form coefficients.

5. The method according to claim 2 , wherein said compressed bit stream is divided into frames, each frame having a corresponding plurality of parametric model data, and wherein said computing normalized signed levels comprises:

computing a spectral envelope of a frame based on said short term filter coefficients;

computing signal levels for said plurality of frequency sub-bands based on said spectral envelope;

calculating a frame gain based on said short term filter coefficients; and

normalizing said computed signal levels based on said overall frame gain and said frame gain based on said short term filter coefficients.

6. The method according to claim 1 , wherein said determining a stability level comprises determining a frequency level of said compressed bit stream above which no voice activity is expected to be present, based on at least one of said parametric model data.

7. The method according to claim 5 , wherein said estimating a background noise level comprises estimating and updating the background noise level present in each frame at each of said plurality of frequency sub-bands.

8. The method according to claim 1 , wherein said identifying the presence of voice data comprises:

deciding if a voice signal is present based on at least one of said background noise estimate and said normalized signal levels; and

indicating the presence of voice activity.

9. A method for determining if a plurality of parametric model data of a compressed bit stream contain voice data, said compressed bit stream being divided into a plurality of frames, each having a corresponding plurality of parametric model data, said method comprising:

computing normalized signal levels for a plurality of frequency sub-bands of said compressed bit stream using at least one of said parametric model data by computing a spectral envelope of a frame based on said short term filter coefficients, computing signal levels for said plurality of frequency sub-bands based on said spectral envelope, calculating a frame gain based on said short term filter coefficients, and normalizing said computed signal levels based on said overall frame gain and said frame gain based on said short term filter coefficients;

determining a stability level for said compressed bit stream using at least one of said parametric model data;

estimating a background noise level for said frequency sub-bands based on at least one of said stability level and said normalized signal levels; and,

identifying the presence of voice data in said compressed bit stream based on said estimation and said normalized signal levels.

10. A method for determining if a plurality of parametric model data of a compressed bit stream contain voice data, comprising:

computing normalized signal levels for a plurality of frequency sub-bands of said compressed bit stream using at least two of said parametric model data;

determining a stability level for said compressed bit stream using at least two of said parametric model data;

estimating a background noise level for said frequency sub-bands based on at least one of said stability level and said normalized signal levels; and

identifying the presence of voice data in said compressed bit stream based on said estimation and said normalized signal levels.