IP Library Granted Patent US 7,319,959
Granted Patent B1
US 7,319,959 · App. 10/439,284 · Granted Jan 15, 2008

Multi-source phoneme classification for noise-robust automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,319,959
App. No.
10/439,284
Granted
Jan 15, 2008
Kind
B1
Abstract

A system and method are disclosed for processing an audio signal including separating the audio signal into a plurality of streams which group sounds from a same source prior to classification and analyzing each separate stream to determine phoneme-level classification. One or more words of the audio signal may then be outputted.

Claims (19)

1. A method of processing an audio signal comprising:

computing 600 spectral values on a logarithmic frequency scale from the audio signal;

separating the 600 spectral values into a plurality of streams which group sounds from a same source prior to classification;

analyzing each separated stream to determine phoneme-level classification; and

outputting one or more words of the audio signal.

2. The method of processing an audio signal as recited in claim 1 wherein phoneme-level classification accuracy is enhanced by providing as an input to a classifier a spectral envelope.

3. The method of processing an audio signal as recited in claim 1 wherein phoneme-level classification accuracy is enhanced by providing as an input to a classifier detected transients.

4. The method of processing an audio signal as recited in claim 1 wherein phoneme-level classification accuracy is enhanced by providing as an input to a classifier pitch and voicing information.

5. The method of processing an audio signal as recited in claim 1 further comprising normalizing for speaker characteristics prior to classification.

6. The method of processing an audio signal as recited in claim 1 further comprising performing noise-threshold tracking.

7. The method of processing an audio signal as recited in claim 6 further comprising adjusting gain and setting noise floor reference levels based on the noise-threshold tracking.

8. The method of processing an audio signal as recited in claim 1 further comprising training with a full phoneme target set.

9. The method of processing an audio signal as recited in claim 1 further comprising incorporating a model of syllabic stress.

10. The method of processing an audio signal as recited in claim 1 wherein the spectral values are computed with a 6 microsecond resolution post-interpolation.

11. The method of processing an audio signal as recited in claim 1 wherein the spectral values are updated every 22 microseconds.

12. The method of processing an audio signal as recited in claim 1 further comprising using the output for automatic voice-dialing for phones.

13. The method of processing an audio signal as recited in claim 1 further comprising using the output for automatic command of a system or device.

14. The method of processing an audio signal as recited in claim 1 further comprising using the output as an interface to a device.

15. The method of processing an audio signal as recited in claim 1 wherein the output comprises a meeting transcription.

Assignments (4)
CHANGE OF NAME Recorded Feb 25, 2016
From: AUDIENCE, INC.
To: AUDIENCE LLC
Reel/Frame 037927/0424 →
MERGER Recorded Feb 25, 2016
From: AUDIENCE LLC
To: KNOWLES ELECTRONICS, LLC
Reel/Frame 037927/0435 →
SECURITY INTEREST Recorded Oct 21, 2003
From: AUDIENCE, INC.
To: VULCON VENTURES INC.
Reel/Frame 014615/0160 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2003
From: WATTS, LLOYD
To: AUDIENCE, INC.
Reel/Frame 014388/0016 →