IP Library Granted Patent US 7,680,662
Granted Patent B2
US 7,680,662 · App. 11/129,254 · Granted Mar 16, 2010

Systems and methods for implementing segmentation in speech recognition systems

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,680,662
App. No.
11/129,254
Granted
Mar 16, 2010
Kind
B2
Abstract

A speech recognition system ( 105 ) includes an acoustic front end ( 115 ) and a processing unit ( 125 ). The acoustic front end ( 115 ) receives frames of acoustic data and determines cepstral coefficients for each of the received frames. The processing unit ( 125 ) determines a number of peaks in the cepstral coefficients for each of the received frames of acoustic data and compares the peaks in the cepstral coefficients of a first one of the received frames with the peaks in the cepstral coefficients of at least a second one of the received frames. The processing unit ( 125 ) then segments the received frames of acoustic data based on the comparison.

Claims (57)

1. A method of segmenting acoustic data for use in a speech recognition process, comprising:

receiving frames of acoustic data;

determining cepstral coefficients for each of the received frames of acoustic data;

determining how many local maxima there are in a plot of the cepstral coefficients for each received frame of acoustic data; and

segmenting the received frames of acoustic data based on results of said local maxima determining.

2. The method of claim 1 , further comprising:

comparing said how many local maxima there are in the plot of the cepstral coefficients of a first one of the received frames with said how many local maxima there are in the plot of the cepstral coefficients of at least a second one of the received frames, wherein the segmenting of the received frames of acoustic data is further based on the comparing.

3. A computerized speech recognition system, comprising:

an acoustic front end configured to:

receive frames of acoustic data,

determine cepstral coefficients for each of the received frames of acoustic data; and

a processing unit configured to:

determine how many local maxima there are in a plot of the cepstral coefficients for each of the received frames of acoustic data,

compare said how many local maxima there are in the plot of the cepstral coefficients of a first one of the received frames with said how many local maxima there are in the plot of the cepstral coefficients of at least a second one of the received frames, and

segment the received frames of acoustic data based on the comparison.

4. A computer-readable medium excluding carrier waves, said medium containing instructions for controlling at least one processing unit to perform a method of segmenting acoustic data for use in a speech recognition process, the method comprising:

receiving cepstral coefficients corresponding to frames of acoustic data;

segmenting the frames of acoustic data based on the received cepstral coefficients; and

determining how many local maxima there are in a plot of the cepstral coefficients corresponding to each of the frames of acoustic data, wherein the segmenting of the frames of acoustic data is further based on the determining.

5. The computer-readable medium of claim 4 , the method further comprising:

comparing how many local maxima there are in a plot of the cepstral coefficients of a first one of the frames with how many local maxima there are in a plot of the cepstral coefficients of at least a second one of the frames, wherein the segmenting of the frames of acoustic data is further based on the comparing.

6. A method of recognizing patterns in acoustic data, comprising:

receiving frames of acoustic data;

determining cepstral coefficients corresponding to the received frames of acoustic data;

determining a number of peaks in a plot of the cepstral coefficients for each received frame of acoustic data;

determining at least one weighting parameter based on the determined number of peaks in the plot of the cepstral coefficients; and

recognizing patterns in the received frames of acoustic data using the at least one weighting parameter.

7. The method of claim 6 , further comprising:

determining, based on the frames of acoustic data, recognition hypothesis scores using a Hidden Markov Model.

8. The method of claim 7 , further comprising:

modifying the recognition hypothesis scores based on the at least one weighting parameter.

9. The method of claim 8 , wherein the recognizing patterns in the frames of acoustic data further uses the modified recognition hypothesis scores.

10. A speech recognition system, comprising:

an acoustic front end configured to receive frames of acoustic data and determine a number of peaks in a plot of cepstral coefficients for each of the received frames,

a processing unit configured to:

determine segmentation information corresponding to the determined number of peaks,

determine at least one weighting parameter based on the determined segmentation information, and

recognize patterns in the received frames of acoustic data using the at least one weighting parameter.

11. The system of claim 10 , the processing unit further configured to:

determine, based on the number of peaks in the plot of the cepstral coefficients in each of the received frames of acoustic data, recognition hypothesis scores using a Hidden Markov Model.

12. The system of claim 11 , the processing unit further configured to:

modify the recognition hypothesis scores based on the at least one weighting parameter.

13. The system of claim 11 , the processing unit further configured to:

recognize patterns in the received frames of acoustic data further using the modified recognition hypothesis scores.

14. A computer readable medium excluding carrier waves, said medium having encoded thereon a data structure comprising:

cepstral coefficient data corresponding to each frame of a plurality of frames of acoustic data, the cepstral coefficient data including how many local maxima there are in a plot of cepstral coefficients corresponding to each frame of acoustic data; and

segmentation data indicating segmentation of the frames of acoustic data based on said how many local maxima there are.

15. An acoustic recognition system, comprising:

an acoustic front end configured to:

receive frames of acoustic data and determine how many local maxima there are in a plot of cepstral coefficients for each of the received frames;

a processing unit configured to generate, based on the determined how many local maxima there are, end frame numbers for each phoneme or Hidden Markov Model (HMM) state contained in the received frames; and

a trainer/HMM decoder configured to use the generated end frame numbers for determining weighted scores that can be used for recognition of acoustic events contained in the received frames of acoustic data.

16. An acoustic recognition system, comprising:

an acoustic front end configured to receive frames of acoustic data and determine a number of peaks of cepstral coefficients for each of the received frames; and

a processing unit configured to:

generate, based on the determined number of peaks, end frame numbers for each phoneme or Hidden Markov Model (HMM) state contained in the received frames of acoustic data, and

determine weighted scores based on the generated end frame numbers that can be used for recognition of acoustic events contained in the received frames of acoustic data.

Assignments (5)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2014
From: BBNT SOLUTIONS LLC
To: BBNT SOLUTIONS LLC; VERIZON CORPORATE SERVICES GROUP INC
Reel/Frame 033524/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2014
From: VERIZON CORPORATE SERVICES GROUP INC.
To: VERIZON PATENT AND LICENSING INC.
Reel/Frame 033421/0403 →
CHANGE OF NAME Recorded Jun 11, 2010
From: BBN TECHNOLOGIES CORP.
To: RAYTHEON BBN TECHNOLOGIES CORP.
Reel/Frame 024523/0625 →
RELEASE OF SECURITY INTEREST Recorded Oct 27, 2009
From: BANK OF AMERICA, N.A. (SUCCESSOR BY MERGER TO FLEET NATIONAL BANK)
To: BBN TECHNOLOGIES CORP. (AS SUCCESSOR BY MERGER TO BBNT SOLUTIONS LLC)
Reel/Frame 023427/0436 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT SUPPLEMENT Recorded Dec 4, 2008
From: BBN TECHNOLOGIES CORP.
To: BANK OF AMERICA, N.A.
Reel/Frame 021926/0017 →