IP Library Granted Patent US 7,797,157
Granted Patent B2
US 7,797,157 · App. 11/032,415 · Granted Sep 14, 2010

Automatic speech recognition channel normalization based on measured statistics from initial portions of speech utterances

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,797,157
App. No.
11/032,415
Granted
Sep 14, 2010
Kind
B2
Abstract

Channel normalization for automatic speech recognition is provided. Statistics are measured from an initial portion of a speech utterance. Feature normalization parameters are estimated based on the measured statistics and a statistically derived mapping relating measured statistics and feature normalization parameters. In some examples, the measured statistics comprise measures of an energy from the initial portion of the speech utterance. In some examples, measures of the energy comprise extreme values of the energy.

Claims (44)

1. A computer-implemented method comprising:

forming a statistically derived mapping based on measured statistics from speech utterances received during off-line processing and feature normalization parameters associated with the speech utterances, wherein each measured statistic is based on portion of a single speech utterance and each feature normalization parameter is based on multiple speech utterances, wherein the speech utterances received during off-line processing are received from a speech database that includes utterances of multiple speakers in multiple acoustic environments and wherein forming the statistically derived mapping comprises determining weights by statistical regression that relate the measured statistics from the speech utterances received during off-line processing to the associated feature normalization parameters;

measuring statistics of energy, from an initial portion of a speech utterance received during on-line processing, where the initial portion of the speech utterance includes a limited number of speech frames having cepstral features and estimated to comprise a minimum and maximum energy to provide estimates of signal and noise levels; and

estimating feature normalization parameters for the speech utterance received during on-line processing based on the measured statistics of the initial portion of a speech utterance by linearly mapping the cepstral features and the signal and noise levels using the statistically derived mapping weights.

2. The method of claim 1 wherein forming the statistically derived mapping comprises:

accepting a plurality of utterances;

measuring statistics from a portion of each of the plurality of utterances; and

forming the statistically derived mapping based on a relationship between inputs that each comprise measured statistics from a portion of a single one of the plurality of utterances and outputs that each comprise feature normalization parameters calculated from multiple ones of the plurality of utterances.

3. The method of claim 2 wherein the portion of each of the plurality of utterances comprises an initial portion of each of the utterances.

4. The method of claim 2 wherein the portion of each of the plurality of utterances comprises an entire portion of each of the utterances.

5. The method of claim 2 wherein the feature normalization parameters corresponding to the plurality of utterances comprise means and variances over time of the plurality of utterances.

6. The method of claim 1 , wherein a first of the measures of the energy consists of a maximum of the energy over an interval of time, and a second of the measures of the energy consists of a minimum of the energy over the interval of time.

7. The method of claim 6 , wherein the interval of time includes at least one speech frame before the utterance has started and includes at least one speech frame within the first vowel.

8. The method of claim 1 , wherein the estimated number of speech frames is determined based on an occurrence of a first vowel spoken in the utterance.

9. The method of claim 1 , wherein measuring statistics includes measuring from an initial portion including a period of silence having close to the minimum energy and a period of speech having close to the maximum energy.

10. The method of claim 1 , further comprising delaying about 100 ms to about 200 ms between measuring a signal to noise ratio (SNR) and applying the normalization parameters to map the cepstral features and the signal and noise levels.

11. The method of claim 1 , further comprising responding quickly to a speech onset and eliminating the need to explicitly detect the time of speech onset by determining minimum energy from the first couple of frames.

12. The method of claim 11 , further comprising determining minimum energy before the speech utterance received during on-line processing has started.

13. The method of claim 1 , wherein the statistically derived mapping is a linear map.

14. The method of claim 13 , wherein a variance of a cepstral feature i at time t is calculated by σ[i,t]=a i+1 (S[t]−N[t])+b +1 ;

where a i and b i are weights of the functional map and S[t] and N[t] are estimates for the signal level and noise level, respectively.

15. A non-transitory computer-readable medium having instructions stored thereon for processing data information, such that the instructions, when executed by a processing device, enable the processing device to perform the operations of:

forming a statistically derived mapping based on measured statistics from speech utterances comprising speech energy and noise and received during off-line processing and feature normalization parameters associated with the speech utterances, wherein each measured statistic is based on portion of a single speech utterance and each feature normalization parameter is based on multiple speech utterances, wherein the speech utterances received during off-line processing are received from a speech database that includes utterances of multiple speakers in multiple acoustic environments and wherein forming the statistically derived mapping comprises determining weights by statistical regression that relate the measured statistics from the speech utterances received during off-line processing to the associated feature normalization parameters;

measuring statistics of energy, from an initial portion of a speech utterance received during on-line processing, where the initial portion of the speech utterance includes a plurality of speech frames estimated to comprise a minimum and maximum energy in the speech utterance; and

estimating feature normalization parameters for the speech utterance received during on-line processing based on the measured statistics of the initial portion and the statistically derived mapping weights.

16. The non-transitory computer-readable medium of claim 15 wherein the estimated number of speech frames is determined based on an occurrence of a first vowel spoken in the utterance.

17. The non-transitory computer-readable medium of claim 15 wherein forming the statistically derived mapping comprises:

accepting a plurality of utterances;

measuring statistics from a portion of each of the plurality of utterances; and

forming the statistically derived mapping based on a relationship between inputs that each comprise measured statistics from a portion of a single of the plurality of utterances and outputs that each comprise feature normalization parameters calculated from multiple of the plurality of utterances.

18. The non-transitory computer-readable medium of claim 17 wherein the portion of each of the plurality of utterances comprises an initial portion of each of the utterances.

19. The non-transitory computer-readable medium of claim 17 wherein the portion of each of the plurality of utterances comprises an entire portion of each of the utterances.

20. A system comprising;

a processor and computer-readable medium associated with modules performing operations executed by the processor and further comprising:

a regression module configured to perform the operations of forming a statistically derived mapping based on measured statistics from speech utterances received during off-line processing and feature normalization parameters associated with the speech utterances, wherein each measured statistic is based on portion of a single speech utterance and each feature normalization parameter is based on multiple speech utterances, wherein the speech utterances received during off-line processing are received from a speech database that includes utterances of multiple speakers in multiple acoustic environments and wherein forming the statistically derived mapping comprises determining weights by statistical regression that relate the measured statistics from the speech utterances received during off-line processing to the associated feature normalization parameters;

an initial processing module configured to perform operations of measuring statistics of energy, from an initial portion of a speech utterance received during on-line processing, where the initial portion of the speech utterance includes a limited number of speech frames based on initial frames before an utterance and an occurrence of a first vowel spoken in the utterance and estimated to comprise a minimum and maximum energy, respectively; and

a mapping module configured to perform operations of estimating feature normalization parameters for the speech utterance received during on-line processing based on the measured statistics of the initial portion and the statistically derived mapping weights.

21. The system of claim 20

wherein the regression module is configured

to accept a plurality of utterances;

measure statistics from a portion of each of the plurality of utterances; and

form the statistically derived mapping based on a relationship between inputs that each comprise measured statistics from a portion of a single of the plurality of utterances and outputs that each comprise feature normalization parameters calculated from multiple of the plurality of utterances.

22. The system of claim 21 wherein the portion of each of the plurality of utterances comprises an initial portion of each of the utterances.

23. The system of claim 21 wherein the portion of each of the plurality of utterances comprises an entire portion of each of the utterances.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
MERGER Recorded Sep 13, 2012
From: VOICE SIGNAL TECHNOLOGIES, INC.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028952/0277 →