IP Library Granted Patent US 8,639,508
Granted Patent B2
US 8,639,508 · App. 13/026,670 · Granted Jan 28, 2014

User-specific confidence thresholds for speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,639,508
App. No.
13/026,670
Granted
Jan 28, 2014
Kind
B2
Abstract

A method of automatic speech recognition includes receiving an utterance from a user via a microphone that converts the utterance into a speech signal, pre-processing the speech signal using a processor to extract acoustic data from the received speech signal, and identifying at least one user-specific characteristic in response to the extracted acoustic data. The method also includes determining a user-specific confidence threshold responsive to the at least one user-specific characteristic, and using the user-specific confidence threshold to recognize the utterance received from the user and/or to assess confusability of the utterance with stored vocabulary.

Claims (39)

1. A method of automatic speech recognition, comprising the steps of:

(a) receiving an utterance from a user via a microphone that converts the utterance into a speech signal;

(b) pre-processing the speech signal using a processor to extract acoustic data from the received speech signal;

(c) identifying at least one user-specific characteristic in response to the extracted acoustic data, wherein the at least one user-specific characteristic comprises a plurality of confidence scores associated with failed attempts of the user to store a nametag; and

(d) determining a user-specific confidence threshold responsive to the at least one user-specific characteristic, wherein the determination is carried out by calculating an average of the plurality of confidence scores and setting the user-specific confidence threshold to a value greater than or equal to the calculated average.

2. The method of claim 1 , further comprising the step of:

(e) using the user-specific confidence threshold to recognize the utterance received from the user, wherein the user-specific confidence threshold is a recognition confidence threshold.

3. The method of claim 1 , further comprising the step of:

(e) using the user-specific confidence threshold to assess confusability of the utterance with stored vocabulary, wherein the user-specific confidence threshold is a confusability confidence threshold.

4. The method of claim 1 , wherein the step (d) determination is also carried out by first verifying that the plurality of confidence scores are within a predetermined range.

5. The method of claim 4 , wherein the predetermined range is plus or minus five percent.

6. The method of claim 1 , wherein the at least one user-specific characteristic includes at least one formant of the utterance.

7. The method of claim 6 , wherein the user-specific confidence threshold is determined using a multiple regression calculation including the at least one formant of the utterance and at least one formant coefficient developed from a plurality of development speakers.

8. The method of claim 1 , wherein the at least one user-specific characteristic includes pitch of the utterance.

9. The method of claim 8 , wherein the user-specific confidence threshold is determined using a multiple regression calculation including the pitch of the utterance and a pitch coefficient developed from a plurality of development speakers.

10. The method of claim 1 , wherein the at least one user-specific characteristic includes pitch and at least one formant of the utterance, and wherein the user-specific confidence threshold is determined using a multiple regression calculation including the pitch and the at least one formant of the utterance and a pitch coefficient and at least one formant coefficient developed from a plurality of development speakers.

11. A method of automatic speech recognition, comprising the steps of:

(a) receiving an utterance from a user via a microphone that converts the utterance into a speech signal;

(b) pre-processing the speech signal using a processor to extract acoustic data from the received speech signal;

(c) identifying at least one user-specific characteristic including pitch and at least one formant in response to the extracted acoustic data;

(d) determining a user-specific confidence threshold responsive to the identified at least one user-specific characteristics, wherein the determination comprises using a multiple regression calculation including the identified user-specific pitch and at least one formant and a pitch coefficient and at least one formant coefficient developed from a plurality of development speakers; and

(e) decoding the acoustic data based on the user-specific confidence threshold to produce a plurality of hypotheses for the received utterance, including calculating confidence scores for the hypotheses.

12. The method of claim 11 , further comprising the step of:

(f) post-processing the plurality of hypotheses, including using the user-specific confidence threshold to identify at least one hypothesis of the plurality of hypotheses as the received utterance, wherein the user-specific confidence threshold is a recognition confidence threshold.

13. The method of claim 11 , further comprising the step of:

(f) post-processing the plurality of hypotheses, including using the user-specific confidence threshold to assess confusability of the utterance with stored vocabulary, wherein the user-specific confidence threshold is a confusability confidence threshold.

14. The method of claim 11 , wherein the at least one user-specific characteristic of step (c) also includes a plurality of confidence scores associated with failed attempts of the user to store a nametag.

15. The method of claim 14 , wherein the step (d) determination is further carried out by calculating an average of the plurality of confidence scores, and setting the user-specific confidence threshold to a value that is relative to the calculated average.

16. The method of claim 15 , wherein the step (d) determination is also carried out by first verifying that the plurality of confidence scores are within a predetermined range.

17. A method of automatic speech recognition, comprising the steps of:

(a) receiving an utterance from a user via a microphone that converts the utterance into a speech signal;

(b) pre-processing the speech signal using a processor to extract acoustic data from the received speech signal;

(c) identifying at least one user-specific characteristic including a plurality of confidence scores associated with failed attempts of the user to store a nametag;

(d) determining a user-specific confidence threshold responsive to the at least one user-specific characteristic;

(e) decoding the acoustic data to produce a plurality of hypotheses for the received utterance, including calculating confidence scores for the hypotheses; and

(f) post-processing the plurality of hypotheses, including using the user-specific confidence threshold to assess confusability of the utterance with stored vocabulary, wherein the user-specific confidence threshold is a confusability confidence threshold.

18. The method of claim 17 , wherein the step (d) determination is carried out by calculating an average of the plurality of confidence scores and setting the user-specific confidence threshold to a value greater than or equal to the calculated average.

19. The method of claim 18 , wherein the at least one user-specific characteristic also includes pitch and at least one formant of the utterance, and wherein the user-specific confidence threshold is also determined using a multiple regression calculation including the pitch and the at least one formant of the utterance, and a pitch coefficient and at least one formant coefficient developed from a plurality of development speakers.

20. The method of claim 19 , wherein the step (d) determination is carried out by setting the user-specific confidence threshold to a value that is relative to the calculated average or the multiple regression calculated value.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2014
From: WILMINGTON TRUST COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 034183/0436 →
SECURITY AGREEMENT Recorded Jun 22, 2012
From: GENERAL MOTORS LLC
To: WILMINGTON TRUST COMPANY
Reel/Frame 028423/0432 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 19, 2011
From: ZHAO, XUFANG; TALWAR, GAURAV
To: GENERAL MOTORS LLC
Reel/Frame 025836/0512 →