IP Library Granted Patent US 9,058,810
Granted Patent B2
US 9,058,810 · App. 13/429,946 · Granted Jun 16, 2015

System and method of performing user-specific automatic speech recognition

Inventors: Bojana Gajic (Trondheim, NO); Shrikanth Sambasivan Narayanan (Los Angeles, CA); Sarangarajan Parthasarathy (New Providence, NJ); Richard Cameron Rose (Watchung, NJ); Aaron Edward Rosenberg (Berkeley Heights, NJ)
Assignee: AT&T Intellectual Property II, L.P.
G10L15/07G10L15/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,058,810
App. No.
13/429,946
Granted
Jun 16, 2015
Kind
B2
Abstract

Speech recognition models are dynamically re-configurable based on user information, application information, background information such as background noise and transducer information such as transducer response characteristics to provide users with alternate input modes to keyboard text entry. Word recognition lattices are generated for each data field of an application and dynamically concatenated into a single word recognition lattice. A language model is applied to the concatenated word recognition lattice to determine the relationships between the word recognition lattices and repeated until the generated word recognition lattices are acceptable or differ from a predetermined value only by a threshold amount. These techniques of dynamic re-configurable speech recognition provide for deployment of speech recognition on small devices such as mobile phones and personal digital assistants as well environments such as office, home or vehicle while maintaining the accuracy of the speech recognition.

Claims (52)

1. A method comprising:

receiving a voice request from a speaker;

sampling an entirety of the voice request at periodic intervals, to yield samples;

identifying a background based on the samples;

modifying a transducer model based on the background generating a modified transducer model and adapting a plurality of speaker independent speech recognition models with the modified transducer model to yield a plurality of modified language models;

receiving a data field selection by the speaker, the data field selection comprising a first data field selection and a second data field selection; and

applying, via a processor, the plurality of modified language models to the voice request for speech recognition based on the data field selection and the background, wherein the speech recognition uses a first language model from the plurality of language models for the first data field selection and a second language model from the plurality of language models for the second data field selection, and wherein the first language model is distinct from the second language model.

2. The method of claim 1 , further comprising:

generating a response to the voice request.

3. The method of claim 2 , wherein the response is one of a voice response, a text response, a tactile response, and a Braille response.

4. The method of claim 1 , wherein the plurality of language models comprises a background model and a transducer model.

5. The method of claim 1 , further comprising:

determining a confidence score for the voice request from the applying of the one of the plurality of modified language models to the voice request.

6. The method of claim 5 , further comprising:

restarting speech recognition when the confidence score is below a threshold.

7. The method of claim 1 , further comprising:

adjusting the periodic intervals based on the samples, to yield adjusted periodic intervals; and

wherein the modifying of the plurality of speaker-independent speech recognition models occurs at the adjusted periodic intervals.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed on the processor, perform operations comprising:

receiving a voice request from a speaker;

sampling an entirety of the voice request at periodic intervals, to yield samples;

identifying a background based on the samples;

modifying a transducer model based on the background generating a modified transducer model and adapting a plurality of speaker independent speech recognition models with the modified transducer model to yield a plurality of modified language models;

receiving a data field selection by the speaker, the data field selection comprising a first data field selection and a second data field selection; and

applying, via a processor, the plurality of modified language models to the voice request for speech recognition based on the data field selection and the background, wherein the speech recognition uses a first language model from the plurality of language models for the first data field selection and a second language model from the plurality of language models for the second data field selection, and wherein the first language model is distinct from the second language model.

9. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in the operations further comprising:

generating a response to the voice request.

10. The system of claim 9 , wherein the response is one of a voice response, a text response, a tactile response, and a Braille response.

11. The system of claim 8 , wherein the plurality of language models comprises a background model and a transducer model.

12. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in the operations further comprising:

determining a confidence score for the voice request from the applying of the one of the plurality of modified language models to the voice request.

13. The system of claim 12 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in the operations further comprising:

restarting speech recognition when the confidence score is below a threshold.

14. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in the operations further comprising:

adjusting the periodic intervals based on the samples.

15. A computer readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving a voice request from a speaker;

sampling an entirety of the voice request at periodic intervals, to yield samples;

identifying a background based on the samples;

modifying a transducer model based on the background generating a modified transducer model and adapting a plurality of speaker independent speech recognition models with the modified transducer model to yield a plurality of modified language models;

receiving a data field selection by the speaker, the data field selection comprising a first data field selection and a second data field selection; and

applying, via a processor, the plurality of modified language models to the voice request for speech recognition based on the data field selection and the background, wherein the speech recognition uses a first language model from the plurality of language models for the first data field selection and a second language model from the plurality of language models for the second data field selection, and wherein the first language model is distinct from the second language model.

16. The computer readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, result in the operations further comprising:

generating a response to the voice request.

17. The computer readable storage device of claim 16 , wherein the response is one of a voice response, a text response, a tactile response, and a Braille response.

18. The computer readable storage device of claim 15 , wherein the plurality of language models comprises a background model and a transducer model.

19. The computer readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, result in the operations further comprising:

determining a confidence score for the voice request from the applying of the one of the plurality of language models to the voice request.

20. The computer readable storage device of claim 15 , having additional instructions stored which, when executed by the computing device, result in the operations further comprising:

restarting speech recognition when the confidence score is below a threshold.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2017
From: AT&T INTELLECTUAL PROPERTY II, L.P.
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 041512/0608 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T CORP.
To: AT&T PROPERTIES, LLC
Reel/Frame 038275/0041 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2016
From: AT&T PROPERTIES, LLC
To: AT&T INTELLECTUAL PROPERTY II, L.P.
Reel/Frame 038275/0130 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2012
From: GAJIC, BOJANA; NARAYANAN, SHRIKANTH SAMBASIVAN; PARTHASARATHY, SARANGARAJAN; ROSE, RICHARD CAMERON; ROSENBERG, AARON EDWARD
To: AT&T CORP.
Reel/Frame 027932/0472 →
Continuity (5)
Continuation 12207175 · Sep 9, 2008
Continuation 11685456 · Mar 13, 2007
Continuation 10091689 · Mar 6, 2002
Provisional Application 60277231 · Mar 20, 2001
Related Publication 20120185237A1 · Jul 19, 2012