IP Library Granted Patent US 9,734,819
Granted Patent B2
US 9,734,819 · App. 13/772,373 · Granted Aug 15, 2017

Recognizing accented speech

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,734,819
App. No.
13/772,373
Granted
Aug 15, 2017
Kind
B2
Abstract

Techniques ( 300, 400, 500 ) and apparatuses ( 100, 200, 700 ) for recognizing accented speech are described. In some embodiments, an accent module recognizes accented speech using an accent library based on device data, uses different speech recognition correction levels based on an application field into which recognized words are set to be provided, or updates an accent library based on corrections made to incorrectly recognized speech.

Claims (49)

1. A computer-implemented method comprising:

receiving, from a user, an utterance that was spoken while focus is set on a field of a form;

determining, from among one or more different field types, a field type associated with the field;

determining, from among different, predefined levels of speech recognition accuracy and from among different, predefined levels of speech recognition latency, a predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type;

selecting, based at least on the predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type, (i) one or more accent libraries that each include phonemes for different pronunciations for words of a language and (ii) a level of correction for a speech recognition system to apply to a transcription;

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system; and

providing the transcription of the utterance in the field of the form.

2. The method of claim 1 , comprising:

selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance.

3. The method of claim 2 , wherein selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance comprises:

identifying countries of addresses stored in an address book; and

selecting the one or more accent libraries based on the countries and on a current location of the computing device.

4. The method of claim 1 , wherein the one or more accent libraries are associated with a default language of the computing device.

5. The method of claim 1 , wherein the speech recognition system generates the transcription of the utterance without accessing a linguistic library.

6. The method of claim 1 , wherein obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system comprises:

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system using the one or more accent libraries.

7. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, from a user, an utterance that was spoken while focus is set on a field of a form;

determining, from among one or more different field types, a field type associated with the field;

determining, from among different, predefined levels of speech recognition accuracy and from among different, predefined levels of speech recognition latency, a predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type;

selecting, based at least on the predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type, (i) one or more accent libraries that each include phonemes for different pronunciations for words of a language and (ii) a level of correction for a speech recognition system to apply to a transcription;

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system; and

providing the transcription of the utterance in the field of the form.

8. The system of claim 7 , wherein the operations further comprise:

selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance.

9. The system of claim 8 , wherein selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance comprises:

identifying countries of addresses stored in an address book; and

selecting the one or more accent libraries based on the countries and on a current location of the computing device.

10. The system of claim 7 , wherein the one or more accent libraries are associated with a default language of the computing device.

11. The system of claim 7 , wherein the speech recognition system generates the transcription of the utterance without accessing a linguistic library.

12. The system of claim 7 , wherein obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system comprises:

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system using the one or more accent libraries.

13. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

receiving, from a user, an utterance that was spoken while focus is set on a field of a form;

determining, from among one or more different field types, a field type associated with the field;

determining, from among different, predefined levels of speech recognition accuracy and from among different, predefined levels of speech recognition latency, a predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type;

selecting, based at least on the predefined level of speech recognition accuracy and a predefined level of speech recognition latency that are indicated as acceptable for the field type, (i) one or more accent libraries that each include phonemes for different pronunciations for words of a language and (ii) a level of correction for a speech recognition system to apply to a transcription;

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system; and

providing the transcription of the utterance in the field of the form.

14. The medium of claim 13 , comprising:

selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance.

15. The medium of claim 14 , wherein selecting the one or more accent libraries based on personal data that is associated with the user and is stored on a computing device that receives the utterance comprises:

identifying countries of addresses stored in an address book; and

selecting the one or more accent libraries based on the countries and on a current location of the computing device.

16. The medium of claim 13 , wherein the one or more accent libraries are associated with a default language of the computing device.

17. The medium of claim 13 , wherein the speech recognition system generates the transcription of the utterance without accessing a linguistic library.

18. The medium of claim 13 , wherein obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system comprises:

obtaining, from the speech recognition system, the transcription of the utterance that is generated by the speech recognition system using the one or more accent libraries.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034244/0014 →