IP Library › Granted Patent US 9,966,060
Granted Patent B2
US 9,966,060 · App. 15/445,863 · Granted May 8, 2018

System and method for user-specified pronunciation of words for speech synthesis and recognition

Inventors: Devang K. Naik (San Jose, CA); Thomas R. Gruber (Emerald Hills, CA); Liam Weiner (San Francisco, CA); Justin G. Binder (Oakland, CA); Charles Srisuwananukorn (San Francisco, CA); Gunnar Evermann (Boston, MA); Shaun Eric Williams (San Jose, CA); Hong Chen (San Jose, CA); Lia T. Napolitano (San Francisco, CA)
Assignee: Apple Inc.
G10L13/027G10L13/04G10L13/08G10L15/063G10L15/22G10L15/265G10L2015/0631G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,966,060
App. No.
15/445,863
Granted
May 8, 2018
Kind
B2
Abstract

The method is performed at an electronic device with one or more processors and memory storing one or more programs for execution by the one or more processors. A first speech input including at least one word is received. A first phonetic representation of the at least one word is determined, the first phonetic representation comprising a first set of phonemes selected from a speech recognition phonetic alphabet. The first set of phonemes is mapped to a second set of phonemes to generate a second phonetic representation, where the second set of phonemes is selected from a speech synthesis phonetic alphabet. The second phonetic representation is stored in association with a text string corresponding to the at least one word.

Claims (63)

1. A method for learning word pronunciations, comprising:

at an electronic device with one or more processors and memory storing one or more programs for execution by the one or more processors:

detecting an error in a speech based interaction with a digital assistant based on detecting a user input other than the speech based interaction;

in response to detecting the error, receiving a speech input from a user, the speech input including a pronunciation of one or more words;

determining, based on the pronunciation of the one or more words, a first set of phonemes from a speech recognition phonetic alphabet and a second set of phonemes from a speech synthesis phonetic alphabet;

updating one or more databases to include the first set of phonemes and the second set of phonemes in association with a text string corresponding to the one or more words; and

performing speech recognition or speech synthesis using the updated one or more databases.

2. The method of claim 1 , wherein the one or more words were received in a prior speech input provided by the user, and wherein the error is an error in speech recognition of the one or more words.

3. The method of claim 1 , wherein the one or more words were output in a speech output by the electronic device, and wherein the error is an error in speech synthesis of the one or more words.

4. The method of claim 1 , further comprising:

receiving the speech input including the one or more words;

performing speech recognition on the speech input to generate the text string corresponding to the one or more words;

determining a confidence metric of the text string; and

detecting the error based on a determination that the confidence metric does not meet a predetermined threshold.

5. The method of claim 1 , further comprising:

synthesizing a speech output including the one or more words; and

detecting the error based on an indication from the user that the one or more words were pronounced incorrectly.

6. The method of claim 1 , further comprising:

performing speech recognition on the speech input to generate the text string corresponding to the one or more words; and wherein updating the one or more databases comprises

updating a speech recognizer to associate the first set of phonemes with the text string.

7. The method of claim 6 , further comprising:

receiving a second speech input including the one or more words;

determining a third set of phonemes for the one or more words;

determining that the one or more words in the second speech input correspond to the text string based on comparing the first set of phonemes and the third set of phonemes.

8. The method of claim 1 , further comprising:

prior to receiving the speech input from the user and after detecting the error, prompting the user to provide the speech input, the speech input including a preferred pronunciation of the one or more words.

9. The method of claim 1 , further comprising:

synthesizing a speech output including the one or more words; and

displaying a textual version of the speech output on a display of the electronic device.

10. A non-transitory computer readable storage medium storing one or more programs, the one or more programs comprising instructions, which when executed by an electronic device, cause the device to:

detect an error in a speech based interaction with a digital assistant based on detecting a user input other than the speech based interaction;

in response to detecting the error, receive a speech input from a user, the speech input including a pronunciation of one or more words;

determine, based on the pronunciation of the one or more words, a first set of phonemes from a speech recognition phonetic alphabet and a second set of phonemes from a speech synthesis phonetic alphabet;

update one or more databases to include the first set of phonemes and the second set of phonemes in association with a text string corresponding to the one or more words; and

perform speech recognition or speech synthesis using the updated one or more databases.

11. The computer readable storage medium of claim 10 , wherein the one or more words were received in a prior speech input provided by the user, and wherein the error is an error in speech recognition of the one or more words.

12. The computer readable storage medium of claim 10 , wherein the one or more words were output in a speech output by the electronic device, and wherein the error is an error in speech synthesis of the one or more words.

13. The computer readable storage medium of claim 10 , wherein the instructions further cause the device to:

receive the speech input including the one or more words;

perform speech recognition on the speech input to generate the text string corresponding to the one or more words;

determine a confidence metric of the text string; and

detect the error based on a determination that the confidence metric does not meet a predetermined threshold.

14. The computer readable storage medium of claim 10 , wherein the instructions further cause the device to:

synthesize a speech output including the one or more words; and

detect the error based on an indication from the user that the one or more words were pronounced incorrectly.

15. An electronic device, comprising:

one or more processors; and

memory storing one or more programs, the one or more programs including instructions, which when executed by the one or more processors, cause the one or more processors to:

detect an error in a speech based interaction with a digital assistant based on detecting a user input other than the speech based interaction;

in response to detecting the error, receive a speech input from a user, the speech input including a pronunciation of one or more words;

determine, based on the pronunciation of the one or more words, a first set of phonemes from a speech recognition phonetic alphabet and a second set of phonemes from a speech synthesis phonetic alphabet;

update one or more databases to include the first set of phonemes and the second set of phonemes in association with a text string corresponding to the one or more words; and

perform speech recognition or speech synthesis using the updated one or more databases.

16. The device of claim 15 , wherein the one or more words were received in a prior speech input provided by the user, and wherein the error is an error in speech recognition of the one or more words.

17. The device of claim 15 , wherein the one or more words were output in a speech output by the electronic device, and wherein the error is an error in speech synthesis of the one or more words.

18. The device of claim 15 , wherein the instructions further cause the one or more processors to:

receive the speech input including the one or more words;

perform speech recognition on the speech input to generate the text string corresponding to the one or more words;

determine a confidence metric of the text string; and

detect the error based on a determination that the confidence metric does not meet a predetermined threshold.

19. The device of claim 15 , wherein the instructions further cause the one or more processors to:

synthesize a speech output including the one or more words; and

detect the error based on an indication from the user that the one or more words were pronounced incorrectly.

Continuity (3)
Division 14298690 · Jun 6, 2014
Provisional Application 61832753 · Jun 7, 2013
Related Publication 20170178619A1 · Jun 22, 2017