IP Library › Granted Patent US 10,102,852
Granted Patent B2
US 10,102,852 · App. 14/686,670 · Granted Oct 16, 2018

Personalized speech synthesis for acknowledging voice actions

Inventors: Fuchun Peng (Cupertino, CA); Jakob Nicolaus Foerster (Zurich, CH); Diego Melendo Casado (San Francisco, CA); Fei Huang (New York, NY); Francoise Beaufays (Mountain View, CA)
Assignee: Google LLC
G10L15/22G10L13/033G10L15/26G10L15/07G10L15/187G10L2015/221G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,102,852
App. No.
14/686,670
Granted
Oct 16, 2018
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for presenting notifications in an enterprise system. In one aspect, a method include actions of obtaining a template that defines (i) trigger criteria for presenting a notification type and (ii) content rules for determining content to include in a notification of the notification type. Additional actions include accessing enterprise resources of an enterprise, the enterprise resources including data describing entities related to the enterprise and relationships among the entities. Further actions include, accessing user information specific to a user and determining that the trigger criteria is satisfied by the enterprise resources and the user information. Additional actions include generating a particular notification of the notification type based at least on the content rules and providing the particular notification to the user.

Claims (75)

1. A computer-implemented method for generating synthesized, spoken acknowledgments of voice queries using personalized pronunciations, the method comprising:

receiving, by a front end component of a voice query processing system that includes (a) the front end component, (b) an automated speech recognizer that is configured to transcribe given audio inputs to terms, (c) a pronunciation generator that is separate from the automated speech recognizer and that is configured to generate sequences of phones that correspond to given audio inputs without mapping the sequences of phones to terms, (d) a text-to speech component, and (e) a voice action engine, audio data encoding a voice query from a user;

providing, by the front end component of the voice query processing system, the audio data to both (i) the pronunciation generator that is separate from the automated speech recognizer and that is configured to generate, while the automated speech recognizer begins transcribing the audio data into terms, sequences of phones that correspond to given audio inputs without mapping the sequences of phone to terms, and (ii) the automated speech recognizer that is configured to transcribe given audio inputs to terms;

obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, a sequence of phones that reflects the user's pronunciation of a particular term of the voice query;

after obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator, obtaining, from the automated speech recognizer, a transcription of the voice query from the audio data, wherein the transcription includes the particular term;

after (i) obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator and (ii) obtaining, from the automated speech recognizer, the transcription of the voice query from the audio data, determining that a spoken acknowledgment that is to be generated for the voice query should use the user's pronunciation for the particular term;

generating, by the text-to-speech component of the voice query processing system, the spoken acknowledgment of the voice query, wherein, when output, the particular term is spoken in accordance with the user's pronunciation for the particular term based at least on the sequence of phones that was generated from the audio data by the pronunciation generator;

providing, by the text-to-speech component of the voice query processing system, the spoken acknowledgment for output; and

after providing the spoken acknowledgement for output, providing the voice query for execution by the voice action engine.

2. The method of claim 1 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

identifying a portion of the audio corresponding to the particular term; and

determining the sequence of phones from the portion of the audio corresponding to the particular term.

3. The method of claim 1 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

mapping the sequence of phones from the audio data corresponding to the particular term to the particular term in the transcription.

4. The method of claim 1 , comprising:

determining that the transcription includes a proper name,

wherein providing the spoken acknowledgment of the voice query for output is in response to determining that the transcription includes a proper name.

5. The method of claim 4 , wherein determining that the transcription includes a proper name comprises:

determining that the transcription includes one or more terms that indicate that the transcription includes a proper name.

6. The method of claim 1 , comprising:

determining, from the audio data, a confidence score for the sequence of phones that reflects a user's pronunciation for the particular term; and

determining that the confidence score for the sequence of phones satisfies a confidence threshold,

wherein providing the spoken acknowledgment of the voice query for output is in response to determining that the confidence score for the sequence of phones satisfies the confidence threshold.

7. The method of claim 1 , wherein obtaining, from the automated speech recognizer, a transcription of the voice query from the audio data comprises:

obtaining the transcription of the voice query from the audio data based at least on canonical pronunciation data associated with the particular term, where the canonical pronunciation data is stored in a pronunciation dictionary and different from the sequence of phones determined from the audio data encoding the voice query from the user.

8. The method of claim 1 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, a sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

obtaining pronunciations for phones within the voice query;

after obtaining the transcription of the voice query from the automated speech recognizer, aligning at least a portion of the obtained pronunciations for the phones within the voice query with the term in the transcription; and

generating custom pronunciation data that reflects the user's pronunciation for the particular term from the portion of the obtained pronunciations for the phones within the voice query aligned with the term in the transcription.

9. The method of claim 1 , wherein generating, by the text-to-speech component of the voice query processing system, a spoken acknowledgment of the voice query comprises:

obtaining text that includes the particular term for the spoken acknowledgement; and

synthesizing the spoken acknowledgement from the text based at least on (i) the sequence of phones for the particular term and (ii) canonical pronunciation data for one or more other terms in the text for the spoken acknowledgement.

10. A system for generating synthesized, spoken acknowledgements of voice queries using personalized pronunciations, the system comprising:

one or more computers; and

one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

receiving, by a front end component of a voice query processing system that includes (a) the front end component, (b) an automated speech recognizer that is configured to transcribe given audio inputs to terms, (c) a pronunciation generator that is separate from the automated speech recognizer and that is configured to generate sequences of phones that correspond to given audio inputs without mapping the sequences of phones to terms, (d) a text-to speech component, and (e) a voice action engine, audio data encoding a voice query from a user;

providing, by the front end component of the voice query processing system, the audio data to both (i) the pronunciation generator that is separate from the automated speech recognizer and that is configured to generate, while the automated speech recognizer begins transcribing the audio data into terms, sequences of phones that correspond to given audio inputs without mapping the sequences of phone to terms, and (ii) the automated speech recognizer that is configured to transcribe given audio inputs to terms;

obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, a sequence of phones that reflects the user's pronunciation of a particular term of the voice query;

after obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator, obtaining, from the automated speech recognizer, a transcription of the voice query from the audio data, wherein the transcription includes the particular term;

after (i) obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator and (ii) obtaining, from the automated speech recognizer, the transcription of the voice query from the audio data, determining that a spoken acknowledgment that is to be generated for the voice query should use the user's pronunciation for the particular term;

generating, by the text-to-speech component of the voice query processing system, the spoken acknowledgment of the voice query, wherein, when output, the particular term is spoken in accordance with the user's pronunciation for the particular term based at least on the sequence of phones that was generated from the audio data by the pronunciation generator; providing, by the text-to-speech component of the voice query processing system, the spoken acknowledgment for output; and

after providing the spoken acknowledgement for output, providing the voice query for execution by the voice action engine.

11. The system of claim 10 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

identifying a portion of the audio corresponding to the particular term; and

determining the sequence of phones from the portion of the audio corresponding to the particular term.

12. The system of claim 10 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

mapping the sequence of phones from the audio data corresponding to the particular term to the particular term in the transcription.

13. The system of claim 10 , the instructions further comprising:

determining that the transcription includes a proper name,

wherein providing the spoken acknowledgment of the voice query for output is in response to determining that the transcription includes a proper name.

14. The system of claim 13 , wherein determining that the transcription includes a proper name comprises:

determining that the transcription includes one or more terms that indicate that the transcription includes a proper name.

15. The system of claim 10 , the instructions further comprising:

determining, from the audio data, a confidence score for the sequence of phones that reflects a user's pronunciation for the particular term; and

determining that the confidence score for the sequence of phones satisfies a confidence threshold,

wherein providing the spoken acknowledgment of the voice query for output is in response to determining that the confidence score for the sequence of phones satisfies the confidence threshold.

16. The system of claim 10 , wherein obtaining, from the automated speech recognizer, a transcription of the voice query from the audio data comprises:

obtaining the transcription of the voice query from the audio data based at least on canonical pronunciation data associated with the particular term, where the canonical pronunciation data is stored in a pronunciation dictionary and different from the sequence of phones determined from the audio data encoding the voice query from the user.

17. A non-transitory computer-readable medium storing instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations for generating synthesized, spoken acknowledgments of voice queries using personalized pronunciations, the operations comprising:

receiving, by a front end component of a voice query processing system that includes (a) the front end component, (b) an automated speech recognizer that is configured to transcribe given audio inputs to terms, (c) a pronunciation generator that is separate from the automated speech recognizer and that is configured to generate sequences of phones that correspond to given audio inputs without mapping the sequences of phones to terms, (d) a text-to speech component, and (e) a voice action engine, audio data encoding a voice query from a user;

providing, by the front end component of the voice query processing system, the audio data to both (i) the pronunciation generator that is separate from the automated speech recognizer and that is configured to generate, while the automated speech recognizer begins transcribing the audio data into terms, sequences of phones that correspond to given audio inputs without mapping the sequences of phone to terms, and (ii) the automated speech recognizer that is configured to transcribe given audio inputs to terms;

obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, a sequence of phones that reflects the user's pronunciation of a particular term of the voice query;

after obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator, obtaining, from the automated speech recognizer, a transcription of the voice query from the audio data, wherein the transcription includes the particular term;

after (i) obtaining the sequence of phones that reflects the user's pronunciation of the particular term of the voice query from the pronunciation generator and (ii) obtaining, from the automated speech recognizer, the transcription of the voice query from the audio data, determining that a spoken acknowledgment that is to be generated for the voice query should use the user's pronunciation for the particular term;

generating, by the text-to-speech component of the voice query processing system, the spoken acknowledgment of the voice query, wherein, when output, the particular term is spoken in accordance with the user's pronunciation for the particular term based at least on the sequence of phones that was generated from the audio data by the pronunciation generator; providing, by the text-to-speech component of the voice query processing system, the spoken acknowledgment for output; and

after providing the spoken acknowledgement for output, providing the voice query for execution by the voice action engine.

18. The medium of claim 17 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

identifying a portion of the audio corresponding to the particular term; and

determining the sequence of phones from the portion of the audio corresponding to the particular term.

19. The medium of claim 17 , wherein obtaining, from the pronunciation generator that is separate from the automated speech recognizer and based on the pronunciation generator processing the audio data before the automated speech recognizer has completed transcribing the audio data, the sequence of phones that reflects the user's pronunciation of a particular term of the voice query comprises:

mapping the sequence of phones from the audio data corresponding to the particular term to the particular term in the transcription.

20. The medium of claim 17 , the instructions further comprising:

determining, from the audio data, a confidence score for the sequence of phones associated with the particular term; and

determining that the confidence score for the sequence of phones satisfies a confidence threshold,

wherein providing the spoken acknowledgment of the voice query for output is in response to determining that the confidence score for the sequence of phones satisfies the confidence threshold.

Assignments (2)
CHANGE OF NAME Recorded Oct 5, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044129/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 18, 2015
From: PENG, FUCHUN; FOERSTER, JAKOB NICOLAUS; CASADO, DIEGO MELENDO; HUANG, FEI; BEAUFAYS, FRANCOISE
To: GOOGLE INC.
Reel/Frame 035659/0550 →
Continuity (1)
Related Publication 20160307569A1 · Oct 20, 2016