IP Library › Granted Patent US 12,462,788
Granted Patent B2
US 12,462,788 · App. 18/312,576 · Granted Nov 4, 2025

Instantaneous learning in text-to-speech during dialog

Inventors: Vijayaditya Peddinti (San Jose, CA); Bhuvana Ramabhadran (Mt. Kisco, NY); Andrew Rosenberg (Brooklyn, NY); Mateusz Golebiewski (Mountain View, CA)
Assignee: Google LLC
G10L13/08G10L15/187
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,788
App. No.
18/312,576
Filed
May 4, 2023
Granted
Nov 4, 2025
Kind
B2
Art Unit
2657
USPC
704/260
Abstract

A method for instantaneous learning in text-to-speech (TTS) during dialog includes receiving a user pronunciation of a particular word present in a query spoken by a user. The method also includes receiving a TTS pronunciation of the same particular word that is present in a TTS input where the TTS pronunciation of the particular word is different than the user pronunciation of the particular word. The method also includes obtaining user pronunciation-related features and TTS pronunciation related features associated with the particular word. The method also includes generating a pronunciation decision selecting one of the user pronunciation or the TTS pronunciation of the particular word that is associated with a highest confidence. The method also include providing the TTS audio that includes a synthesized speech representation of the response to the query using the user pronunciation or the TTS pronunciation for the particular word.

Claims (66)

1 . A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

processing, using a trained automatic speech recognition model, audio data representing a query spoken by a user to identify a particular word present in the query;

receiving a user pronunciation of the particular word present in the query spoken by the user;

receiving a text-to-speech (TTS) pronunciation of the same particular word, the TTS pronunciation of the particular word comprising a verified preferred pronunciation of the particular word that is verified by a person as being an accurate pronunciation of the particular word;

determining, by processing the user pronunciation and the TTS pronunciation using a trained pronunciation decision model, that the verified preferred pronunciation of the particular word is different than the user pronunciation of the particular word;

generating, using the trained pronunciation decision model, a pronunciation decision selecting the verified preferred pronunciation of the particular word, the pronunciation decision comprising:

an indication for the user that the user pronunciation of the particular word is different from the verified preferred pronunciation of the particular word; and

the verified preferred pronunciation of the particular word; and

providing the pronunciation decision to a user device associated with the user.

2 . The method of claim 1 , wherein the operations further comprise, after receiving the TTS pronunciation of the same particular word:

obtaining user pronunciation-related features associated with the user pronunciation of the particular word; and

obtaining TTS pronunciation-related features associated with the TTS pronunciation of the particular word.

3 . The method of claim 2 , wherein selecting the verified preferred pronunciation of the particular word comprises determining that the verified preferred pronunciation of the particular word is associated with a highest confidence for use in TTS audio.

4 . The method of claim 2 , wherein the user pronunciation-related features associated with the user pronunciation of the particular word comprise at least one of:

a geographical area of the user when the query was spoken by the user;

linguistic demographic information associated with the user; or

a frequency of using the user pronunciation when pronouncing the particular word in previous queries spoken by the user and/or other users.

5 . The method of claim 2 , wherein receiving the TTS pronunciation of the particular word comprises:

receiving, as input to a trained TTS system, a TTS input comprising a textual representation of a response to the query;

generating, as output from the trained TTS system, an initial sample of TTS audio comprising an initial synthesized speech representation of the response to the query; and

extracting a TTS acoustic representation of the particular word from the initial sample of the TTS audio, the TTS acoustic representation conveying the TTS pronunciation of the particular word.

6 . The method of claim 2 , wherein receiving the TTS pronunciation of the particular word comprises processing a textual representation of a response to the query to generate a TTS phoneme representation that conveys the TTS pronunciation of the particular word.

7 . The method of claim 1 , wherein the operations further comprise:

receiving audio data corresponding to the query spoken by the user; and

processing, using an automated speech recognizer (ASR), the audio data to generate a transcription of the query.

8 . The method of claim 7 , wherein receiving the user pronunciation of the particular word comprises at least one of:

extracting the user pronunciation of the particular word from an intermittent state of the ASR while using the ASR to process the audio data;

extracting a user acoustic representation of the particular word from the audio data, the user acoustic representation conveying the user pronunciation of the particular word; or

processing the audio data to generate a user phoneme representation that conveys the user pronunciation of the particular word.

9 . The method of claim 1 , wherein the operations further comprise, after generating the pronunciation decision selecting the verified preferred pronunciation of the particular word:

receiving explicit feedback from the user indicating which one of the user pronunciation of the particular word or the TTS pronunciation of the particular word the user prefers for pronouncing the particular word in subsequent TTS outputs; and

updating the trained pronunciation decision model based on the explicit feedback from the user.

10 . A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that, when executed on the data processing hardware, cause the data processing hardware to perform operations comprising:

processing, using an automatic speech recognition (ASR) model, audio data representing a query spoken by a user to identify a particular word present in the query;

receiving a user pronunciation of the particular word present in the query spoken by the user;

receiving a text-to-speech (TTS) pronunciation of the same particular word, the TTS pronunciation of the particular word comprising a verified preferred pronunciation of the particular word that is verified by a person as being an accurate pronunciation of the particular word;

determining, by processing the user pronunciation and the TTS pronunciation using a trained pronunciation decision model, that the verified preferred pronunciation of the particular word is different than the user pronunciation of the particular word;

generating, using the trained pronunciation decision model, a pronunciation decision selecting the verified preferred pronunciation of the particular word, the pronunciation decision comprising:

an indication for the user that the user pronunciation of the particular word is different from the verified preferred pronunciation of the particular word; and

the verified preferred pronunciation of the particular word; and

providing the pronunciation decision to a user device associated with the user.

11 . The system of claim 10 , wherein the operations further comprise, after receiving the TTS pronunciation of the same particular word:

obtaining user pronunciation-related features associated with the user pronunciation of the particular word; and

obtaining TTS pronunciation-related features associated with the TTS pronunciation of the particular word.

12 . The system of claim 11 , wherein selecting the verified preferred pronunciation of the particular word comprises determining that the verified preferred pronunciation of the particular word is associated with a highest confidence for use in TTS audio.

13 . The system of claim 11 , wherein the user pronunciation-related features associated with the user pronunciation of the particular word comprise at least one of:

a geographical area of the user when the query was spoken by the user;

linguistic demographic information associated with the user; or

a frequency of using the user pronunciation when pronouncing the particular word in previous queries spoken by the user and/or other users.

14 . The system of claim 11 , wherein receiving the TTS pronunciation of the particular word comprises:

receiving, as input to a trained TTS system, a TTS input comprising a textual representation of a response to the query;

generating, as output from the trained TTS system, an initial sample of TTS audio comprising an initial synthesized speech representation of the response to the query; and

extracting a TTS acoustic representation of the particular word from the initial sample of the TTS audio, the TTS acoustic representation conveying the TTS pronunciation of the particular word.

15 . The system of claim 11 , wherein receiving the TTS pronunciation of the particular word comprises processing a textual representation of a response to the query to generate a TTS phoneme representation that conveys the TTS pronunciation of the particular word.

16 . The system of claim 10 , wherein the operations further comprise:

receiving audio data corresponding to the query spoken by the user; and

processing, using an automated speech recognizer (ASR), the audio data to generate a transcription of the query.

17 . The system of claim 16 , wherein receiving the user pronunciation of the particular word comprises at least one of:

extracting the user pronunciation of the particular word from an intermittent state of the ASR while using the ASR to process the audio data;

extracting a user acoustic representation of the particular word from the audio data, the user acoustic representation conveying the user pronunciation of the particular word; or

processing the audio data to generate a user phoneme representation that conveys the user pronunciation of the particular word.

18 . The system of claim 10 , wherein the operations further comprise, after generating the pronunciation decision selecting the verified preferred pronunciation of the particular word:

receiving explicit feedback from the user indicating which one of the user pronunciation of the particular word or the TTS pronunciation of the particular word the user prefers for pronouncing the particular word in subsequent TTS outputs; and

updating the pronunciation decision model based on the explicit feedback from the user.

Continuity (2)
Continuation 17190456 · Mar 3, 2021
Related Publication 20230274727A1 · Aug 31, 2023
References Cited (40)
US 5766015A · Shpiro · 1998 [cited by examiner]
US 6249763B1 · Minematsu · 2001 [cited by examiner]
US 7280963B1 · Beaufays · 2007 [cited by examiner]
US 10319365B1 · Nicolis · 2019 [cited by examiner]
US 10339920B2 · Adams · 2019 [cited by examiner]
US 10403291B2 · Moreno · 2019 [cited by examiner]
US 11546390B1 · Boodaei · 2023 [cited by examiner]
US 20020082831A1 · Hwang · 2002 [cited by examiner]
US 20020118846A1 · Narusawa · 2002 [cited by examiner]
US 20020146669A1 · Bender · 2002 [cited by examiner]
US 20040230431A1 · Gupta · 2004 [cited by examiner]
US 20050273337A1 · Erell et al. · 2005 [cited by applicant]
US 20070233487A1 · Cohen · 2007 [cited by examiner]
US 20100153115A1 · Klee · 2010 [cited by examiner]
US 20110218806A1 · Alewine et al. · 2011 [cited by applicant]
US 20130277915A1 · Garrett · 2013 [cited by examiner]
US 20140234809A1 · Colvard · 2014 [cited by examiner]
US 20140365216A1 · Gruber · 2014 [cited by examiner]
US 20140379709A1 · Mack · 2014 [cited by examiner]
US 20150161985A1 · Peng · 2015 [cited by examiner]
US 20150206539A1 · Campbell · 2015 [cited by examiner]
US 20150243278A1 · Kibre · 2015 [cited by examiner]
US 20150364141A1 · Lee · 2015 [cited by examiner]
US 20160188727A1 · Waibel · 2016 [cited by examiner]
US 20160307569A1 · Peng et al. · 2016 [cited by applicant]
US 20170130465A1 · Claudin · 2017 [cited by examiner]
US 20170178619A1 · Naik et al. · 2017 [cited by applicant]
US 20180130465A1 · Kim · 2018 [cited by examiner]
US 20180174483A1 · Bhunachet · 2018 [cited by examiner]
US 20180190269A1 · Lokeswarappa · 2018 [cited by examiner]
US 20200228336A1 · Streit · 2020 [cited by examiner]
US 20200357390A1 · Bromand · 2020 [cited by examiner]
US 20210099759A1 · Armstrong · 2021 [cited by examiner]
US 20230083096A1 · Trehan · 2023 [cited by examiner]
US 20230162731A1 · Wu · 2023 [cited by examiner]
US 20240193193A1 · Hattori · 2024 [cited by examiner]
JP H11175082A · 1999 [cited by applicant]
JP 2001159865A · 2001 [cited by applicant]
International Search Report and Written Opinion for the related Application No. PCT/US2021/018219, dated May 31, 2022, 68 pages. [cited by applicant]
Japanese Office Action for the related Application No. 2024-151468 dated Sep. 2, 2025. [cited by applicant]