IP Library Granted Patent US 12,417,757
Granted Patent B2
US 12,417,757 · App. 18/656,197 · Granted Sep 16, 2025

On-device speech synthesis of textual segments for training of on-device speech recognition model

Inventors: Françoise Beaufays (Mountain View, CA); Johan Schalkwyk (Scarsdale, NY); Khe Chai Sim (Dublin, CA)
Assignee: GOOGLE LLC
G10L13/047G10L15/063G10L2015/0635
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,417,757
App. No.
18/656,197
Granted
Sep 16, 2025
Kind
B2
Abstract

Processor(s) of a client device can: identify a textual segment stored locally at the client device; process the textual segment, using a speech synthesis model stored locally at the client device, to generate synthesized speech audio data that includes synthesized speech of the identified textual segment; process the synthesized speech, using an on-device speech recognition model that is stored locally at the client device, to generate predicted output; and generate a gradient based on comparing the predicted output to ground truth output that corresponds to the textual segment. In some implementations, the generated gradient is used, by processor(s) of the client device, to update weights of the on-device speech recognition model. In some implementations, the generated gradient is additionally or alternatively transmitted to a remote system for use in remote updating of global weights of a global speech recognition model.

Claims (44)

1. A method of personalizing an end-to-end speech recognition model for a user, the method implemented by one or more processors and the method comprising:

identifying a textual segment;

generating synthesized speech audio data that includes synthesized speech of the identified textual segment, wherein generating the synthesized speech audio data comprises processing the textual segment using a speech synthesis model adapted to one or more voice characteristics of the user;

processing, using the end-to-end speech recognition model, the synthesized speech audio data to generate a predicted textual segment;

generating a gradient based on comparing the predicted textual segment to the textual segment; and

updating one or more weights of the end-to-end speech recognition model based on the generated gradient.

2. The method of claim 1 , wherein the one or more processors are of a client device and further comprising:

transmitting, to a remote system, the generated gradient without transmitting any of: the textual segment, the synthesized speech audio data, and the predicted textual segment;

wherein the remote system utilizes the generated gradient, and additional gradients from additional client devices, to update global weights of a global end-to-end speech recognition model.

3. The method of claim 1 , wherein the identified textual segment is a name and wherein identifying the textual segment is based on the name being saved as a contact in a contacts list.

4. The method of claim 1 , wherein the identified textual segment is a name of a media item and wherein identifying the textual segment includes identifying the textual segment from a media playlist.

5. The method of claim 1 , wherein identifying the textual segment is based on determining that the textual segment is new.

6. The method of claim 1 , wherein the one or more processors are of a client device and further comprising:

determining, based on sensor data from one or more sensors of the client device, that a current state of the client device satisfies one or more conditions;

wherein generating the synthesized speech audio data, processing the synthesized speech audio data to generate the predicted textual segment, generating the gradient, and/or updating the one or more weights are performed responsive to determining that the current state of the client device satisfies the one or more conditions.

7. The method of claim 6 , wherein the one or more conditions include at least one of: the client device is charging, the client device has at least a threshold state of charge, or the client device is not being carried.

8. The method of claim 6 , wherein the one or more conditions include that the client device has at least a threshold state of charge and the client device is not being carried.

9. The method of claim 1 , wherein the speech synthesis model is adapted to the one or more voice characteristics of the user based on prior audio data that captures a prior utterance of the user, the prior utterance of the user being provided prior to generating the synthesized speech audio.

10. A system used in personalizing an end-to-end speech recognition model for a user, the system comprising:

memory storing instructions;

one or more processors operable to execute the instructions to:

identify a textual segment;

generate synthesized speech audio data that includes synthesized speech of the identified textual segment, wherein in generating the synthesized speech audio data one or more of the processors are to process the textual segment using a speech synthesis model adapted to one or more voice characteristics of the user;

process, using the end-to-end speech recognition model, the synthesized speech audio data to generate a predicted textual segment;

generate a gradient based on comparing the predicted textual segment to the textual segment; and

update one or more weights of the end-to-end speech recognition model based on the generated gradient.

11. The system of claim 10 , wherein the one or more processors are of a client device and wherein one or more of the processors are further operable to:

transmit, to a remote system, the generated gradient without transmitting any of: the textual segment, the synthesized speech audio data, and the predicted textual segment;

wherein the remote system utilizes the generated gradient, and additional gradients from additional client devices, to update global weights of a global end-to-end speech recognition model.

12. The system of claim 10 , wherein the identified textual segment is a name and wherein in identifying the textual segment one or more of the processors are to identify the textual segment based on the name being saved as a contact in a contacts list.

13. The system of claim 10 , wherein the identified textual segment is a name of a media item and wherein in identifying the textual segment one or more of the processors are to identify the textual segment from a media playlist.

14. The system of claim 10 , wherein in identifying the textual segment one or more of the processors are to identify the textual segment based on determining that the textual segment is new.

15. The system of claim 10 , wherein the one or more processors are of a client device and wherein one or more of the processors are further operable to:

determine, based on sensor data from one or more sensors of the client device, that a current state of the client device satisfies one or more conditions;

wherein generating the synthesized speech audio data, processing the synthesized speech audio data to generate the predicted textual segment, generating the gradient, and/or updating the one or more weights are performed responsive to determining that the current state of the client device satisfies the one or more conditions.

16. The system of claim 15 , wherein the one or more conditions include at least one of: the client device is charging, the client device has at least a threshold state of charge, or the client device is not being carried.

17. The system of claim 15 , wherein the one or more conditions include that the client device has at least a threshold state of charge and the client device is not being carried.

18. The system of claim 10 , wherein the speech synthesis model is adapted to the one or more voice characteristics of the user based on prior audio data that captures a prior utterance of the user, the prior utterance of the user being provided prior to generating the synthesized speech audio.

19. One or more non-transitory computer readable storage media storing computer instructions that are executable by one or more processors to personalize an end-to-end speech recognition model for a user by:

identifying a textual segment;

generating synthesized speech audio data that includes synthesized speech of the identified textual segment, wherein generating the synthesized speech audio data comprises processing the textual segment using a speech synthesis model adapted to one or more voice characteristics of the user;

processing, using the end-to-end speech recognition model, the synthesized speech audio data to generate a predicted textual segment;

generating a gradient based on comparing the predicted textual segment to the textual segment; and

updating one or more weights of the end-to-end speech recognition model based on the generated gradient.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2024
From: BEAUFAYS, FRANÇOISE; SCHALKWYK, JOHAN; SIM, KHE CHAI
To: GOOGLE LLC
Reel/Frame 067512/0129 →
Continuity (5)
Continuation 18204324 · May 31, 2023
Continuation 17479285 · Sep 20, 2021
Continuation 16959546
Provisional Application 62872140 · Jul 9, 2019
Related Publication 20240290317A1 · Aug 29, 2024
References Cited (61)
US 6970820B2 · Junqua · 2005 [cited by examiner]
US 7277855B1 · Acker · 2007 [cited by examiner]
US 7668718B2 · Kahn et al. · 2010 [cited by applicant]
US 7957969B2 · Alewine et al. · 2011 [cited by applicant]
US 8818793B1 · Bangalore · 2014 [cited by applicant]
US 9508338B1 · Kaszczuk et al. · 2016 [cited by applicant]
US 9697822B1 · Naik et al. · 2017 [cited by applicant]
US 9792900B1 · Kaskari · 2017 [cited by applicant]
US 9911437B2 · Melamed · 2018 [cited by examiner]
US 10388272B1 · Thomson et al. · 2019 [cited by applicant]
US 10789956B1 · Dube · 2020 [cited by applicant]
US 11127392B2 · Beaufays et al. · 2021 [cited by applicant]
US 11705106B2 · Beaufays et al. · 2023 [cited by applicant]
US 20030069729A1 · Bickley et al. · 2003 [cited by applicant]
US 20050182629A1 · Coorman et al. · 2005 [cited by applicant]
US 20060136205A1 · Song · 2006 [cited by examiner]
US 20060149558A1 · Kahn et al. · 2006 [cited by applicant]
US 20070055526A1 · Eide et al. · 2007 [cited by applicant]
US 20070208570A1 · Bhardwaj · 2007 [cited by applicant]
US 20080235024A1 · Goldberg · 2008 [cited by examiner]
US 20110153620A1 · Coifman · 2011 [cited by examiner]
US 20110288863A1 · Rasmussen · 2011 [cited by applicant]
US 20120310642A1 · Cao et al. · 2012 [cited by applicant]
US 20130262096A1 · Wilhelms-Tricarico et al. · 2013 [cited by applicant]
US 20130325446A1 · Levien et al. · 2013 [cited by applicant]
US 20150025891A1 · Goldberg · 2015 [cited by examiner]
US 20150081293A1 · Hsu et al. · 2015 [cited by applicant]
US 20150161983A1 · Yassa · 2015 [cited by applicant]
US 20160247521A1 · Melamed · 2016 [cited by examiner]
US 20160379626A1 · Deisher et al. · 2016 [cited by applicant]
US 20170069311A1 · Grost · 2017 [cited by examiner]
US 20170206889A1 · Lev-Tov et al. · 2017 [cited by applicant]
US 20170301347A1 · Fuhrman · 2017 [cited by applicant]
US 20190005947A1 · Kim et al. · 2019 [cited by applicant]
US 20190005952A1 · Kruse et al. · 2019 [cited by applicant]
US 20200175961A1 · Thomson · 2020 [cited by examiner]
US 20210006920A1 · Beaufays et al. · 2021 [cited by applicant]
US 20210104223A1 · Beaufays et al. · 2021 [cited by applicant]
US 20220005458A1 · Beaufays et al. · 2022 [cited by applicant]
US 20220005462A1 · Hwang · 2022 [cited by applicant]
US 20230306955A1 · Beaufays et al. · 2023 [cited by applicant]
CN 108133705 · 2018 [cited by applicant]
CN 108182936 · 2018 [cited by applicant]
CN 109215637 · 2019 [cited by applicant]
CN 109887484 · 2019 [cited by applicant]
JP H0389294 · 1991 [cited by applicant]
JP H11338489 · 1999 [cited by applicant]
JP 2001013983 · 2001 [cited by applicant]
JP 2005043461 · 2005 [cited by applicant]
JP 2005301097 · 2005 [cited by applicant]
JP 2013218095 · 2013 [cited by applicant]
JP 2019101291 · 2019 [cited by applicant]
China National Intellectual Property Administration; Notice of Allowance issued in Application No. 201980091350.4; 4 pages; dated Jun. 18, 2024. [cited by applicant]
Intellectual Property India; Hearing Notice issued in Application No. IN202127030140; 2 pages; dated Aug. 19, 2024. [cited by applicant]
China National Intellectual Property Administration; Notification of First Office Action issued in Application No. 201980091350.4; 15 pages; dated Jan. 9, 2024. [cited by applicant]
Japanese Patent Office; Notice of Allowance issued in Application No. 2021-541637, 3 pages, dated Jun. 13, 2022. [cited by applicant]
Intellectual Property India; Examination Report issued in Application No. IN202127030140; 6 pages; dated Mar. 14, 2022. [cited by applicant]
The Korean Intellectual Property Office; Allowance of Patent issued in Application No. 10-2021-7024199, 3 pages, dated Mar. 25, 2022. [cited by applicant]
The Korean Intellectual Property Office; Notice of Office Action issued in Application No. 10-2021-7024199; 10 pages; dated Dec. 15, 2021. [cited by applicant]
Ueno, S. et al. , “Multi-Speaker Sequence-to-Sequence Speech Synthesis for Data Augmentation in Acoustic-to-Word Speech Recognition;” 2019 IEEE International Conference on Acoustics, Speech and Signal Processing; pp. 61… [cited by applicant]
European Patent Office; International Search Report and Written Opinion of PCT Ser. No. PCT/US2019/054314; 17 pages; dated Feb. 21, 2020. [cited by applicant]