IP Library › Granted Patent US 12,573,402
Granted Patent B2
US 12,573,402 · App. 17/710,137 · Granted Mar 10, 2026

Generating and/or utilizing unintentional memorization measure(s) for automatic speech recognition model(s)

Inventors: Om Dipakbhai Thakkar (San Jose, CA); Hakim Sidahmed (Washington, DC); W. Ronny Huang (Kensington, MD); Rajiv Mathews (Sunnyvale, CA); Françoise Beaufays (Mountain View, CA); Florian Tramèr (Zurich, CH)
Assignee: GOOGLE LLC
G10L15/26G10L13/02G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,402
App. No.
17/710,137
Granted
Mar 10, 2026
Kind
B2
Abstract

An unintentional memorization measure can be used to determine whether an automatic speech recognition (ASR) model has unintentionally memorized one or more phrases during training of the ASR model. Various implementations include generating one or more candidate transcripts based on the vocabulary of the ASR model. For example, the system can generate a candidate transcript by appending a token of the vocabulary to a previous candidate transcript. Various implementations include processing the candidate transcript using a speech synthesis model to generate synthesized speech audio data that includes synthesized speech of the candidate transcript. Additionally or alternatively, the synthesized speech audio data can be processed using the ASR model to generate ASR output. Various implementations can include generating a loss based on comparing the ASR output and the candidate transcript.

Claims (59)

1 . A method implemented by one or more processors, the method comprising:

generating a candidate transcript using a vocabulary corresponding to an automatic speech recognition (“ASR”) model;

generating, based on processing the candidate transcript using a speech synthesis model, synthesized speech audio data that includes synthesized speech of the candidate transcript;

processing the synthesized speech audio data using the ASR model to generate ASR output that reflects a predicted text representation of the synthesized speech;

generating a loss based on comparison of the ASR output to the candidate transcript;

determining, based on the loss that is based on the comparison of the ASR output to the candidate transcript, whether the ASR model unintentionally memorized one or more occurrences, in training data used to train the ASR model, of a corresponding human speaking the candidate transcript;

generating an overall memorization measure as a function of the determination of whether the ASR model unintentionally memorized the one or more occurrences;

determining, based on the overall memorization measure, whether to transmit the ASR model to a plurality of client devices; and

in response to determining, based on the overall memorization measure, to transmit the ASR model to the plurality of client devices:

transmitting, to the plurality of client devices, the ASR model or weights of the ASR model.

2 . The method of claim 1 , wherein generating the candidate transcript using the vocabulary corresponding to the ASR model comprises:

prior to generating the candidate transcript:

generating a prior candidate transcript using the vocabulary corresponding to the ASR model, wherein the prior candidate transcript includes a beginning portion of words of the candidate transcript but lacks one or more ending words of the candidate transcript;

generating, based on processing the prior candidate transcript using the speech synthesis model, prior synthesized speech audio data that includes prior synthesized speech of the prior candidate transcript;

processing the prior synthesized speech audio data using the ASR model to generate prior ASR output that reflects a prior predicted text representation of the prior synthesized speech;

generating a prior loss based on comparison of the prior ASR output to the prior candidate transcript; and

selecting, based on the prior loss, the prior candidate transcript for augmentation.

3 . The method of claim 2 , further comprising:

prior to generating the candidate transcript:

generating an additional prior candidate transcript using the vocabulary corresponding to the ASR model, wherein the additional prior candidate transcript includes a beginning portion of words of the prior candidate transcript but lacks one or more ending words of the prior candidate transcript;

generating, based on processing the additional prior candidate transcript using the speech synthesis model, additional prior synthesized speech audio data that includes synthesized speech of the additional prior candidate transcript;

processing the additional prior synthesized speech audio data using the ASR model to generate additional prior ASR output that reflects an additional prior predicted text representation of the additional prior synthesized speech;

generating an additional prior loss based on comparison of the additional prior ASR output to the additional prior candidate transcript; and

selecting, based on (a) the prior loss of the prior candidate transcript and (b) the additional prior loss of the additional prior candidate transcript, the prior candidate transcript for augmentation instead of the additional prior candidate transcript.

4 . The method of claim 1 , wherein generating the loss based on comparison of the ASR output to the candidate transcript comprises:

generating the loss based on comparing a transcript portion of the ASR output with the candidate transcript.

5 . The method of claim 1 , wherein generating the loss based on comparison of the ASR output to the candidate transcript comprises:

for each word in the candidate transcript,

comparing a probability portion of the ASR output of a corresponding portion of the ASR output with the word in the candidate transcript; and

generating the loss based on the comparing.

6 . The method of claim 1 , further comprising:

determining, based on the overall memorization measure for the ASR model, whether to utilize a particular technique in federated training of the ASR model and/or of an additional machine learning model.

7 . The method of claim 1 , further comprising:

transmitting, in response to a request by a third party, the overall memorization measure.

8 . The method of claim 1 , further comprising:

transmitting, in response to a request by a third party, the candidate transcript.

9 . The method of claim 1 , wherein the vocabulary corresponding to the ASR model comprises a set of tokens, a set of characters, a set of letters, a set of words, a set of word-pieces, and/or a set of phonemes.

10 . The method of claim 1 , wherein determining, based on the overall memorization measure, whether to transmit the ASR model to a plurality of client devices comprises:

comparing the overall memorization measure to a prior memorization measure generated based on prior weights of the ASR model, wherein the weights of the ASR model are updated relative to the prior weights of the ASR model.

11 . The method of claim 10 , wherein the weights of the ASR model are updated based on client model updates from the plurality of client devices, the client model updates from the plurality of client devices being based on client gradients locally generated at the client devices.

12 . A method implemented by one or more processors, the method comprising:

receiving an automatic speech recognition (“ASR”) model and a vocabulary corresponding to the ASR model;

generating, based on the vocabulary of the ASR model, a set of candidate transcripts;

for each candidate transcript in the set of candidate transcripts:

generating, based on processing the candidate transcript using a speech synthesis model, synthesized speech audio data that includes synthesized speech of the candidate transcript;

processing the synthesized speech audio data using the ASR model to generate ASR output that reflects a predicted text representation of the synthesized speech;

generating a loss based on comparison of the ASR output to the candidate transcript;

determining, based on the loss that is based on the comparison of the ASR output to the candidate transcript, whether the ASR model unintentionally memorized one or more occurrences, in training data used to train the ASR model, of a corresponding human speaking the candidate transcript;

determining, based on whether the ASR model unintentionally memorized the one or more occurrences of the corresponding human speaking the candidate transcript, whether to transmit the ASR model to a plurality of client devices; and

in response to determining to transmit the ASR model to the plurality of client devices:

transmitting, to the plurality of client devices, the ASR model or weights of the ASR model.

13 . The method of claim 12 , further comprising:

generating an overall memorization measure as a function of the determination of whether the ASR model unintentionally memorized the one or more occurrences corresponding to one or more candidate transcripts in the set of candidate transcripts.

14 . The method of claim 13 , further comprising:

identifying a subset of candidate transcripts, wherein the loss corresponding to each of the candidate transcripts in the subset of candidate transcripts indicates a probability indicating the ASR model unintentionally memorized the corresponding one or more occurrences satisfies a threshold value.

15 . The method of claim 14 , wherein receiving the ASR model and the vocabulary corresponding to the ASR model comprises receiving the ASR model and the vocabulary from a third party, and further comprising transmitting the subset of candidate transcripts to the third party.

16 . The method of claim 12 , wherein the training data used to train the ASR model includes one or more targeted training instances and further comprising:

determining whether the set of candidate transcripts includes the one or more targeted training instances; and

determining an overall memorization measure as a function of whether the set of candidate transcripts includes the one or more targeted training instances.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 16, 2022
From: THAKKAR, OM DIPAKBHAI; SIDAHMED, HAKIM; HUANG, W. RONNY; MATHEWS, RAJIV; BEAUFAYS, FRANCOISE; TRAMÈR, FLORIAN
To: GOOGLE LLC
Reel/Frame 059921/0543 →
Continuity (1)
Related Publication 20230317082A1 · Oct 5, 2023
References Cited (20)
US 20110016110A1 · Egi · 2011 [cited by examiner]
US 20140163981A1 · Cook · 2014 [cited by examiner]
US 20220068255A1 · Chen · 2022 [cited by examiner]
US 20230169954A1 · Thomas · 2023 [cited by examiner]
JP 2011022705 · 2011 [cited by applicant]
WO 2020229684 · 2020 [cited by applicant]
WO 2021006920 · 2021 [cited by applicant]
WO 2021081061 · 2021 [cited by applicant]
WO 2021101501 · 2021 [cited by applicant]
WO 2021234839 · 2021 [cited by applicant]
Thakkar et al. “Understanding unintended memorization in language models under federated learning.” Proceedings of the Third Workshop on Privacy in Natural Language Processing. 2021 (https://aclanthology.org/2021.privat… [cited by examiner]
Huang, W.R. et al., “Detecting Unintended Memorization in Language-Model-Fused ASR”; arXiv.org, Cornell University Library; arXiv:2204.09606; 6 pages; dated Apr. 20, 2022. [cited by applicant]
Carlini, N. et al., “The Secret Sharer: Evaluating and Testing Unintended Memorization in Neural Networks”; arXiv.org, Cornell University Library; arXiv:1802.08232; 18 pages; dated Feb. 22, 2018. [cited by applicant]
European Patent Office; International Search Report and Written Order issued in Application No. PCT/US2022/047275; 13 pages; dated Dec. 21, 2022. [cited by applicant]
European Patent Office, Intention to Grant issued in Application No. 22809262.3; 59 pages; dated Feb. 14, 2025. [cited by applicant]
Giulia Garau “Speaker Normalisation for Large Vocabulary Multiparty Conversational Speech Recognition” University of Edinburgh. 2009. 190 pages. [cited by applicant]
Cho et al., “Learning Speaker Embedding from Text-to-Speech” arXiv:2010.11221v1 [eess.AS] 5 pages, dated Oct. 21, 2020. [cited by applicant]
Dang et al., “A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter It” arXiv:2104.07815v1 [cs.CL] 16 pages, dated Apr. 15, 2021. [cited by applicant]
Yu Zhang “Exploring Neural Network Architectures For Acoustic Modeling” Massachusetts Institute of Technology. 132 pages, dated Sep. 2017. [cited by applicant]
Japanese Patent Office, Notice of Refusal issued in Application No. 2024556755, 10 pages, dated Aug. 26, 2025. [cited by applicant]