IP Library › Granted Patent US 12,260,875
Granted Patent B2
US 12,260,875 · App. 18/609,362 · Granted Mar 25, 2025

Phrase extraction for ASR models

Inventors: Ehsan Amid (Mountain View, CA); Om Dipakbhai Thakkar (Sunnyvale, CA); Rajiv Mathews (Sunnyvale, CA); Francoise Beaufays (Mountain View, CA)
Assignee: Google LLC
G10L21/0332G10L15/063G10L15/08G10L21/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,260,875
App. No.
18/609,362
Granted
Mar 25, 2025
Kind
B2
Abstract

A method of phrase extraction for ASR models includes obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance and modifying the audio data to obfuscate a particular phrase recited in the utterance. The method also includes processing, using a trained ASR model, the modified audio data to generate a predicted transcription of the utterance, and determining whether the predicted transcription includes the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance. When the predicted transcription includes the particular phrase, the method includes generating an output indicating that the trained ASR model leaked the particular phrase from a training data set used to train the ASR model.

Claims (38)

1. A computer-implemented method executed on data processing hardware that causes the data processing hardware to perform operations comprising:

obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance;

processing the ground-truth transcription to identify a particular phrase included in the ground-truth transcription that is associated with sensitive data;

modifying the audio data to obfuscate the particular phrase recited in the utterance;

processing, using a trained automated speech recognition (ASR) model, the modified audio data to generate a predicted transcription of the utterance;

determining whether the predicted transcription includes the particular phrase or another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance; and

when the predicted transcription includes the other phrase substituted for the particular phrase, generating an output indicating that the trained ASR model leaked the other phrase from a training data set used to train the ASR model.

2. The computer-implemented method of claim 1 , wherein the operations further comprise,

when the predicted transcription includes the particular phrase, generating an output indicating that the trained ASR model leaked the particular phrase from the training data set used to train the ASR model.

3. The computer-implemented method of claim 2 , wherein the operations further comprise, when the predicted transcription does not include the particular phrase or the other phrase substituted for the particular phrase from the ground-truth transcription that is associated with the same category of information as the particular phrase, generating an output indicating that the trained ASR model has not leaked any information from the training data set used to train the ASR model.

4. The computer-implemented method of claim 1 , wherein the audio data comprises an audio waveform.

5. The computer-implemented method of claim 4 , wherein the audio waveform corresponds to human speech.

6. The computer-implemented method of claim 4 , wherein the audio waveform corresponds to synthesized speech.

7. The computer-implemented method of claim 1 , wherein modifying the audio data comprises:

based on the ground-truth transcription, identifying a segment of the audio data that aligns with the particular phrase in the ground-truth transcription; and

performing data augmentation on the identified segment of the audio data to obfuscate the particular phrase recited in the utterance.

8. The computer-implemented method of claim 7 , wherein performing data augmentation on the identified segment of the audio data comprises adding noise to the identified segment of the audio data.

9. The computer-implemented method of claim 7 , wherein performing data augmentation on the identified segment of the audio data comprises removing the identified segment of the audio data.

10. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance;

processing the ground-truth transcription to identify a particular phrase included in the ground-truth transcription that is associated with sensitive data;

modifying the audio data to obfuscate the particular phrase recited in the utterance;

processing, using a trained automated speech recognition (ASR) model, the modified audio data to generate a predicted transcription of the utterance;

determining whether the predicted transcription includes the particular phrase or another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance; and

when the predicted transcription includes the other phrase substituted for the particular phrase, generating an output indicating that the trained ASR model leaked the other phrase from a training data set used to train the ASR model.

11. The system of claim 10 , wherein the operations further comprise,

when the predicted transcription includes the particular phrase, generating an output indicating that the trained ASR model leaked the particular phrase from the training data set used to train the ASR model.

12. The system of claim 11 , wherein the operations further comprise, when the predicted transcription does not include the particular phrase or the other phrase substituted for the particular phrase from the ground-truth transcription that is associated with the same category of information as the particular phrase, generating an output indicating that the trained ASR model has not leaked any information from the training data set used to train the ASR model.

13. The system of claim 10 , wherein the audio data comprises an audio waveform.

14. The system of claim 13 , wherein the audio waveform corresponds to human speech.

15. The system of claim 13 , wherein the audio waveform corresponds to synthesized speech.

16. The system of claim 10 , wherein modifying the audio data comprises:

based on the ground-truth transcription, identifying a segment of the audio data that aligns with the particular phrase in the ground-truth transcription; and

performing data augmentation on the identified segment of the audio data to obfuscate the particular phrase recited in the utterance.

17. The system of claim 16 , wherein performing data augmentation on the identified segment of the audio data comprises adding noise to the identified segment of the audio data.

18. The system of claim 16 , wherein performing data augmentation on the identified segment of the audio data comprises removing the identified segment of the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 21, 2024
From: AMID, EHSAN; THAKKAR, OHM; MATHEWS, RAJIV; BEAUFAYS, FRANCOISE
To: GOOGLE LLC
Reel/Frame 066859/0582 →
Continuity (3)
Continuation 17643848 · Dec 13, 2021
Provisional Application 63264836 · Dec 2, 2021
Related Publication 20240221772A1 · Jul 4, 2024
References Cited (23)
US 9123343B2 · Kurki-Suonio · 2015 [cited by applicant]
US 11538467B1 · Feyisetan · 2022 [cited by applicant]
US 11955134B2 · Amid · 2024 [cited by examiner]
US 20150287401A1 · Lee et al. · 2015 [cited by applicant]
US 20170139905A1 · Na · 2017 [cited by applicant]
US 20200034663A1 · Michiels et al. · 2020 [cited by applicant]
US 20200082259A1 · Gu et al. · 2020 [cited by applicant]
US 20200311540A1 · Chakraborty · 2020 [cited by examiner]
US 20210064760A1 · Sharma et al. · 2021 [cited by applicant]
US 20210150269A1 · Choudhury · 2021 [cited by examiner]
US 20210303724A1 · Goshen · 2021 [cited by examiner]
US 20210357508A1 · Elovici et al. · 2021 [cited by applicant]
US 20210377035A1 · Walheim · 2021 [cited by examiner]
Kwon et al. “Selective Audio Adversarial Example in Evasion Attack on Speech Recognition System”. IEEE Transactions on Information Forensics and Security, vol. 15, 2020. Published Jun. 27, 2019, pp. 526-538 (Year: 2019). [cited by examiner]
Pal et al. “Synthetic speech detection using fundamental frequency variation and spectral features”. Computer Speech & Language 48 (2018) 31-50 (Year: 2018). [cited by examiner]
Zhang et al. “Privacy-preserving Machine Learning through Data Obfuscation”. arXiv:1807.01860v2 [cs.CR] Jul. 13, 2018 (Year: 2018). [cited by examiner]
Giuseppe Ateniese et al, “Hacking smart machines with smarter ones: How to extract meaningful data from machine learning classifiers”, International Journal of Security and Networks (IJSN), vol. {0} 10, No. {0} 3, Jun. … [cited by applicant]
Trung Dang et al, “A Method to Reveal Speaker Identity in Distributed ASR Training, and How to Counter It”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Apr. 15, 2021 (Apr.… [cited by applicant]
Trung Dang et al, “Revealing and Protecting Labels in Distributed Training”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853,Oct. 31, 2021 (Oct. 31, 2021). [cited by applicant]
International Search Report and Written Opinion issued in related PCT Application No. PCT/US2021/062998, dated Sep. 9, 2022. [cited by applicant]
Shah et al. “Evaluating the vulnerability of end-to-end automatic speech recognition models to membership inference attacks.” Interspeech 2021, Aug. 30-Sep. 3, 2021, Brno Czech Republic (Year: 2021). [cited by applicant]
Khanna et al. “Identifying Privacy Vulnerabilities in Key Stages of Computer Vision, Natural Language Processing, and Voce Processing Systems”. International Journal of Business Intelligence and Big Data Analytics, pp. … [cited by applicant]
Tseng et al. “Membership Inference Attacks Against Self-supervised Speech Models”. ArXiv:2111.05113v1 [cs.CR] Nov. 9, 2021 (Year: 2021). [cited by applicant]