IP Library › Granted Patent US 11,955,134
Granted Patent B2
US 11,955,134 · App. 17/643,848 · Granted Apr 9, 2024

Phrase extraction for ASR models

Inventors: Ehsan Amid (Mountain View, CA); Om Thakkar (Freemont, CA); Rajiv Mathews (Mountain View, CA); Francoise Beaufays (Mountain View, CA)
Assignee: Google LLC
G10L21/0332G10L15/063G10L15/08G10L21/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,955,134
App. No.
17/643,848
Granted
Apr 9, 2024
Kind
B2
Abstract

A method of phrase extraction for ASR models includes obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance and modifying the audio data to obfuscate a particular phrase recited in the utterance. The method also includes processing, using a trained ASR model, the modified audio data to generate a predicted transcription of the utterance, and determining whether the predicted transcription includes the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance. When the predicted transcription includes the particular phrase, the method includes generating an output indicating that the trained ASR model leaked the particular phrase from a training data set used to train the ASR model.

Claims (42)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance;

modifying the audio data to obfuscate a particular phrase recited in the utterance;

processing, using a trained automated speech recognition (ASR) model, the modified audio data to generate a predicted transcription of the utterance;

determining whether the predicted transcription includes the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance; and

when the predicted transcription includes the particular phrase, generating an output indicating that the trained ASR model leaked the particular phrase from a training data set used to train the ASR model.

2. The computer-implemented method of claim 1 , wherein the operations further comprise, when the predicted transcription includes another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase, generating an output indicating that the trained ASR model leaked the other phrase from the training data set used to train the ASR model.

3. The computer-implemented method of claim 1 , wherein the operations further comprise, when the predicted transcription does not include the particular phrase or another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase, generating an output indicating that the trained ASR model has not leaked any information from the training data set used to train the ASR model.

4. The computer-implemented method of claim 1 , wherein the audio data comprises an audio waveform.

5. The computer-implemented method of claim 4 , wherein the audio waveform corresponds to human speech.

6. The computer-implemented method of claim 4 , wherein the audio waveform corresponds to synthesized speech.

7. The computer-implemented method of claim 1 , wherein modifying the audio data comprises:

based on the ground-truth transcription, identifying a segment of the audio data that aligns with the particular phrase in the ground-truth transcription; and

performing data augmentation on the identified segment of the audio data to obfuscate the particular phrase recited in the utterance.

8. The computer-implemented method of claim 7 , wherein performing data augmentation on the identified segment of the audio data comprises adding noise to the identified segment of the audio data.

9. The computer-implemented method of claim 7 , wherein performing data augmentation on the identified segment of the audio data comprises replacing the identified segment of the audio data with noise.

10. The computer-implemented method of claim 1 , wherein the operations further comprise:

processing the ground-truth transcription of the utterance to identify any phrases included in the ground-truth transcription that are associated with a specific category of information,

wherein modifying the audio data occurs in response to identifying that the particular phrase included in the ground-truth transcription is associated with the specific category of information.

11. The computer-implemented method of claim 10 , wherein the specific category of information comprises names, addresses, dates, zip codes, patient diagnosis, account numbers, or telephone numbers.

12. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

obtaining audio data characterizing an utterance and a corresponding ground-truth transcription of the utterance;

modifying the audio data to obfuscate a particular phrase recited in the utterance;

processing, using a trained automated speech recognition (ASR) model, the modified audio data to generate a predicted transcription of the utterance;

determining whether the predicted transcription includes the particular phrase by comparing the predicted transcription of the utterance to the ground-truth transcription of the utterance; and

when the predicted transcription includes the particular phrase, generating an output indicating that the trained ASR model leaked the particular phrase from a training data set used to train the ASR model.

13. The system of claim 12 , wherein the operations further comprise, when the predicted transcription includes another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase, generating an output indicating that the trained ASR model leaked the other phrase from the training data set used to train the ASR model.

14. The system of claim 12 , wherein the operations further comprise, when the predicted transcription does not include the particular phrase or another phrase substituted for the particular phrase from the ground-truth transcription that is associated with a same category of information as the particular phrase, generating an output indicating that the trained ASR model has not leaked any information from the training data set used to train the ASR model.

15. The system of claim 12 , wherein the audio data comprises an audio waveform.

16. The system of claim 15 , wherein the audio waveform corresponds to human speech.

17. The system of claim 15 , wherein the audio waveform corresponds to synthesized speech.

18. The system claim 12 , wherein modifying the audio data comprises:

based on the ground-truth transcription, identifying a segment of the audio data that aligns with the particular phrase in the ground-truth transcription; and

performing data augmentation on the identified segment of the audio data to obfuscate the particular phrase recited in the utterance.

19. The system of claim 18 , wherein performing data augmentation on the identified segment of the audio data comprises adding noise to the identified segment of the audio data.

20. The system of claim 18 , wherein performing data augmentation on the identified segment of the audio data comprises replacing the identified segment of the audio data with noise.

21. The system of claim 12 , wherein the operations further comprise:

processing the ground-truth transcription of the utterance to identify any phrases included in the ground-truth transcription that are associated with a specific category of information,

wherein modifying the audio data occurs in response to identifying that the particular phrase included in the ground-truth transcription is associated with the specific category of information.

22. The system of claim 21 , wherein the specific category of information comprises names, addresses, dates, zip codes, patient diagnosis, account numbers, or telephone numbers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 24, 2022
From: AMID, EHSAN; THAKKAR, OM; MATHEWS, RAJIV; BEAUFAYS, FRANCOISE
To: GOOGLE LLC
Reel/Frame 058747/0348 →
Continuity (2)
Provisional Application 63264836 · Dec 2, 2021
Related Publication 20230178094A1 · Jun 8, 2023
Cited By (2)
US 12,260,875 US 12,499,309