IP Library › Granted Patent US 12,334,059
Granted Patent B2
US 12,334,059 · App. 18/619,684 · Granted Jun 17, 2025

Contrastive Siamese network for semi-supervised speech recognition

Inventors: Jaeyoung Kim (Cupertino, CA); Soheil Khorram (Redwood City, CA); Hasim Sak (Santa Clara, CA); Anshuman Tripathi (Mountain View, CA); Han Lu (Redmond, WA); Qian Zhang (Mountain View, CA)
Assignee: Google LLC
G10L15/16G06N3/088G10L15/1815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,059
App. No.
18/619,684
Granted
Jun 17, 2025
Kind
B2
Abstract

A method includes receiving a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions. At a target branch of a contrastive Siamese network, the method also includes generating a sequence of encoder outputs for the plurality of unlabeled audio samples and modifying time characteristics of the encoder outputs to generate a sequence of target branch outputs. At an augmentation branch of a contrastive Siamese network, the method also includes performing augmentation on the unlabeled audio samples, generating a sequence of augmented encoder outputs for the augmented unlabeled audio samples, and generating predictions of the sequence of target branch outputs generated at the target branch. The method also includes determining an unsupervised loss term based on target branch outputs and predictions of the sequence of target branch outputs. The method also includes updating parameters of the audio encoder based on the unsupervised loss term.

Claims (64)

1. A computer-implemented method when executed on data processing hardware causes the data processing hardware to perform operations comprising:

receiving training samples comprising:

a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions; and

a plurality of labeled audio samples corresponding to spoken utterances paired with corresponding transcriptions:

executing a semi-supervised training process for training a speech recognition model, the semi-supervised training process comprising an unsupervised subnetwork training process and a supervised subnetwork training process;

during execution of the unsupervised subnetwork training process:

performing augmentation on the unlabeled audio samples;

generating, using an audio encoder of the speech recognition model, a sequence of augmented encoder outputs for the augmented unlabeled audio samples;

generating, using a prediction network configured to receive the sequence of augmented encoder outputs, predictions of a sequence of target branch outputs; and

determining an unsupervised loss term based on the sequence target branch outputs and the predictions of the sequence of target branch outputs;

during execution of the supervised subnetwork training process:

generating, using the speech recognition model, speech recognition results for the labeled audio samples; and

determining a supervised loss term based on the speech results for the labeled audio samples and the corresponding transcriptions of the labeled audio samples; and

updating parameters of the speech recognition model based on the unsupervised loss term and the supervised loss term.

2. The computer-implemented method of claim 1 , wherein the speech recognition results generated for the labeled audio samples using the speech recognition model comprises a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step.

3. The computer-implemented method of claim 1 , wherein the operations further comprise updating parameters of the speech recognition model based on the supervised loss term independently of updating parameters of the speech recognition model based on the unsupervised loss term.

4. The computer-implemented method of claim 1 , wherein the operations further comprise applying data augmentation to at least one of the labeled audio samples.

5. The computer-implemented method of claim 4 , wherein applying data augmentations comprises at least one of adding noise, adding reverberation, or manipulating timing.

6. The computer-implemented method of claim 1 , wherein the unsupervised loss term comprises a contrastive loss term.

7. The computer-implemented method of claim 1 , wherein performing augmentation on the unlabeled audio samples comprises performing time modification and masking on the unlabeled audio samples.

8. The computer-implemented method of claim 1 , wherein the operations further comprise generating, as output from the audio encoder, a higher order feature representation for the plurality of unlabeled audio samples.

9. The computer-implemented method of claim 1 , wherein the speech recognition model comprises a Transformer-Transducer (T-T) model and the operations further comprise:

receiving, as input to the audio encoder of the T-T model, a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions;

generating, by the audio encoder, at each of a plurality of time steps, a sequence of acoustic frames extracted from audio data characterizing a spoken utterance;

receiving, as input to a label encoder of the T-T model, a sequence of non-blank symbols output by a final softmax layer;

generating, by the label encoder, at each of the plurality of time steps, a dense representation;

receiving, as input to a joint network of the T-T model, the higher order feature representation generated by the audio encoder at each of the plurality of time steps and the dense representation generated by the label encoder at each of the plurality of time steps; and

generating, by the joint network, at each of the plurality of time steps, a probability distribution over possible speech recognition hypothesis at the corresponding time step.

10. The computer-implemented method of claim 1 , wherein the audio encoder comprises:

a stack of strided convolutional layers and transformer layers; or

a stack of conformer layers.

11. A system comprising:

data processing hardware; and

memory hardware storing instructions that when executed on the data processing hardware causes the data processing hardware to perform operations comprising:

receiving training samples comprising:

a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions; and

a plurality of labeled audio samples corresponding to spoken utterances paired with corresponding transcriptions:

executing a semi-supervised training process for training a speech recognition model, the semi-supervised training process comprising an unsupervised subnetwork training process and a supervised subnetwork training process;

during execution of the unsupervised subnetwork training process:

performing augmentation on the unlabeled audio samples;

generating, using an audio encoder of the speech recognition model, a sequence of augmented encoder outputs for the augmented unlabeled audio samples;

generating, using a prediction network configured to receive the sequence of augmented encoder outputs, predictions of a sequence of target branch outputs; and

determining an unsupervised loss term based on the sequence target branch outputs and the predictions of the sequence of target branch outputs;

during execution of the supervised subnetwork training process:

generating, using the speech recognition model, speech recognition results for the labeled audio samples; and

determining a supervised loss term based on the speech results for the labeled audio samples and the corresponding transcriptions of the labeled audio samples; and

updating parameters of the speech recognition model based on the unsupervised loss term and the supervised loss term.

12. The system of claim 11 , wherein the speech recognition results generated for the labeled audio samples using the speech recognition model comprises a probability distribution over possible speech recognition hypotheses for the labeled audio sample at the corresponding output step.

13. The system of claim 11 , wherein the operations further comprise updating parameters of the speech recognition model based on the supervised loss term independently of updating parameters of the speech recognition model based on the unsupervised loss term.

14. The system of claim 11 , wherein the operations further comprise applying data augmentation to at least one of the labeled audio samples.

15. The system of claim 14 , wherein applying data augmentations comprises at least one of adding noise, adding reverberation, or manipulating timing.

16. The system of claim 11 , wherein the unsupervised loss term comprises a contrastive loss term.

17. The system of claim 11 , wherein performing augmentation on the unlabeled audio samples comprises performing time modification and masking on the unlabeled audio samples.

18. The system of claim 11 , wherein the operations further comprise generating, as output from the audio encoder, a higher order feature representation for the plurality of unlabeled audio samples.

19. The system of claim 11 , wherein the speech recognition model comprises a Transformer-Transducer (T-T) model and the operations further comprise:

receiving, as input to the audio encoder of the T-T model, a plurality of unlabeled audio samples corresponding to spoken utterances not paired with corresponding transcriptions;

generating, by the audio encoder, at each of a plurality of time steps, a sequence of acoustic frames extracted from audio data characterizing a spoken utterance;

receiving, as input to a label encoder of the T-T model, a sequence of non-blank symbols output by a final softmax layer;

generating, by the label encoder, at each of the plurality of time steps, a dense representation;

receiving, as input to a joint network of the T-T model, the higher order feature representation generated by the audio encoder at each of the plurality of time steps and the dense representation generated by the label encoder at each of the plurality of time steps; and

generating, by the joint network, at each of the plurality of time steps, a probability distribution over possible speech recognition hypothesis at the corresponding time step.

20. The system of claim 11 , wherein the audio encoder comprises:

a stack of strided convolutional layers and transformer layers; or

a stack of conformer layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: KIM, JAEYOUNG; KHORRAM, SOHEIL; SAK, HASIM; TRIPATHI, ANSHUMAN; LU, HAN; ZHANG, QIAN
To: GOOGLE LLC
Reel/Frame 066933/0515 →
Continuity (3)
Continuation 17644337 · Dec 14, 2021
Provisional Application 63261895 · Sep 30, 2021
Related Publication 20240242712A1 · Jul 18, 2024
References Cited (13)
US 10332509B2 · Catanzaro et al. · 2019 [cited by applicant]
US 11961515B2 · Kim · 2024 [cited by examiner]
US 20190325275A1 · Lee · 2019 [cited by examiner]
US 20200027444A1 · Prabhavalkar et al. · 2020 [cited by applicant]
US 20210056417A1 · Zhang · 2021 [cited by examiner]
US 20210089964A1 · Zhang · 2021 [cited by examiner]
US 20210133623A1 · Amrani · 2021 [cited by examiner]
US 20230096805A1 · Kim · 2023 [cited by examiner]
International Search Report and Written Opinion for the related application No. PCT/US2021/063417, dated May 31, 2022, 56 pages. [cited by applicant]
Chen Xinlei et al: “Exploring Simple Siamese Representation Learning”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 20, 2021 (Jun. 20, 2021), pp. 15745-15753, XP034006641, DOI: … [cited by applicant]
Herman Kamper et al: “Improved acoustic word embeddings for zer•• resource languages using multilingual transfer”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 2, 202… [cited by applicant]
Kamper Herman et al: “Deep convolutional acoustic word embeddings using word-pair side information”, 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, Mar. 20, 2016 (Mar. 20, … [cited by applicant]
Lain et al., “Speech Emotion Recognition via Contrastive Loss under Siamese Networks”; Oct. 26, 2018; ASMMC-MMAC'18; pp. 21-26 (Year: 2018). [cited by applicant]