IP Library › Granted Patent US 12,548,559
Granted Patent B1
US 12,548,559 · App. 18/341,159 · Granted Feb 10, 2026

Training neural network components

Inventors: I-Fan Chen (Sammamish, WA); Satya Venkata Phani Sankar Nidadavolu (Redmond, WA); Brian King (Bellingham, WA); Pegah Ghahremani (Seattle, WA); Pin-Jui Ku (Brookhaven, GA)
Assignee: Amazon Technologies, Inc.
G10L15/16G10L15/063G10L15/1815
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,559
App. No.
18/341,159
Granted
Feb 10, 2026
Kind
B1
Abstract

A machine learning model may be configured for training using an associated learning technique. A model configured for end-to-end backpropagation may adapted for associated learning by introducing functions for projecting hidden vectors and labels to a shared representation space and for reconstructing labels from representation vectors. An associated learning loss may be calculated at each layer, with the resulting gradients backpropagated locally through that layer rather than all layers. A reconstruction loss may be calculated using each layer's output including the predicted label. Training by associated learning may be parallelized (e.g., layer by layer) to yield efficiency gains. In addition, associated learning training may be more robust to training label errors. The resulting model may be used to, for example, predict data sequences in an autoregressive manner in which subsequent portions of the output data sequence are predicted in part based on previous predicted portions of the output data sequence.

Claims (87)

1 . A computer-implemented method for training a machine learning model, the method comprising:

receiving first input data representing a first sequence of tokens from a training dataset;

processing the first input data using a first plurality of components corresponding to a first layer of a first machine learning model to determine a first latent representation of a first portion of the first input data and a second latent representation of a second portion of the first input data, the second portion following the first portion in the first sequence;

determining, using a second plurality of components corresponding to a second layer of the first machine learning model, a third latent representation of the first latent representation and a fourth latent representation of the second latent representation;

processing the fourth latent representation using a first neural network component of the second plurality of components to determine a predicted second latent representation;

processing the second latent representation and the predicted second latent representation using a second neural network component of the first plurality of components to determine a predicted second portion of the first input data;

configuring a third plurality of components for a second machine learning model using:

a first associated learning loss determined using the first latent representation and the second latent representation, and

a first reconstruction loss determined using the second portion and the predicted second portion; and

configuring a fourth plurality of components for the second machine learning model using:

a second associated learning loss determined using the third latent representation and the fourth latent representation, and

a second reconstruction loss determined using the second latent representation and the predicted second latent representation, wherein the second machine learning model represents an update of the first machine learning model.

2 . The computer-implemented method of claim 1 , further comprising:

receiving second input data representing a second sequence of frames of audio data including speech;

processing the second input data using an encoder component to generate first data representing a hidden state representation of the second input data; and

processing the first data using the third plurality of components and the fourth plurality of components to generate first output data representing a transcript of the speech.

3 . The computer-implemented method of claim 1 , further comprising:

determining a third machine learning model that includes the third plurality of components but not the fourth plurality of components;

sending, to a device, first data representing the third machine learning model; and

causing the device to process second input data representing a second sequence of frames of audio data including speech using the third machine learning model to generate first output data representing a transcript of the speech.

4 . The computer-implemented method of claim 3 , further comprising:

determining, using a third component of the first plurality of components, a first stochastic term; and

determining, using a fourth component of the second plurality of components, a second stochastic term, wherein:

the second latent representation includes the first stochastic term, and

the fourth latent representation includes the second stochastic term.

5 . A computer-implemented method comprising:

receiving first input data representing speech;

processing the first input data using an encoder component to generate first encoded data representing a hidden state representation of the first input data;

processing, using first plurality of components of a first machine learning model, the first encoded data and a first portion of first output data to predict a second portion of the first output data, the first output data representing a transcript of the speech, wherein:

the first plurality of components are trained according to a first loss calculated using first data representing a latent representation of a first portion of training data and second data representing a latent representation of a second portion of the training data following the first portion, and

the first machine learning model includes a second plurality of components trained according to a second loss calculated using second third data representing a latent representation of the first data and fourth data representing a latent representation of the second data; and

sending the first output data to a first system component.

6 . The computer-implemented method of claim 5 , further comprising:

receiving, by the first system component, the first output data; and

performing, by the first system component, natural language understanding (NLU) processing to determine semantic content of the speech.

7 . The computer-implemented method of claim 5 , wherein the first input data represents a first natural language, and the first output data represents a second natural language, the method further comprising:

receiving, by the first system component, the first output data;

performing, by the first system component, text-to-speech processing to generate synthesized speech in the second natural language; and

causing a device to output the synthesized speech.

8 . The computer-implemented method of claim 5 , further comprising, prior to receiving the first input data:

processing, using a third plurality of components corresponding to a second machine learning model, the training data to determine the first data and the second data;

determining the first loss using the first data and the second data; and

determining the first plurality of components for the first machine learning model using the first loss, the first machine learning model representing an update of the second machine learning model.

9 . The computer-implemented method of claim 8 , further comprising:

determining, using a fourth plurality of components corresponding to the second machine learning model, the third data and the fourth data;

determining the second loss using the second data and the fourth data; and

determining the second plurality of components for the first machine learning model using the second loss.

10 . The computer-implemented method of claim 8 , further comprising:

determining, using the third plurality of components, a stochastic term, wherein the first data is additionally determined using the stochastic term.

11 . The computer-implemented method of claim 5 , further comprising, prior to receiving the first input data:

processing the second data using a first component of a second machine learning model to determine a predicted second portion of the training data;

determining a first reconstruction loss using the predicted second portion of the training data and the second portion of the training data; and

determining the first plurality of components using the first reconstruction loss, the first machine learning model representing an update of the second machine learning model.

12 . The computer-implemented method of claim 11 , further comprising:

processing the fourth data using a second component of the second machine learning model to determine fifth data representing a latent representation of the fourth data, wherein determining the predicted second portion of the training data additionally includes processing the fifth data using the first component of the second machine learning model.

13 . A system, comprising:

at least one processor; and

at least one memory comprising instructions that, when executed by the at least one processor, cause the system to:

receive first input data representing speech;

process the first input data using an encoder component to generate first encoded data representing a hidden state representation of the first input data;

process, using first plurality of components of a first machine learning model, the first encoded data and a first portion of first output data to predict a second portion of the first output data, the first output data representing a transcript of the speech, wherein:

the first plurality of components are trained according to a first loss calculated using first data representing a latent representation of a first portion of training data and second data representing a latent representation of a second portion of the training data following the first portion, and

the first machine learning model includes a second plurality of components trained according to a second loss calculated using third data representing a latent representation of the first data and fourth data representing a latent representation of the second data; and

send the first output data to a first system component.

14 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive, by the first system component, the first output data; and

perform, by the first system component, natural language understanding (NLU) processing to determine semantic content of the speech.

15 . The system of claim 13 , wherein the first input data represents a first natural language, the first output data represents a second natural language, and the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

receive, by the first system component, the first output data;

perform, by the first system component, text-to-speech processing to generate synthesized speech in the second natural language; and

cause a device to output the synthesized speech.

16 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

process, using a third plurality of components corresponding to a second machine learning model, the training data to determine the first data and the second data;

determine the first loss using the first data and the second data; and

determine the first plurality of components for the first machine learning model using the first loss, the first machine learning model representing an update of the second machine learning model.

17 . The system of claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determine, using a fourth plurality of components corresponding to the second machine learning model, the third data and the fourth data;

determine the second loss using the second data and the fourth data; and

determine the second plurality of components for the first machine learning model using the second loss.

18 . The system of claim 16 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

determining, using the third plurality of components, a stochastic term, wherein the first data is additionally determined using the stochastic term.

19 . The system of claim 13 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to, prior to receiving the first input data:

process the second data using a first component of a second machine learning model to determine a predicted second portion of the training data;

determine a first reconstruction loss using the predicted second portion of the training data and the second portion of the training data; and

determine the first plurality of components using the first reconstruction loss, the first machine learning model representing an update of the second machine learning model.

20 . The system of claim 19 , wherein the at least one memory further comprises instructions that, when executed by the at least one processor, further cause the system to:

process the fourth data using a second component of the second machine learning model to determine fifth data representing a latent representation of the fourth data, wherein determining the predicted second portion of the training data additionally includes processing the fifth data using the first component of the second machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 26, 2023
From: CHEN, I-FAN; NIDADAVOLU, SATYA VENKATA PHANI SANKAR; KING, BRIAN; GHAHREMANI, PEGAH; KU, PIN JUI
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 064058/0617 →
Continuity (1)
Provisional Application 63490332 · Mar 15, 2023
References Cited (28)
US 11735171B2 · Qian · 2023 [cited by examiner]
US 12112752B1 · Gupta · 2024 [cited by examiner]
US 12190862B2 · Rosenberg · 2025 [cited by examiner]
US 20210350786A1 · Chen · 2021 [cited by examiner]
US 20220122581A1 · Chen · 2022 [cited by examiner]
US 20220189456A1 · Pang · 2022 [cited by examiner]
US 20220246132A1 · Zhang · 2022 [cited by examiner]
US 20220366898A1 · Qian · 2022 [cited by examiner]
US 20230104228A1 · Li · 2023 [cited by examiner]
US 20230169281A1 · Zheng · 2023 [cited by examiner]
US 20230317059A1 · Rosenberg · 2023 [cited by examiner]
US 20230376734A1 · Liu · 2023 [cited by examiner]
US 20240062064A1 · Puzovic · 2024 [cited by examiner]
US 20240096077A1 · Luo · 2024 [cited by examiner]
Jinyu Li, et al. “An Overview of Noise-Robust Automatic Speech Recognition,” IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 22, No. 4, pp. 745-777, Apr. 2014. [cited by applicant]
Shane Settle, et al. “End-to-End Multi-Speaker Speech Recognition.” In ICASSP, Apr. 2018, 7 pages. Retrieved on Aug. 1, 2023 from https://www.merl.com/publications/docs/TR2018-001.pdf. [cited by applicant]
Tobias Menne, et al. “Analysis of Deep Clustering as Preprocessing for Automatic Speech Recognition of Sparsely Overlapping Speech.” In Proceedings of Interspeech (2019), 5 pages, https://arxiv.org/abs/1905.03500v2, Sep… [cited by applicant]
I-Fan Chen, et al. “Investigation of Training Label Error Impact on RNN-T,” (Dec. 2021), 8 pages, https://arxiv.org/abs/2112.00350v1. [cited by applicant]
Kartik Audhkhasi, et al. “Mixture Model Attention: Flexible Streaming and Non-Streaming Automatic Speech Recognition.” In Proceedings of Interspeech (2021), pp. 1812-1816. [cited by applicant]
Aleksei Kalinov, et al. “Carnelinet: Neural Mixture Model for Automatic Speech Recognition,” (Jul. 2021), 8 pages, https://arxiv.org/abs/2107.10708v1. [cited by applicant]
Katrin Tomanek, et al. “Residual Adapters for Parameter-Efficient ASR Adaptation to Atypical and Accented Speech,” In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, (2021), pp. 6… [cited by applicant]
Heng-Jui Chang, et al. “Towards Lifelong Learning of End-to-end ASR,” In Proceedings of Interspeech (2021), pp. 2552-2555. [cited by applicant]
Dennis YH Wu, et al. “Associated Learning: A Methodology to Decompose End-To-End Backpropagation on CNN, RNN, and Transformer,” Published as a Conference Paper at ICLR 2022, (2022), 18 pages. Retrieved from https://open… [cited by applicant]
Yu-Wei Kao and Hung-Hsuan Chen, “Associated Learning: Decomposing End-to-end Backpropagation based on Auto-encoders and Target Propagation,” (2021), 34 pages, https://arxiv.org/abs/1906.05560v4. [cited by applicant]
David Rumelhart, et al. “Learning representations by back-propagating errors.” Nature, vol. 323, pp. 533-536 (Oct. 1986). https://doi.org/10.1038/323533a0. [cited by applicant]
Yu-An Chung, et al. “Vector-Quantized Autoregressive Predictive Coding,” In Proceedings of Interspeech (2020), pp. 3760-3764. [cited by applicant]
Wei-Ning Hsu, et al. “Hubert: How Much Can a Bad Teacher Benefit ASR Pre-Training?,” ICASSP 2021—2021 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), (2021), pp. 6533-6537, doi: 10.110… [cited by applicant]
William Chan, et al. “Listen, Attend and Spell: A Neural Network for Large Vocabulary Conversational Speech Recognition,” 2016 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), 2016, pp.… [cited by applicant]