IP Library › Granted Patent US 12,321,861
Granted Patent B2
US 12,321,861 · App. 18/303,179 · Granted Jun 3, 2025

Initialization of parameters for machine-learned transformer neural network architectures

Inventors: Maksims Volkovs (Toronto, CA); Xiao Shi Huang (Toronto, CA); Juan Felipe Perez Vallejo (Toronto, CA)
Assignee: The Toronto-Dominion Bank
G06N3/084G06F9/30036G06F9/3555G06F9/463G06F18/214G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,321,861
App. No.
18/303,179
Granted
Jun 3, 2025
Kind
B2
Abstract

An online system trains a transformer architecture by an initialization method which allows the transformer architecture to be trained without normalization layers of learning rate warmup, resulting in significant improvements in computational efficiency for transformer architectures. Specifically, an attention block included in an encoder or a decoder of the transformer architecture generates the set of attention representations by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the set of attention representations. The initialization method may be performed by scaling the parameters of the value matrix and the output matrix with a factor that is inverse to a number of the set of encoders or a number of the set of decoders.

Claims (40)

1. A system comprising:

a processor configured to execute instructions;

a computer-readable medium containing instructions for execution on the processor, the instructions causing the processor to perform steps of:

accessing a machine-learned model including:

a set of blocks that each perform a respective operation on one or more inputs to the block to generate one or more outputs for the block, wherein one or more of the blocks are configured as attention blocks and an attention block is coupled to receive an input key, an input query, an input value and generate an attention by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the attention, and wherein at least a subset of the blocks are configured as residual blocks in which an input to a residual block is added to an output of the residual block;

initializing parameters of the machine-learned model including parameters of the value matrix and the output matrix for each attention block;

scaling the parameters of the value matrix and the output matrix by multiplying a scaling factor that depends on a number of the subset of residual blocks;

obtaining a set of training data, each training data instance including an ordered set of training input embeddings and an ordered set of training output embeddings;

determining a loss function indicating a difference between one or more estimated output embeddings and the ordered set of training output embeddings, the estimated output embeddings generated by applying the machine-learned model to the ordered set of training input embeddings; and

updating the parameters of the machine-learned model to reduce the loss function.

2. The system of claim 1 , wherein at least one attention block does not include a normalization layer coupled to receive a set of inputs and normalize the set of inputs.

3. The system of claim 1 , wherein the set of blocks are arranged as a set of decoders, and at least one decoder includes a first attention block, a second attention block, and a multi-layer perceptron (MLP) block placed after the first and second attention blocks,

wherein the first attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder, and wherein the second attention block is a cross-attention block coupled to receive the input query from the sequence of output embeddings or an output of a previous decoder, and the input key and the input value from an encoder.

4. The system of claim 3 , wherein the scaling factor is a number inverse to 3N+1, wherein N is the number of decoders in the set.

5. The system of claim 1 , wherein the set of blocks are arranged as a set of encoders, and at least one encoder includes an attention block and a multi-layer perceptron (MLP) block placed after the attention block,

wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder.

6. The system of claim 5 , wherein the scaling factor is a number inverse to 2N+1, where N is the number of encoders in the set.

7. The system of claim 1 , wherein the scaling factor is inverse to the number of residual blocks.

8. The system of claim 1 , wherein initializing the parameters of the machine-learned model comprises initializing values of at least a portion of the parameters of the machine-learned model using a Xavier initialization method.

9. The system of claim 8 , wherein initializing the parameters of the machine-learned model further comprises sampling the values of the at least the portion of the parameters from a uniform distribution with a range of [−1/sqrt (n), 1/sqrt (n)], where n is a size of a previous neural network layer.

10. The system of claim 1 , wherein the subset of blocks configured as residual blocks include the one or more blocks configured as attention blocks.

11. A method, comprising:

accessing, by a processor when executing instructions, a machine-learned model including:

a set of blocks that each perform a respective operation on one or more inputs to the block to generate one or more outputs for the block, wherein one or more of the blocks are configured as attention blocks and an attention block is coupled to receive an input key, an input query, an input value and generate an attention by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the attention, and wherein at least a subset of the blocks are configured as residual blocks in which an input to a residual block is added to an output of the residual block;

initializing, by the processor, parameters of the machine-learned model including parameters of the value matrix and the output matrix for each attention block;

scaling, by the processor, the parameters of the value matrix and the output matrix by multiplying a scaling factor that depends on a number of the subset of residual blocks;

obtaining, by the processor, a set of training data, each training data instance including an ordered set of training input embeddings and an ordered set of training output embeddings;

determining, by the processor, a loss function indicating a difference between the one or more estimated output embeddings and the ordered set of training output embeddings; and

updating, by the processor, the parameters of the machine-learned model to reduce the loss function for the training data in the set.

12. The method of claim 11 , wherein at least one attention block does not include a normalization layer coupled to receive a set of inputs and normalize the set of inputs.

13. The method of claim 11 , wherein the set of blocks are arranged as a set of decoders, and at least one decoder includes a first attention block, a second attention block, and a multi-layer perceptron (MLP) block placed after the first and second attention blocks,

wherein the first attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder, and wherein the second attention block is a cross-attention block coupled to receive the input query from the sequence of output embeddings or an output of a previous decoder, and the input key and the input value from an encoder.

14. The method of claim 13 , wherein the scaling factor is a number inverse to 3N+1, wherein N is the number of decoders in the set.

15. The method of claim 11 , wherein the set of blocks are arranged as a set of encoders, and at least one encoder includes an attention block and a multi-layer perceptron (MLP) block placed after the attention block,

wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder.

16. The method of claim 15 , wherein the scaling factor is a number inverse to 2N+1, where N is the number of encoders in the set.

17. The method of claim 11 , wherein the scaling factor is inverse to the number of residual blocks.

18. The method of claim 11 , wherein initializing the parameters of the machine-learned model comprises initializing values of at least a portion of the parameters of the machine-learned model using a Xavier initialization method.

19. The method of claim 18 , wherein initializing the parameters of the machine-learned model further comprises sampling the values of the at least the portion of the parameters from a uniform distribution with a range of [−1/sqrt (n), 1/sqrt (n)], where n is a size of a previous neural network layer.

20. The method of claim 11 , wherein the subset of blocks configured as residual blocks include the one or more blocks configured as attention blocks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 19, 2023
From: VOLKOVS, MAKSIMS; HUANG, XIAO SHI; VALLEJO, JUAN FELIPE PEREZ
To: THE TORONTO-DOMINION BANK
Reel/Frame 063381/0641 →
Continuity (3)
Continuation 17169211 · Feb 5, 2021
Provisional Application 62976040 · Feb 13, 2020
Related Publication 20230252301A1 · Aug 10, 2023
References Cited (32)
US 11126660B1 · Sen et al. · 2021 [cited by applicant]
US 20180075343A1 · van den Oord · 2018 [cited by examiner]
US 20180300400A1 · Paulus · 2018 [cited by applicant]
US 20180329897A1 · Kalchbrenner · 2018 [cited by examiner]
US 20190046068A1 · Ceccaldi et al. · 2019 [cited by applicant]
US 20190130213A1 · Shazeer et al. · 2019 [cited by applicant]
US 20190287012A1 · Celikyilmaz et al. · 2019 [cited by applicant]
US 20210381992A1 · Aguiar · 2021 [cited by applicant]
WO WO2018126325A1 · 2018 [cited by applicant]
WO WO2019220128A1 · 2019 [cited by applicant]
Bilkhu, M. et al., “Attention is All You Need for Videos: Self-Attention Based Video Summarization Using Universal Transformers,” arXiv: 1906.02792, Jun. 6, 2019, pp. 1-15. [cited by applicant]
Chen, M.X. et al., “The Best of Both Worlds: Combining Recent Advances in Neural Machine Translation,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Jul. 2018, pp. 76-86. [cited by applicant]
Chen, Q. et al., “Behavior Sequence Transformer for E-commerce Recommendation in Alibaba,” arXiv:1905.06874, May 15, 2019, pp. 1-4. [cited by applicant]
Devlin, J. et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv:1810.04805, Oct. 11, 2018, pp. 1-14. [cited by applicant]
Gehring, J. et al., “Convolutional Sequence to Sequence Learning,” Proceedings of the 34th International Conference on Machine Learning, Aug. 2017, pp. 1243-1252. [cited by applicant]
Glorot, X. et al., “Understanding the Difficulty of Training Deep Feedforward Neural Networks,” Proceedings of the 13th International Conference on Artificial Intelligence and Statistics (AISTATS), May 2010, pp. 249-256. [cited by applicant]
Lan, Z. et al., “ALBERT: A Lite BERT for Self-supervised Learning of Language Representations,” ICLR, Apr. 2020, pp. 1-17. [cited by applicant]
Liu, L. et al., “On the Variance of the Adaptive Learning Rate and Beyond,” arXiv:1908.03265, Aug. 8, 2019, pp. 1-14. [cited by applicant]
Ma, J. et al., “On the Adequacy of Untuned Warmup for Adaptive Optimization,” arXiv:1910.04209, Oct. 9, 2019, pp. [cited by applicant]
Nguyen, T.Q. et al., “Transformers without Tears: Improving the Normalization of Self-Attention,” arXiv:1910.05895, Dec. 30, 2019, pp. [cited by applicant]
Ott, M. et al., “Scaling Neural Machine Translation,” Proceedings of the Third Conference on Machine Translation (WMT), vol. 1: Research Papers, Oct. 2018, pp. 1-9. [cited by applicant]
PCT International Search Report and Written Opinion, PCT Application No. PCT/CA2021/050130, Apr. 26, 2021, 8 pages. [cited by applicant]
Popel, M. et al., “Training Tips for the Transformer Model,” The Prague Bulletin of Mathematical Linguistics, No. 110, Apr. 2018, pp. 43-70. [cited by applicant]
Rikters, M., “Impact of Corpora Quality on Neural Machine Translation,” arXiv:1810.08392, Oct. 19, 2018, pp. 1-8. [cited by applicant]
Stanojević, M. et al., “BEER: Better Evaluation as Ranking,” Proceedings of the Ninth Workshop on Statistical Machine Translation, Jun. 2014, pp. 414-419. [cited by applicant]
Vaswani, A. et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Dec. 2017, pp. 1-11. [cited by applicant]
Wang, Q. et al., “Learning Deep Transformer Models for Machine Translation,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Jul. 2019, pp. 1810-1822. [cited by applicant]
Xiong, R. et al., “On Layer Normalization in the Transformer Architecture,” Proceedings of the 37 [cited by applicant]
Yang, Z. et al., “XLNet: Generalized Autoregressive Pretraining for Language Understanding,” 33 [cited by applicant]
Zhang, H. et al., “Fixup Initialization: Residual Learning Without Normalization,” International Conference on Learning Representations, May 2019, pp. 1-16. [cited by applicant]
Zhang, W. et al., “Bridging the Gap between Training and Inference for Neural Machine Translation,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics, Aug. 2019, pp. 4334-4343. [cited by applicant]
United States Office Action, U.S. Appl. No. 17/169,211, filed Oct. 6, 2022, 9 pages. [cited by applicant]
Cited By (24)
US 12,499,241 US 12,517,812 US 12,536,264 US 12,541,894 US 12,585,435 US 12,592,301 US 12,625,680 US 12,641,178 US 12,645,429 US 12,645,689 US 12,645,838 US 12,646,051 US 12,650,836 US 12,657,566 US 12,670,334 US 12,670,640 US 12,682,179 US 12,693,842 US 12,699,556 US 12,705,398 US 12,711,683 US 12,725,152 US 12,750,363 US 12,750,443