IP Library › Granted Patent US 11,663,488
Granted Patent B2
US 11,663,488 · App. 17/169,211 · Granted May 30, 2023

Initialization of parameters for machine-learned transformer neural network architectures

Inventors: Maksims Volkovs (Toronto, CA); Xiao Shi Huang (Toronto, CA); Juan Felipe Perez Vallejo (Toronto, CA)
Assignee: THE TORONTO-DOMINION BANK
G06N3/084G06F9/30036G06F9/3555G06F9/463G06F18/214G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,663,488
App. No.
17/169,211
Filed
Feb 5, 2021
Granted
May 30, 2023
Kind
B2
Art Unit
2181
USPC
712/7
Abstract

An online system trains a transformer architecture by an initialization method which allows the transformer architecture to be trained without normalization layers of learning rate warmup, resulting in significant improvements in computational efficiency for transformer architectures. Specifically, an attention block included in an encoder or a decoder of the transformer architecture generates the set of attention representations by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the set of attention representations. The initialization method may be performed by scaling the parameters of the value matrix and the output matrix with a factor that is inverse to a number of the set of encoders or a number of the set of decoders.

Claims (48)

1. A system comprising:

a processor configured to execute instructions;

a computer-readable medium containing instructions for execution on the processor, the instructions causing the processor to perform steps of:

accessing a machine-learned model including:

a set of encoders and a set of decoders, the set of encoders coupled to receive a sequence of input embeddings and generate an encoded output, and

the set of decoders coupled to receive a sequence of output embeddings and the encoded output and generate a prediction for a next word,

wherein at least one encoder of the set of encoders or at least one decoder of the set of decoders includes an attention block, the attention block coupled to receive an input key, an input query, an input value and generate an attention by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the attention;

initializing parameters of the machine-learned model including parameters of the value matrix and the output matrix;

scaling the parameters of the value matrix and the output matrix by multiplying a scaling factor that is inverse to a number of the set of encoders or a number of the set of decoders;

obtaining a set of training text, each training text including an ordered set of training input embeddings and an ordered set of training output embeddings; and

for each training text in the set:

generating one or more estimated output embeddings by applying the set of encoders and the set of decoders to the ordered set of training input embeddings, and

determining a loss function indicating a difference between the one or more estimated output embeddings and the ordered set of training output embeddings; and

updating the parameters of the machine-learned model to reduce the loss function for the training text in the set.

2. The system of claim 1 , wherein the at least one encoder or the at least one decoder does not include a normalization layer coupled to receive a set of inputs and normalize the set of inputs.

3. The system of claim 1 , wherein the at least one decoder includes the attention block, a second attention block placed after the attention block, and a multi-layer perceptron (MLP) block placed after the attention block,

wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder, and

wherein the second attention block is an encoder-decoder attention block coupled to receive a second input key that corresponds to the encoded output, a second input query that corresponds to an output generated from at least the attention block, and a second input value that corresponds to the encoded output.

4. The system of claim 3 , further comprising scaling initialized parameters of the MLP block with the scaling factor.

5. The system of claim 1 , wherein the at least one decoder includes the attention block, a second attention block placed after the attention block, and a multi-layer perceptron (MLP) block placed after the second attention block, and wherein the scaling factor is inverse to a number of residual blocks in the set of decoders.

6. The system of claim 1 , wherein the at least one encoder includes the attention block and a multi-layer perceptron (MLP) block placed after the attention block, and wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of input embeddings or an output of a previous encoder.

7. The system of claim 6 , further comprising scaling initialized parameters of the MLP block with the scaling factor.

8. The system of claim 1 , wherein the at least one encoder includes the attention block and a multi-layer perceptron (MLP) block placed after the attention block, and wherein the scaling factor is inverse to a number of residual blocks in the set of encoders.

9. The system of claim 1 , wherein initializing the parameters of the machine-learned model comprises initializing values of at least a portion of the parameters of the machine-learned model using a Xavier initialization method.

10. The system of claim 9 , wherein initializing the parameters of the machine-learned model further comprises sampling the values of the at least the portion of the parameters from a uniform distribution with a range of [−1/sqrt(n), 1/sqrt(n)], where n is a size of a previous neural network layer.

11. A method, comprising:

accessing a machine-learned model including:

a set of encoders and a set of decoders, the set of encoders coupled to receive a sequence of input embeddings and generate an encoded output, and

the set of decoders coupled to receive a sequence of output embeddings and the encoded output and generate a prediction for a next word,

wherein at least one encoder of the set of encoders or at least one decoder of the set of decoders includes an attention block, the attention block coupled to receive an input key, an input query, an input value and generate an attention by applying a key matrix to the input key, a query matrix to the input query, a value matrix to the input value to generate an output, and applying an output matrix to the output to generate the attention;

initializing parameters of the machine-learned model including parameters of the value matrix and the output matrix;

scaling the parameters of the value matrix and the output matrix by multiplying a scaling factor that is inverse to a number of the set of encoders or a number of the set of decoders;

obtaining a set of training text, each training text including an ordered set of training input embeddings and an ordered set of training output embeddings; and

for each training text in the set:

generating one or more estimated output embeddings by applying the set of encoders and the set of decoders to the ordered set of training input embeddings, and

determining a loss function indicating a difference between the one or more estimated output embeddings and the ordered set of training output embeddings; and

updating the parameters of the machine-learned model to reduce the loss function for the training text in the set.

12. The method of claim 11 , wherein the at least one encoder or the at least one decoder does not include a normalization layer coupled to receive a set of inputs and normalize the set of inputs.

13. The method of claim 11 , wherein the at least one decoder includes the attention block, a second attention block placed after the attention block, and a multi-layer perceptron (MLP) block placed after the attention block,

wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of output embeddings or an output of a previous decoder, and

wherein the second attention block is an encoder-decoder attention block coupled to receive a second input key that corresponds to the encoded output, a second input query that corresponds to an output generated from at least the attention block, and a second input value that corresponds to the encoded output.

14. The method of claim 13 , further comprising scaling initialized parameters of the MLP block with the scaling factor.

15. The method of claim 11 , wherein the at least one decoder includes the attention block, a second attention block placed after the attention block, and a multi-layer perceptron (MLP) block placed after the second attention block, and wherein the scaling factor is inverse to a number of residual blocks in the set of decoders.

16. The method of claim 11 , wherein the at least one encoder includes the attention block and a multi-layer perceptron (MLP) block placed after the attention block, and wherein the attention block is a self-attention block coupled to receive the input key, the input query, and the input value that each corresponds to the sequence of input embeddings or an output of a previous encoder.

17. The method of claim 16 , further comprising scaling initialized parameters of the MLP block with the scaling factor.

18. The method of claim 11 , wherein the at least one encoder includes the attention block and a multi-layer perceptron (MLP) block placed after the attention block, and wherein the scaling factor is inverse to a number of residual blocks in the set of encoders.

19. The method of claim 11 , wherein initializing the parameters of the machine-learned model comprises initializing values of at least a portion of the parameters of the machine-learned model using a Xavier initialization method.

20. The method of claim 19 , wherein initializing the parameters of the machine-learned model further comprises sampling the values of the at least the portion of the parameters from a uniform distribution with a range of [−1/sqrt(n), 1/sqrt(n)], where n is a size of a previous neural network layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2023
From: VOLKOVS, MAKSIMS; HUANG, XIAO SHI; VALLEJO, JUAN FELIPE PEREZ
To: THE TORONTO-DOMINION BANK
Reel/Frame 062803/0703 →
Continuity (2)
Provisional Application 62976040 · Feb 13, 2020
Related Publication 20210255862A1 · Aug 19, 2021
Cited By (30)
US 12,306,906 US 12,316,715 US 12,399,687 US 12,499,241 US 12,517,812 US 12,536,264 US 12,541,544 US 12,541,894 US 12,566,541 US 12,585,435 US 12,591,559 US 12,592,301 US 12,625,680 US 12,641,178 US 12,645,429 US 12,645,689 US 12,645,838 US 12,646,051 US 12,650,836 US 12,657,566 US 12,670,334 US 12,670,640 US 12,688,620 US 12,693,842 US 12,699,556 US 12,705,398 US 12,711,683 US 12,725,152 US 12,750,363 US 12,750,443