IP Library › Granted Patent US 11,403,486
Granted Patent B2
US 11,403,486 · App. 17/095,257 · Granted Aug 2, 2022

Methods and systems for training convolutional neural network using built-in attention

Inventors: Niamul Quader (Toronto, CA); Md Ibrahim Khalil (Toronto, CA); Juwei Lu (North York, CA); Peng Dai (Markham, CA); Wei Li (Markham, CA)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06K9/6256G06N3/04G06N3/08G06V30/194
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,403,486
App. No.
17/095,257
Granted
Aug 2, 2022
Kind
B2
Abstract

Methods and systems for updating the weights of a set of convolution kernels of a convolutional layer of a neural network are described. A set of convolution kernels having attention-infused weights is generated by using an attention mechanism based on characteristics of the weights. For example, a set of location-based attention multipliers is applied to weights in the set of convolution kernels, a magnitude-based attention function is applied to the weights in the set of convolution kernels, or both. An output activation map is generated using the set of convolution kernels with attention-infused weights. A loss for the neural network is computed, and the gradient is back propagated to update the attention-infused weights of the convolution kernels.

Claims (124)

1. A method for updating weights of a set of convolution kernels of a convolutional layer of a neural network during training of the neural network, the method comprising:

obtaining the set of convolution kernels of the convolutional layer;

generating a set of convolution kernels having attention-infused weights by performing at least one of:

applying a set of location-based attention multipliers to weights in the set of convolution kernels; or

applying a magnitude-based attention function to the weights in the set of convolution kernels;

performing convolution on an input activation map using the set of convolution kernels with attention-infused weights to generate an output activation map; and

updating the attention-infused weights in the set of convolution kernels using a back propagated gradient of a loss computed for the neural network.

2. The method of claim 1 , wherein the set of location-based attention multipliers is applied to the weights in the set of convolution kernels, to obtain a set of location-excited weights, and wherein the magnitude-based attention function is applied to the set of location-excited weights.

3. The method of claim 1 , further comprising:

prior to computing the loss for the neural network, applying a channel-based attention function to the output activation map.

4. The method of claim 1 , wherein applying the set of location-based attention multipliers further comprises:

learning the set of location-based attention multipliers.

5. The method of claim 4 , wherein learning the set of location-based attention multipliers comprises:

performing average pooling to obtain an averaged weight for each convolution kernel;

feeding the averaged weights of the convolution kernels through one or more fully connected layers, to learn the attention multiplier for each convolution kernel; and

expanding the attention multiplier across all weights in each respective convolution kernel to obtain the set of location-based attention multipliers.

6. The method of claim 5 , wherein feeding the averaged weights of the convolution kernels through the one or more fully connected layers comprises:

feeding the averaged weights of the convolution kernels through a first fully connected layer;

applying, to an output of the first fully connected layer, a first activation function;

feeding an output of the first activation function to a second fully connected layer; and

applying, to an output of the second fully connected layer, a second activation function.

7. The method of claim 1 , wherein the magnitude-based attention function applies greater attention to weights of greater magnitude, and lesser attention to weights of lesser magnitude.

8. The method of claim 7 , wherein the magnitude-based attention function is

w

A

=

f

A

⁡

(

w

m

)

=

M

A

*

0.5

*

ln

⁢

1

+

w

m

/

M

A

1

-

w

m

/

M

A

where w m is a weight for a convolution kernel, W A is the weight after applying magnitude-based attention, M A =(1+∈ A )*M, M is the maximum of all W m in a convolutional layer and ∈ A is a hyperparameter with a selected small value.

9. The method of claim 1 , further comprising:

prior to applying the set of location-based attention multipliers or the magnitude-based attention function, standardizing the weights in the set of convolution kernels.

10. A processing system comprising a processing device and a memory storing instructions which, when executed by the processing device, cause the processing system to update weights of a set of convolution kernels of a convolutional layer of a convolutional neural network during training of the neural network by:

obtaining the set of convolution kernels of the convolutional layer;

generating a set of convolution kernels having attention-infused weights by performing at least one of:

applying a set of location-based attention multipliers to weights in the set of convolution kernels; or

applying a magnitude-based attention function to the weights in the set of convolution kernels;

performing convolution on an input activation map using the set of convolution kernels with attention-infused weights to generate an output activation map; and

updating the attention-infused weights in the set of convolution kernels using a back propagated gradient of a loss computed for the neural network.

11. The processing system of claim 10 , wherein the set of location-based attention multipliers is applied to the weights in the set of convolution kernels, to obtain a set of location-excited weights, and wherein the magnitude-based attention function is applied to the set of location-excited weights.

12. The processing system of claim 10 , wherein the instructions further cause the processing system to:

prior to computing the loss for the neural network, apply a channel-based attention function to the output activation map.

13. The processing system of claim 10 , wherein the instructions further cause the processing system to apply the set of location-based attention multipliers further by:

learning the set of location-based attention multipliers.

14. The processing system of claim 13 , wherein the instructions further cause the processing system to learn the set of location-based attention multipliers by:

performing average pooling to obtain an averaged weight for each convolution kernel;

feeding the averaged weights of the convolution kernels through one or more fully connected layers, to learn the attention multiplier for each convolution kernel; and

expanding the attention multiplier across all weights in each respective convolution kernel to obtain the set of location-based attention multipliers.

15. The processing system of claim 14 , wherein the instructions further cause the processing system to feed the averaged weights of the convolution kernels through the one or more fully connected layers by:

feeding the averaged weights of the convolution kernels through a first fully connected layer;

applying, to an output of the first fully connected layer, a first activation function;

feeding an output of the first activation function to a second fully connected layer; and

applying, to an output of the second fully connected layer, a second activation function.

16. The processing system of claim 10 , wherein the magnitude-based attention function applies greater attention to weights of greater magnitude, and lesser attention to weights of lesser magnitude.

17. The processing system of claim 16 , wherein the magnitude-based attention function is

w

A

=

f

A

⁡

(

w

m

)

=

M

A

*

0.5

*

ln

⁢

1

+

w

m

/

M

A

1

-

w

m

/

M

A

where w m is a weight for a convolution kernel, w A is the weight after applying magnitude-based attention, M A =(1+∈ A )*M, M is the maximum of all w m in a convolutional layer and ∈ A is a hyperparameter with a selected small value.

18. The processing system of claim 10 , wherein the instructions further cause the processing system to:

prior to applying the set of location-based attention multipliers or the magnitude-based attention function, standardize the weights in the set of convolution kernels.

19. A non-transitory computer-readable medium having instructions tangibly stored thereon, wherein the instructions, when executed by a processing device of a processing system, causes the processing system to update weights of a set of convolution kernels of a convolutional layer of a convolutional neural network during training of the neural network by:

obtaining the set of convolution kernels of the convolutional layer;

generating a set of convolution kernels having attention-infused weights by performing at least one of:

applying a set of location-based attention multipliers to weights in the set of convolution kernels; or

applying a magnitude-based attention function to the weights in the set of convolution kernels;

performing convolution on an input activation map using the set of convolution kernels with attention-infused weights to generate an output activation map; and

updating the attention-infused weights in the set of convolution kernels using a back propagated gradient of a loss computed for the neural network.

20. The non-transitory computer-readable medium of claim 19 , wherein the set of location-based attention multipliers is applied to the weights in the set of convolution kernels, to obtain a set of location-excited weights, and wherein the magnitude-based attention function is applied to the set of location-excited weights.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2021
From: QUADER, NIAMUL; KHALIL, MD IBRAHIM; LU, JUWEI; DAI, PENG; LI, WEI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 057725/0873 →
Continuity (2)
Provisional Application 62934744 · Nov 13, 2019
Related Publication 20210142106A1 · May 13, 2021
Cited By (1)
US 12,210,942