IP Library Granted Patent US 12,165,032
Granted Patent B2
US 12,165,032 · App. 16/586,813 · Granted Dec 10, 2024

Neural networks with area attention

Inventors: Yang Li (Palo Alto, CA); Lukasz Mieczyslaw Kaiser (Mountain View, CA); Samuel Bengio (Los Altos, CA); Si Si (San Jose, CA)
Assignee: Google LLC
G06N3/045G06F16/903G06N3/10G06F16/00G06F40/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,165,032
App. No.
16/586,813
Granted
Dec 10, 2024
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for implementing an area attention layer in a neural network system. The area attention layer area implements a way for a neural network model to attend to areas in the memory, where each area contains a group of items that are structurally adjacent.

Claims (62)

1. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to implement a neural network system configured to receive a neural network input and to generate a neural network output, the neural network system comprising:

an area attention layer, wherein the area attention layer is configured to, during processing of the neural network input:

receive data specifying a memory comprising a set of multiple items and, for each item, a respective key and a respective value;

determine a plurality of areas within the memory, wherein: (i) each area includes one or more items in the memory, (ii) one or more of the areas include multiple adjacent items in the memory, and (iii) the plurality of areas within the memory comprises one or more given areas that each include a respective number of items that is greater than one but less than a total number of items in the set of items;

determine, for each of the plurality of areas, a respective area key and a respective area value from at least the keys and values of the items in the area, comprising for each of the one or more given areas;

combining the keys for the items in the given area to generate the respective area key for the given area, wherein combining the keys comprises determining the area key based on a mean of the keys of the items in the given area; and

combining the values for the items the given area to generate the respective area value for the given area;

receive an attention query;

apply an attention mechanism between the attention query and the area keys for each area of the plurality of areas to generate a respective attention weight for each area of the plurality of areas, comprising, for each of the one or more given areas;

determining the respective attention weight for the given area based on a product between the attention query and the respective area key generated by combining the keys for the items in the given area; and

generate an area attention layer output by combining the area values for each area of the plurality of areas in accordance with the attention weights, comprising, for each of the one or more given areas;

combining the respective attention weight for the given area and the respective area value generated by combining the values for items in the given area.

2. The system of claim 1 , wherein the area attention layer is further configured to:

provide the area attention layer output to another component of the neural network system.

3. The system of claim 1 , wherein the memory is arranged as a sequence of items, and wherein determining a plurality of areas comprises:

identifying, as a different area, each combination of adjacent items that includes no more than a maximum number of items.

4. The system of claim 1 , wherein the memory is arranged as a two-dimensional grid of items, and wherein determining a plurality of areas comprises:

identifying, as a different area, each rectangular region of items within the two-dimensional grid that has no more than a maximum height and no more than a maximum width.

5. The system of claim 1 , wherein for each of the given areas, combining the values for the items in the given area comprises:

determining the area value for the given area to be a sum of the values of the items in the given area.

6. The system of claim 1 , wherein for each of the one or more given areas, combining the keys for the items in the given area comprises:

determining the area key for the given area to be the mean of the keys of the items in the given area.

7. The system of claim 1 , wherein the area key and area value for each area that includes only a single item are the key and value for the single item.

8. The system of claim 1 , wherein each item corresponds to a respective portion of the neural network input, and wherein the value for each item is an encoded representation of the respective portion of the neural network input.

9. The system of claim 1 , wherein (i) the values and keys for the items in the memory and (ii) the attention query are provided as input to the area attention layer by respective other components of the neural network system.

10. The system of claim 1 , wherein the key is the same as the value for each item in the memory.

11. The system of claim 1 , wherein the key and the value are different for each item in the memory.

12. The system of claim 1 , wherein for each of the one or more given areas, combining the keys for the items in the given area to generate the area key for the given area further comprises:

determining a variance of the keys for the items in the given area.

13. One or more computer storage media storing instructions that, when executed by one or more computers, cause the one or more computers to implement an area attention layer configured to perform operations comprising:

receiving data specifying a memory comprising a set of multiple of items and, for each item, a respective key and a respective value;

determining a plurality of areas within the memory, wherein: (i) each area includes one or more items in the memory, (ii) one or more of the areas include multiple adjacent items in the memory, and (iii) the plurality of areas within the memory comprises one or more given areas that each include a respective number of items that is greater than one but less than a total number of items in the set of items;

determining, for each of the plurality of areas, a respective area key and a respective area value from at least the keys and values of the items in the area, comprising for each of the one or more given areas;

combining the keys for the items in the given area to generate the respective area key for the given area, wherein combining the keys comprises determining the area key based on a mean of the keys of the items in the given area; and

combining the values for the items in the memory within the given area to generate the respective area value for the given area;

receiving an attention query;

applying an attention mechanism between the attention query and the area keys for each area of the plurality of areas to generate a respective attention weight for each area of the plurality of areas, comprising, for each of the one or more given areas;

determining the respective attention weight for the given area based on a product between the attention query and the respective area key generated by combining the keys for the items in the given area; and

generate an area attention layer output by combining the area values for each area of the plurality of areas in accordance with the attention weights, comprising, for each of the one or more given areas;

combining the respective attention weight for the given area and the respective area value generated by combining the values for items in the given area.

14. The one or more computer storage media of claim 13 , wherein the operations comprise:

providing the area attention layer output to another component of a neural network system.

15. The one or more computer storage media of claim 13 , wherein the memory is arranged as a sequence of items, and wherein determining a plurality of areas comprises:

identifying, as a different area, each combination of adjacent items that includes no more than a maximum number of items.

16. A method performed by one or more computers, the method comprising:

receive data specifying a memory comprising a set of multiple items and, for each item, a respective key and a respective value;

determine a plurality of areas within the memory, wherein: (i) each area includes one or more items in the memory, (ii) one or more of the areas include items in the memory, and (iii) the plurality of areas within the memory comprises one or more given areas that each include a respective number of items that is greater than one but less than a total number of items in the set of items;

determine, for each of the plurality of areas, a respective area key and a respective area value from at least the keys and values of the items in the area, comprising for each of the one or more given areas;

combining the keys for the items in the given area to generate the respective area key for the given area, wherein combining the keys comprises determining the area key based on a mean of the keys of the items in the given area; and

combining the values for the items in the given area to generate the respective area value for the given area;

receive an attention query;

apply an attention mechanism between the attention query and the area keys for each area of the plurality of areas to generate a respective attention weight for each area of the plurality of areas, comprising, for each of the one or more given areas;

determining the respective attention weight for the given area based on a product between the attention query and the respective area key generated by combining the keys for the items in the given area; and

generate an area attention layer output by combining the area values for each area of the plurality of areas in accordance with the attention weights, comprising, for each of the one or more given areas;

combining the respective attention weight for the given area and the respective area value generated by combining the values for items in the given area.

17. The method of claim 16 , wherein the memory is arranged as a sequence of items, and wherein determining a plurality of areas comprises:

identifying, as a different area, each combination of adjacent items that includes no more than a maximum number of items.

18. The method of claim 16 , wherein the memory is arranged as a two-dimensional grid of items, and wherein determining a plurality of areas comprises:

identifying, as a different area, each rectangular region of items within the two-dimensional grid that has no more than a maximum height and no more than a maximum width.

19. The method of claim 16 , further comprising providing the area attention layer output to another component of a neural network system.

20. The method of claim 16 , wherein for each of the given areas, combining the values for the items in the given area comprises:

determining the area value for the given area to be a sum of the values of the items in the given area.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2019
From: LI, YANG; KAISER, LUKASZ MIECZYSLAW; BENGIO, SAMUEL; SI, SI
To: GOOGLE LLC
Reel/Frame 050837/0842 →
Continuity (2)
Provisional Application 62737913 · Sep 27, 2018
Related Publication 20200104681A1 · Apr 2, 2020
Cited By (1)
US 12,511,486