IP Library › Granted Patent US 12,705,489
Granted Patent B2
US 12,705,489 · App. 17/306,813 · Granted Aug 11, 2026

Attention-based neural networks with branching blocks

Inventors: David Martin Dohan (San Francisco, CA); David Richard So (San Francisco, CA); Chen Liang (Stanford, CA); Quoc V. Le (Sunnyvale, CA)
Assignee: Google LLC
G06N3/086G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,705,489
App. No.
17/306,813
Filed
May 3, 2021
Granted
Aug 11, 2026
Kind
B2
Art Unit
2128
USPC
706/15
Abstract

A method for receiving training data for training a neural network to perform a machine learning task and for searching for, using the training data, an optimized neural network architecture for performing the machine learning task is described. Searching for the optimized neural network architecture includes: maintaining population data; maintaining threshold data; and repeatedly performing the following operations: selecting one or more candidate architectures from the population data; generating a new architecture from the one or more selected candidate architectures; for the new architecture: training a neural network having the new architecture until termination criteria for the training are satisfied; and determining a final measure of fitness of the neural network having the new architecture after the training; and adding data defining the new architecture and the final measure of fitness for the neural network having the new architecture to the population data.

Claims (46)

1 . A system for performing a machine learning task on a network input to generate a network output, the system comprising one or more computers and one or more storage devices storing instructions that, when executed by the one or more computers, cause the one or more computers to implement:

an attention neural network configured to perform the machine learning task, the attention neural network comprising a plurality of operation blocks, the plurality of operation blocks comprising a plurality of branching attention blocks, each respective branching attention block comprising a respective first branch that comprises a respective first attention layer that comprises a respective first number of attention heads and a respective second branch that comprises a respective second attention layer that comprises a respective second number of attention heads, each respective branching attention block configured to:

receive a branching attention block input;

provide the branching attention block input to the respective first branch included in the respective branching attention block and to the respective second branch included in the respective branching attention block;

generate, by the respective first attention layer included in the respective first branch included in the respective branching attention block, a first attention layer output by applying an attention mechanism on a first attention layer input using each attention head in the respective first number of attention heads included in the respective first attention layer, wherein the first attention layer input is derived from the branching attention block input;

generate, by the respective second attention layer included in the respective second branch included in the respective branching attention block, a second attention layer output by applying an attention mechanism on a second attention layer input using each attention head in the respective second number of attention heads included in the second attention layer, wherein the second attention layer input is derived from the branching attention block input;

determine a branching attention block output based on determining a combination of the first attention layer output and the second attention layer output; and

provide the branching attention block output as input to a subsequent block in the plurality of operation blocks included in the attention neural network,

wherein, within each respective branching attention block, the respective first number of attention heads of the respective first attention layer included in the respective first branch is different from the respective second number of attention heads of the respective second attention layer included in the respective second branch, and

wherein the respective first number of attention heads of the respective first attention layer included in the respective first branch in each respective branching attention block of the plurality of branching attention blocks is a same as the respective first number of attention heads of the respective first attention layer included in the respective first branch in another respective branching attention block of the plurality of branching attention blocks.

2 . The system of claim 1 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more depth-wise separable convolutional neural network layers.

3 . The system of claim 1 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more gated linear unit (GLU) neural network layers.

4 . The system of claim 1 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more swish activation neural network layers.

5 . The system of claim 1 , wherein the respective first attention layer included in the respective first branch is different than the respective second attention layer included in the respective second branch.

6 . The system of claim 1 , wherein determining the combination comprises computing an addition between the first attention layer output and the second attention layer output.

7 . The system of claim 1 , wherein the respective first attention layer is configured to generate the first attention layer output at least in part by applying a multi-head attention mechanism that uses the respective first number of attention heads to the first attention input layer, wherein the respective second attention layer is configured to generate the second attention layer output at least in part by applying a multi-head attention mechanism that uses the respective second number of attention heads to the second attention layer input, and wherein the first number is greater than the second number.

8 . The system of claim 1 , wherein the network input is an input sequence that includes a respective input at each of multiple positions in an input order, the network output is an output sequence that includes a respective output at each of multiple positions in an output order, or both.

9 . The system of claim 1 , wherein the network input comprises text data, image data, or audio data.

10 . A computer-implemented method comprising, at each respective branching attention block in a plurality of branching attention blocks in an attention neural network:

receiving, at the respective branching attention block, a branching attention block input, wherein the respective branching attention block comprises a respective first branch that comprises a respective first attention layer that comprises a respective first number of attention heads and a respective second branch that comprises a respective second attention layer that comprises a respective second number of attention heads;

providing, by the respective branching attention block, the branching attention block input to the respective first branch included in the respective branching attention block and to the respective second branch included in the respective branching attention block;

generating, by the respective first attention layer included in the respective first branch included in the respective branching attention block, a first attention layer output by applying an attention mechanism on a first attention layer input using each attention head in the respective first number of attention heads included in the respective first attention layer, wherein the first attention layer input is derived from the branching attention block input;

generating, by the respective second attention layer included in the respective second branch included in the respective branching attention block, a second attention layer output by applying an attention mechanism on a second attention layer input using each attention head in the respective second number of attention heads included in the respective second attention layer, wherein the second attention layer input is derived from the branching attention block input;

determining, by the respective branching attention block, a branching attention block output based on determining a combination of the first attention layer output and second attention layer output; and

providing, by the branching attention block, the branching attention block output as input to a subsequent block in a plurality of operation blocks included in the attention neural network,

wherein, within each respective branching attention block, the respective first number of attention heads of the respective first attention layer included in the respective first branch is different from the respective second number of attention heads of the respective second attention layer included in the respective second branch, and

wherein the respective first number of attention heads of the respective first attention layer included in the respective first branch in each respective branching attention block of the plurality of branching attention blocks is a same as the respective first number of attention heads of the respective first attention layer included in the respective first branch in another respective branching attention block of the plurality of branching attention blocks.

11 . The method of claim 10 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more depth-wise separable convolutional neural network layers.

12 . The method of claim 10 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more gated linear unit (GLU) neural network layers.

13 . The method of claim 10 , wherein at least one of the plurality of operation blocks included in the attention neural network comprises one or more swish activation neural network layers.

14 . The method of claim 10 , wherein the respective first attention layer included in the respective first branch is different than the respective second attention layer included in the respective second branch.

15 . The method of claim 10 , wherein determining the combination comprises computing an addition between the first attention layer output and the second attention layer output.

16 . The method of claim 14 , wherein the first attention layer is configured to generate the first attention layer output at least in part by applying a multi-head attention mechanism that uses the first number of attention heads to the first attention layer input, wherein the second attention layer is configured to generate the second attention layer output at least in part by applying a multi-head attention mechanism that uses the second number of attention heads to the second attention layer input, and wherein the first number is greater than the second number.

17 . The method of claim 10 , wherein the network input is an input sequence that includes a respective input at each of multiple positions in an input order, the network output is an output sequence that includes a respective output at each of multiple positions in an output order, or both.

18 . The method of claim 10 , wherein the network input comprises text data, image data, or audio data.

19 . One or more non-transitory computer storage media storing instructions that when executed by one or more computers cause the one or more computers to implement:

an attention neural network configured to perform the machine learning task, the attention neural network comprising a plurality of operation blocks, the plurality of operation blocks comprising a plurality of branching attention blocks, each respective branching attention block comprising a respective first branch that comprises a respective first attention layer that comprises a respective first number of attention heads and a respective second branch that comprises a respective second attention layer that comprises a respective second number of attention heads, each respective branching attention block configured to:

receive a branching attention block input;

provide the branching attention block input to the respective first branch included in the respective branching attention block and to the respective second branch included in the respective branching attention block;

generate, by the respective first attention layer included in the respective first branch included in the respective branching attention block, a first attention layer output by applying an attention mechanism on a first attention layer input using each attention head in the respective first number of attention heads included in the respective first attention layer, wherein the first attention layer input is derived from the branching attention block input;

generate, by the respective second attention layer included in the respective second branch included in the respective branching attention block, a second attention layer output by applying an attention mechanism on a second attention layer input using each attention head in the respective second number of attention heads included in the second attention layer, wherein the second attention layer input is derived from the branching attention block input;

determine a branching attention block output based on determining a combination of the first attention layer output and the second attention layer output; and

provide the branching attention block output as input to a subsequent block in the plurality of operation blocks included in the attention neural network,

wherein, within each respective branching attention block, the respective first number of attention heads of the respective first attention layer included in the respective first branch is different from the respective second number of attention heads of the respective second attention layer included in the respective second branch, and

wherein the respective first number of attention heads of the respective first attention layer included in the respective first branch in each respective branching attention block of the plurality of branching attention blocks is a same as the respective first number of attention heads of the respective first attention layer included in the respective first branch in another respective branching attention block of the plurality of branching attention blocks.

20 . The computer storage media of claim 19 , wherein the network input comprises text data, image data, or audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2021
From: DOHAN, DAVID MARTIN; SO, DAVID RICHARD; LIANG, CHEN; LE, QUOC V.
To: GOOGLE LLC
Reel/Frame 056131/0228 →
Continuity (2)
Continuation 16447866 · Jun 20, 2019
Related Publication 20210256390A1 · Aug 19, 2021
References Cited (117)
US 10997503B2 · Dohan · 2021 [cited by examiner]
US 20020046143A1 · Eder · 2002 [cited by applicant]
US 20080154809A1 · Stockwell · 2008 [cited by applicant]
US 20110029467A1 · Spehr · 2011 [cited by applicant]
US 20110087625A1 · Tanner · 2011 [cited by applicant]
US 20140243913A1 · Lineaweaver · 2014 [cited by applicant]
US 20160071520A1 · Hayakawa · 2016 [cited by applicant]
US 20170154260A1 · Hamada et al. · 2017 [cited by applicant]
US 20180137219A1 · Goldfarb · 2018 [cited by applicant]
US 20180144248A1 · Lu · 2018 [cited by examiner]
US 20180322801A1 · Dey · 2018 [cited by applicant]
US 20190050973A1 · Bernal · 2019 [cited by applicant]
US 20190073591A1 · Andoni et al. · 2019 [cited by applicant]
US 20190180187A1 · Rawal · 2019 [cited by applicant]
US 20190180188A1 · Liang et al. · 2019 [cited by applicant]
US 20190207814A1 · Jain · 2019 [cited by applicant]
US 20190258714A1 · Zhong · 2019 [cited by examiner]
US 20200151573A1 · Das · 2020 [cited by applicant]
Xiong et al. (ANTNets: Mobile Convolutional Neural Networks for Resource Efficient Image Classification, Apr. 2019, pp. 1-9) (Year: 2019). [cited by examiner]
Balestriero et al. (From Hard to Soft: Understanding Deep Network Nonlinearities via Vector Quantization and Statistical Inference, Oct. 2018, pp. 1-14) (Year: 2018). [cited by examiner]
Medina et al. (Parallel Attention Mechanisms in Neural Machine Translation, Dec. 2018, pp. 547-552) (Year: 2018). [cited by examiner]
Park et al. (BAM: Bottleneck Attention Module, Jul. 2018, pp. 1-14) (Year: 2018). [cited by examiner]
Voita et al. (Analyzing Multi-Head Self-Attention: Specialized Heads Do the Heavy Lifting, the Rest Can Be Pruned, Jun. 7, 2019, pp. 1-12) (Year: 2019). [cited by examiner]
Gu et al. (Improving Multi-Head Attention with Capsule Networks, Dec. 2018, pp. 314-326) (Year: 2018). [cited by examiner]
Ahmed et al. (Weighted Transformer Network for Machine Translation, Nov. 2017, pp. 1-10) (Year: 2017). [cited by examiner]
Ahmed et al, “Weighted Transformer Network for Machine Translation” arXiv, Nov. 2017, 10 pages. [cited by applicant]
Angeline et al, “An evolutionary algorithm that constructs recurrent neural networks” IEEE Transactions on Neural Networks, Jul. 1993, 28 pages. [cited by applicant]
Bahdanau et al, “Neural machine translation by jointly learning to align and translate” arXiv, Apr. 2015, 15 pages. [cited by applicant]
Baker et al, “Accelerating neural architecture search using performance prediction” arXiv, Nov. 2017, 14 pages. [cited by applicant]
Baker et al, “Designing neural network architectures using reinforcement learning” arXiv, Nov. 2016, 16 pages. [cited by applicant]
Bergstra et al, “Random search for hyperparameter optimization” Journal of Machine Learning Research, Feb. 2012, 25 pages. [cited by applicant]
Brock et al, “SMASH: One-shot model architecture search through hypernetworks” arXiv, Aug. 2017, 21 pages. [cited by applicant]
Cai et al, “Efficient architecture search by network transformation” arXiv, Nov. 2017, 8 pages. [cited by applicant]
Chelba et al, “One billion word benchmark for measuring progress in statistical language modeling” arXiv, Mar. 2014, 6 pages. [cited by applicant]
Chen et al, “Dual path networks” arXiv, Aug. 2017, 11 pages. [cited by applicant]
Chen et al, “The best of both worlds: Combining recent advances in neural machine translation” arXiv, Apr. 2018, 12 pages. [cited by applicant]
Cireşan et al, “Multi-col. deep neural networks for image classification” arXiv, Feb. 2012, 20 pages. [cited by applicant]
Coleman et al, “Analysis of dawnbench, a time-to-accuracy machine learning performance benchmark” arXiv, Jun. 2018, 11 pages. [cited by applicant]
Cortes et al, “Adanet: Adaptive structural learning of artificial neural networks” arXiv, Feb. 2017, 14 pages. [cited by applicant]
Cubuk et al, “Autoaugment: Learning augmentation policies from data” arXiv, Apr. 2019, 14 pages. [cited by applicant]
Dai et al, “Semi-supervised sequence learning” arXiv, Nov. 2015, 10 pages. [cited by applicant]
Dauphin et al, “Language modeling with gated convolutional networks” arXiv, Sep. 2017, 9 pages. [cited by applicant]
Deng et al, “Imagenet: A large-scale hierarchical image database” IEEE Conference on Computer Vision and Pattern Recognition, 2009, 8 pages. [cited by applicant]
Devlin et al, “BERT: pre-training of deep bidirectional transformers for language understanding” arXiv, May 2019, 16 pages. [cited by applicant]
Domhan et al, “Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves” IJCAI, 2017, 9 pages. [cited by applicant]
Elfwing et al, “Sigmoid-weighted linear units for neural network function approximation in reinforcement learning” Neural Networks, Jan. 2018, 9 pages. [cited by applicant]
Elsken et al, “Neural architecture search: A survey” arXiv, Apr. 2019, 21 pages. [cited by applicant]
Elsken et al, “Simple and efficient architecture search for convolutional neural networks” arXiv, Nov. 2017, 14 pages. [cited by applicant]
Fahlman et al, “The cascade-correlation learning architecture” NIPS, 1990, 9 pages. [cited by applicant]
Feurer et al, “Efficient and robust automated machine learning” NIPS, 2015, 9 pages. [cited by applicant]
Floreano et al, “Neuroevolution: from architectures to learning” Evolutionary Intelligence, 2008, 16 pages. [cited by applicant]
Gehring et al, “Convolutional sequence to sequence learning” arXiv, Jul. 2017, 15 pages. [cited by applicant]
Goldberg et al, “A comparative analysis of selection schemes used in genetic algorithms” Foundations of Genetic Algorithms, 1991, 27 pages. [cited by applicant]
He et al, “Deep residual learning for image recognition” arXiv, Dec. 2015, 12 pages. [cited by applicant]
Henderson et al, “Deep reinforcement learning that matters” arXiv, Jan. 2019, 26 pages. [cited by applicant]
Hochreiter et al, “Long short-term memory” Neural Computation, 1997, 32 pages. [cited by applicant]
Hornby et al, “Alps: the age-layered population structure for reducing the problem of premature convergence” GECCO, Jul. 2006, 8 page. [cited by applicant]
Hu et al, “Squeeze-and-excitation networks” arXiv, May 2019, 13 pages. [cited by applicant]
Huang et al, “Densely connected convolutional networks” arXiv, Jan. 2018, 9 pages. [cited by applicant]
Huang et al, “Gpipe: Efficient training of giant neural networks using pipeline parallelism” arXiv, Jul. 2019, 11 pages. [cited by applicant]
Jamieson et al, “Hyperband: a novel bandit-based approach to hyperparameter optimization” arXiv, Jun. 2018, 52 pages. [cited by applicant]
Jamieson et al, “Non-stochastic best arm identification and hyperparameter optimization” Proceedings of the 19th International Conference on Artificial Intelligence on Statistics (AISTATS), 2016, 9 pages. [cited by applicant]
Klein et al, “Learning curve prediction with bayesian neural networks” ICLR, 2017, 16 pages. [cited by applicant]
Krizhevsky et al, “Imagenet classification with deep convolutional neural networks” NIPS, 2012, 9 pages. [cited by applicant]
Krizhevsky et al, “Learning multiple layers of features from tiny images” Apr. 2009, 60 pages. [cited by applicant]
Lei Ba et al, “Layer Normalization” arXiv, Jul. 2016, 14 pages. [cited by applicant]
Liu et al, “Hierarchical representations for efficient architecture search” arXiv, Feb. 2018, 13 pages. [cited by applicant]
Liu et al, “Progressive neural architecture search” arXiv, Jul. 2018, 20 pages. [cited by applicant]
Loshchilov et al, “Sgdr: Stochastic gradient descent with warm restarts” arXiv, May 2017, 16 pages. [cited by applicant]
Maas et al, “Rectifier non-linearities improve neural network acoustic models” International Conference on Machine Learning, 2013, 6 pages. [cited by applicant]
Meissner et al. (Optimized Particle Swarm Optimization (OPSO) and its application to artificial neural network training, Mar. 2006, pp. 1-11) (Year: 2006). [cited by applicant]
Mendoza et al, “Towards automatically-tuned neural networks” JMLR: Workshop and Conference Proceedings, 2016, 8 pages. [cited by applicant]
Merrienboer et al, “On the properties of neural machine translation” Encoderdecoder approaches arXiv, Oct. 2014, 9 pages. [cited by applicant]
Miikkulainen et al, “Evolving deep neural networks” arXiv, Mar. 2017, 8 pages. [cited by applicant]
Miller et al, “Designing neural networks using genetic algorithms” ICGA, 1989, 6 pages. [cited by applicant]
Negrinho et al, “Deeparchitect: Automatically designing and training deep architectures” arXiv, Apr. 2017, 12 pages. [cited by applicant]
Ott et al, “Scaling neural machine translation” arXiv, Sep. 2018, 9 pages. [cited by applicant]
paracrawl.com [online] “Borader Web-Scale Provision of Parallel Corpora for European Languages,” Sep. 2017 [retrieved on Aug. 12, 2019] retreieved from: URL <https://paracrawl.eu/download.html>, 5 page. [cited by applicant]
Peters et al, “Deep contextualized word representations” arXiv, Mar. 2018, 15 pages. [cited by applicant]
Pham et al, “Efficient neural architecture search via parameter sharing” arXiv, Feb. 2018, 11 pages. [cited by applicant]
Pham et al, “Faster discovery of neural architectures by searching for paths in a large model” ICLR workshop, 2018, 15 pages. [cited by applicant]
Post, “A call for clarity in reporting BLEU scores” arXiv, Sep. 2018, 6 pages. [cited by applicant]
Radford et al, “Improving language understanding by generative pre training” Open AI, 2018, 12 pages. [cited by applicant]
Ramachandran et al, “Searching for activation functions” arXiv, Oct. 2017, 13 pages. [cited by applicant]
Real et al, “Large-scale evolution of image classifiers” arXiv, Jun. 2017, 18 pages. [cited by applicant]
Real et al, “Regularized Evolution for Image Classifier Architecture Search,” arXiv, Feb. 16, 16 pages. [cited by applicant]
Salimans et al, “Evolution strategies as a scalable alternative to reinforcement learning” arXiv, Sep. 2017, 13 pages. [cited by applicant]
Saxena et al, “Convolutional neural fabrics” NIPS, 2016, 9 pages. [cited by applicant]
Shaw et al, “Self-attention with relative position representations” arXiv, Apr. 2018, 5 pages. [cited by applicant]
Shazeer et al, “Adafactor: adaptive learning rates with sublinear memory cost” arXiv, Apr. 2018, 9 pages. [cited by applicant]
Simmons et al, “False-positive psychology: Undisclosed flexibility in data collection and analysis allows presenting anything as significant” Psychology Science, Oct. 2011, 9 pages. [cited by applicant]
So et al, “The evolved transformer,” arXiv, May 17, 2019, 14 pages. [cited by applicant]
Srivastava et al, “Dropout: A simple way to prevent neural networks from overfitting” Journal of Machine Learning, Jun. 2014, 30 pages. [cited by applicant]
Stanley et al, “Evolving neural networks through augmenting topologies” Massachusetts Institute of Technology, 2002, 29 pages. [cited by applicant]
Stanley et al, “Real-time neuroevolution in the nero video game” IEEE Transactions on Evolutionary Computation, Dec. 2005, 41 pages. [cited by applicant]
Stanley et al. (Evolving Neural Networks through Augmenting Topologies, Jun. 2001, pp. 1-27) (Year: 2001). [cited by applicant]
Suganuma et al, “A genetic programming approach to designing convolutional neural network architectures” arXiv, Aug. 2017, 9 pages. [cited by applicant]
Sutskever et al, “Sequence to sequence learning with neural networks” Advances in Neural Information Processing Systems, 2014, 9 pages. [cited by applicant]
Szegedy et al, “Going deeper with convolutions” arXiv, Sep. 2014, 12 pages. [cited by applicant]
Szegedy et al, “Inception-v4, inception-resnet and the impact of residual connections on learning” arXiv, Aug. 2016, 12 pages. [cited by applicant]
Van Den Oord et al, “Wavenet: A generative model for raw audio” arXiv, Sep. 2016, 15 pages. [cited by applicant]
Vaswani et al, “Attention Is All You Need”, arXiv, Dec. 6, 2017, 15 pages. [cited by applicant]
Vaswani et al, “Tensor2tensor for neural machine translation” arXiv, Mar. 2018, 9 pages. [cited by applicant]
Wan et al, “Regularization of neural networks using dropconnect” Proceedings of the 30th International Conference on Machine Learning, 2013, 12 pages. [cited by applicant]
Wu et al, “Google's neural machine translation system: Bridging the gap between human and machine translation” arXiv, Oct. 2016, 23 pages. [cited by applicant]
Wu et al, “Pay less attention with lightweight and dynamic convolutions” arXiv, Feb. 2019, 14 pages. [cited by applicant]
Xie et al, “Aggregated residual transformations for deep neural networks” arXiv, Apr. 2017, 10 pages. [cited by applicant]
Xie et al, “Genetic cnn” IEEE International Conference on Computer Vision (ICCV), 2017, 10 pages. [cited by applicant]
Xie et al, “SNAS: stochastic neural architecture search” ICLR, 2019, 17 pages. [cited by applicant]
Yao, “Evolving artificial neural networks” IEEE, 1999, 25 pages. [cited by applicant]
Yu et al, “Fast and accurate reading comprehension by combining self-attention and convolution” ICLR, 2018, 15 pages. [cited by applicant]
Zagoruyko et al, “Wide residual networks” arXiv, 2016, 15 pages. [cited by applicant]
Zela et al, “Towards automated deep learning: efficient joint neural architecture and hyperparameter search” arXiv, Jul. 2018, 11 pages. [cited by applicant]
Zhang et al, “Polynet: A pursuit of structural diversity in very deep networks” arXiv, Jul. 2017, 9 pages. [cited by applicant]
Zhong et al, “Practical network blocks design with q-learning” arXiv, Aug. 2017, 11 pages. [cited by applicant]
Zoph et al, “Learning transferable architectures for scalable image recognition” arXiv, Apr. 2018, 14 pages. [cited by applicant]
Zoph et al, “Neural architecture search with reinforcement learning” arXiv, Feb. 2017, 16 pages. [cited by applicant]