IP Library › Granted Patent US 12,346,817
Granted Patent B2
US 12,346,817 · App. 17/232,803 · Granted Jul 1, 2025

Neural architecture search

Inventors: Barret Zoph (Sunnyvale, CA); Yun Jia Guan (Stanford, CA); Hieu Hy Pham (Redwood City, CA); Quoc V. Le (Sunnyvale, CA)
Assignee: Google LLC
G06N3/082G06N3/04G06N3/044G06N3/045G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,346,817
App. No.
17/232,803
Granted
Jul 1, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for determining neural network architectures. One of the methods includes generating, using a controller neural network, a batch of output sequences, each output sequence in the batch specifying a respective subset of a plurality of components of a large neural network that should be active during the processing of inputs by the large neural network; for each output sequence in the batch: determining a performance metric of the large neural network on the particular neural network task (i) in accordance with current values of the large network parameters and (ii) with only the subset of components specified by the output sequences active; and using the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters of the controller neural network.

Claims (46)

1. A method of determining an architecture for a neural network for performing a particular neural network task, the method comprising:

generating, in accordance with current values of a plurality of controller parameters, a batch of output sequences, each output sequence in the batch specifying a respective subset of a plurality of components of a large neural network, wherein the large neural network includes the respective subset of the plurality of components and one or more other components of the plurality of components that are not in the respective subset, wherein the large neural network has a plurality of large network parameters, wherein the large neural network comprises a plurality of layers, and wherein the respective subset of the plurality of components of the large neural network forms a smaller neural network that (i) includes only the respective subset of the plurality of components, (ii) does not include the one or more other components of the plurality of components of the large neural network that are not in the respective subset and (iii) has, for each component in the respective subset, current values of the large neural network parameters for that component;

for each output sequence in the batch:

determining a performance metric of the smaller neural network on the particular neural network task in accordance with current values of the large network parameters for the components in the smaller neural network; and

using, by an updating engine, the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters;

generating, in accordance with the adjusted values of the controller parameters, a new output sequence that specifies a new subset of the plurality of components of the large neural network, the new subset of the plurality of components forming a new neural network; and

training, by a training engine in collaboration with the updating engine, the new neural network with only the components in the new subset specified by the new output sequence on training data to determine adjusted values of the large network parameters for the components in the new subset.

2. The method of claim 1 , further comprising:

generating, in accordance with the adjusted values of the controller parameters, a new output sequence that specifies a new subset of the plurality of components of the large neural network, the new subset of the plurality of components forming a new neural network; and

training the new neural network with only the components in the new subset specified by the new output sequence on training data to determine adjusted values of the large network parameters for the components in the new subset.

3. The method of claim 1 , wherein using the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters comprises:

adjusting the current values of the controller parameters to cause generated output sequences to have increased performance metrics using a reinforcement learning technique.

4. The method of claim 3 , wherein the reinforcement learning technique is a policy gradient technique.

5. The method of claim 4 , wherein the reinforcement learning technique is a REINFORCE technique.

6. The method of claim 1 , wherein the current values of the large network parameters are fixed while determining the performance of the smaller neural network.

7. The method of claim 1 , wherein each output sequence comprises respective outputs at each of a plurality of time steps, wherein each time step corresponds to a respective node in a directed acyclic graph (DAG) that represents the large neural network, wherein the DAG comprises a plurality of edges connecting nodes in the DAG, and wherein the output sequence defines, for each node, an input received by the node and a computation performed by the node.

8. The method of claim 7 , wherein generating the batch of output sequences comprises:

generating, for each particular node of a plurality of nodes in the DAG, at a first time step corresponding to the node, a probability distribution over nodes that are connected to the particular node by an incoming edge in the DAG.

9. The method of claim 7 wherein generating the batch of output sequences comprises:

generating, for each particular node of a plurality of nodes in the DAG, at a first time step corresponding to the node, a respective independent probability for each node that is connected to the particular node by an incoming edge in the DAG that defines a likelihood that the edge will be designated as active.

10. The method of claim 8 , for each particular node of the plurality of nodes in the DAG, at a second time step corresponding to the node, generating a probability distribution over possible computations performed by the particular node.

11. The method of claim 1 , wherein the large neural network is a recurrent neural network.

12. The method of claim 1 , wherein the large neural network is a convolutional neural network.

13. The method of claim 1 , further comprising:

generating, in accordance with the adjusted values of the controller parameters, a final output sequence that defines a final set of components.

14. The method of claim 13 , further comprising:

performing the particular neural network task for received network inputs by processing the received network inputs with only the final set of components being active.

15. The method of claim 1 , wherein each of training engine and the updating engine is implemented as one or more software or components installed on one or more computers in one or more locations.

16. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform operations for determining an architecture for a neural network for performing a particular neural network task, the operations comprising:

generating, in accordance with current values of a plurality of controller parameters, a batch of output sequences, each output sequence in the batch specifying a respective subset of a plurality of components of a large neural network, wherein the large neural network includes the respective subset of the plurality of components and one or more other components of the plurality of components that are not in the respective subset, wherein the large neural network has a plurality of large network parameters, wherein the large neural network comprises a plurality of layers, and wherein the respective subset of the plurality of components of the large neural network forms a smaller neural network that (i) includes only the respective subset of the plurality of components, (ii) does not include one or more other components of the plurality of components of the large neural network that are not in the respective subset and (iii) has, for each component in the respective subset, current values of the large neural network parameters for that component;

for each output sequence in the batch:

determining a performance metric of the smaller neural network on the particular neural network task in accordance with current values of the large network parameters for the components in the smaller neural network;

using, by an updating engine, the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters;

generating, in accordance with the adjusted values of the controller parameters, a new output sequence that specifies a new subset of the plurality of components of the large neural network, the new subset of the plurality of components forming a new neural network; and

training, by a training engine in collaboration with the updating engine, the new neural network with only the components in the new subset specified by the new output sequence on training data to determine adjusted values of the large network parameters for the components in the new subset.

17. One or more non-transitory computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform operations for determining an architecture for a neural network for performing a particular neural network task, the operations comprising:

generating, in accordance with current values of a plurality of controller parameters, a batch of output sequences, each output sequence in the batch specifying a respective subset of a plurality of components of a large neural network, wherein the large neural network includes the respective subset of the plurality of components and one or more other components of the plurality of components that are not in the respective subset, wherein the large neural network has a plurality of large network parameters, wherein the large neural network comprises a plurality of layers, and wherein the respective subset of the plurality of components of the large neural network forms a smaller neural network that (i) includes only the respective subset of the plurality of components, (ii) does not include one or more other components of the plurality of components of the large neural network that are not in the respective subset and (iii) has, for each component in the respective subset, current values of the large neural network parameters for that component;

for each output sequence in the batch:

determining a performance metric of the smaller neural network on the particular neural network task in accordance with current values of the large network parameters for the components in the smaller neural network;

using, by an updating engine, the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters;

generating, in accordance with the adjusted values of the controller parameters, a new output sequence that specifies a new subset of the plurality of components of the large neural network, the new subset of the plurality of components forming a new neural network; and

training, by a training engine in collaboration with the updating engine, the new neural network with only the components in the new subset specified by the new output sequence on training data to determine adjusted values of the large network parameters for the components in the new subset.

18. The system of claim 16 , wherein using the performance metrics for the output sequences in the batch to adjust the current values of the controller parameters comprises:

adjusting the current values of the controller parameters to cause generated output sequences to have increased performance metrics using a reinforcement learning technique.

19. The system of claim 16 , wherein the current values of the large network parameters are fixed while determining the performance of the smaller neural network.

20. The system of claim 16 , wherein each output sequence comprises respective outputs at each of a plurality of time steps, wherein each time step corresponds to a respective node in a directed acyclic graph (DAG) that represents the large neural network, wherein the DAG comprises a plurality of edges connecting nodes in the DAG, and wherein the output sequence defines, for each node, an input received by the node and a computation performed by the node.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 16, 2021
From: ZOPH, BARRET; GUAN, YUN JIA; PHAM, HIEU HY; LE, QUOC V.
To: GOOGLE LLC
Reel/Frame 055947/0013 →
Continuity (4)
Continuation 16859781 · Apr 27, 2020
Continuation PCTUS2018058041 · Oct 29, 2018
Provisional Application 62578361 · Oct 27, 2017
Related Publication 20210232929A1 · Jul 29, 2021
References Cited (68)
US 20160180200A1 · Vijayanarasimhan et al. · 2016 [cited by applicant]
US 20160232440A1 · Gregor et al. · 2016 [cited by applicant]
US 20170046614A1 · Golovashkin et al. · 2017 [cited by applicant]
CN 106372648 · 2017 [cited by applicant]
CN 106650922 · 2017 [cited by applicant]
JP 2017004509A · 2017 [cited by applicant]
JP 2019533257A · 2019 [cited by applicant]
Shuai, Bing, et al. “Dag-recurrent neural networks for scene labeling.” Proceedings of the IEEE conference on computer vision and pattern recognition. 2016. (Year: 2016). [cited by examiner]
Ando, Ei, Toshio Nakata, and Masafumi Yamashita. “Approximating the longest path length of a stochastic DAG by a normal distribution in linear time.” Journal of Discrete Algorithms 7.4 (2009): 420-438. (Year: 2009). [cited by examiner]
Zoph, Barret, and Quoc V. Le. “Neural architecture search with reinforcement learning.” arXiv preprint arXiv:1611.01578 (2016). ( Year: 2016). [cited by examiner]
Dutta, Bhaskar, Anders Wallqvist, and Jaques Reifman. “PathNet: a tool for pathway analysis using topological information.” Source code for biology and medicine 7 (2012): 1-12. (Year: 2012). [cited by examiner]
Srivastava, Rupesh K., Klaus Greff, and Jürgen Schmidhuber. “Training very deep networks.” Advances in neural information processing systems 28 (2015). (Year: 2015). [cited by examiner]
Tishby, Naftali, and Noga Zaslavsky. “Deep learning and the information bottleneck principle.” 2015 IEEE information theory workshop (ITW). IEEE, 2015. (Year: 2015). [cited by examiner]
Liu, Bo, et al. “Deep Neural Networks for High Dimension, Low Sample Size Data.” IJCAI. vol. 2017. 2017. (Year: 2017). [cited by examiner]
Ioannou, Yani Andrew. Structural priors in deep neural networks. Diss. 2018. (Year: 2018). [cited by examiner]
Imai, Shunsuke, and Hajime Nobuhara. “Stepwise pathnet: Transfer learning algorithm to improve network structure versatility.” 2018 IEEE International Conference on Systems, Man, and Cybernetics (SMC). IEEE, 2018. (Year… [cited by examiner]
IN Office Action in Indian Application No. 202027017927, dated Jun. 23, 2021, 6 pages (with English translation). [cited by applicant]
JP Office Action in Japanese Application No. 2020-523808, dated Jun. 21, 2021, 9 pages (with English translation). [cited by applicant]
Anonymous, “Faster Discovery of Neural Architectures by Searching for Paths within a Large Model” ICLR, 2018, 15 pages. [cited by applicant]
Brock et al, “SMASH: One-Shot Model Architecture Search through HyperNetworks” arXiv, Aug. 2017, 21 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Application No. PCT/US2018/058041, dated May 7, 2020, 9 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Application No. PCT/US2018/058041, dated Feb. 13, 2019, 15 pages. [cited by applicant]
Pham et al., “Efficient Neural Architecture Search via Paramater Sharing” arXiv, Feb. 2018, 11 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning” arXiv, Feb. 2017, 16 pages. [cited by applicant]
EP Office Action in European Appln. No. 18807483.5, dated Sep. 15, 2022, 6 pages. [cited by applicant]
Office Action in Japanese Appln. No. 2022-41344, dated Apr. 24, 2023, 9 pages (with—English Translation). [cited by applicant]
Office Action in Indian Appln. 202027017927, dated Aug. 24, 2023, 3 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201880075801.0, dated Apr. 30, 2024, 7 pages (with English translation). [cited by applicant]
Office Action in Chinese Appln. No. 201880075801.0, dated Feb. 2, 2024, 24 pages (with English translation). [cited by applicant]
Ahmed et al., “Connectivity learning in multi-branch networks,” CoRR, Submitted on Dec. 2017, arXiv:1709.09582v2, 17 pages. [cited by applicant]
Bahdanau et al., “Neural machine translation by jointly learning to align and translate” CoRR, Submitted on May 2016, arXiv:1409.0473v7, 15 pages. [cited by applicant]
Baker et al., “Designing Neural Network Architectures Using Reinforcement Learning,” Presented in ICLR, 2017, 18 pages. [cited by applicant]
Bello et al., “Neural Optimizer Search With Reinforcement Learning,” Presented at ICML, Jul. 2017, 10 pages. [cited by applicant]
Bottou, “A theoretical approach to connectionnist learning; Application to automatic speech recognition” PhD thesis, 1991, 239 pages (with English abstract). [cited by applicant]
Cai et al., “Reinforcement learning for architecture search by network transformation,” CoRR, Submitted on Jul. 2017, arXiv:1707.04873v1, 10 pages. [cited by applicant]
Devries et al., “Improved Regularizatyion of Convolutional Neural networks with cutout,” CoRR, Submitted on Nov. 2017, arXiv:1708.04552v2, 8 pages. [cited by applicant]
Gal et al., “A Theoretically Grounded Application of Dropout in Recurrent Neural networks,” CoRR, Submitted on Oct. 2016, arXiv:1512.05287v5, 14 pages. [cited by applicant]
Gastaldi, “Shake-shake regularization of 3-branch residual networks,” Presented at ICLR Workshop Track, 2017, 5 pages. [cited by applicant]
Ha et al., “Hypernetworks,” CoRR, Submitted on Dec. 2016, arXiv:1609.09106v4, 29 pages. [cited by applicant]
He et al., “Deep Residual learning for Image Recognition,” Presented at CPVR, 2016, 770-778. [cited by applicant]
He et al., “Delving deep into rectifiers: Surpassing human-level performance on imagenet classification,” Presented at CVPR, 2015, 1026-1034. [cited by applicant]
Hochreiter et al., “Long short-term memory,” In Neural computations, 1997, 32 pages. [cited by applicant]
Huang et al., “Densly connected convolutional networks,” CoRR, Submitted on Jan. 2018, arXiv:1608.06993v5, 9 pages. [cited by applicant]
Inan et al., “Tying word vectors and word classifiers: a loss framework for language modeling,” CoRR, Submitted on Mar. 2017, arXiv:1611.01462v3, 13 pages. [cited by applicant]
Ioffe et al., “Batch normalization: Accelerating deep network training by reducing internal covariate shift,” CoRR, Submitted on Mar. 2015, arXiv:1502.03167v3, 11 pages. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” CoRR, Submitted on Jan. 2017, arXiv:1412.6980v9, 15 pages. [cited by applicant]
Krause et al., “Dynamic evaluation of neural sequence models,” CoRR, Submitted on Oct. 2017, arXiv:1709.07432v2, 10 pages. [cited by applicant]
Krizahevsky, “Learning multiple layers of features from tiny images,” Technical report, Apr. 8, 2009, 60 pages. [cited by applicant]
Larsson et al., “Fractalnet: Ultra-deep neural networks without residuals,” CoRR, Submitted on May 2017, arXiv:1605.07648v4, 11 pages. [cited by applicant]
Lin et al., “Network in network,” CoRR, Submitted on Mar. 2013, arXiv:1312.4400v3, 10 pages. [cited by applicant]
Loshchilov et al., “SGDR: Stochastic gradient descent with warm restarts,” CoRR, Submitted on May 2017, arXiv:1608.03983v5, 16 pages. [cited by applicant]
Melis et al., “On the state of the art of evaluation in neural language models,” CoRR, Submitted on Nov. 2017, arXiv:1707.05589v2, 10 pages. [cited by applicant]
Merity et al., “Regularizing and optimizing LSTM language models,” CoRR, Submitted on Aug. 2017, arXiv:1708.02182v1, 10 pages. [cited by applicant]
Negrinho et al., “Deeparchitect: Automatically designing and training deep architectures,” CoRR, Submitted on Apr. 2017, arXiv:1704.08792v1, 12 pages. [cited by applicant]
Nesterov et al., “A method for solving the convex programming problem with convergence rate o(1/k2),” Soviet Mathematics Doklady, 1983, 544-547 (with English abstract). [cited by applicant]
Saxena et al., “Convolutional neural fabrics,” Advances in Neural Information Processing Systems, Barcelona, Spain, Dec. 5-10, 2016; Advances in Neural Information Processing Systems 29, Dec. 2016, 9 pages. [cited by applicant]
Schulman et al., “Proximal Policy optimization algorithms,” CoRR, Submitted on Aug. 2017, arXiv:1707.06347v2, 12 pages. [cited by applicant]
Sutskever et al., “Sequence to sequence learning with neural networks,” NIPS, 2014, 9 pages. [cited by applicant]
Szegedy et al., “Re-thinking the inception architecture for computer vision,” Presented in CPVR, 2016, 2818-2826. [cited by applicant]
Veniat et al., “Learning time-efficient deep architectures with budgeted super networks,” CoRR, Submitted on May 2017, arXiv:1706.00046v1, 11 pages. [cited by applicant]
Williams et al., “Function optimization using connectionist reinforcement learning algorithms.,” Connection Science, 1991, 3(3): 29 pages. [cited by applicant]
Williams et al., “Simple statistical gradient-following algorithms for connectionist reinforcement learning,” Machine Learning, 1992, 5-32. [cited by applicant]
Xie et al., “Aggregated residual transformations for deep neural networks,” In CVPR, 2017, 1492-1500. [cited by applicant]
Zagoruyko et al., “Wide Residual Networks,” CoRR, Submitted on Jun. 2017, arXiv:1605.07146v4, 15 pages. [cited by applicant]
Zaremba et al., “Recurrent neural network regularization,” CoRR, Submitted on Dec. 2014, arXiv:1409.2329v4, 8 pages. [cited by applicant]
Zhong et al., “Practical Network Blocks Design with q-learning,” CoRR, submitted on Aug. 2017, arXiv:1708.05552v1, 11 pages. [cited by applicant]
Zilly et al., “Re-current highway networks,” In ICML, 2017, 10 pages. [cited by applicant]
Zoph et al., “Learning transferable architectures for scalable image recognition” CoRR, Submitted on Dec. 2017, 1707.07012v3, 14 pages. [cited by applicant]
Cited By (1)
US 12,633,105