IP Library › Granted Patent US 12,293,276
Granted Patent B2
US 12,293,276 · App. 18/430,483 · Granted May 6, 2025

Neural architecture search with factorized hierarchical search space

Inventors: Mingxing Tan (Newark, CA); Quoc Le (Sunnyvale, CA); Bo Chen (Pasadena, CA); Vijay Vasudevan (Los Altos Hills, CA); Ruoming Pang (New York, NY)
Assignee: GOOGLE LLC
G06N3/04G06F17/15G06N3/044G06N3/084G06N20/10G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,293,276
App. No.
18/430,483
Granted
May 6, 2025
Kind
B2
Abstract

The present disclosure is directed to an automated neural architecture search approach for designing new neural network architectures such as, for example, resource-constrained mobile CNN models. In particular, the present disclosure provides systems and methods to perform neural architecture search using a novel factorized hierarchical search space that permits layer diversity throughout the network, thereby striking the right balance between flexibility and search space size. The resulting neural architectures are able to be run relatively faster and using relatively fewer computing resources (e.g., less processing power, less memory usage, less power consumption, etc.), all while remaining competitive with or even exceeding the performance (e.g., accuracy) of current state-of-the-art mobile-optimized models.

Claims (48)

1. A computer-implemented method, the method comprising:

defining, by one more computing devices, an initial network structure for an artificial neural network, the initial network structure comprising a plurality of blocks;

associating, by the one or more computing devices, a plurality of sub-search spaces respectively with the plurality of blocks, wherein the sub-search space for each block has one or more searchable parameters associated therewith, wherein the one or more searchable parameters included in the sub-search space associated with at least one of the plurality of blocks comprise a number of layers included in the block; and

for each of one or more iterations:

modifying, by one or more computing devices, at least one of the searchable parameters in the sub-search space associated with at least one of the plurality of blocks to generate a new network structure for the artificial neural network.

2. The computer-implemented method of claim 1 , wherein the plurality of sub-search spaces are independent from each other such that modification of at least one of the searchable parameters in one of the sub-search spaces does not necessitate modification of the searchable parameters of any other of the sub-search spaces.

3. The computer-implemented method of claim 1 , wherein, for the at least one of the plurality of blocks, the number of layers comprise a number of identical layers and the searchable parameters for such block are uniformly applied to the number of identical layers included in such block.

4. The computer-implemented method of claim 1 , wherein the one or more searchable parameters included in the sub-search space associated with at least one of the plurality of blocks comprise an operation to be performed by each of one or more layers included in the block.

5. The computer-implemented method of claim 4 , wherein a set of available operations for the searchable parameter of the operation to be performed comprise one or more of:

a convolution;

a depthwise convolution;

an inverted bottleneck convolution; or

a group convolution.

6. The computer-implemented method of claim 1 , wherein the one or more searchable parameters included in the sub-search space associated with at least one of the plurality of blocks comprise one or more of:

a kernel size;

a skip operation to be performed; or

an output filter size.

7. The computer-implemented method of claim 1 , wherein all of the plurality of sub-search spaces share a same set of searchable parameters.

8. The computer-implemented method of claim 1 , wherein at least two of the plurality of sub-search spaces have different sets of searchable parameters associated therewith.

9. The computer-implemented method of claim 1 , further comprising, for each iteration:

measuring, by the one or more computing devices, one or more performance characteristics of the new network structure for the artificial neural network.

10. The computer-implemented method of claim 9 , further comprising, for each iteration:

determining, by the one or more computing devices, a reward to provide to a controller in a reinforcement learning scheme based at least in part on the one or more performance characteristics.

11. The computer-implemented method of claim 9 , wherein the one or more performance characteristics comprise a real-world latency associated with implementation of the new network structure on a real-world mobile device.

12. A computing system, comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

defining an initial network structure for an artificial neural network, the initial network structure comprising a plurality of blocks, wherein a plurality of sub-search spaces are respectively associated with the plurality of blocks, the sub-search space for each block having one or more searchable parameters associated therewith; and

for each of a plurality of iterations:

modifying at least one of the searchable parameters in the sub-search space associated with at least one of the plurality of blocks to generate a new network structure for the artificial neural network.

13. The computing system of claim 12 , wherein the plurality of sub-search spaces comprise a plurality of independent sub-search spaces such that modification of at least one of the searchable parameters in one of the sub-search spaces does not necessitate modification of the searchable parameters of any other of the sub-search spaces.

14. The computing system of claim 12 , wherein the one or more searchable parameters included in the sub-search space associated with at least one of the plurality of blocks comprise a number of layers included in the block.

15. The computing system of claim 12 , wherein the one or more searchable parameters for each block are uniformly applied to all layers included in such block.

16. One or more non-transitory computer-readable media that store instructions that, when executed by one or more processors of a computing system, cause the computing system to perform operations, the operations comprising:

defining, by one more computing devices, an initial network structure for an artificial neural network, the initial structure comprising a plurality of blocks, wherein a plurality of sub-search spaces are respectively associated with the plurality of blocks, the sub-search space for each block having a plurality of searchable parameters associated therewith, the plurality of searchable parameters for each block comprising at least a number of identical layers included in the block and an operation to be performed by each of the number of identical layers included in the block; and

for each of a plurality of iterations:

modifying, by one or more computing devices, at least one of the searchable parameters in the sub-search space associated with at least one of the plurality of blocks to generate a new network structure for the artificial neural network, wherein the number of identical layers included in at least one of the plurality of blocks comprises two or more identical layers.

17. The one or more non-transitory computer-readable media of claim 16 , wherein the plurality of sub-search spaces comprise a plurality of independent sub-search spaces such that modification of at least one of the searchable parameters in one of the sub-search spaces does not necessitate modification of the searchable parameters of any other of the sub-search spaces.

18. The one or more non-transitory computer-readable media of claim 16 , wherein a set of available operations for the searchable parameter of the operation to be performed by each of the number of identical layers comprise one or more of:

a convolution;

a depthwise convolution;

a mobile inverted bottleneck convolution; or

a group convolution.

19. The one or more non-transitory computer-readable media of claim 16 , wherein the plurality of searchable parameters included in the sub-search space associated with at least one of the plurality of blocks additionally comprise one or more of:

a kernel size;

a skip operation to be performed; or

an output filter size.

20. The one or more non-transitory computer-readable media of claim 16 , wherein the plurality searchable parameters included in the sub-search space associated with at least one of the plurality of blocks additionally comprise an input size.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2024
From: TAN, MINGXING; LE, QUOC; CHEN, BO; VASUDEVAN, VIJAY; PANG, RUOMING
To: GOOGLE LLC
Reel/Frame 066693/0864 →
Continuity (5)
Continuation 18154321 · Jan 13, 2023
Continuation 17495398 · Oct 6, 2021
Continuation 16258927 · Jan 28, 2019
Provisional Application 62756254 · Nov 6, 2018
Related Publication 20240273336A1 · Aug 15, 2024
References Cited (43)
US 8682820B2 · Robinson et al. · 2014 [cited by applicant]
US 9729639B2 · Sustaeta et al. · 2017 [cited by applicant]
US 20140143193A1 · Zheng et al. · 2014 [cited by applicant]
US 20190123974A1 · Georgios · 2019 [cited by examiner]
US 20200065961A1 · Zhao et al. · 2020 [cited by applicant]
US 20200327367A1 · Ma et al. · 2020 [cited by applicant]
US 20210124990A1 · Lian · 2021 [cited by applicant]
Baker et al., “Designing Neural Network Architectures Using Reinforcement Learning”, arXiv:1611.02167v3, Mar. 22, 2017, 18 pages. [cited by applicant]
Cai et al., “Reinforcement Learning for Architecture Search by Network Transformation”, arXiv:1707.04873v1, Jul. 16, 2017, 10 pages. [cited by applicant]
Dong et al., “DPP-Net: Device-Aware Progressive Search for Pareto-Optimal Neural Architectures”, arXiv:1806.08198v2, Jul. 25, 2018, 15 pages. [cited by applicant]
Dong et al., “PPP-Net: Platform-Aware Progressive Search for Pareto-Optimal Neural Architectures”, International Conference on Learning Representations Workshop, Vancouver, Canada, Apr. 30-May 3, 2018, 4 pages. [cited by applicant]
Elsken et al., “Efficient Multi-Objective Neural Architecture Search via Lamarckian Evolution”, arXiv:1804.09081v3, Dec. 22, 2018, 23 pages. [cited by applicant]
Elsken et al., “Multi-Objective Architecture Search for CNNs”, arXIV:1804.09081v3, Dec. 22, 2018, 23 pages. [cited by applicant]
Gholami et al., “SqueezeNext: Hardware-Aware Neural Network Design”, arXiv:1803.10615v2, Aug. 27, 2018, 12 pages. [cited by applicant]
Gordon et al., “Morphnet: Fast & Simple Resource-Constrained Structure Learning of Deep Networks”, arXiv:1711.06798v3, Apr. 17, 2018, 15 pages. [cited by applicant]
Goyal et al., “Accurate, Large Minibatch SGD: Training Imagenet in 1 Hour”, arXiv:1706.02677v2, Apr. 30, 2018, 12 pages. [cited by applicant]
Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning Trained Quantization and Huffman Coding”, arXiv:1510.00149v5, Feb. 15, 2016, 14 pages. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition”, Computer Vision and Pattern Recognition, Las Vegas, Nevada, United States, Jun. 26-Jul. 1, 2016, pp. 770-778. [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications”, arXiv:1704.04861v1, Apr. 17, 2017, 9 pages. [cited by applicant]
Hsu et al., “MONAS: Multi-Objective Neural Architecture Search Using Reinforcement Learning”, arXiv:1806.10332v1, Jun. 27, 2018, 8 pages. [cited by applicant]
Hu et al., “Squeeze-and-Excitation Networks”, arXiv:1709.01507v3, Oct. 25, 2018, 14 pages. [cited by applicant]
Huang et al., “CondenseNet: An Efficient DenseNet Using Learned Group Convolutions”, arXiv:1711.09224v2, Jun. 7, 2018, 10 pages. [cited by applicant]
Iandola et al., “SqueezeNet: Alex Net-Level Accuracy with 50x Fewer Parameters and <0.5 mb Model Size”. arXiv:160207360v4, Nov. 4, 2016, 13 pages. [cited by applicant]
Jacob et al., “Quantization and Training of Neural Networks for Efficient Integer-Arithmetic-Only Inference”, arXiv:1712.05877v1, Dec. 15, 2017, 14 pages. [cited by applicant]
Khosla et al., “Novel Dataset for Fine-Grained Image Categorization: Stanford Dogs”, First Workshop on Fine-Grained Visual Categorization Computer Vision and Pattern Recognition, Colorado Springs, Colorado, United State… [cited by applicant]
Kim et al, “NEMO: Neuro-Evolution with Multiobjective Optimization of Deep Neural Network for Speed and Accuracy”, Journal of Machine Learning Research: Workshop and Conference Proceedings, vol. 1, 2017, pp. 1-8. [cited by applicant]
Lin et al., Microsoft COCO: Common Objects in Context, arXiv:1405.0312v3, Feb. 21, 2015. 15 pages. [cited by applicant]
Liu et al., “DARTS: Differentiable Architecture Search”, arXiv:1806.09055v1, Jun. 24, 2018, 12 pages. [cited by applicant]
Liu et al., “Hierarchical Representations for Efficient Architecture Search”, arXiv:1711.00436v2. Feb. 22, 2018, 13 pages. [cited by applicant]
Liu et al., “Progressive Neural Architecture Search” arXiv:1712.00559v1, Dec. 2, 2017, 11 pages. [cited by applicant]
Liu et al., “SSD: Single Shot Multibox Detector”, arXiv:1512.02325v5, Dec. 29, 2016, 17 pages. [cited by applicant]
Pham et al., “Efficient Neural Architecture Search via Parameter Sharing”, arXiv:1802.03268v2, Feb. 12, 2018, 11 pages. [cited by applicant]
Real et al., “Regularized Evolution for Image Classifier Architecture Search”, arXiv:1802.01548v8, Oct. 26, 2018, 16 pages. [cited by applicant]
Redmon et al., “YOLO9000: Better, Faster, Stronger”, arXiv:1612.08242v1, Dec. 25, 2016, 9 pages. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge”, arXiv:1409.0575v3, Jan. 30, 2015, 43 pages. [cited by applicant]
Sandler et al., “MobileNetV2: Inverted Residuals and Linear Bottlenecks”, arXiv:1801.04381v3, Apr. 2, 2018, 14 pages. [cited by applicant]
Schulman et al., “Proximal Policy Optimization Algorithms”, arXiv:1707.06347v2, Aug. 28, 2017, 12 pages. [cited by applicant]
Szegedy et al., “Inception-v4, Inception-ResNet and the Impact of Residual Connections on Learning”, arXiv:1602.07261v2, Aug. 23, 2016, 12 pages. [cited by applicant]
Yang et al., “NetAdapt: Platform-Aware Neural Network Adaptation for Mobile Applications”, arXiv:180403230v2, Sep. 28, 2018, 16 pages. [cited by applicant]
Zhang et al., “ShuffleNet: An Extremely Efficient Convolutional Neural Network for Mobile Devices”, arxiv:1707.01083v2, Dec. 7, 2017, 9 pages. [cited by applicant]
Zhou et al., “Resource-Efficient Neural Architect”, arXiv 1806.07912v1, Jun. 12, 2018, 14 pages. [cited by applicant]
Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition”, arXiv:1707.07012v2, Apr. 11, 2018, 14 pages. [cited by applicant]
Zoph et al., “Neural Architecture Search with Reinforcement Learning”, arXiv:1611.01578v2, Feb. 15, 2017, 16 pages. [cited by applicant]