IP Library › Granted Patent US 12,579,150
Granted Patent B2
US 12,579,150 · App. 17/721,873 · Granted Mar 17, 2026

Hybrid and hierarchical multi-trial and OneShot neural architecture search on datacenter machine learning accelerators

Inventors: Sheng Li (Cupertino, CA); Garrett Axel Andersen (Austin, TX); Norman Paul Jouppi (Palo Alto, CA); Quoc V. Le (Sunnyvale, CA); Liqun Cheng (Palo Alto, CA); Parthasarathy Ranganathan (San Jose, CA); Julian Paul Grady (Pittsburgh, PA); Yang Li (Palo Alto, CA); Martin Wicke (San Francisco, CA); Yifeng Lu (Palo Alto, CA); Yun Ni (San Mateo, CA); Kun Wang (Pittsburgh, PA)
Assignee: Google LLC
G06F16/2457G06F16/24554G06N3/063G06N3/0985
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,579,150
App. No.
17/721,873
Granted
Mar 17, 2026
Kind
B2
Abstract

According to various implementations, generally disclosed herein is a hybrid and hierarchical neural architecture search (NAS) approach. The approach includes performing a search space partitioning scheme to divide the search space into sub-search spaces. The approach further includes performing a first type of NAS, such as a Multi-trial NAS, to cover a search across the sub-search spaces. The approach also includes performing a second type of NAS, such as a One-Shot NAS, to cover each sub-search space. The approach further includes automatically stopping the second type of NAS based on one or more early stopping criteria.

Claims (41)

1 . A method for performing a neural architecture search (NAS), the method comprising:

partitioning, with one or more processors, a search space for a model architecture into a plurality of sub-search spaces corresponding to different compiler configurations for the model architecture;

performing, with the one or more processors, a first type of NAS across the plurality of sub-search spaces, wherein the first type of NAS is a multi-trial search involving searching for different compiler flags;

performing, with the one or more processors, a second type of NAS within each sub-search space based on the performance of the first type of NAS, wherein the second type of NAS is a One-Shot search involving searching for different hyperparameters;

stopping, with the one or more processors, the second type of NAS based on an early stopping criterion; and

selecting, with the one or more processors, a compiler configuration for the model architecture based on results of the first type of NAS and the second type of NAS.

2 . The method of claim 1 , wherein partitioning the search space further comprises partitioning the search space based on one or more hyperparameters that influence the search space.

3 . The method of claim 2 , wherein:

partitioning the search space further comprises selecting a principal model architecture parameter as a dimension for the first type of NAS; and

performing a second type of NAS further comprises searching for a remainder of model architecture dimensions.

4 . The method of claim 1 , wherein partitioning the search space further comprises:

computing a size of the search space based on a capacity influenced by the search space; and

automatically partitioning the search space based on the computed size.

5 . The method of claim 4 , wherein the capacity influenced by the search space further comprises one of a machine learning hardware memory capacity, compute throughput, memory or network bandwidth, or power.

6 . The method of claim 1 , wherein partitioning the search space further comprises partitioning the search space based on one or more hyperparameters that influence at least one of a quality or efficiency of machine learning model results.

7 . The method of claim 1 , further comprising monitoring, with the one or more processors, the early stopping criterion.

8 . The method of claim 1 , wherein the early stopping criterion comprises one or more of an architecture searchable parameter approaching a convergence, a quality threshold, a threshold amount of data consumed, or a convergence rate threshold.

9 . A system comprising:

one or more processors; and

one or more storage devices coupled to the one or more processors and storing instructions, when performed by the one or more processors, causes the one or more processors to perform operations for performing a neural architecture search (NAS), the operations comprising:

partitioning a search space for a model architecture into a plurality of sub-search spaces corresponding to different compiler configurations for the model architecture;

performing a first type of NAS across the plurality of sub-search spaces, wherein the first type of NAS is a multi-trial search involving searching for different compiler flags;

performing a second type of NAS within each sub-search space based on the performance of the first type of NAS, wherein the second type of NAS is as One-Shot search involving searching for different hyperparameters;

stopping the second type of NAS based on an early stopping criterion; and

selecting a compiler configuration for the model architecture based on results of the first type of NAS and the second type of NAS.

10 . The system of claim 9 , wherein partitioning the search space further comprises partitioning the search space based on one or more hyperparameters that influence the search space.

11 . The system of claim 10 , wherein:

partitioning the search space further comprises selecting a principal model architecture parameter as a dimension for the first type of NAS; and

performing a second type of NAS further comprises search for a remainder of model architecture dimensions.

12 . The system of claim 9 , wherein partitioning the search space further comprises:

computing a size of the search space based on a capacity influenced by the search space; and

automatically partitioning the search space based on the computed size.

13 . The system of claim 12 , wherein the capacity influenced by the search space further comprises one of a machine learning hardware memory capacity, compute throughput, memory or network bandwidth, or power.

14 . The system of claim 9 , wherein partitioning the search space further comprises partitioning the search space based on one or more hyperparameters that influence at least one of quality or efficiency of machine learning model results.

15 . The system of claim 9 , wherein the early stopping criterion comprises one or more of an architecture searchable parameter approaching a convergence, a quality threshold, a threshold amount of data consumed, or a convergence rate threshold.

16 . A non-transitory computer readable medium for storing instructions that, when executed by one or more processors, causes the one or more processors to perform operations for performing a neural architecture search (NAS), the operations comprising:

partitioning a search space for a model architecture into a plurality of sub-search spaces corresponding to different compiler configurations for the model architecture;

performing a first type of NAS across the plurality of sub-search spaces, wherein the first type of NAS is a multi-trial search involving searching for different compiler flags;

performing a second type of NAS within each sub-search space based on the performance of the first type of NAS, wherein the second type of NAS is a One-Shot search involving searching for different hyperparameters;

stopping the second type of NAS based on an early stopping criterion; and

selecting a compiler configuration for the model architecture based on results of the first type of NAS and the second type of NAS.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2022
From: LI, SHENG; ANDERSEN, GARRETT AXEL; JOUPPI, NORMAN PAUL; LE, QUOC V.; CHENG, LIQUN; RANGANATHAN, PARTHASARATHY; GRADY, JULIAN PAUL; LI, YANG; WICKE, MARTIN; LU, YIFENG; NI, YUN; WANG, KUN
To: GOOGLE LLC
Reel/Frame 059620/0027 →
Continuity (2)
Provisional Application 63320880 · Mar 17, 2022
Related Publication 20230297580A1 · Sep 21, 2023
References Cited (42)
US 20170160706A1 · Düll et al. · 2017 [cited by applicant]
US 20200104687A1 · Gesmundo · 2020 [cited by applicant]
US 20200143227A1 · Tan et al. · 2020 [cited by applicant]
US 20200265301A1 · Burger et al. · 2020 [cited by applicant]
US 20210034928A1 · Oh · 2021 [cited by examiner]
US 20210383223A1 · Tan et al. · 2021 [cited by applicant]
US 20220019890A1 · Staffler et al. · 2022 [cited by applicant]
US 20220035878A1 · Sarah · 2022 [cited by examiner]
US 20220036136A1 · Muehlberg et al. · 2022 [cited by applicant]
US 20220108054A1 · Akhauri et al. · 2022 [cited by applicant]
US 20220147680A1 · Zehngut · 2022 [cited by examiner]
US 20220229960A1 · Nath et al. · 2022 [cited by applicant]
US 20220405450A1 · Bunandar et al. · 2022 [cited by applicant]
US 20230064692A1 · Chen · 2023 [cited by examiner]
US 20230096654A1 · Salameh · 2023 [cited by examiner]
US 20230153506A1 · Park et al. · 2023 [cited by applicant]
CN 112116156A · 2020 [cited by applicant]
WO 2022072890A1 · 2022 [cited by applicant]
WO 2022076933A1 · 2022 [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2022/033520 dated Dec. 8, 2022. 17 pages. [cited by applicant]
Li et al. Searching for Fast Model Families on Datacenter Accelerators. 2021. Computer Vision Foundation, pp. 8085-8095. [cited by applicant]
Lin et al. NAAS: Neural Accelerator Architecture Search. May 27, 2021. 7 pages. [cited by applicant]
Lin et al. Neural-Hardware Architecture Search. 2019. 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), Vancouver, Canada, 5 pages. [cited by applicant]
Parashar et al. Timeloop: A Systematic Approach to DNN Accelerator Evaluation. Apr. 25, 2019. 2019 IEEE International Symposium on Performance Analysis of Systems and Software (ISPASS). 12 pages. [cited by applicant]
Tang et al. NeuroMeter: An Integrated Power, Area, and Timing Modeling Framework for Machine Learning Accelerators. Apr. 22, 2021. 2021 IEEE International Symposium on High-Performance Computer Architecture (HPCA). pp. … [cited by applicant]
Yang et al. Co-Exploration of Neural Architectures and Heterogeneous ASIC Accelerator Designs Targeting Multiple Tasks. Feb. 10, 2020. 7 pages. [cited by applicant]
Zhang et al. A Full-Stack Search Technique for Domain Optimized Deep Learning Accelerators. Feb. 1, 2022. ASPLOS '22, Feb. 28-Mar. 4, 2022, Lausanne, Switzerland. 16 pages. [cited by applicant]
International Search Report and Written Opinion for International Application No. PCT/US2023/019338 dated Jul. 12, 2023. 15 pages. [cited by applicant]
Penney et al. A Survey of Machine Learning Applied to Computer Architecture Design. arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 26, 2019 (Sep. 26, 2019), 14 pages. [cited by applicant]
Robine et al. Smaller World Models for Reinforcement Learning. arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 2, 2021 (Mar. 2, 2021), 9 pages. [cited by applicant]
Sadasivam et al. Invited: Efficient Reinforcement Learning for Automating Human Decision-Making in SOC Design. 2018 55th ACM/ESDA/IEEE Design Automation Conference (DAC), IEEE, Jun. 24, 2018 (Jun. 24, 2018), pp. 1-6. [cited by applicant]
Shi et al. Reinforcement Learning Based Test Case Prioritization for Enhancing the Security of Software. 2020 IEEE 7th International Conference On Data Science and Advanced Analytics (DSAA) , IEEE, Oct. 6, 2020 (Oct. 6,… [cited by applicant]
Bender et al. Can weight sharing outperform random architecture search? An investigation with TuNAS. Aug. 13, 2020. 13 pages. [cited by applicant]
Bender et al. Understanding and Simplifying One-Shot Architecture Search. 2018. Proceedings of the 35 th International Conference on Machine Learning, Stockholm, Sweden, 10 pages. [cited by applicant]
Bergstra et al. Algorithms for Hyper-Parameter Optimization. 2011. Advances in Neural Information Processing Systems 24 (NIPS 2011), pp. 1-9. [cited by applicant]
Cho et al. B2EA: An Evolutionary Algorithm Assisted by Two Bayesian Optimization Modules for Neural Architecture Search. Feb. 17, 2022. 25 pages. [cited by applicant]
Golovin et al. Google Vizier: A Service for Black-Box Optimization. Aug. 13-17, 2017. Halifax, NS, Canada. KDD 2017 Applied Data Science Paper, pp. 1487-1496. [cited by applicant]
Li et al. Searching for Fast Model Families on Datacenter Accelerators. Feb. 10, 2021, pp. 8085-8095. [cited by applicant]
Real et al. Regularized Evolution for Image Classifier Architecture Search. Feb. 16, 2019. AAAI 2019, the Thirty-Third AAAI Conference on Artificial Intelligence.16 pages. [cited by applicant]
Tan et al. MnasNet: Platform-Aware Neural Architecture Search for Mobile. May 29, 2019. 9 pages. [cited by applicant]
Office Action for European Patent Application No. 22738246.2 dated Feb. 3, 2026. 9 pages. [cited by applicant]
Zhang et al. Fast Hardware-Aware Neural Architecture Search. Jun. 14, 2020. 2020 IEEE/CVF Conference On Computer Vision and Pattern Recognition Workshops (CVPRW), IEEE, pp. 2959-2967, DOI: 10.1109/CVPRW50498.2020.00354. [cited by applicant]