IP Library › Granted Patent US 12,596,914
Granted Patent B2
US 12,596,914 · App. 17/516,025 · Granted Apr 7, 2026

Generative adversarial neural architecture search

Inventors: Seyed Saeed Changiz Rezaei (Vancouver, CA); Fred Xuefei Han (Edmonton, CA); Di Niu (Edmonton, CA)
Assignee: Huawei Technologies Co., Ltd.
G06N3/0475G06N3/045G06N3/088
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,914
App. No.
17/516,025
Granted
Apr 7, 2026
Kind
B2
Abstract

A method and system for neural architectural search (NAS) for performing a task. A generative adversarial network comprising a generator and a discriminator receives, from a user device, a query for neural network architecture, the query including a search space. The generator of the generative adversarial network generates a plurality of generated neural network architectures responsive to the received search space. The discriminator of the generative adversarial network selects an optimal neural network architecture from among the plurality of generated neural network architectures. The optimal generated neural network architecture is transmitted to the user device.

Claims (192)

1 . A method for neural architectural search (NAS) for performing a task, the method comprising:

receiving, by a generative adversarial network comprising a generator and a discriminator, from a user device, a query for neural network architecture, the query including a search space, wherein the discriminator of the generative adversarial network is trained by minimizing a cross-entropy loss between a truth class and predicted probabilities generated from candidate architectures from a set of true architectures or a set of fake architectures respectively;

generating, via the generator of the generative adversarial network, a plurality of generated neural network architectures responsive to the received search space;

selecting, via the discriminator of the generative adversarial network, an optimal neural network architecture from among the plurality of generated neural network architectures; and

transmitting the optimal generated neural network architecture to the user device.

2 . The method of claim 1 , further comprising:

training the generative adversarial network by:

generating, by the generator component of the generative adversarial network, generated neural network architectures based on training data,

evaluating, by the discriminator of the generative adversarial network, the generated neural network architectures,

ranking, by the discriminator, the generated neural network architectures, and

adjusting one or more parameters of the generator based on the ranking of the generated neural network architectures.

3 . The method of claim 2 , wherein the training data comprises the set of true architectures.

4 . The method of claim 2 , wherein the discriminator applies a Siamese scheme in which a pair of architectures in which one architecture is from the set of true architectures and the other architecture is from either the set of true architectures or the set of fake architectures comprising of generated neural network architectures.

5 . The method of claim 2 , wherein the generator and the discriminator are trained by playing a two-player minimax game whose corresponding parameters, φ t and θ t respectively, are optimized using:

(

θ

t

,

ϕ

t

)

←

arg

⁢

min

θ

t

⁢

max

ϕ

t

⁢

V

⁢

(

G

θ

t

,

D

ϕ

t

)

=

𝔼

x

∼

p

𝒥

[

log

⁢

D

ϕ

t

(

x

)

]

+

𝔼

x

∼

p

𝒥

(

x

;

ϕ

t

)

[

log

⁢

(

1

-

D

ϕ

t

(

x

)

)

]

where G denotes the generator and D denotes the discriminator,

x is a Directed Acyclic Graph (DAG) connecting one or more operations, and

is a distribution of currently top performing architectures .

6 . The method of claim 1 , wherein the generator comprises an encoder and a decoder.

7 . The method of claim 6 , wherein the encoder is a graph neural network (GNN)-based, auto-regressive architecture encoder comprising a recurrent neural network (RNN) and GNN.

8 . The method of claim 6 , wherein the decoder comprises a multi-layer perceptron (MLP) that outputs an operator probability distribution and a Gated Recurrent Unit (GRU) that recursively determines edge connections to previous operators.

9 . The method of claim 1 , wherein the generator represents each architecture as a Directed Acyclic Graph (DAG) of operators and connections therebetween.

10 . The method of claim 1 , wherein the discriminator comprises a GNN followed by a MLP classifier.

11 . A system for neural architectural search (NAS) using a generative adversarial network, comprising:

a processing device; and

a memory coupled to the processing device, the memory storing computer-executable instructions that, in response to execution by the processing device, cause the system to:

receive, by a generator of the generative adversarial network, from a user device, a query for neural network architecture, the query including a search space; and

generate, via the generator of the generative adversarial network, a plurality of generated neural network architectures responsive to the received search space; and

select, via a discriminator of the generative adversarial network, an optimal neural network architecture from among the plurality of generated neural network architectures;

wherein the optimal generated neural network architecture is transmitted to the user device, and

wherein the discriminator of the generative adversarial network is trained by minimizing a cross-entropy loss between a truth class and predicted probabilities generated from candidate architectures from a set of true architectures or a set of fake architectures respectively.

12 . The system of claim 11 , wherein the generative adversarial network is trained by:

generating, by the generator component of the generative adversarial network, generated neural network architectures based on training data,

evaluating, by the discriminator of the generative adversarial network, the generated neural network architectures,

ranking, by the discriminator, the generated neural network architectures, and

adjusting one or more parameters of the generator based on the ranking of the generated neural network architectures.

13 . The system of claim 12 , wherein the discriminator applies a Siamese scheme in which a pair of architectures in which one architecture is from the set of true architectures and the other architecture is from either the set of true architectures or the set of fake architectures comprising of generated neural network architectures.

14 . The system of claim 11 , wherein the generator and the discriminator are trained by playing a two-player minimax game whose corresponding parameters, φ t and θ t , respectively, are optimized using:

(

θ

t

,

ϕ

t

)

←

arg

⁢

min

θ

t

⁢

max

ϕ

t

⁢

V

⁢

(

G

θ

t

,

D

ϕ

t

)

=

𝔼

x

∼

p

𝒥

[

log

⁢

D

ϕ

t

(

x

)

]

+

𝔼

x

∼

p

𝒥

(

x

;

ϕ

t

)

[

log

⁢

(

1

-

D

ϕ

t

(

x

)

)

]

where G denotes the generator and D denotes the discriminator,

x is a Directed Acyclic Graph (DAG) connecting one or more operations, and

is a distribution of currently top performing architectures .

15 . The system of claim 11 , wherein the generator comprises an encoder and a decoder.

16 . The system of claim 15 , wherein the encoder is a graph neural network (GNN)-based, auto-regressive architecture encoder comprising a recurrent neural network (RNN) and GNN.

17 . The system of claim 11 , wherein the decoder comprises a multi-layer perceptron (MLP) that outputs an operator probability distribution and a Gated Recurrent Unit (GRU) that recursively determines edge connections to previous operators.

18 . The system of claim 11 , wherein the generator represents each architecture as a Directed Acyclic Graph (DAG) of operators and connections therebetween.

19 . The system of claim 11 , wherein the discriminator comprises a GNN followed by a MLP classifier.

20 . A non-transitory machine-readable storage medium having tangibly stored thereon executable instructions for execution by one or more processors that, in response to execution by the one or more processors, cause the one or more processors to:

receive, by a generative adversarial network comprising a generator and a discriminator, from a user device, a query for neural network architecture, the query including a search space, wherein the discriminator of the generative adversarial network is trained by minimizing a cross-entropy loss between a truth class and predicted probabilities generated from candidate architectures from a set of true architectures or a set of fake architectures respectively;

generate, via the generator of the generative adversarial network, a plurality of generated neural network architectures responsive to the received search space;

select, via the discriminator of the generative adversarial network, an optimal neural network architecture from among the plurality of generated neural network architectures; and

transmit the optimal generated neural network architecture to the user device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 10, 2021
From: CHANGIZ REZAEI, SEYED SAEED; HAN, FRED XUEFEI; NIU, DI
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 058075/0396 →
Continuity (1)
Related Publication 20230140142A1 · May 4, 2023
References Cited (56)
US 10152970B1 · Olabiyi · 2018 [cited by examiner]
US 20190114348A1 · Gao et al. · 2019 [cited by applicant]
US 20190147320A1 · Mattyus · 2019 [cited by examiner]
US 20190258984A1 · Rehman et al. · 2019 [cited by applicant]
US 20190317739A1 · Turek · 2019 [cited by examiner]
US 20210256392A1 · Chen et al. · 2021 [cited by applicant]
US 20220138473A1 · Kwatra · 2022 [cited by examiner]
WO 2021063476A1 · 2021 [cited by applicant]
WO WO2022065771A1 · 2022 [cited by examiner]
David Laredo, “Automatic Model Selection for Neural Networks”, 2019 arXiv (Year: 2019). [cited by examiner]
Zhilong Lu, “Leveraging Graph Neural Network With LSTM For Traffic Speed Prediction”, 2019 IEEE (Year: 2019). [cited by examiner]
Goodfellow et al., “Generative Adversarial Nets”, arxiv:1496.2661v1, 2014. [cited by examiner]
Huang et al., “Searching by Generating: flexible and efficient one-shot NAS with architecture generator”, CVPR, 2021. [cited by examiner]
Lukasik et al., “Learning Where to Look—Generative NAS is surprisingly efficient”, Computer Vision—ECCV, 2022. [cited by examiner]
Elsken, Thomas, Jan Hendrik Metzen, and Frank Hutter. “Neural architecture search: A survey.” arXiv preprint arXiv:1808.05377 2018. [cited by applicant]
Kaicheng Yu, Christian Sciuto, Martin Jaggi, Claudiu Musat, and Mathieu Salzmann. Evaluating the search phase of neural architecture search 2019. [cited by applicant]
Liam Li and Ameet Talwalkar. Random search and reproducibility for neural architecture search. In Uncertainty in Artificial Intelligence, PMLR 2020. [cited by applicant]
Hanxiao Liu, Karen Simonyan, and Yiming Yang. Darts: Differentiable architecture search. arXiv preprint arXiv:1806.09055 2018. [cited by applicant]
Sirui Xie, Hehui Zheng, Chunxiao Liu, and Liang Lin. Snas: stochastic neural architecture search. arXiv preprint arXiv:1812.09926 2018. [cited by applicant]
Kirthevasan Kandasamy, Willie Neiswanger, Jeff Schneider, Barnabas Poczos, and Eric P Xing. Neural architecture search with bayesian optimisation and optimal transport. In Advances in neural information processing syste… [cited by applicant]
Hieu Pham, Melody Y Guan, Barret Zoph, Quoc V Le, and Jeff Dean, Efficient neural architecture search via parameter sharing, arXiv preprint arXiv:1802.03268 2018. [cited by applicant]
Barret Zoph, Vijay Vasudevan, Jonathon Shlens, and Quoc V Le, Learning transferable architectures for scalable mage recognition, In Proceedings of the IEEE conference on computer vision and pattern recognition 2018. [cited by applicant]
Antoine Yang, Pedro M Esperança, and Fabio M Carlucci, Nas evaluation is frustratingly hard, arXiv preprint arXiv:1912.12522 2019. [cited by applicant]
Chris Ying, Aaron Klein, Eric Christiansen, Esteban Real, Kevin Murphy, and Frank Hutter, Nasbench-101: Towards reproducible neural architecture search, In International Conference on Machine Learning 2019. [cited by applicant]
Xuanyi Dong and Yi Yang, Nas-bench-201: Extending the scope of reproducible neural architecture search, In International Conference on Learning Representations 2020. [cited by applicant]
Julien Siems, Lucas Zimmer, Arber Zela, Jovita Lukasik, Margret Keuper, and Frank Hutter, Nasbench-301 and the case for surrogate benchmarks for neural architecture search, arXiv preprint arXiv:2008.09777 2020. [cited by applicant]
Xin Chen, Lingxi Xie, Jun Wu, and Qi Tian, Progressive differentiable architecture search: Bridging the depth gap between search and evaluation, In Proceedings of the IEEE International Conference on Computer Vision 201… [cited by applicant]
Yuhui Xu, Lingxi Xie, Xiaopeng Zhang, Xin Chen, Guo-Jun Qi, Qi Tian, and Hongkai Xiong, Pc-darts: Partial channel connections for memory-efficient architecture search, In International Conference on Learning Representat… [cited by applicant]
Xiangning Chen and Cho-Jui Hsieh, Stabilizing differentiable architecture search via perturbation based regularization, arXiv preprint arXiv:2002.05283 2020. [cited by applicant]
Liam Li, Mikhail Khodak, Maria-Florina Balcan, and Ameet Talwalkar, Geometry-aware gradient algorithms for neural architecture search, arXiv preprint arXiv:2004.07802 2020. [cited by applicant]
Yao Shu, Wei Wang, and Shaofeng Cai, Understanding architectures learnt by cell-based neural architecture search, In International Conference on Learning Representations 2019. [cited by applicant]
Han Cai, Ligeng Zhu, and Song Han, Proxylessnas: Direct neural architecture search on target task and hardware, arXiv preprint arXiv:1812.00332 2018. [cited by applicant]
Bichen Wu, Xiaoliang Dai, Peizhao Zhang, Yanghan Wang, Fei Sun, Yiming Wu, Yuandong Tian, Peter Vajda, Yangqing Jia, and Kurt Keutzer, Fbnet: Hardware-aware efficient convnet design via differentiable neural architectur… [cited by applicant]
Mingxing Tan, Bo Chen, Ruoming Pang, Vijay Vasudevan, Mark Sandler, Andrew Howard, and Quoc V Le, Mnasnet: Platform-aware neural architecture search for mobile, In Proceedings of the IEEE Conference on Computer Vision a… [cited by applicant]
Mingxing Tan and Quoc V Le, Efficientnet: Rethinking model scaling for convolutional neural networks, arXiv preprint arXiv:1905.11946 2019. [cited by applicant]
Gabriel Bender, Hanxiao Liu, Bo Chen, Grace Chu, Shuyang Cheng, Pieter-Jan Kindermans, and Quoc V Le, Can weight sharing outperform random architecture search? an investigation with tunas, In Proceedings of the IEEE/CVF… [cited by applicant]
Linnan Wang, Yiyang Zhao, Yuu Jinnai, Yuandong Tian, and Rodrigo Fonseca, Neural architecture search using deep neural networks and monte carlo tree search, In Proceedings of the AAAI Conference on Artificial Intelligen… [cited by applicant]
Zewei Chen, Fengwei Zhou, George Trimponias, and Zhenguo Li, Multi-objective neural architecture search via non-stationary policy gradient, arXiv preprint arXiv:2001.08437 2020. [cited by applicant]
Jiahui Yu, Pengchong Jin, Hanxiao Liu, Gabriel Bender, Pieter-Jan Kindermans, Mingxing Tan, Thomas Huang, Kiaodan Song, Ruoming Pang, and Quoc Le, Bignas: Scaling up neural architecture search with big single-stage mode… [cited by applicant]
Renqian Luo, Xu Tan, Rui Wang, Tao Qin, Enhong Chen, and Tie-Yan Liu, Semi-supervised neural architecture search, arXiv preprint arXiv:2002.10389 2020. [cited by applicant]
Reuven Y Rubinstein and Dirk P Kroese, The cross-entropy method: a unified approach to combinatorial optimization, Monte-Carlo simulation and machine learning, Springer Science & Business Media 2013. [cited by applicant]
John Schulman, Filip Wolski, Prafulla Dhariwal, Alec Radford, and Oleg Klimov, Proximal policy optimization algorithms, arXiv preprint arXiv:1707.06347 2017. [cited by applicant]
Goodfellow, Ian, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio, Generative adversarial nets, In Advances in neural information processing systems 2014. [cited by applicant]
Kingma, Diederik P., and Max Welling, Auto-encoding variational bayes, arXiv preprint arXiv:1312.6114 2013. [cited by applicant]
Sutton, Richard S., David McAllester, Satinder Singh, and Yishay Mansour, Policy gradient methods for reinforcement earning with function approximation, Advances in neural information processing systems 12 1999. [cited by applicant]
Renqian Luo, Fei Tian, Tao Qin, Enhong Chen, Tie- Yan Liu, Neural architecture optimization 2018. [cited by applicant]
Xiaoliang Dai, Alvin Wan, Peizhao Zhang, Bichen Wu, Zijian He, Zhen Wei, Kan Chen, Yuandong Tian, Matthew Yu, and Peter Vajda, bnetv3: Joint architecture-recipe search using neural acquisition function, arXiv preprint a… [cited by applicant]
Alexia Jolicoeur-Martineau, The relativistic discriminator: a key element missing from standard gan, arXivpreprint arXiv:1807.00734 Sep. 10, 2018. [cited by applicant]
Christopher Morris, Martin Ritzert, Matthias Fey, William L, Hamilton, Jan Eric Lenssen, Gau-rav Rattan, and Martin Grohe, Weisfeiler and Leman go neural: Higher-order Graph Neural Networks, arXiv:1810.02244 Oct. 4, 201… [cited by applicant]
Colin White, Willie Neiswanger, and Yash Savani, Bananas: Bayesian optimization with neural architectures for neural architecture search, arXiv preprint arXiv:1910.11858 Oct. 25, 2018. [cited by applicant]
Jiaxuan You, Bowen Liu, Zhitao Ying, Vijay Pande, and Jure Leskovec, Graph convolutional policy network for goal-directed molecular graph generation, In Advances in neural information processing systems 2018. [cited by applicant]
Seyed Saeed Changiz Rezaei, Fred X. Han, Di Niu, Mohammad Salameh, Keith Mills, Shuo Lian, Wei Lu, Shangling Jui, Generative Adversarial Neural Architecture Search, arXiv:2105.09356 Jun. 23, 2021. [cited by applicant]
Chen, Xiangning, and Cho-Jui Hsieh, Stabilizing Differentiable Architecture Search via Perturbation-based Regularization, arXiv:2002.05283 Feb. 12, 2020. [cited by applicant]
Arber Zela, Thomas Elsken, Tonmoy Saikia, Yassine Marrakchi, Thomas Brox and Frank Hutter, Understanding and robustifying differentiable architecture search, arXiv:1909.09656 Sep. 2, 2019. [cited by applicant]
Dong, Xuanyi, and Yi Yang, “Searching for a robust neural architecture in four GPU hours”, Proceedings of the IEEE Conference on computer vision and pattern recognition, arXiv:1910.04465 Oct. 10, 2019. [cited by applicant]
Yunchen Pum Zhe Gan, Ricardo Henao, Xin Yuan, Chunyuan Li, Andrew Stevens and Lawrence Carin, Variational Autoencoder for Deep Learning of Images, Labels and Captions, arXiv:1609.08976 Sep. 28, 2016. [cited by applicant]