IP Library › Granted Patent US 12,530,560
Granted Patent B2
US 12,530,560 · App. 17/072,485 · Granted Jan 20, 2026

System and method for differential architecture search for neural networks

Inventors: Pan Zhou (Singapore, SG); Chu Hong Hoi (Singapore, SG)
Assignee: Salesforce, Inc.
G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,560
App. No.
17/072,485
Granted
Jan 20, 2026
Kind
B2
Abstract

A method for generating a neural network, including initializing the neural network including a plurality of cells, each cell corresponding to a graph including one or more nodes, each node corresponding to a latent representation of a dataset. A plurality of gates are generated, wherein each gate independently determines whether an operation between two nodes is used. A first regularization is performed using the plurality of gates. The first regularization is one of a group-structured sparsity regularization and a path-depth-wised regularization. An optimization is performed on the neural network by adjusting its network parameters and gate parameters based on the regularization of the sparsity.

Claims (76)

1 . A method for providing a multi-layer neural network, comprising:

initializing the multi-layer neural network including a plurality of cells, each cell corresponding to a graph including a plurality of nodes, each node corresponding to a latent representation of a dataset;

generating a plurality of gates for a first cell, wherein each gate independently determines whether an operation between two nodes in the first cell is used, wherein the first cell includes:

two input nodes corresponding to outputs of two previous cells respectively,

one or more intermediate nodes connecting with previous nodes in the cell, and

an output node provided by concatenating the one or more intermediate nodes;

performing a first regularization using the plurality of gates, wherein the first regularization includes a group-structured sparsity regularization and a path-depth-wise regularization,

wherein the group-structured sparsity regularization includes:

assigning a first regularization constant to a skip connection group; and

assigning a second regularization constant to a non-skip connection group;

wherein the first regularization constant is larger than the second regularization constant; and

wherein the path-depth-wise regularization is configured for correcting cell-selection bias caused by the plurality of gates in group-structured sparsity regularization, wherein the cell-selection bias is generated after intermediate nodes are connected with input nodes in the multi-layer neural network; and

performing an optimization on the multi-layer neural network by adjusting its network parameters and gate parameters based on the first regularization using the first regularization constant and the second regularization constant.

2 . The method of claim 1 , wherein the performing the group-structured sparsity regularization includes:

generating two or more groups from the plurality of gates based on a corresponding operation type of each gate; and

performing a regularization of sparsity of each of the two or more groups.

3 . The method of claim 2 , wherein the performing the group-structured sparsity regularization includes:

generating a first group loss based on activation probability of a first group of the two or more groups; and

generating a second group loss based on activation probability of a second group of the two or more groups;

wherein the optimization is performed based on the first group loss and the second group loss.

4 . The method of claim 1 , wherein the performing the path-depth-wise regularization includes:

generating a path-depth-wise loss using probabilities based on a depth of a path of the multi-layer neural network;

wherein the optimization is performed based on the path-depth-wise loss.

5 . The method of claim 1 , wherein the operation is one of skip connection, zero operation, and pooling operation, separable convolutions, dilated separable convolutions, average pooling and max pooling.

6 . The method of claim 1 , wherein the optimization uses gradient descent.

7 . A non-transitory machine-readable medium comprising a plurality of machine-readable instructions which, when executed by one or more processors, are adapted to cause the one or more processors to perform a method comprising:

initializing a multi-layer neural network including a plurality of cells, each cell corresponding to a graph including a plurality of nodes, each node corresponding to a latent representation of a dataset;

generating a plurality of gates for a first cell, wherein each gate independently determines whether an operation between two nodes in the first cell is used, wherein the first cell includes:

two input nodes corresponding to outputs of two previous cells respectively,

one or more intermediate nodes connecting with previous nodes in the cell, and

an output node provided by concatenating the one or more intermediate nodes;

performing a first regularization using the plurality of gates, wherein the first regularization includes a group-structured sparsity regularization and a path-depth-wise regularization,

wherein performing the group-structured sparsity regularization includes:

assigning a first regularization constant to a skip connection group; and

assigning a second regularization constant to a non-skip connection group;

wherein the first regularization constant is larger than the second regularization constant; and

wherein the path-depth-wise regularization is configured for correcting cell-selection bias caused by the plurality of gates in group-structured sparsity regularization, wherein the cell-selection bias is generated after intermediate nodes are connected with input nodes in the multi-layer neural network via skip connections and after connections between neighboring nodes are cut via zero operation; and

performing an optimization on the multi-layer neural network by adjusting its network parameters and gate parameters based on the first regularization using the first regularization constant and the second regularization constant.

8 . The non-transitory machine-readable medium of claim 7 , wherein the performing the group-structured sparsity regularization includes:

generating two or more groups from the plurality of gates based on a corresponding operation type of each gate; and

performing a regularization of sparsity of each of the two or more groups.

9 . The non-transitory machine-readable medium of claim 8 , wherein the performing the group-structured sparsity regularization includes:

generating a first group loss based on activation probability of a first group of the two or more groups; and

generating a second group loss based on activation probability of a second group of the two or more groups;

wherein the optimization is performed based on the first group loss and the second group loss.

10 . The non-transitory machine-readable medium of claim 7 , wherein the performing the path-depth-wise regularization includes:

generating a path-depth-wise loss using probabilities based on a depth of a path of the multi-layer neural network;

wherein the optimization is performed based on the path-depth-wise loss.

11 . The non-transitory machine-readable medium of claim 7 , wherein the operation is one of skip connection, zero operation, and pooling operation, separable convolutions, dilated separable convolutions, average pooling and max pooling.

12 . The non-transitory machine-readable medium of claim 7 , wherein the optimization uses gradient descent.

13 . A system, comprising:

a non-transitory memory; and

one or more hardware processors coupled to the non-transitory memory and configured to read instructions from the non-transitory memory to cause the system to perform a method comprising:

initializing a multi-layer neural network including a plurality of cells, each cell corresponding to a graph including a plurality of nodes, each node corresponding to a latent representation of a dataset;

generating a plurality of gates for a first cell, wherein each gate independently determines whether an operation between two nodes in the first cell is used, wherein the first cell includes:

two input nodes corresponding to outputs of two previous cells respectively,

one or more intermediate nodes connecting with previous nodes in the cell, and

an output node provided by concatenating the one or more intermediate nodes;

performing a first regularization using the plurality of gates, wherein the first regularization includes a group-structured sparsity regularization and a path-depth-wise regularization,

wherein performing the group-structured sparsity regularization includes:

assigning a first regularization constant to a skip connection group; and

assigning a second regularization constant to a non-skip connection group;

wherein the first regularization constant is larger than the second regularization constant; and

wherein the path-depth-wise regularization is configured for correcting cell-selection bias caused by the plurality of gates in group-structured sparsity regularization, wherein the cell-selection bias is generated after intermediate nodes are connected with input nodes in the multi-layer neural network via skip connections and after connections between neighboring nodes are cut via zero operation; and

performing an optimization on the multi-layer neural network by adjusting its network parameters and gate parameters based on the first regularization using the first regularization constant and the second regularization constant.

14 . The system of claim 13 , wherein the performing the group-structured sparsity regularization includes:

generating two or more groups from the plurality of gates based on a corresponding operation type of each gate; and

performing a regularization of sparsity of each of the two or more groups.

15 . The system of claim 14 , wherein the performing the group-structured sparsity regularization includes:

generating a first group loss based on activation probability of a first group of the two or more groups; and

generating a second group loss based on activation probability of a second group of the two or more groups;

wherein the optimization is performed based on the first group loss and the second group loss.

16 . The system of claim 13 , wherein the performing the path-depth-wise regularization includes:

generating a path-depth-wise loss using probabilities based on a depth of a path of the multi-layer neural network;

wherein the optimization is performed based on the path-depth-wise loss.

17 . The system of claim 13 , wherein the operation is one of skip connection, zero operation, and pooling operation, separable convolutions, dilated separable convolutions, average pooling and max pooling.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2020
From: ZHOU, PAN; HOI, CHU HONG
To: SALESFORCE.COM, INC.
Reel/Frame 054114/0481 →
Continuity (2)
Provisional Application 63034269 · Jun 3, 2020
Related Publication 20210383188A1 · Dec 9, 2021
References Cited (135)
US 10282663B2 · Socher et al. · 2019 [cited by applicant]
US 10346721B2 · Albright et al. · 2019 [cited by applicant]
US 10474709B2 · Paulus · 2019 [cited by applicant]
US 10521465B2 · Paulus · 2019 [cited by applicant]
US 10542270B2 · Zhou et al. · 2020 [cited by applicant]
US 10546217B2 · Albright et al. · 2020 [cited by applicant]
US 10558750B2 · Lu et al. · 2020 [cited by applicant]
US 10565305B2 · Lu et al. · 2020 [cited by applicant]
US 10565306B2 · Lu et al. · 2020 [cited by applicant]
US 10565318B2 · Bradbury · 2020 [cited by applicant]
US 10565493B2 · Merity et al. · 2020 [cited by applicant]
US 10573295B2 · Zhou et al. · 2020 [cited by applicant]
US 10592767B2 · Trott et al. · 2020 [cited by applicant]
US 10699060B2 · Mccann · 2020 [cited by applicant]
US 10747761B2 · Zhong et al. · 2020 [cited by applicant]
US 10776581B2 · Mccann et al. · 2020 [cited by applicant]
US 10783875B2 · Hosseini-Asl et al. · 2020 [cited by applicant]
US 10817650B2 · Mccann et al. · 2020 [cited by applicant]
US 10839284B2 · Hashimoto et al. · 2020 [cited by applicant]
US 10846478B2 · Lu et al. · 2020 [cited by applicant]
US 10902289B2 · Gao et al. · 2021 [cited by applicant]
US 10909157B2 · Paulus et al. · 2021 [cited by applicant]
US 10929607B2 · Zhong et al. · 2021 [cited by applicant]
US 10958925B2 · Zhou et al. · 2021 [cited by applicant]
US 10963652B2 · Hashimoto et al. · 2021 [cited by applicant]
US 10963782B2 · Xiong et al. · 2021 [cited by applicant]
US 10970486B2 · Machado et al. · 2021 [cited by applicant]
US 11663481B2 · Liu · 2023 [cited by examiner]
US 20160350653A1 · Socher et al. · 2016 [cited by applicant]
US 20170024645A1 · Socher et al. · 2017 [cited by applicant]
US 20170032280A1 · Socher · 2017 [cited by applicant]
US 20170140240A1 · Socher et al. · 2017 [cited by applicant]
US 20180096219A1 · Socher · 2018 [cited by applicant]
US 20180121788A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180121799A1 · Hashimoto et al. · 2018 [cited by applicant]
US 20180129931A1 · Bradbury et al. · 2018 [cited by applicant]
US 20180129937A1 · Bradbury et al. · 2018 [cited by applicant]
US 20180268287A1 · Johansen et al. · 2018 [cited by applicant]
US 20180268298A1 · Johansen et al. · 2018 [cited by applicant]
US 20180336453A1 · Merity et al. · 2018 [cited by applicant]
US 20180373987A1 · Zhang et al. · 2018 [cited by applicant]
US 20190130248A1 · Zhong et al. · 2019 [cited by applicant]
US 20190130249A1 · Bradbury et al. · 2019 [cited by applicant]
US 20190130273A1 · Keskar et al. · 2019 [cited by applicant]
US 20190130312A1 · Xiong et al. · 2019 [cited by applicant]
US 20190130896A1 · Zhou et al. · 2019 [cited by applicant]
US 20190188568A1 · Keskar et al. · 2019 [cited by applicant]
US 20190213482A1 · Socher et al. · 2019 [cited by applicant]
US 20190251431A1 · Keskar et al. · 2019 [cited by applicant]
US 20190258939A1 · Min et al. · 2019 [cited by applicant]
US 20190286073A1 · Asl et al. · 2019 [cited by applicant]
US 20190355270A1 · Mccann et al. · 2019 [cited by applicant]
US 20190362246A1 · Lin et al. · 2019 [cited by applicant]
US 20200005765A1 · Zhou et al. · 2020 [cited by applicant]
US 20200065651A1 · Merity et al. · 2020 [cited by applicant]
US 20200090033A1 · Ramachandran et al. · 2020 [cited by applicant]
US 20200090034A1 · Ramachandran et al. · 2020 [cited by applicant]
US 20200103911A1 · Ma et al. · 2020 [cited by applicant]
US 20200104643A1 · Hu et al. · 2020 [cited by applicant]
US 20200104699A1 · Zhou et al. · 2020 [cited by applicant]
US 20200105272A1 · Wu et al. · 2020 [cited by applicant]
US 20200117854A1 · Lu et al. · 2020 [cited by applicant]
US 20200117861A1 · Bradbury · 2020 [cited by applicant]
US 20200142917A1 · Paulus · 2020 [cited by applicant]
US 20200175305A1 · Trott et al. · 2020 [cited by applicant]
US 20200234113A1 · Liu · 2020 [cited by applicant]
US 20200272940A1 · Sun et al. · 2020 [cited by applicant]
US 20200285704A1 · Rajani et al. · 2020 [cited by applicant]
US 20200285705A1 · Zheng et al. · 2020 [cited by applicant]
US 20200285706A1 · Singh et al. · 2020 [cited by applicant]
US 20200285993A1 · Liu et al. · 2020 [cited by applicant]
US 20200302178A1 · Gao et al. · 2020 [cited by applicant]
US 20200330026A1 · Zhu · 2020 [cited by examiner]
US 20200334334A1 · Keskar et al. · 2020 [cited by applicant]
US 20200364299A1 · Niu et al. · 2020 [cited by applicant]
US 20200364542A1 · Sun · 2020 [cited by applicant]
US 20200364580A1 · Shang et al. · 2020 [cited by applicant]
US 20200372116A1 · Gao et al. · 2020 [cited by applicant]
US 20200372319A1 · Sun et al. · 2020 [cited by applicant]
US 20200372339A1 · Che et al. · 2020 [cited by applicant]
US 20200372341A1 · Asai et al. · 2020 [cited by applicant]
US 20200372632A1 · Chauhan · 2020 [cited by examiner]
US 20200380213A1 · Mccann et al. · 2020 [cited by applicant]
US 20210042604A1 · Hashimoto et al. · 2021 [cited by applicant]
US 20210049236A1 · Nguyen et al. · 2021 [cited by applicant]
US 20210073459A1 · Mccann et al. · 2021 [cited by applicant]
US 20210089588A1 · Le et al. · 2021 [cited by applicant]
US 20210089882A1 · Sun et al. · 2021 [cited by applicant]
US 20210089883A1 · Li et al. · 2021 [cited by applicant]
Louizos et al., “Learning Sparse Neural Networks Through L0 Regularization,” Jun. 22, 2018, pp. 1-39. (Year: 2018). [cited by examiner]
Zhou et al., “Theory-Inspired Path-Regularized Differential Network Architecture Search,” 2020, pp. 1-12 (Year: 2020). [cited by examiner]
Chu et al., “Fair DARTS: Eliminating Unfair Advantages in Differentiable Architecture Search,” arXiv preprint arXiv: 1911.12126, 2019, pp. 1-36. Provided by applicant on May 3, 2021 (Year: 2019). [cited by examiner]
Yu et al. “Exploiting sparseness in deep neural networks for large vocabulary speech recognition,” 2012, 2012 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 4409-4412. (Year: 2012… [cited by examiner]
Alessandro, Rinaldo, “Advanced statistics theory I,” UC Berkeley Lecture, Available Online at <http://www.stat.cmu.edu/˜arinaldo/Teaching/36755/F17/Scribed_Lectures/F17_0911.pdf>, Lecture 3: Sep. 11, 2017, 5 pages. [cited by applicant]
Allen-Zhu et al., “A convergence theory for deep learning via over-parameterization,” In Proc. Int'l Conf. Machine Learning, Jun. 17, 2019, 53 pages. [cited by applicant]
Balduzzi et al., “The shattered gradients problem: If resnets are the answer, then what is the question?,” In Proceedings of the 34th International Conference on Machine Learning, Sydney, Australia, PMLR 70, 2017, pp. 3… [cited by applicant]
Cai et al., “Path-level network transformation for efficient architecture search,” In Proceedings of the 35th International Conference on Machine Learning, Stockholm, Sweden, PMLR 80, 2018, 12 pages. [cited by applicant]
Cai et al., “Proxylessnas: Direct neural architecture search on target task and hardware,” In Int'l Conf. Learning Representations, 2018, 13 pages. [cited by applicant]
Chen et al., “Progressive Differentiable Architecture Search: Bridging the Depth Gap between Search and Evaluation,” In IEEE International Conference on Computer Vision, 2019, 10 pages. [cited by applicant]
Chu et al., “Fair DARTS: Eliminating unfair advantages in differentiable architecture search,” arXiv preprint arXiv:1911.12126, 2019, 36 pages. [cited by applicant]
DeVries et al., “Improved regularization of convolutional neural networks with cutout,” arXiv preprint arXiv:1708.04552, 2017, 8 pages. [cited by applicant]
Dong et al., “Searching for A Robust Neural Architecture in Four GPU Hours,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 1761-1770. [cited by applicant]
Du et al., “Gradient Descent Finds Global Minima of Deep Neural Networks,” In Proceedings of the 36th International Conference on Machine Learning, Long Beach, California, PMLR 97, 2019, 45 pages. [cited by applicant]
Du et al., “Gradient descent provably optimizes over-parameterized neural networks,” In Int'l Conf. Learning Representations, 2018, 19 pages. [cited by applicant]
Dyson, “Statistical theory of the energy levels of complex systems I,” Journal of Mathematical Physics, vol. 3, No. 1, 1962, pp. 140-156. [cited by applicant]
He et al., “Deep Residual Learning for Image Recognition,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2016, 12 pages. [cited by applicant]
He et al., “Identity Mappings in Deep Residual Networks,” In Proc. European Conf. Computer Vision, 2016, pp. 630-645. [cited by applicant]
Howard et al., “MobileNets: Efficient Convolutional Neural Networks for Mobile Vision Applications,” arXiv preprint arXiv:1704.04861, Apr. 17, 2017, 9 pages. [cited by applicant]
Huang et al., “Densely Connected Convolutional Networks,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2017, 9 pages. [cited by applicant]
Hwang, “Cauchy's interlace theorem for eigenvalues of hermitian matrices,” The American Mathematical Monthly, vol. 111, No. 2, Feb. 2004, pp. 157-159. [cited by applicant]
Kingma et al., “ADAM: A method for stochastic optimization,” Int'l Conf Learning Representations, 2015, 15 pages. [cited by applicant]
Krizhevsky et al., “Learning multiple layers of features from tiny images,” Apr. 8, 2009, 63 pages. [cited by applicant]
Liang et al., “DARTS+: Improved differentiable architecture search with early stopping,” arXiv preprint arXiv:1909.06035, 2019, 15 pages. [cited by applicant]
Liu et al., “DARTS: Differentiable architecture search,” In Int'l Conf. Learning Representations, 2018, 13 pages. [cited by applicant]
Liu et al., “Progressive Neural Architecture Search,” In Proc. European Conf. Computer Vision, 2018, 16 pages. [cited by applicant]
Loshchilov et al., “SGDR: Stochastic gradient descent with warm restarts,” arXiv preprint arXiv:1608.03983, In Int'l Conf. Learning Representations, 2017, 16 pages. [cited by applicant]
Louizos, et al., “Learning sparse neural networks through Lo regularization,” In Int'l Conf. Learning Representations, 2018, 13 pages. [cited by applicant]
Ma et al., “ShuffleNet V2: Practical Guidelines for Efficient CNN Architecture Design,” In Proc. European Conf. Computer Vision, 2018, pp. 116-131. [cited by applicant]
Maddison et al., “A* sampling,” In Proc. Conf. Neural Information Processing Systems, 2014, 19, pages. [cited by applicant]
Orhan et al., “Skip connections eliminate singularities,” In Int'l Conf. Learning Representations 2018, arXiv preprint arXiv:1701.09175, 2018, 22 pages. [cited by applicant]
Pham et al., “Efficient neural architecture search via parameter sharing,” In Proc. Int'l Conf. Machine Learning, 2018, 11 pages. [cited by applicant]
Real et al., “Regularized evolution for image classifier architecture search,” In AAAI Conference on Artificial Intelligence, vol. 33, 2019, pp. 4780-4789. [cited by applicant]
Russakovsky et al., “ImageNet Large Scale Visual Recognition Challenge,” Int' l. J. Computer Vision, vol. 115, No. 3, 2015, 43 pages. [cited by applicant]
Saw et al., “Chebyshev Inequality with Estimated Mean and Variance,” The American Statistician, vol. 38, No. 2, 1984, pp. 130-132. [cited by applicant]
Shu et al., “Understanding architectures learnt by cell-based neural architecture search,” In Int'l Conf. Learning Representations, 2020, 21 pages. [cited by applicant]
Srivastava et al., “Dropout: a simple way to prevent neural networks from overfitting,” J. of Machine Learning Research, vol. 15, No. 1, 2014, pp. 1929-1958. [cited by applicant]
Tan et al., “Mnasnet: Platform-aware neural architecture search for mobile,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 2820-2828. [cited by applicant]
Tian, “An analytical formula of population gradient for two-layered relu network and its applications in convergence and critical point analysis,” In Proc. Int'l Conf Machine Learning, 2017, pp. 3404-3413. [cited by applicant]
Wu et al., “FBnet: Hardware-aware efficient convnet design via differentiable neural architecture search,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2019, pp. 10734-10742. [cited by applicant]
Xie et al., “SNAS: stochastic neural architecture search,” In Int'l Conf. Learning Representations, 2019, arXiv:1812.09926v3, 2019, 17 pages. [cited by applicant]
Xu et al., “PC-DARTS: Partial channel connections for memory-efficient architecture search,” In Int'l Conf Learning Representations, 2019, 13 pages. [cited by applicant]
Zela et al., “Understanding and robustifying differentiable architecture search,” In Int'l Conf. Learning Representations, 2020, 28 pages. [cited by applicant]
Zhou et al., “Bayesnas: A bayesian approach for neural architecture search,” In Proc. Int'l Conf. Machine Learning, 2019, 25 pages. [cited by applicant]
Zoph et al., “Learning transferable architectures for scalable image recognition,” In Proc. IEEE Conf. Computer Vision and Pattern Recognition, 2018, 14, pages. [cited by applicant]
Zoph et al., “Neural architecture search with reinforcement learning,” In Int'l Conf. Learning Representations, 2017, 16 pages. [cited by applicant]