IP Library › Granted Patent US 12,314,856
Granted Patent B2
US 12,314,856 · App. 18/612,917 · Granted May 27, 2025

Population based training of neural networks

Inventors: Maxwell Elliot Jaderberg (London, GB); Wojciech Czarnecki (London, GB); Timothy Frederick Goldie Green (London, GB); Valentin Clement Dalibard (London, GB)
Assignee: DeepMind Technologies Limited
G06N3/08G06N3/048
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,856
App. No.
18/612,917
Granted
May 27, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for training a neural network. A method includes: training a neural network having a plurality of network parameters to perform a particular neural network task and to determine trained values of the network parameters using an iterative training process having a plurality of hyperparameters, the method comprising: maintaining a plurality of candidate neural networks and, for each of the candidate neural networks, data specifying: (i) respective values of the network parameters for the candidate neural network, (ii) respective values of the hyperparameters for the candidate neural network, and (iii) a quality measure that measures a performance of the candidate neural network on the particular neural network task; and for each of the plurality of candidate neural networks, repeatedly performing additional training operations.

Claims (65)

1. A method of training a neural network having a plurality of network parameters to perform a particular neural network task and to determine trained values of the network parameters using an iterative training process having a plurality of hyperparameters, the method comprising:

maintaining a plurality of candidate neural networks and, for each of the plurality of candidate neural networks, data specifying: (i) values of the network parameters of the candidate neural network, (ii) values of the hyperparameters of the candidate neural network, and (iii) a quality measure that measures a performance of the candidate neural network on the particular neural network task;

for each of the plurality of candidate neural networks, repeatedly performing the following training operations, comprising:

repeatedly updating the values of the network parameters of the candidate neural network in accordance with the maintained values of the hyperparameters of the candidate neural network until a termination criterion is satisfied, wherein the maintained values of hyperparameters remain unchanged;

updating the quality measure of the candidate neural network based on the updated values of the network parameters of the candidate neural network;

updating the respective values of the hyperparameters of the candidate neural network based on the updated quality measure of the candidate neural network, the updating comprising:

sampling another candidate neural network from the plurality of candidate neural networks, the other candidate neural network having respective values of the network parameters and a respective quality measure for the respective values of the network parameters; and

updating the hyperparameters of the candidate neural network based on a result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network;

updating the maintained data of the candidate neural network to specify the updated values of the hyperparameters, the updated values of the network parameters, and the updated value of the quality measure; and

selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures for the plurality of candidate neural networks after the training operations have repeatedly been performed.

2. The method of claim 1 , wherein selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures of the plurality of candidate neural networks comprises:

selecting the maintained parameter values of the candidate neural network having a best maintained quality measure of any of the plurality of candidate neural networks after the training operations have repeatedly been performed.

3. The method of claim 1 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the maintained values of the hyperparameters of the other candidate neural network.

4. The method of claim 3 , further comprising: setting new values of the network parameters of the candidate neural network to the respective values of the network parameters of the other candidate neural network.

5. The method of claim 1 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is not greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the updated values of the hyperparameters of the candidate neural network.

6. The method of claim 5 , further comprising: perturbing the new values of the hyperparameters of the candidate neural network according to a predetermined factor or probability distribution.

7. The method of claim 1 , further comprising:

providing the trained values of the network parameters for use in processing new inputs to the neural network.

8. A system comprising:

one or more computers and one or more storage devices on which are stored instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations for training a neural network having a plurality of network parameters to perform a particular neural network task and to determine trained values of the network parameters using an iterative training process having a plurality of hyperparameters, the operations comprising:

maintaining a plurality of candidate neural networks and, for each of the plurality of candidate neural networks, data specifying: (i) values of the network parameters of the candidate neural network, (ii) values of the hyperparameters of the candidate neural network, and (iii) a quality measure that measures a performance of the candidate neural network on the particular neural network task;

for each of the plurality of candidate neural networks, repeatedly performing the following training operations, comprising:

repeatedly updating the values of the network parameters of the candidate neural network in accordance with the maintained values of the hyperparameters of the candidate neural network until a termination criterion is satisfied, wherein the maintained values of hyperparameters remain unchanged;

updating the quality measure of the candidate neural network based on the updated values of the network parameters of the candidate neural network;

updating the respective values of the hyperparameters of the candidate neural network based on the updated quality measure of the candidate neural network, the updating comprising:

sampling another candidate neural network from the plurality of candidate neural networks, the other candidate neural network having respective values of the network parameters and a respective quality measure for the respective values of the network parameters; and

updating the hyperparameters of the candidate neural network based on a result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network;

updating the maintained data of the candidate neural network to specify the updated values of the hyperparameters, the updated values of the network parameters, and the updated value of the quality measure; and

selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures for the plurality of candidate neural networks after the training operations have repeatedly been performed.

9. The system of claim 8 , wherein selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures of the plurality of candidate neural networks comprises:

selecting the maintained parameter values of the candidate neural network having a best maintained quality measure of any of the plurality of candidate neural networks after the training operations have repeatedly been performed.

10. The system of claim 8 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the maintained values of the hyperparameters of the other candidate neural network.

11. The system of claim 10 , wherein the operations further comprise: setting new values of the network parameters of the candidate neural network to the respective values of the network parameters of the other candidate neural network.

12. The system of claim 8 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is not greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the updated values of the hyperparameters of the candidate neural network.

13. The system of claim 12 , wherein the operations further comprise: perturbing the new values of the hyperparameters of the candidate neural network according to a predetermined factor or probability distribution.

14. The system of claim 8 , wherein the operations further comprise:

providing the trained values of the network parameters for use in processing new inputs to the neural network.

15. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more computers, cause the one or more computers to perform operations for training a neural network having a plurality of network parameters to perform a particular neural network task and to determine trained values of the network parameters using an iterative training process having a plurality of hyperparameters, the operations comprising:

maintaining a plurality of candidate neural networks and, for each of the plurality of candidate neural networks, data specifying: (i) values of the network parameters of the candidate neural network, (ii) values of the hyperparameters of the candidate neural network, and (iii) a quality measure that measures a performance of the candidate neural network on the particular neural network task;

for each of the plurality of candidate neural networks, repeatedly performing the following training operations, comprising:

repeatedly updating the values of the network parameters of the candidate neural network in accordance with the maintained values of the hyperparameters of the candidate neural network until a termination criterion is satisfied, wherein the maintained values of hyperparameters remain unchanged;

updating the quality measure of the candidate neural network based on the updated values of the network parameters of the candidate neural network;

updating the respective values of the hyperparameters of the candidate neural network based on the updated quality measure of the candidate neural network, the updating comprising:

sampling another candidate neural network from the plurality of candidate neural networks, the other candidate neural network having respective values of the network parameters and a respective quality measure for the respective values of the network parameters; and

updating the hyperparameters of the candidate neural network based on a result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network;

updating the maintained data of the candidate neural network to specify the updated values of the hyperparameters, the updated values of the network parameters, and the updated value of the quality measure; and

selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures for the plurality of candidate neural networks after the training operations have repeatedly been performed.

16. The one or more non-transitory computer-readable storage media of claim 15 , wherein selecting the trained values of the network parameters from the parameter values in the maintained data based on the maintained quality measures of the plurality of candidate neural networks comprises:

selecting the maintained parameter values of the candidate neural network having a best maintained quality measure of any of the plurality of candidate neural networks after the training operations have repeatedly been performed.

17. The one or more non-transitory computer-readable storage media of claim 15 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the maintained values of the hyperparameters of the other candidate neural network.

18. The one or more non-transitory computer-readable storage media of claim 17 , wherein the operations further comprise: setting new values of the network parameters of the candidate neural network to the respective values of the network parameters of the other candidate neural network.

19. The one or more non-transitory computer-readable storage media of claim 15 , wherein updating the hyperparameters of the candidate neural network based on the result of comparing the updated quality measure of the candidate neural network and the respective quality measure of the other candidate neural network comprises:

determining whether the respective quality measure of the other candidate neural network is greater than the updated quality measure of the candidate neural network; and

in response to determining that the respective quality measure of the other candidate neural network is not greater than the updated quality measure of the candidate neural network, setting new values of the hyperparameters of the candidate neural network to the updated values of the hyperparameters of the candidate neural network.

20. The one or more non-transitory computer-readable storage media of claim 19 , wherein the operations further comprise: perturbing the new values of the hyperparameters of the candidate neural network according to a predetermined factor or probability distribution.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 071498/0210 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2024
From: JADERBERG, MAXWELL ELLIOT; CZARNECKI, WOJCIECH; GREEN, TIMOTHY FREDERICK GOLDIE; DALIBARD, VALENTIN CLEMENT
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 067303/0656 →
Continuity (4)
Continuation 18120715 · Mar 13, 2023
Continuation 16766631
Provisional Application 62590177 · Nov 22, 2017
Related Publication 20240346310A1 · Oct 17, 2024
References Cited (82)
US 6523016B1 · Michalski · 2003 [cited by applicant]
US 11132609B2 · Pascanu · 2021 [cited by examiner]
US 20190034762A1 · Hashimoto · 2019 [cited by applicant]
US 20190370648A1 · Zoph et al. · 2019 [cited by applicant]
CN 107229972 · 2017 [cited by applicant]
CN 107301459 · 2017 [cited by applicant]
EP 1557788 · 2005 [cited by applicant]
Abadi et al, “TensorFlow: A system for large-scale machine learning”, 12th USENIX Symposium on Operating Systems Design and Implementation, Nov. 2016, 21 pages. [cited by applicant]
Agarwal et al, “Oracle inequalities for computationally budgeted model selection,” Proceedings of the 24th Annual Conference on Learning Theory, 2011, pp. 69-86. [cited by applicant]
Back et al, “An overview of parameter control methods by self-adaptation in evolutionary algorithms,” Fundamenta Informaticae, 1998, 35(1-4):51-66. [cited by applicant]
Beattie et al, “Deepmind lab,” arXiv preprint arXiv:1612.03801, Dec. 2016, 11 pages. [cited by applicant]
Bellemare et al, “The arcade learning environment: An evaluation platform for general agents,” J. Artif. Intell. Res.(JAIR), Jun. 2013, pp. 47:253-279. [cited by applicant]
Bengio, “Gradient-Based Optimization of Hyperparameters,” Neural Computation, Sep. 1999, 18 pages. [cited by applicant]
Bergstra et al, “Algorithms for hyper-parameter optimization,” Advances in Neural Information Processing Systems, Dec. 2011, pp. 2546-2554. [cited by applicant]
Bergstra et al, “Random search for hyper-parameter optimization,” Journal of Machine Learning Research, Feb. 2012, 281-305. [cited by applicant]
Castillo et al, “G-prop-III: Global optimization of multilayer perceptrons using an evolutionary algorithm,” Proceedings of the 1st Annual Conference on Genetic and Evolutionary Computation, Jul. 1999, 8 pages. [cited by applicant]
Castillo et al, “Lamarckian evolution and the Baldwin effect in evolutionary neural networks,” arXiv preprint cs/0603004, Mar. 2006, 5 pages. [cited by applicant]
Claesen et al., “Hyperparameter Search in Machine Learning,” https://arxiv.org/abs/1502.02127v1, Feb. 2015, 5 pages. [cited by applicant]
Clune et al, “Natural selection fails to optimize mutation rates for long-term adaptation on rugged fitness landscapes,” PLoS Computational Biology, Sep. 2008, 4(9):e1000187: 8 pages. [cited by applicant]
Domhan et al, “Speeding up automatic hyperparameter optimization of deep neural networks by extrapolation of learning curves,” IJCAI, Jun. 2015, pp. 3460-3468. [cited by applicant]
Extended Search Report in European Appln. No. 24172554.8, dated Aug. 5, 2024, 14 pages. [cited by applicant]
Fernando et al, “Convolution by Evolution: Differentiable Pattern Producting Networks,” arxiv.org, Cornell University Library, Jun. 2016, 109-116. [cited by applicant]
Friedrichs et al, “Evolutionary tuning of multiple SVM parameters” Elxevier Science, Oct. 2004, 13 pages. [cited by applicant]
Gagliolo et al., “Learning dynamic algorithm portfolios,” Annals of Mathematics and Artificial Intelligence, Aug. 2006, 47(3):295-328. [cited by applicant]
Gloger, “Self-adaptive evolutionary algorithms,” Universität Paderborn, Jan. 2004, 17 pages. [cited by applicant]
Golovin et al, “Google vizier: A service for black-box optimization,” Proceedings of the 23rd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, ACM, Aug. 2017, 1487-1496. [cited by applicant]
Gonzalez et al, “Batch Bayesian optimization via local penalization,” Artificial Intelligence and Statistics, May 2016, pp. 648-657. [cited by applicant]
Goodfellow et al, “Generative adversarial nets,” Advances in Neural Information Processing Systems, Dec. 2014, 9 pages. [cited by applicant]
Gruau et al., “Adding Learning to the Cellular Development of Neural Networks: Evolution and the Baldwin Effect,” Evolutionary Computation, 1993, 1(3):213-233. [cited by applicant]
Gulrajani et al, “Improved training of Wasserstein GANs,” Advances in Neural Information Processing Systems, Mar. 2017, 11 pages. [cited by applicant]
György et al., “Efficient multi-start strategies for local search algorithms,” Journal of Artificial Intelligence Research, Jul. 2011, pp. 407-444. [cited by applicant]
Hansen et al, “Completely Derandomized Self-Adaptation in Evolution Strategies” Evolutionary Computation, Jun. 2001, 9(2):39 pages. [cited by applicant]
Hinton et al., “Lecture 6.5-rmsprop: Divide the gradient by a running average of its recent magnitude,” COURSERA: Neural networks for machine learning, retrieved from URL <https://www.cs.toronto.edu/˜tijmen/csc321/slide… [cited by applicant]
Houck et al, “Empirical investigation of the benefits of partial lamarckianism,” Evolutionary Computation, Mar. 1997, 5(1):31-60. [cited by applicant]
Hüsken et al, “Optimization for problem classes-neural networks that learn to learn,” Combinations of Evolutionary Computation and Neural Networks, May 2000, 12 pages. [cited by applicant]
Hutter et al, “Sequential model-based optimization for general algorithm configuration,” LION, 2011, pp. 5:507-523. [cited by applicant]
Igel et al, “Evolutionary optimization of neural systems: The use of strategy adaptation,” Trends and Applications in Constructive Approximation, 2005, 23 pages. [cited by applicant]
Intention to Grant Patent in European Appln. No. 18807623.6, dated Oct. 23, 2023, 18 pages. [cited by applicant]
Jaderberg et al, “Reinforcement learning with unsupervised auxiliary tasks,” arXiv preprint arXiv:1611.05397, Nov. 2016, 14 pages. [cited by applicant]
Jaderberg et al., “Population Based Training of Neural Networks,” https://arxiv.org/abs/1711.09846v2, Nov. 2017, 21 pages. [cited by applicant]
Kingma et al., “Adam: A method for stochastic optimization,” https://arxiv.org/abs/1412.6980v8, Jul. 2015, 15 pages. [cited by applicant]
Klein et al, “Fast Bayesian optimization of machine learning hyperparameters on large datasets,” arXiv preprint arXiv:1605.07079, May 2016, 17 pages. [cited by applicant]
Ku et al., “Exploring the effects of Lamarckian and Baldwinian learning in evolving recurrent neural networks,” Proceedings of 1997 IEEE International Conference on Evolutionary Computation, Apr. 1997, 5 pages. [cited by applicant]
Li et al, “Hyperband: A Novel Bandit-Based Approach to Hyperparameter Optimization,” https://arxiv.org/abs/1603.06560v3, Nov. 2016, 48 pages. [cited by applicant]
Liu et al, “Hierarchical representations for efficient architecture search,” https://arxiv.org/abs/1711.00436v1, Nov. 2017, 13 pages. [cited by applicant]
Loshchilov et al., “CMA-ES for Hyperparameter Optimization of Deep Neural Networks,” https://arxiv.org/abs/1604.07269, Apr. 2016, 8 pages. [cited by applicant]
Loshchilov et al., “SGDR: Stochastic Gradient Descent with Warm Restarts,” https://arxiv.org/abs/1608.03983, 2016, 16 pages. [cited by applicant]
Masse et al., “Speed learning on the fly,” arXiv preprint arXiv:1511.02540, Nov. 2015, 25 pages. [cited by applicant]
Mnih et al., “Asynchronous methods for deep reinforcement learning,” International Conference on Machine Learning, Jun. 2016, 10 pages. [cited by applicant]
Office Action in European Appln. No. 18807623.6, dated Jul. 13, 2021, 11 pages. [cited by applicant]
Office Action in European Appln. No. 18807623.6, dated Sep. 18, 2023, 19 pages. [cited by applicant]
Panayotov et al, “Librispeech: An ASR corpus based on public domain audio books,” 2015 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2015, 5 pages. [cited by applicant]
PCT International Preliminary Report on Patentability in International Appln. No. PCT/EP2018/082162, mailed Jun. 4, 2020, 12 pages. [cited by applicant]
PCT International Search Report and Written Opinion in International Appln. No. PCT/EP2018/082162, mailed Feb. 27, 2019, 19 pages. [cited by applicant]
Radford et al, “Unsupervised representation learning with deep convolutional generative adversarial networks,” https://arxiv.org/abs/1511.06434, last revised Jan. 2016, 16 pages. [cited by applicant]
Rasley et al, “Hyperdrive: Exploring hyperparameters with POP scheduling,” Proceedings of the 18th International Middleware Conference, Dec. 2017, 13 pages. [cited by applicant]
Real et al, “Large-scale evolution of image classifiers,” arXiv preprint arXiv:1703.01041, Jun. 2017, 18 pages. [cited by applicant]
Real et al, “Regularized Evolution for Image Classifier Architecture Search” arXiv, Oct. 2018, 16 pages. [cited by applicant]
Renders et al., “Hybrid methods using genetic algorithms for global optimization,” IEEE Transactions on Systems, Man, and Cybernetics, Part B (Cybernetics), Apr. 1996, 26(2):243-258. [cited by applicant]
Rosca et al, “Variational approaches for auto-encoding generative adversarial networks,” arXiv preprint arXiv:1706.04987, Oct. 2017, 21 pages. [cited by applicant]
Sabharwal et al, “Selecting Near-Optimal Learners via Incremental Data Allocation,” AAAI, Mar. 2016, 9 pages. [cited by applicant]
Salimans et al, “Improved techniques for training GANs,” Advances in Neural Information Processing Systems, Dec. 2016, 9 pages. [cited by applicant]
Salustowicz et al., “Probabilistic Incremental Program Evolution: Stochastic Search Through Program Space,” European Conference on Machine Learning, 1997, 213-220. [cited by applicant]
Shah et al., “Parallel predictive entropy search for batch global optimization of expensive objective functions,” Advances in Neural Information Processing Systems, Dec. 2015, 9 pages. [cited by applicant]
Smith, “Cyclical learning rates for training neural networks,” https://arxiv.org/abs/1506.01186, Apr. 2017, 10 pages. [cited by applicant]
Snoek et al, “Practical Bayesian optimization of machine learning algorithms,” Advances in neural information processing systems, 2012, 9 pages. [cited by applicant]
Snoek et al, “Scalable Bayesian optimization using deep neural networks,” International Conference on Machine Learning, Jun. 2015, 10 pages. [cited by applicant]
Spears, “Adapting crossover in evolutionary algorithms,” Proceedings of the Fourth Annual Conference on Evolutionary Programming, Mar. 1995, 18 pages. [cited by applicant]
Springenberg et al, “Bayesian optimization with robust Bayesian neural networks,” Advances in Neural Information Processing Systems, Dec. 2016, pp. 9 pages. [cited by applicant]
Srinivas et al, “Gaussian process optimization in the bandit setting: No regret and experimental design,” arXiv preprint arXiv:0912.3995, 2009, 17 pages. [cited by applicant]
Swersky et al, “Freeze-thaw Bayesian optimization,” arXiv preprint arXiv:1406.3896, Jun. 2014, 12 pages. [cited by applicant]
Swersky et al, “Multi-task Bayesian optimization,” Advances in neural information processing systems, Jan. 2013, 9 pages. [cited by applicant]
Van den Oord et al, “Wavenet: A Generative Model for Raw Audio” https://arxiv.org/abs/1609.03499, Sep. 2016, 15 pages. [cited by applicant]
Vaswani et al, “Attention is all you need,” arXiv preprint arXiv:1706.03762, 2017, 11 pages. [cited by applicant]
Vezhnevets et al, “Feudal networks for hierarchical reinforcement learning,” arXiv preprint arXiv:1703.01161, Mar. 2017, 12 pages. [cited by applicant]
Vinyals et al., “Starcraft II: A new challenge for reinforcement learning,” arXiv preprint arXiv:1708.04782, Aug. 2017, 20 pages. [cited by applicant]
Welch, “The generalization of ‘student's’ problem when several different population variances are involved,” Biometrika, Jan. 1947, 34(1/2):28-35. [cited by applicant]
Wu et al., “The parallel knowledge gradient method for batch Bayesian optimization,” Advances in Neural Information Processing Systems, Dec. 2016, 9 pages. [cited by applicant]
Xue et al, “A survey on evolutionary computation approaches to feature selection,” IEEE Transactions on Evolutionary Computation, Aug. 2016, 20(4):606-625. [cited by applicant]
Yang et al, “LR-GAN: Layered recursive generative adversarial networks for image generation,” https://arxiv.org/abs/1703.01560, Aug. 2017, 21 pages. [cited by applicant]
Young et al., “Optimizing deep learning hyper-parameters through an evolutionary algorithm,” Proceedings of the Workshop on Machine Learning in High-Performance Computing Environments, Nov. 2015, 5 pages. [cited by applicant]
Zhang et al., “Evolutionary computation meets machine learning: A survey,” IEEE Computational Intelligence Magazine, Nov. 2011, 6(4):68-75. [cited by applicant]