Automated generation of machine learning models
This document relates to automated generation of machine learning models, such as neural networks. One example system includes a hardware processing unit and a storage resource. The storage resource can store computer-readable instructions cause the hardware processing unit to perform an iterative model-growing process that involves modifying parent models to obtain child models. The iterative model-growing process can also include selecting candidate layers to include in the child models based at least on weights learned in an initialization process of the candidate layers. The system can also output a final model selected from the child models.
1 . A method performed on a computing device, the method comprising:
outputting a graphical user interface having a first graphical element for designating an operation search space of operations available to be performed by candidate layers in an iterative model-growing process;
receiving first user input via the first graphical element, the first user input identifying a set of multiple operations to include in the operation search space;
adding an initial parent model to a parent model pool;
performing two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:
selecting a particular parent model having a plurality of layers from the parent model pool;
inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the first user input;
initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;
based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;
training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;
evaluating the trained child models using one or more criteria; and
based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and
after the two or more iterations, selecting at least one trained child model as a final model and outputting the final model.
2 . The method of claim 1 , further comprising:
receiving second user input via a second graphical element of the graphical user interface, wherein the second user input designates a default model, designates a randomly-generated model, or navigates to an existing model to use as the initial parent model.
3 . The method of claim 1 , wherein the first user input received via the first graphical element of the graphical user interface selects at least two different convolution operations, and the respective operations performed by the respective candidate layers are selected randomly from the operation search space.
4 . The method of claim 1 , wherein the first user input received via the first graphical element of the graphical user interface selects at least two different pooling operations, and the respective operations performed by the respective candidate layers are selected randomly from the operation search space.
5 . The method of claim 1 , further comprising:
receiving second user input directed to a second graphical element of the graphical user interface, the second user input identifying a specified amount of computational resources to use for the iterative model-growing process; and
responsive to expending the specified amount of computational resources, stopping the iterative model-growing process and selecting the final model.
6 . The method of claim 5 , wherein the second user input directed to the second graphical element designates a number of GPU-days to expend for the iterative model-growing process.
7 . The method of claim 5 , wherein the second user input directed to the second graphical element designates a length of time to expend for the iterative model-growing process.
8 . The method of claim 1 , further comprising:
receiving second user input directed to a second graphical element of the graphical user interface, the second user input directed to the second graphical element identifying model size as a particular criterion for evaluating the trained child models; and
designating the individual trained child model as the new parent model based at least on a model size of the individual trained child model.
9 . The method of claim 1 , further comprising:
receiving second user input directed to a second graphical element of the graphical user interface, the second user input directed to the second graphical element specifying connectivity parameters for the child models; and
generating the child models according to the connectivity parameters.
10 . The method of claim 9 , wherein the connectivity parameters specified by the second user input indicate a number of previous layers that contribute to the one or more selected candidate layers of the child models.
11 . The method of claim 9 , wherein the connectivity parameters specified by the second user input directed to the second graphical element indicate whether skip connections are employed in the child models.
12 . The method of claim 1 , wherein the first user input selects a group of at least two different convolutional kernel sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.
13 . The method of claim 1 , wherein the first user input selects a group of at least two different convolutional stride sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.
14 . The method of claim 1 , wherein the first user input selects a group of pooling operations available to be performed by the respective candidate layers when performing the iterative model-growing process, the group including at least a max pooling operation and an average pooling operation.
15 . The method of claim 1 , wherein the first user input selects a group of at least two different pooling window sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.
16 . A system comprising:
a hardware processing unit; and
a storage resource storing computer-readable instructions which, when executed by the hardware processing unit, cause the hardware processing unit to:
receive first user input via a first graphical element of a graphical user interface, the first user input identifying a set of multiple operations to include in an operation search space for an iterative model-growing process;
add an initial parent model to a parent model pool;
perform two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:
selecting a particular parent model having a plurality of layers from the parent model pool;
inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the first user input;
initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;
based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;
training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;
evaluating the trained child models using one or more criteria; and
based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and
after the two or more iterations, select at least one trained child model as a final model and outputting the final model.
17 . The system of claim 16 , wherein the operation search space includes multiple convolution operations and multiple pooling operations designated by the first user input that is received via the first graphical element of the graphical user interface.
18 . The system of claim 16 , wherein the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to:
configure the respective candidate layers to perform respective operations randomly selected from the set of multiple operations identified by the first user input.
19 . The system of claim 18 , wherein the first user input directed to the first graphical element identifies at least two different convolution operations and at least two different pooling operations from which the respective operations are randomly selected.
20 . The system of claim 19 , wherein the first user input directed to the first graphical element specifies at least two different window sizes and at least two different strides for the at least two different convolution operations.
21 . A computer-readable storage medium storing instructions which, when executed by a processing device, cause the processing device to perform acts comprising:
outputting a graphical user interface having a first graphical element for designating an initial parent model and a second graphical element for designating an operation search space for an iterative model-growing process;
receiving first user input via the first graphical element, the first user input identifying a particular machine learning model as the initial parent model;
receiving second user input via the second graphical element, the second user input identifying a set of multiple operations to include in the operation search space;
configuring the operation search space of the iterative model-growing process to be restricted to the set of multiple operations identified by the second user input received via the second graphical element of the graphical user interface;
adding the initial parent model to a parent model pool;
performing two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:
selecting a particular parent model having a plurality of layers from the parent model pool;
inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the second user input;
initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;
based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;
training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;
evaluating the trained child models using one or more criteria; and
based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and
after the two or more iterations, selecting at least one trained child model as a final model and outputting the final model.