IP Library › Granted Patent US 12,626,141
Granted Patent B2
US 12,626,141 · App. 18/080,407 · Granted May 12, 2026

Automated generation of machine learning models

Inventors: Debadeepta Dey (Kenmore, WA); Hanzhang Hu (Pittsburg, PA); Richard A. Caruana (Woodinville, WA); John C. Langford (Scarsdale, NY); Eric J. Horvitz (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
G06N3/086G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,141
App. No.
18/080,407
Granted
May 12, 2026
Kind
B2
Abstract

This document relates to automated generation of machine learning models, such as neural networks. One example system includes a hardware processing unit and a storage resource. The storage resource can store computer-readable instructions cause the hardware processing unit to perform an iterative model-growing process that involves modifying parent models to obtain child models. The iterative model-growing process can also include selecting candidate layers to include in the child models based at least on weights learned in an initialization process of the candidate layers. The system can also output a final model selected from the child models.

Claims (68)

1 . A method performed on a computing device, the method comprising:

outputting a graphical user interface having a first graphical element for designating an operation search space of operations available to be performed by candidate layers in an iterative model-growing process;

receiving first user input via the first graphical element, the first user input identifying a set of multiple operations to include in the operation search space;

adding an initial parent model to a parent model pool;

performing two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:

selecting a particular parent model having a plurality of layers from the parent model pool;

inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the first user input;

initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;

based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;

training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;

evaluating the trained child models using one or more criteria; and

based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and

after the two or more iterations, selecting at least one trained child model as a final model and outputting the final model.

2 . The method of claim 1 , further comprising:

receiving second user input via a second graphical element of the graphical user interface, wherein the second user input designates a default model, designates a randomly-generated model, or navigates to an existing model to use as the initial parent model.

3 . The method of claim 1 , wherein the first user input received via the first graphical element of the graphical user interface selects at least two different convolution operations, and the respective operations performed by the respective candidate layers are selected randomly from the operation search space.

4 . The method of claim 1 , wherein the first user input received via the first graphical element of the graphical user interface selects at least two different pooling operations, and the respective operations performed by the respective candidate layers are selected randomly from the operation search space.

5 . The method of claim 1 , further comprising:

receiving second user input directed to a second graphical element of the graphical user interface, the second user input identifying a specified amount of computational resources to use for the iterative model-growing process; and

responsive to expending the specified amount of computational resources, stopping the iterative model-growing process and selecting the final model.

6 . The method of claim 5 , wherein the second user input directed to the second graphical element designates a number of GPU-days to expend for the iterative model-growing process.

7 . The method of claim 5 , wherein the second user input directed to the second graphical element designates a length of time to expend for the iterative model-growing process.

8 . The method of claim 1 , further comprising:

receiving second user input directed to a second graphical element of the graphical user interface, the second user input directed to the second graphical element identifying model size as a particular criterion for evaluating the trained child models; and

designating the individual trained child model as the new parent model based at least on a model size of the individual trained child model.

9 . The method of claim 1 , further comprising:

receiving second user input directed to a second graphical element of the graphical user interface, the second user input directed to the second graphical element specifying connectivity parameters for the child models; and

generating the child models according to the connectivity parameters.

10 . The method of claim 9 , wherein the connectivity parameters specified by the second user input indicate a number of previous layers that contribute to the one or more selected candidate layers of the child models.

11 . The method of claim 9 , wherein the connectivity parameters specified by the second user input directed to the second graphical element indicate whether skip connections are employed in the child models.

12 . The method of claim 1 , wherein the first user input selects a group of at least two different convolutional kernel sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.

13 . The method of claim 1 , wherein the first user input selects a group of at least two different convolutional stride sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.

14 . The method of claim 1 , wherein the first user input selects a group of pooling operations available to be performed by the respective candidate layers when performing the iterative model-growing process, the group including at least a max pooling operation and an average pooling operation.

15 . The method of claim 1 , wherein the first user input selects a group of at least two different pooling window sizes available to be performed by the respective candidate layers when performing the iterative model-growing process.

16 . A system comprising:

a hardware processing unit; and

a storage resource storing computer-readable instructions which, when executed by the hardware processing unit, cause the hardware processing unit to:

receive first user input via a first graphical element of a graphical user interface, the first user input identifying a set of multiple operations to include in an operation search space for an iterative model-growing process;

add an initial parent model to a parent model pool;

perform two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:

selecting a particular parent model having a plurality of layers from the parent model pool;

inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the first user input;

initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;

based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;

training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;

evaluating the trained child models using one or more criteria; and

based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and

after the two or more iterations, select at least one trained child model as a final model and outputting the final model.

17 . The system of claim 16 , wherein the operation search space includes multiple convolution operations and multiple pooling operations designated by the first user input that is received via the first graphical element of the graphical user interface.

18 . The system of claim 16 , wherein the computer-readable instructions, when executed by the hardware processing unit, cause the hardware processing unit to:

configure the respective candidate layers to perform respective operations randomly selected from the set of multiple operations identified by the first user input.

19 . The system of claim 18 , wherein the first user input directed to the first graphical element identifies at least two different convolution operations and at least two different pooling operations from which the respective operations are randomly selected.

20 . The system of claim 19 , wherein the first user input directed to the first graphical element specifies at least two different window sizes and at least two different strides for the at least two different convolution operations.

21 . A computer-readable storage medium storing instructions which, when executed by a processing device, cause the processing device to perform acts comprising:

outputting a graphical user interface having a first graphical element for designating an initial parent model and a second graphical element for designating an operation search space for an iterative model-growing process;

receiving first user input via the first graphical element, the first user input identifying a particular machine learning model as the initial parent model;

receiving second user input via the second graphical element, the second user input identifying a set of multiple operations to include in the operation search space;

configuring the operation search space of the iterative model-growing process to be restricted to the set of multiple operations identified by the second user input received via the second graphical element of the graphical user interface;

adding the initial parent model to a parent model pool;

performing two or more iterations of the iterative model-growing process, the iterative model-growing process comprising:

selecting a particular parent model having a plurality of layers from the parent model pool;

inserting a plurality of candidate layers into the particular parent model, respective candidate layers being configured to perform respective operations selected from the set of multiple operations identified by the second user input;

initializing the plurality of candidate layers to obtain learned weights for the candidate layers during training while the plurality of candidate layers are connected to the particular parent model, the plurality of candidate layers being initialized while maintaining weights of the plurality of layers of the particular parent model;

based at least on the learned weights of the plurality of candidate layers, selecting less than all of the plurality of candidate layers as selected candidate layers to include in each child model of a plurality of child models for subsequent training, each respective child model including the plurality of layers of the particular parent model and one or more of the selected candidate layers;

training the plurality of child models having the one or more selected candidate layers, the training resulting in trained child models;

evaluating the trained child models using one or more criteria; and

based at least on the evaluating, designating an individual trained child model as a new parent model and adding the new parent model to the parent model pool; and

after the two or more iterations, selecting at least one trained child model as a final model and outputting the final model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2022
From: DEY, DEBADEEPTA; CARUANA, RICHARD A.; HORVITZ, ERIC J.; HU, HANZHANG; LANGFORD, JOHN C.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 062072/0398 →
Continuity (2)
Continuation 16213470 · Dec 7, 2018
Related Publication 20230115700A1 · Apr 13, 2023
References Cited (61)
US 10282864B1 · Kim · 2019 [cited by examiner]
US 20070094168A1 · Ayala · 2007 [cited by examiner]
US 20150206048A1 · Talathi · 2015 [cited by applicant]
US 20180024510A1 · Matsushima · 2018 [cited by examiner]
US 20180365557A1 · Kobayashi · 2018 [cited by examiner]
US 20190251439A1 · Zoph · 2019 [cited by examiner]
US 20200097847A1 · Convertino · 2020 [cited by examiner]
US 20220292357A1 · Xu · 2022 [cited by applicant]
US 20220343165A1 · Hu · 2022 [cited by applicant]
US 20240144051A1 · Kirshenboim · 2024 [cited by applicant]
CN 108764292A · 2022 [cited by applicant]
JP 2018195314A · 2018 [cited by applicant]
KR 100845230B1 · 2008 [cited by applicant]
KR 20180068292A · 2018 [cited by applicant]
KR 20180084969A · 2018 [cited by applicant]
WO 2017154284A1 · 2017 [cited by applicant]
WO 2018167885A1 · 2018 [cited by applicant]
Prellberg et al., “Lamarckian Evolution of Convolutional Neural Networks,” Jun. 21, 2018, arXiv:1806.08099v1 [cs.NE], 12 pages (Year: 2018). [cited by examiner]
Garg et al., “Fabrik: An online collaborative neural network editor,” Oct. 27, 2018, arXiv preprint arXiv:1810.11649v1 [cs.LG], 12 pages (Year: 2018). [cited by examiner]
Li et al., “Learning without Forgetting,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 12, pp. 2935-2947, 2018 (Year: 2018). [cited by examiner]
Nayman, et al., “XNAS: Neural Architecture Search with Expert Advice”, In Proceedings of 33rd Conference on Neural Information Processing Systems, 2019, pp. 1-11. [cited by applicant]
“Archai Documentation”, Retrieved from: https://microsoft.github.io/archai/, Retrieved on Nov. 18, 2022, 2 Pages. [cited by applicant]
“Office Action Issued in Indian Patent Application No. 202117024539”, Mailed Date: Jan. 9, 2023, 6 Pages. [cited by applicant]
“30-Minute Tutorial”, Retrieved from: https://microsoft.github.io/archai/user-guide/tutorial.html#network-architecture-search-nas, Retrieved on Nov. 18, 2022, 9 Pages. [cited by applicant]
“Office Action Issued in Russian Patent Application No. 2021119674”, Mailed Date: Mar. 17, 2023, 14 Pages. [cited by applicant]
“Office Action Issued in Israel Patent Application No. 283463”, Mailed Date: Aug. 2, 2023, 3 Pages. [cited by applicant]
“Written Opinion Issued in Singaporean Patent Application No. 11202105300T”, Mailed Date: Jul. 10, 2023, 5 Pages. [cited by applicant]
“Office Action Issued in Japanese Patent Application No. 2021-525023”, Mailed Date: Oct. 2, 2023, 7 Pages. [cited by applicant]
Benmeziane, et al., “A Comprehensive Survey on Hardware-Aware Neural Architecture Search”, Arxiv.org, Cornell University Library, Jan. 22, 2021, 30 pages. [cited by applicant]
Gupta, et al., “Accelerator-aware Neural Network Design using AutoML”, Arxiv.org, Cornell University Library, Mar. 5, 2020, 5 pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2023/033653, mailed on Jan. 25, 2024, 20 pages. [cited by applicant]
Loni, et al., “FastStereoNet: A Fast Neural Architecture Search for Improving the Inference of Disparity Estimation on Resource-Limited Platforms”, IEEE Transactions on Systems, Man, And Cybernetics: Systems, vol. 52, N… [cited by applicant]
Office Action received for Indonesian Application No. P00202103943, mailed on Nov. 20, 2023, 6 Pages (English Translation Provided). [cited by applicant]
Office Action Received for Russian Application No. 2021119674, mailed on Oct. 30, 2023, 24 pages (English Translation Provided). [cited by applicant]
Tan, et al., “MnasNet: Platform-Aware Neural Architecture Search for Mobile”, Arxiv.org, Cornell University Library, Jul. 31, 2018, 9 pages. [cited by applicant]
Office Action received for Indonesian Application No. P00202103943, mailed on May 20, 2024, 4 Pages. [cited by applicant]
EPO Notification Rule 94(3) received in European Application No. 19809290.0, mailed on May 27, 2024, 11 pages (English Translation Provided). [cited by applicant]
Decision to Grant Received for Japanese Application No. 2021-525023, mailed on Mar. 21, 2024, 5 pages (English Translation Provided). [cited by applicant]
First Office Action Received for Chinese Application No. 201980080971.2, mailed on Jul. 15, 2024, 31 pages. (English Translation Provided). [cited by applicant]
Office Action Issued in Australian Patent Application No. 2019394750, Mailed Date: Aug. 12, 2024, 03 Pages. [cited by applicant]
Office Action Received for Mexican Application No. MX/a/2021/006555, mailed on Jul. 23, 2024, 9 pages. (English Translation Provided). [cited by applicant]
Office Action Issued in Australian Patent Application No. 2019394750, Mailed Date: Jun. 28, 2024, 03 Pages. [cited by applicant]
Liu, et al., “Progressive Neural Architecture Search”, In Proceedings of the European Conference on Computer Vision, Oct. 6, 2018, pp. 19-35. [cited by applicant]
Office Action Received for Canadian Application No. 3119027, mailed on Jan. 27, 2025, 4 pages. [cited by applicant]
Summons to attend oral proceedings pursuant to Rule 115(1) received in European Application No. 19809290.0, mailed on Feb. 20, 2025, 11 pages. [cited by applicant]
Non-Final Office Action mailed on Jun. 27, 2025, in U.S. Appl. No. 17/978,587 43 Pages. [cited by applicant]
Communication under Rule 71(3) Received for European Application No. 19809290.0, mailed on Jul. 7, 2025, 8 pages. [cited by applicant]
Notice of Allowance for Chinese Application No. 201980080971.2, mailed on Jun. 30, 2025, 4 pages. (English Translation Provided). [cited by applicant]
Notice of Allowance Received for Korea Application No. 10-2021-7017204, mailed on Jul. 7, 2025, 06 pages. (English Translation Provided). [cited by applicant]
Prellberg et al., “Lamarckian Evolution of Convolutional Neural Network”, arXiv:1806.08099v1, Jun. 21, 2018, 12 pages. [cited by applicant]
Intimation of Grant Received for Indian Application No. 202117024539, mailed on Mar. 13, 2025, 01 Page. [cited by applicant]
Second Office Action Received for Chinese Application No. 201980080971.2, mailed on Mar. 18, 2025, 24 pages. (English Translation Provided). [cited by applicant]
Office Action Received for Japanese Application No. 2024067927, mailed on Mar. 28, 2025, 8 pages. (English Translation is Provided). [cited by applicant]
International Preliminary Report on Patentability received for PCT Application No. PCT/US23/033653, mailed on May 15, 2025, 14 pages. [cited by applicant]
Notice of Allowance Received for Japanese Application No. 2024-067927, mailed on Jul. 3, 2025, 05 pages. (English Translation is Provided). [cited by applicant]
Examination Report Received for Singaporean Application No. 11202105300T, mailed on Oct. 16, 2025, 4 pages. [cited by applicant]
Notice of allowance Received for Israel Application No. 283463, mailed on Aug. 27, 2024, 4 pages. [cited by applicant]
Notice of Refusal Received for Israel Application No. 283463, mailed on Jan. 7, 2025, 1 page. [cited by applicant]
Notification of Grant received for Singaporean Application No. 11202105300T, mailed on Jan. 2, 2026, 2 pages. [cited by applicant]
Decision to grant a European patent pursuant to Article 97(1) received in European Application No. 19809290.0, mailed on Nov. 27, 2025, 2 pages. [cited by applicant]
Non-Final Office Action mailed on Jan. 23, 2026 in U.S. Appl. No. 17/978,587, 65 Pages. [cited by applicant]