IP Library › Granted Patent US 12,450,479
Granted Patent B2
US 12,450,479 · App. 17/508,665 · Granted Oct 21, 2025

Systems and methods for tuning hyperparameters of a model and advanced curtailment of a training of the model

Inventors: Michael McCourt (San Francisco, CA); Taylor Jackle-Spriggs (San Francisco, CA); Ben Hsu (San Francisco, CA); Simon Howey (San Francisco, CA); Halley Nicki Vance (San Francisco, CA); James Blomo (San Francisco, CA); Patrick Hayes (San Francisco, CA); Scott Clark (San Francisco, CA)
Assignee: Intel Corporation
G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,450,479
App. No.
17/508,665
Granted
Oct 21, 2025
Kind
B2
Abstract

A system and method for tuning hyperparameters and training a model includes implementing a hyperparameter tuning service that tunes hyperparameters of a model that includes receiving, via an API, a tuning request that includes: (i) a first part comprising tuning parameters for generating tuned hyperparameter values for hyperparameters of the model; and (ii) a second part comprising model training control parameters for monitoring and controlling a training of the model, wherein the model training control parameters include criteria for generating instructions for curtailing a training run of the model; monitoring the training run for training the model based on the second part of the tuning request, wherein the monitoring of the training run includes periodically collecting training run data; and computing an advanced training curtailment instruction based on the training run data that automatically curtails the training run prior to a predefined maximum training schedule of the training run.

Claims (37)

1. An apparatus comprising:

at least one memory;

instructions in the apparatus; and

processor circuitry to execute the instructions to:

compute an average metric value for a plurality of first training runs;

compute a metric value for a second training run at an evaluation interval;

evaluate the metric value for the evaluation interval of the second training run relative to the average metric value based on a metric goal; and

stop the second training run based on an early termination policy and the evaluation of the metric value relative to the average metric value.

2. The apparatus of claim 1 , wherein the metric value corresponds to a performance metric for tuning a hyperparameter of a model.

3. The apparatus of claim 2 , wherein the processor circuitry is to execute the instructions to define the metric goal, the metric goal to optimize the performance metric of the model based on the plurality of first training runs.

4. The apparatus of claim 2 , wherein the metric value corresponds to an accuracy metric of the model and the metric goal is to maximize accuracy of the model, the processor to execute the instructions to stop the second training run in response to the metric value being less than the average metric value.

5. The apparatus of claim 1 , wherein the evaluation interval defines a frequency of applying the early termination policy, the processor circuitry to execute the instructions to apply the early termination policy based on the evaluation interval.

6. The apparatus of claim 1 , wherein the processor circuitry is to execute the instructions to select an evaluation delay, the evaluation delay to cause a minimum number of evaluation intervals to complete before applying the early termination policy at the evaluation interval.

7. The apparatus of claim 1 , wherein the processor circuitry is to execute the instructions to use the early termination policy to stop the second training run based on the second training run being a low-performance run.

8. The apparatus of claim 1 , wherein the processor circuitry is to execute the instructions to stop the second training run based on the early termination policy prior to completion of the second training run.

9. At least one non-transitory computer readable medium comprising instructions that, when executed, cause at least one processor to at least:

compute an average metric value for a plurality of first training runs;

compute a metric value for a second training run at an evaluation interval;

evaluate the metric value for the evaluation interval of the second training run relative to the average metric value based on a metric goal; and

stop the second training run based on an early termination policy and the evaluation of the metric value relative to the average metric value.

10. The at least one non-transitory computer readable medium of claim 9 , wherein the metric value and the average metric value are based on a performance metric corresponding to a tuning of a hyperparameter of a model.

11. The at least one non-transitory computer readable medium of claim 10 , wherein the instructions cause the at least one processor to define the metric goal, the metric goal to optimize the performance metric of the model based on the plurality of first training runs.

12. The at least one non-transitory computer readable medium of claim 10 , wherein the metric value corresponds to an accuracy metric of the model and the metric goal is to maximize accuracy of the model, the instructions to cause the at least one processor to stop the second training run in response to the metric value being less than the average metric value.

13. The at least one non-transitory computer readable medium of claim 9 , wherein the evaluation interval defines a frequency of applying the early termination policy, the instructions to cause the at least one processor to apply the early termination policy based on the evaluation interval.

14. The at least one non-transitory computer readable medium of claim 9 , wherein the instructions cause the at least one processor to select an evaluation delay, the evaluation delay to cause a minimum number of evaluation intervals to complete before applying the early termination policy at the evaluation interval.

15. The at least one non-transitory computer readable medium of claim 9 , wherein the instructions cause the at least one processor to use the early termination policy to stop the second training run based on the second training run being a low-performance run.

16. The at least one non-transitory computer readable medium of claim 9 , wherein the instructions cause the at least one processor to stop the second training run based on the early termination policy prior to completion of the second training run.

17. A method to tune a hyperparameter of a model, the method comprising:

computing an average metric value for a plurality of first training runs;

computing a metric value for a second training run at an evaluation interval;

evaluating the metric value for the evaluation interval of the second training run relative to the average metric value based on a metric goal; and

stopping the second training run based on an early termination policy and the evaluation of the metric value relative to the average metric value.

18. The method of claim 17 , further including defining the metric goal, the metric goal to optimize a metric of the model based on the plurality of first training runs, the metric corresponding to a performance of the model.

19. The method of claim 17 , wherein the metric value corresponds to an accuracy metric of a model, the metric goal to maximize accuracy of the model, the method further including stopping the second training run in response to the metric value being less than the average metric value.

20. The method of claim 17 , wherein the evaluation interval defines a frequency of applying the early termination policy, the method further including applying the early termination policy based on the evaluation interval.

21. The method of claim 17 , wherein the evaluation interval is a first evaluation interval, the method further including selecting an evaluation delay, the evaluation delay to cause a minimum number of evaluation intervals to complete before the first evaluation interval.

22. The method of claim 17 , further including stopping the second training run based on the early termination policy prior to completion of the second training run.

Continuity (3)
Continuation 16849422 · Apr 15, 2020
Provisional Application 62833895 · Apr 15, 2019
Related Publication 20220114450A1 · Apr 14, 2022
References Cited (53)
US 7363281B2 · Jin et al. · 2008 [cited by applicant]
US 8364613B1 · Lin et al. · 2013 [cited by applicant]
US 9786036B2 · Annapureddy · 2017 [cited by applicant]
US 9858529B2 · Adams et al. · 2018 [cited by applicant]
US 10217061B2 · Hayes et al. · 2019 [cited by applicant]
US 10282237B1 · Johnson et al. · 2019 [cited by applicant]
US 10379913B2 · Johnson et al. · 2019 [cited by applicant]
US 10445150B1 · Johnson et al. · 2019 [cited by applicant]
US 10528891B1 · Cheng et al. · 2020 [cited by applicant]
US 10558934B1 · Cheng et al. · 2020 [cited by applicant]
US 10565025B2 · Johnson et al. · 2020 [cited by applicant]
US 10607159B2 · Hayes et al. · 2020 [cited by applicant]
US 10621514B1 · Cheng et al. · 2020 [cited by applicant]
US 10740695B2 · Cheng et al. · 2020 [cited by applicant]
US 11157812B2 · McCourt et al. · 2021 [cited by applicant]
US 20070019065A1 · Mizes · 2007 [cited by applicant]
US 20080183648A1 · Goldberg et al. · 2008 [cited by applicant]
US 20090244070A1 · Mattikalli et al. · 2009 [cited by applicant]
US 20100083196A1 · Liu · 2010 [cited by applicant]
US 20150288573A1 · Baughman et al. · 2015 [cited by applicant]
US 20160110657A1 · Gibiansky et al. · 2016 [cited by applicant]
US 20160132787A1 · Drevo et al. · 2016 [cited by applicant]
US 20160232540A1 · Gao et al. · 2016 [cited by applicant]
US 20170124487A1 · Szeto et al. · 2017 [cited by applicant]
US 20180121797A1 · Prabhu et al. · 2018 [cited by applicant]
US 20180129892A1 · Bahl et al. · 2018 [cited by applicant]
US 20180240041A1 · Koch et al. · 2018 [cited by applicant]
US 20180356949A1 · Wang et al. · 2018 [cited by applicant]
US 20190114537A1 · Wesolowski et al. · 2019 [cited by applicant]
US 20190156229A1 · Tee et al. · 2019 [cited by applicant]
US 20190220755A1 · Carbune et al. · 2019 [cited by applicant]
US 20200019888A1 · McCourt et al. · 2020 [cited by applicant]
US 20200050968A1 · Lee et al. · 2020 [cited by applicant]
US 20200111018A1 · Golovin et al. · 2020 [cited by applicant]
US 20200151029A1 · Johnson et al. · 2020 [cited by applicant]
US 20200202254A1 · Hayes et al. · 2020 [cited by applicant]
US 20200302342A1 · Cheng et al. · 2020 [cited by applicant]
WO 2018213119A1 · 2018 [cited by applicant]
Rasley et al., “Hyperdrive: Exploring hyperparameters with pop scheduling.” In Proceedings of the 18th ACM/IFIP/USENIX Middleware Conference, pp. 1-13. 2017. (Year: 2017). [cited by applicant]
Abadi et al., “Tensorflow: A system for large-scale machine learning.” In 12th {USENIX} symposium on operating systems design and implementation ({OSDI} 16), pp. 265-283. 2016. (Year: 2016). [cited by applicant]
Bergstra et al.. “Making a science of model search: Hyperparameter optimization in hundreds of dimensions for vision architectures.” In International conference on machine learning, pp. 115-123. 2013. (Year: 2013). [cited by applicant]
Golovin et al., “Google vizier: A service for black-box optimization.” In Proceedings of the 23rd ACM SIGKDD International conference on knowledge discovery and data mining, pp. 1487-1495. 2017. (Year: 2017). [cited by applicant]
Johnson et al., “Orchestrate: Infrastructure for Enabling Parallelism during Hyperparameter Optimization.” arXiv preprint arXiv:1812.07751 (2018). (Year: 2018). [cited by applicant]
Dewancker et al., “A stratified analysis of Bayesian optimization methods.” arXiv preprint arXiv: 1603.09441 (2016). (Year: 2016). [cited by applicant]
Bergstra et al., “Hyperopt: a Python Library for Model Selection and Hyperparameter Optimization,” Computational Science & Discovery, 2015, 25 pages. [cited by applicant]
Diaz et al., “An Effective Algorithm for Hyperparameter Optimization of Neural Networks,” IBM Journal of Research and Development, vol. 61, No. 4/5, 2017, 20 pages. [cited by applicant]
Golovin et al., “Google Vizier: A Service for Black-Box Optimization,” KDD '17, Aug. 12-17, Halifax, NS, Canada, 2017, 10 pages. [cited by applicant]
Gardner et al., “Bayesian Optimization with Inequality Constraints,” Proceedings of the 31st International Conference on Machine Learning, Beijing, China, 2014, 10 pages. [cited by applicant]
Zou et al., “Regularization and Variable Selection via the Elastic Net,” Journal of the Royal Statistical Society, vol. 37, Issue 2, 2005, 20 pages. [cited by applicant]
Zhou et al., “Combining Global and Local Surrogate Models to Accelerate Evolutionary Optimization,” IEEE Transactions on Systems, Man, and Cybernetics—Part C: Applications and Reviews, vol. 37, No. 1, Jan. 2007, 11 page… [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 16/849,422, dated Oct. 13, 2020, 18 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 16/849,422, dated Mar. 3, 2021, 7 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 16/849,422, dated Jun. 23, 2021, 7 pages. [cited by applicant]