IP Library Granted Patent US 12,367,249
Granted Patent B2
US 12,367,249 · App. 17/505,568 · Granted Jul 22, 2025

Framework for optimization of machine learning architectures

Inventors: Anthony Sarah (San Diego, CA); Daniel Cummings (Austin, TX); Juan Pablo Munoz (Folsom, CA); Tristan Webb (San Diego, CA)
Assignee: Intel Corporation
G06F16/953G06N5/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,249
App. No.
17/505,568
Granted
Jul 22, 2025
Kind
B2
Abstract

The present disclosure is related to framework for automatically and efficiently finding machine learning (ML) architectures that are optimized to one or more specified performance metrics and/or hardware platforms. This framework provides ML architectures that are applicable to specified ML domains and are optimized for specified hardware platforms in significantly less time than could be done manually and in less time than existing ML model searching techniques. Furthermore, a user interface is provided that allows a user to search for different ML architectures based on modified search parameters, such as different hardware platform aspects and/or performance metrics. Other embodiments may be described and/or claimed.

Claims (49)

1. An apparatus for identifying machine learning (ML) architectures, the apparatus comprising:

interface circuitry to obtain an ML configuration, the ML configuration including a current set of input search parameters;

machine-readable instructions; and

at least one processor circuit to be programmed based on the machine-readable instructions to:

identify a set of previous ML architectures based on previous searches using respective previous sets of input search parameters, the previous sets of input search parameters different from the current set of input search parameters but having at least one of an ML task, an ML domain or hardware platform information in common with the current set of search parameters;

initialize a set of candidate ML architectures to include the set of previous ML architectures;

search the set of candidate ML architectures based on the current set of input search parameters;

determine, based on the search, an output set of ML architectures from the set of candidate ML architectures, the output set of ML architectures to satisfy one or more of the current set of search parameters; and

evaluate performance of ones of the ML architectures in the output set of ML architectures.

2. The apparatus of claim 1 , wherein the ML configuration includes a super-network, and one or more of the at least one processor circuit is to generate sub-networks of the super-network to include in the set of candidate ML architectures, the sub-networks to have fewer ML parameters than the super-network.

3. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to search the set of candidate ML architectures based on one or more constraints between two or more ML parameters of the candidate ML architectures in the initialized set of candidate ML architectures.

4. The apparatus of claim 1 , wherein the search is based on an algorithm, and one or more of the at least one processor circuit is to select the algorithm from a plurality of algorithms based on at least one of hardware platform information or a performance metric included in the ML configuration.

5. The apparatus of claim 4 , wherein the algorithms include one or more of grid search, random search, Bayesian optimization, an evolutionary algorithm, a tree-structured Parzen estimator, or a user-defined optimization algorithm.

6. The apparatus of claim 5 , wherein the algorithms include the evolutionary algorithm, and the evolutionary algorithm is one of Strength Pareto Evolutionary Algorithm 2 (SPEA-2) or Nondominated Sorting Genetic Algorithm-II.

7. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to determine the set of optimal ML architectures to be a set of Pareto optimal solutions.

8. The apparatus of claim 7 , wherein one or more of the at least one processor circuit is to:

for successive iterations until convergence is reached:

rank ones of the candidate ML architectures in the set of candidate ML architectures;

perform crowding distance sorting to select individual candidate ML architectures from the ranked set of candidate ML architectures; and

carry the selected individual candidate ML architectures into a next optimization iteration.

9. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

measure one or more performance metrics based on operation of ones of the ML architectures in the output set of ML architectures using a test dataset.

10. The apparatus of claim 1 , wherein one or more of the at least one processor circuit is to:

predict one or more performance metrics of ones of the ML architectures in the output set of ML architectures using one or more proxy functions.

11. The apparatus of claim 10 , wherein the one or more proxy functions include one or more of associative arrays, mapping functions, dictionaries, hash tables, look-up tables (LUTs), linked lists, ML classifiers, parameter counting, computational throughput metrics, Jacobian covariance functions, saliency pruning functions, channel pruning functions, heuristic functions, or hyper-heuristic functions.

12. The apparatus of claim 1 , wherein the ML configuration includes an identifier of a hardware platform on which a selected one of the output set of ML architectures is to be deployed.

13. The apparatus of claim 1 , wherein the ML configuration includes data of one or more hardware components of a hardware platform.

14. One or more non-transitory computer readable media (NTCRM) comprising instructions to cause at least one processor circuit to at least:

access a machine learning (ML) configuration from a client device;

determine a set of candidate ML architectures based on sub-networks included in a super-network indicated by the ML configuration;

determine, based on the set of candidate ML architectures, an output set of ML architectures that satisfy at least one of an ML parameter or hardware platform information (HPI) included in the ML configuration;

determine performance metrics for the set of optimal ML architectures; and

cause presentation of information corresponding to the output set of ML architectures and the determined performance metrics at the client device.

15. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

obtain a selection of one of the MP architectures from among the output set of ML architectures; and

cause the selected one of the ML architectures to be sent to the client device.

16. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to determine the output set of ML architectures to be a Pareto frontier of a multi-objective optimization problem, and the information includes a graphical representation of the Pareto frontier.

17. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

determine the set of candidate ML architectures to include first ones of the sub-networks that were previously found to satisfy the at least one of the ML parameter or the HPI; or

determine the set of candidate ML architectures to include second ones of the sub-networks that satisfy one or more heuristics derived from at least one previously generated output ML architecture, the one or more heuristics to indicate a relationship between individual ML parameters of the at least one previously generated output ML architecture.

18. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

search the set of candidate ML architectures based on one or more constraints between two or more ML parameters of the candidate ML architectures in the set of candidate ML architectures.

19. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

operate one or more algorithms to determine the output set of ML architectures based on the set of candidate ML architectures, the one or more optimization algorithms including one or more of grid search, random search, Bayesian optimization, a genetic algorithm, a tree-structured Parzen estimator, Strength Pareto Evolutionary Algorithm 2 (SPEA-2), Nondominated Sorting Genetic Algorithm-II, or a user-defined optimization algorithm.

20. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

measure the performance metrics based on operation of ones of the ML architectures in the output set of ML architectures using a test dataset.

21. The one or more NTCRM of claim 14 , wherein the instructions are to cause one or more of the at least one processor circuit to:

predict one or more of the performance metrics based on one or more proxy functions, the one or more proxy functions including one or more of associative arrays, mapping functions, dictionaries, hash tables, look-up tables (LUTs), linked lists, ML classifiers, parameter counting, computational throughput metrics, Jacobian covariance functions, saliency pruning functions, channel pruning functions, heuristic functions, or hyper-heuristic functions.

22. The one or more NTCRM of claim 14 , wherein one or more of the at least one processor circuit is to determine the output set of ML architectures to satisfy at least the HPI, and the HPI includes technical data of a hardware platform or technical data of one or more hardware components of the hardware platform.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 20, 2021
From: SARAH, ANTHONY; CUMMINGS, DANIEL; MUNOZ, JUAN PABLO; WEBB, TRISTAN
To: INTEL CORPORATION
Reel/Frame 057848/0730 →
Continuity (1)
Related Publication 20220035878A1 · Feb 3, 2022
References Cited (180)
US 10528891B1 · Cheng · 2020 [cited by examiner]
US 10744372B2 · Meyerson · 2020 [cited by examiner]
US 10803392B1 · Khan et al. · 2020 [cited by applicant]
US 11347995B2 · Bender · 2022 [cited by examiner]
US 11443162B2 · Yang · 2022 [cited by examiner]
US 11527308B2 · Shahrzad · 2022 [cited by examiner]
US 11531861B2 · Tan · 2022 [cited by examiner]
US 11620515B2 · Liu · 2023 [cited by applicant]
US 20070208548A1 · McConaghy · 2007 [cited by applicant]
US 20100325072A1 · Truemper · 2010 [cited by examiner]
US 20130185234A1 · Arora · 2013 [cited by examiner]
US 20150095253A1 · Lim · 2015 [cited by examiner]
US 20160328644A1 · Lin et al. · 2016 [cited by applicant]
US 20170220928A1 · Hajizadeh · 2017 [cited by examiner]
US 20170261949A1 · Hoffmann et al. · 2017 [cited by applicant]
US 20180137219A1 · Goldfarb · 2018 [cited by examiner]
US 20190005390A1 · Stitt et al. · 2019 [cited by applicant]
US 20190019096A1 · Yoshida · 2019 [cited by examiner]
US 20190279102A1 · Cataltepe · 2019 [cited by applicant]
US 20190286984A1 · Vasudevan · 2019 [cited by examiner]
US 20190347548A1 · Amizadeh · 2019 [cited by examiner]
US 20190354837A1 · Zhou · 2019 [cited by examiner]
US 20190370659A1 · Dean · 2019 [cited by examiner]
US 20200044955A1 · Pugaczewski · 2020 [cited by applicant]
US 20200057965A1 · Howard · 2020 [cited by applicant]
US 20200104710A1 · Vasudevan · 2020 [cited by examiner]
US 20200104715A1 · Denolf · 2020 [cited by examiner]
US 20200125545A1 · Idicula · 2020 [cited by examiner]
US 20200143227A1 · Tan et al. · 2020 [cited by applicant]
US 20200143243A1 · Liang · 2020 [cited by examiner]
US 20200167593A1 · Kim · 2020 [cited by applicant]
US 20200193266A1 · Scheidegger · 2020 [cited by examiner]
US 20200272909A1 · Parmentier · 2020 [cited by examiner]
US 20200311561A1 · Naaman et al. · 2020 [cited by applicant]
US 20200320399A1 · Huang et al. · 2020 [cited by applicant]
US 20200349050A1 · Ghobadi · 2020 [cited by applicant]
US 20200357486A1 · Kok · 2020 [cited by examiner]
US 20200401899A1 · Dohan · 2020 [cited by examiner]
US 20210034924A1 · McCourt · 2021 [cited by examiner]
US 20210056378A1 · Yang · 2021 [cited by examiner]
US 20210064627A1 · Kleiner et al. · 2021 [cited by applicant]
US 20210081763A1 · Abdelfattah · 2021 [cited by examiner]
US 20210097383A1 · Kaur · 2021 [cited by applicant]
US 20210097443A1 · Li · 2021 [cited by examiner]
US 20210110140A1 · Munoz et al. · 2021 [cited by applicant]
US 20210110276A1 · Chu · 2021 [cited by examiner]
US 20210110302A1 · Nam · 2021 [cited by examiner]
US 20210158929A1 · Sjolund · 2021 [cited by applicant]
US 20210303967A1 · Bender · 2021 [cited by applicant]
US 20210304061A1 · Kolar et al. · 2021 [cited by applicant]
US 20210350233A1 · Saboori et al. · 2021 [cited by applicant]
US 20210357744A1 · Mittal · 2021 [cited by examiner]
US 20210357959A1 · Cella · 2021 [cited by applicant]
US 20210390376A1 · Byrne · 2021 [cited by examiner]
US 20220012089A1 · Nasr-Azadani et al. · 2022 [cited by applicant]
US 20220019880A1 · Dasgupta · 2022 [cited by examiner]
US 20220027792A1 · Cummings et al. · 2022 [cited by applicant]
US 20220035877A1 · Nittur Sridhar · 2022 [cited by examiner]
US 20220035878A1 · Sarah et al. · 2022 [cited by applicant]
US 20220036128A1 · Levanony · 2022 [cited by applicant]
US 20220036194A1 · Sundaresan · 2022 [cited by examiner]
US 20220172038A1 · Chen · 2022 [cited by examiner]
US 20220180125A1 · Shen et al. · 2022 [cited by applicant]
US 20220198217A1 · Dong · 2022 [cited by examiner]
US 20220198260A1 · Xue · 2022 [cited by examiner]
US 20220230048A1 · Li · 2022 [cited by examiner]
US 20220269835A1 · Yang · 2022 [cited by examiner]
US 20220277231A1 · Ostergaard et al. · 2022 [cited by applicant]
US 20220284582A1 · Yang · 2022 [cited by examiner]
US 20220328128A1 · Kok · 2022 [cited by applicant]
US 20220348903A1 · Ranganathan et al. · 2022 [cited by applicant]
US 20220398500A1 · Singhal · 2022 [cited by applicant]
US 20220414534A1 · Ananthanarayanan · 2022 [cited by examiner]
US 20230032748A1 · Sodhi · 2023 [cited by applicant]
US 20230274151A1 · Xu · 2023 [cited by examiner]
US 20230325711A1 · Haraldson · 2023 [cited by examiner]
US 20240007414A1 · Jain · 2024 [cited by examiner]
US 20240046148A1 · Bega · 2024 [cited by examiner]
US 20240289687A1 · Kumar et al. · 2024 [cited by applicant]
US 20240362472A1 · Fu · 2024 [cited by examiner]
EP 3975060A1 · 2022 [cited by applicant]
WO WO2021158313A1 · 2021 [cited by applicant]
Dzmitry Bahdanau et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, May 19, 2016, 15 pages, arXiv:1409.0473v7 [cs.CL]. [cited by applicant]
Irwan Bello et al., “Attention Augmented Convolutional Networks”, 2019, pp. 3286-3295, Proceedings of the IEEE/CVF Int'l Conference on Computer Vision. [cited by applicant]
Cristian Buciluâ et al., “Model Compression”, Aug. 20, 2006, pp. 535-541, Proceedings of the 12th Assn. for Computing Machinery (ACM) Special Interest Group on Knowledge Discovery in Data (SIGKDD) Int'l Conference on Kn… [cited by applicant]
Liang-Chieh Chen et al., “Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs”, Apr. 27, 2017, pp. 834-848, IEEE transactions on pattern analysis and machine i… [cited by applicant]
Yu Cheng et al., “A Survey of Model Compression and Acceleration for Deep Neural Networks”, Jun. 14, 2020, 10 pages, arXiv:1710.09282v9 [cs.LG]. [cited by applicant]
Aston Zhang et al., “Dive into Deep Learning” Jul. 25, 2021, 323 pages, Release 0.17.0, chs. 4-10 (Jul. 25, 2021), available at: https://d2l.ai/. [cited by applicant]
Tim Dettmers et al., “Sparse Networks from Scratch: Faster Training without Losing Performance”, Aug. 23, 2019, 14 pages, arXiv:1907.04840v2 [cs.LG]. [cited by applicant]
Yu Cheng et al., “A Survey of Model Compression and Acceleration for Deep Neural Networks”, Oct. 23, 2017, 10 pages, IEEE Signal Processing Magazine, Special Issue on Deep Learning for Image Understanding. [cited by applicant]
Alex Krizhevsky et al., “Learning Multiple Layers of Features from Tiny Images”, Master's Thesis, U. of Toronto, Citeseer, 60 pages (Apr. 8, 2009), https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf. [cited by applicant]
Souvik Kundu et al., “A Tunable Robust Pruning Framework Through Dynamic Network Rewiring of DNNS”, arXiv:2011.03083v2 [cs.CV], 8 pages (Nov. 24, 2020). [cited by applicant]
Namhoon Lee et al., “SNIP: Single-Shot Network Pruning based on Connection Sensitivity”, Int'l Conference on Learning Representations (ICLR) 2019, 15 pages (May 6, 2019). [cited by applicant]
Tsung-Yi Lin et al., “Feature Pyramid Networks for Object Detection”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017, pp. 2117-2125 (2017). [cited by applicant]
Zhuang Liu et al., “Rethinking the Value of Network Pruning”, arXiv:1810.05270v2 [cs.LG], 21 pages (Mar. 5, 2019). [cited by applicant]
Marc E. McDill, “Forest Resource Management”, Penn State Univ., ch. 11, pp. 203-233 (Jun. 22, 1999), https://faculty.washington.edu/toths/Presentations/Lecture%202/Ch11_LPIntro.pdf. [cited by applicant]
Hesham Mostafa et al., “Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization”, Proceedings of the 36th Int'l Conference on Machine Learning (PMLR), pp. 4646-4655 (May 2… [cited by applicant]
Junki Park et al., “OPTIMUS: Optimized Matrix Multiplication Structure for Transformer Neural Network Accelerator”, Mar. 15, 2020, pp. 363-378, Proceedings of Machine Learning and Systems, vol. 2. [cited by applicant]
Prajit Ramachandran et al., “Stand-Alone Self-Attention in Vision Models”, arXiv:1906.05909v1 [cs.CV], 15 pages (Jun. 13, 2019), https://arxiv.org/pdf/1906.05909.pdf. [cited by applicant]
Christian Szegedy et al., “Going Deeper with Convolutions”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015), doi: 10.1109/CVPR.2015.7298594. [cited by applicant]
James Bergstra et al., “Algorithms for Hyper-Parameter Optimization”, 2011, 9 pages, Advances in Neural Information Processing Systems 24 (NIPS 2011). [cited by applicant]
James Bergstra et al., “Random Search for Hyper-Parameter Optimization”, Feb. 1, 2012, 25 pages, J. of Machine Learning Research, vol. 13, No. 2. [cited by applicant]
Han Cai et al., “Once-For-All: Train One Network and Specialize for Efficient Deployment”, Apr. 29, 2020, 15 pages, arXiv:1908.09791v5 [cs.LG]. [cited by applicant]
Kalyanmoy Deb et al., “A Fast and Elitist Multiobjective Genetic Algorithm: NSGA-II”, Apr. 2002, 16 pages, IEEE Transactions on Evolutionary Computation, vol. 6, No. 2. [cited by applicant]
Ian Dewancker et al., “Bayesian Optimization Primer”, 2020, 4 pages. Retrieved from the Internet: http://static.sigopt.com/b/20a144d208ef255d3b981ce419667ec25d8412e2/static/pdf/SigOpt_Bayesian_Optimization_Primer.pdf. [cited by applicant]
Ian Dewancker et al., “Bayesian Optimization for Machine Learning: A Practical Guidebook”, Dec. 14, 2016, 15 pages, arXiv preprint arXiv:1612.04858. [cited by applicant]
A.E. Eiben et al., “Introduction to Evolutionary Computing”, 2015,295 pages, second edition, Springer Heidelberg New York Dordrecht London. [cited by applicant]
Wenlan Huang et al., “Survey on Multi-Objective Evolutionary Algorithms”, Aug. 1, 2019, 8 pages, IOP Conf. Series: J. of Physics: Conf. Series, vol. 1288, No. 1, p. 012057. [cited by applicant]
Christian Igel et al., “Covariance Matrix Adaptation for Multi-objective Optimization”, 2007, pp. 1-28, Evolutionary Computation, vol. 15, No. 1. [cited by applicant]
Danilo Vasconcellos Vargas et al., “General Subpopulation Framework and Taming the Conflict Inside Populations”, Jan. 2, 2019, 37 pages, arXiv:1901.00266v1 [cs.NE]. [cited by applicant]
Extended European Search Report mailed Jan. 2, 2023 for European Patent Application No. 22186944.9, 14 pages. [cited by applicant]
Tang et al., “A Semi-Supervised Assessor of Neural Architectures”, arXiv:2005.06821v1 [cs.CV], arxiv.org, Cornell Univ. Library, Ithaca, NY, 10 pages (May 14, 2020). [cited by applicant]
Kokiopoulou et al., “Fast Task-Aware Architecture Inference”, arXiv:1902.05781v1 [cs.LG], arxiv.org, Cornell Univ. Library, Ithaca, NY, 10 pages (Feb. 15, 2019). [cited by applicant]
Zichao Lu et al., “NSGANetV2: Evolutionary Multi-objective Surrogate-Assisted Neural Architecture Search”, Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, pp. 35-51 (Aug. 2020). [cited by applicant]
Extended European Search Report mailed Jan. 10, 2023 for European Patent Application No. 22186932.4, 9 pages. [cited by applicant]
“Intel Agilex FSeries 027 FPGA R25A Product Specifications” retrieved on Oct. 6, 2021, 3 pages. Retrieved from Internet at: https://www.intel.com/content/www/us/en/products/sku/208599/intel-agilex-fseries-027-fpga-r25a/… [cited by applicant]
“Intel Atom x6413E Processor Product Specifications”, retrieved on Oct. 6, 2021, 4 pages. Retrieved from Internet at https://ark.intel.com/content/www/us/en/ark/products/207908/intel-atom-x6413e-processor-1-5m-cache-up-… [cited by applicant]
“Intel® Xeon® Platinum 8362 Processor Product Specifications” retrieved on Oct. 6, 2021, 4 pages. Retrieved from internet at: https://www.intel.com/content/www/us/en/products/sku/217216/intel-xeon-platinum-8362-processo… [cited by applicant]
Mahrokh Javadi et al., “Combining Manhattan and Crowding distances in Decision Space for Multimodal Multi- objective Optimization Problems”, Sep. 12, 2019, 6 pages, Eurogen 2019. [cited by applicant]
Will Koehrsen, “A Conceptual Explanation of Bayesian Hyperparameter Optimization for Machine Learning”, Jun. 28, 2018. 17 pages, Toward Data Science, available at: https://towardsdatascience.com/a-conceptual-explanation… [cited by applicant]
Hanxiao Liu et al., “DARTS: Differentiable Architecture Search”, Apr. 23, 2019, 13 pages, arXiv:1806.09055v2 [cs.LG]. [cited by applicant]
Joseph Mellor et al., “Neural Architecture Search without Training”, Jul. 1, 2021, pp. 7588-7598, Int'l Conference on Machine Learning, PMLR. [cited by applicant]
Jasper Snoek et al., “Practical Bayesian Optimization of Machine Learning Algorithms”, Aug. 29, 2012, 9 pages, Advances in Neural Information Processing Systems 25 (NIPS 2012). [cited by applicant]
Eckart Zitzler et al., “SPEA2: Improving the Strength Pareto Evolutionary Algorithm”, May 2001, 21 pages, Computer Engineering and Communication Networks Lab (TIK), Swiss Fed. Inst. of Tech. (ETH), Zurich, Ch, TIK-Repor… [cited by applicant]
Hanrui Wang et al., “HAT: Hardware-Aware Transformers for Efficient Natural Language Processing”, May 28, 2020, 14 pages, arXiv:2005.14187v1 [cs.CL]. [cited by applicant]
Hanrui Wang et al., “HAT: Hardware-Aware Transformers for Efficient Natural Language Processing”, 2020, 37 pages, ACL 2020 Slides, available at: https://hat.mit.edu/assets/ACL20_HAT_HanruiWang.pdf. [cited by applicant]
Kalyanmoy Deb, “Multi-Objective Optimization Using Evolutionary Algorithms”, Indian Institute of Technology—Kanpur, Dept. of Mechanical Engineering, Kanpur, India, KanGAL Report No. 2011003, 24 pages (Feb. 10, 2011), ht… [cited by applicant]
Zhichao Lu et al., “NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm”, arXiv:1810.03522v2 [cs.CV], 13 pages (Apr. 18, 2019). [cited by applicant]
Yonglong Tian et al., “Contrastive Representation Distillation”, arXiv:1910.10699v2 [cs.LG], 19 pages (Jan. 18, 2020). [cited by applicant]
Tianyun Zhang et al., “A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers”, Proceedings of the European Conference on Computer Vision (ECCV) 2018, in Lecture Notes in Computer Sc… [cited by applicant]
Ashish Vaswani et al., “Attention Is All You Need”, Advances in Neural Information Processing Systems 30 (NIPS 2017), pp. 5998-6008 (2017), https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa… [cited by applicant]
Chunnan Wang et al., “Multi-Objective Neural Architecture Search Based on Diverse Structures and Adaptive Recommendation”, arXiv:2007.02749v2 [cs.CV], 11 pages (Aug. 13, 2020), https://arxiv.org/pdf/2007.02749.pdf. [cited by applicant]
Sergey Zagoruyko et al., “Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer”, arXiv:1612.03928v3 [cs.CV], 13 pages (Feb. 12, 2017), https://arxiv.org/p… [cited by applicant]
Chenzhuo Zhu et al., “Trained Ternary Quantization”, arXiv:1612.01064v3 [cs.LG], 10 pages (Feb. 23, 2017), https://arxiv.org/pdf/1612.01064.pdf. [cited by applicant]
Barret Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition 2018, pp. 8697-8710 (2018), https://openaccess.thec… [cited by applicant]
Mohamed S. Abdelfattah et al., “Zero-Cost Proxies for Lightweight NAS”, Mar. 19, 2021, 17 pages, arXiv:2101.08134v2 [cs.LG]. [cited by applicant]
Edward Alibekov et al., “Proxy Functions for Approximate Reinforcement Learning”, 2019, pp. 224-229, IFAC PapersOnLine 52-11. [cited by applicant]
Bowen Baker et al., “Designing Neural Network Architectures using Reinforcement Learning”, Mar. 22, 2017, 18 pages, arXiv:1611.02167v3. [cited by applicant]
Han Cai et al., “Proxylessnas: Direct neural architecture search on target task and hardware”, Feb. 23, 2019, 13 pages, arXiv:1812.00332v2 [cs.LG]. [cited by applicant]
Olivier Chapelle et al., “Semi-Supervised Learning”, 2006, 524 pages, MIT Press, Cambridge MA, London England. [cited by applicant]
Thomas Elsken et al., “Neural architecture search: A survey”, Mar. 19, 2019, 21 pages, J. of Machine Learning Research, vol. 20, No. 1. [cited by applicant]
Thomas N. Kipf et al., “Semi-Supervised Classification with Graph Convolutional Networks”, Feb. 22, 2017, 14 pages, arXiv:1609.02907v4. [cited by applicant]
M. Z. Naser et al., “Insights into Performance Fitness and Error Metrics for Machine Learning”, May 17, 2020, 25 pages, arXiv:2006.00887v1. [cited by applicant]
Kenneth O. Stanley et al., “Evolving Neural Networks through Augmenting Topologies”, Jun. 2002, pp. 99-127, Evolutionary Computation, vol. 10, No. 2. [cited by applicant]
Mingxing Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, May 24, 2019, pp. 6105-6114 , Int'l Conference on Machine Learning, PMLR. [cited by applicant]
Jack Turner et al., “BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget”, Jan. 23, 2020, 15 pages, arXiv:1906.04113v2. [cited by applicant]
Jesper E. Van Engelen et al., “A survey on semi-supervised learning”, Nov. 15, 2019, pp. 373-440, Machine Learning, vol. 109, No. 2. [cited by applicant]
Barret Zoph et al., “Neural Architecture Search with Reinforcement Learning”, Feb. 15, 2017, 16 pages, arXiv: 1611.01578v2. [cited by applicant]
Song Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, arXiv:1510.00149v5 [cs.CV], 14 pages (Feb. 15, 2016). [cited by applicant]
Andrew Howard et al., “Searching for MobileNetV3”, Proceedings of the IEEE/Computer Vision Foundation (CVF) Int'l Conference on Computer Vision (ICCV 2019), pp. 1314-1324 (Oct. 2019). [cited by applicant]
Liam Li et al., “Geometry-Aware Gradient Algorithms for Neural Architecture Search”, arXiv:2004.07802v5 [cs.LG], 25 pages (Mar. 18, 2021). [cited by applicant]
Mu Li et al., “Scaling Distributed Machine Learning with the Parameter Server”, 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI '14), pp. 583-598 (Oct. 2014). [cited by applicant]
Misha Khodak et al., “In defense of weight-sharing for neural architecture search: an optimization perspective”, Machine Learning Blog Carnegie Mellon University (ML@CMU), 12 pages (Jul. 17, 2020), https://blog.ml.cmu.e… [cited by applicant]
Hieu Pham et al. “Efficient Neural Architecture Search via Parameters Sharing”, Proceedings of the 35th Int'l Conference on Machine Learning (PMLR), vol. 80, pp. 4095-4104 (2018). [cited by applicant]
Aurick Qiao et al., “Litz: Elastic Framework for High-Performance DistributedMachine Learning”, 2018 USENIX Annual Technical Conference (USENIX ATC '18), pp. 631-643 (Jul. 2018). [cited by applicant]
Mohammad Rastegari et al. “XNOR-Net: Imagenet Classification Using Binary Convolutional Neural Networks”, European Conference on Computer Vision (ECCV), Springer, Cham., pp. 525-542 (Oct. 8, 2016). [cited by applicant]
Lingxi Xie et al., “Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap”, arXiv:2008.01475v2 [cs.CV], 24 pages (Aug. 5, 2020). [cited by applicant]
Yuge Zhang et al., “Deeper Insights Into Weight Sharing in Neural Architecture Search”, arXiv:2001.01431v1 [cs.LG], 16 pages (Jan. 6, 2020). [cited by applicant]
Jonathan Frankle et al., “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks”, arXiv:1803.03635v5 [cs.LG], 42 pages (Mar. 4, 2019), https://arxiv.org/pdf/1803.03635.pdf. [cited by applicant]
Jianping Gou et al., “Knowledge distillation: A survey”, Int'l J. of Comp. Vision, vol. 129, No. 6, pp. 1789-1819 (Jun. 2021), https://arxiv.org/pdf/2006.05525.pdf. [cited by applicant]
Lucas Hansen, “Tiny ImageNet Challenge Submission”, CS231n: Convolutional Neural Networks for Visual Recognition, Stanford University, 6 pages (2015), http://cs231n.stanford.edu/reports/2015/pdfs/lucash_final.pdf. [cited by applicant]
Kaiming He et al., “Deep Residual Learning for Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition 2016, pp. 770-778 (2016), https:/openaccess.thecvf.com/content_cvpr_2016/p… [cited by applicant]
Adrián Hernández et al., “Attention Mechanisms and Their Applications to Complex Systems”, Entropy, vol. 23, No. 3, p. 283, 18 pages (Feb. 26, 2021), https://www.mdpi.com/1099-4300/23/3/283/htm. [cited by applicant]
Geoffrey Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv preprint arXiv:1503.02531, 9 pages (Mar. 9, 2015), https://arxiv.org/pdf/1503.02531.pdf. [cited by applicant]
Le Hou et al., “High Resolution Medical Image Analysis with Spatial Partitioning”, arXiv:1909.03108v3 [eess.IV], 5 pages (Sep. 12, 2019), https://arxiv.org/pdf/1909.03108.pdf. [cited by applicant]
Cheng-Zhi Anna Huang et al., “Music Transformer: Generating Music with Long-Term Structure”, arXiv:1809.04281v3 [cs.LG], 14 pages (Dec. 12, 2018), https://arxiv.org/pdf/1809.04281.pdf. [cited by applicant]
Judit Ács, “Masking attention weights in PyTorch”, Judit Ács's blog, 4 pages (Dec. 27, 2018; last visited Oct. 2, 2021), http://juditacs.github.io/2018/12/27/masked-attention.html. [cited by applicant]
Salman Khan et al., “Transformers in Vision: A Survey”, arXiv:2101.01169v2 [cs.CV], 28 pages (Feb. 22, 2021), https://arxiv.org/pdf/2101.01169v2.pdf. [cited by applicant]
Taehyeon Kim et al., “Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation”, 2021, pp. 2628-2635, (Proceedings of the Thirtieth International Joint Conference on Artificial Intelli… [cited by applicant]
Taehyeon Kim et al., “Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation”, May 19, 2021, 11 pages, arXiv:2105.08919 [cs.LG]. [cited by applicant]
Salman Khan et al., “Transformers in Vision: A Survey”, ACM Computing Surveys, 38 pages (Accepted Dec. 2021; available online Jan. 6, 2022), https://dl.acm.org/doi/abs/10.1145/3505244. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Sep. 16, 2024, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/504,996, dated Mar. 8, 2024, 19 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Oct. 25, 2024, 8 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Patent Application No. 17/506, 161, dated Jan. 30, 2025, 34 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Mar. 12, 2025, 9 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/506,161, dated Jan. 30, 2025, 34 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowability,” issued in connection with U.S. Appl. No. 17/504,996, dated Apr. 28, 2025, 2 pages. [cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/497,736 dated May 7, 2025, 36 pages. [cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/506,161, dated May 19, 2025, 5 pages. [cited by applicant]
Cited By (1)
US 12,555,573