US 20130185234A1
· Arora
· 2013
[cited by examiner]
US 20150095253A1
· Lim
· 2015
[cited by examiner]
US 20160328644A1
· Lin et al.
· 2016
[cited by applicant]
US 20170220928A1
· Hajizadeh
· 2017
[cited by examiner]
US 20180137219A1
· Goldfarb
· 2018
[cited by examiner]
US 20190019096A1
· Yoshida
· 2019
[cited by examiner]
US 20190286984A1
· Vasudevan
· 2019
[cited by examiner]
US 20190370659A1
· Dean
· 2019
[cited by examiner]
US 20200104710A1
· Vasudevan
· 2020
[cited by examiner]
US 20200104715A1
· Denolf
· 2020
[cited by examiner]
US 20200272909A1
· Parmentier
· 2020
[cited by examiner]
US 20200349050A1
· Ghobadi
· 2020
[cited by applicant]
US 20200357486A1
· Kok
· 2020
[cited by examiner]
US 20210081763A1
· Abdelfattah
· 2021
[cited by examiner]
US 20210110140A1
· Munoz et al.
· 2021
[cited by applicant]
US 20210350233A1
· Saboori et al.
· 2021
[cited by applicant]
US 20210357959A1
· Cella
· 2021
[cited by applicant]
US 20220012089A1
· Nasr-Azadani et al.
· 2022
[cited by applicant]
US 20220027792A1
· Cummings et al.
· 2022
[cited by applicant]
US 20220036194A1
· Sundaresan
· 2022
[cited by examiner]
US 20220198217A1
· Dong
· 2022
[cited by examiner]
US 20220198260A1
· Xue
· 2022
[cited by examiner]
US 20220284582A1
· Yang
· 2022
[cited by examiner]
US 20220328128A1
· Kok
· 2022
[cited by applicant]
US 20220348903A1
· Ranganathan et al.
· 2022
[cited by applicant]
US 20230274151A1
· Xu
· 2023
[cited by examiner]
US 20230325711A1
· Haraldson
· 2023
[cited by examiner]
US 20240046148A1
· Bega
· 2024
[cited by examiner]
US 20240289687A1
· Kumar et al.
· 2024
[cited by applicant]
US 20240362472A1
· Fu
· 2024
[cited by examiner]
EP 3975060A1
· 2022
[cited by applicant]
WO WO2021158313A1
· 2021
[cited by applicant]
Dzmitry Bahdanau et al., “Neural Machine Translation by Jointly Learning to Align and Translate”, May 19, 2016, 15 pages, arXiv:1409.0473v7 [cs.CL].
[cited by applicant]
Irwan Bello et al., “Attention Augmented Convolutional Networks”, 2019, pp. 3286-3295, Proceedings of the IEEE/CVF Int'l Conference on Computer Vision.
[cited by applicant]
Cristian Buciluâ et al., “Model Compression”, Aug. 20, 2006, pp. 535-541, Proceedings of the 12th Assn. for Computing Machinery (ACM) Special Interest Group on Knowledge Discovery in Data (SIGKDD) Int'l Conference on Kn…
[cited by applicant]
Liang-Chieh Chen et al., “Deeplab: Semantic Image Segmentation with Deep Convolutional Nets, Atrous Convolution, and Fully Connected CRFs”, Apr. 27, 2017, pp. 834-848, IEEE transactions on pattern analysis and machine i…
[cited by applicant]
Yu Cheng et al., “A Survey of Model Compression and Acceleration for Deep Neural Networks”, Jun. 14, 2020, 10 pages, arXiv:1710.09282v9 [cs.LG].
[cited by applicant]
Aston Zhang et al., “Dive into Deep Learning” Jul. 25, 2021, 323 pages, Release 0.17.0, chs. 4-10 (Jul. 25, 2021), available at: https://d2l.ai/.
[cited by applicant]
Tim Dettmers et al., “Sparse Networks from Scratch: Faster Training without Losing Performance”, Aug. 23, 2019, 14 pages, arXiv:1907.04840v2 [cs.LG].
[cited by applicant]
Yu Cheng et al., “A Survey of Model Compression and Acceleration for Deep Neural Networks”, Oct. 23, 2017, 10 pages, IEEE Signal Processing Magazine, Special Issue on Deep Learning for Image Understanding.
[cited by applicant]
Alex Krizhevsky et al., “Learning Multiple Layers of Features from Tiny Images”, Master's Thesis, U. of Toronto, Citeseer, 60 pages (Apr. 8, 2009), https://www.cs.toronto.edu/˜kriz/learning-features-2009-TR.pdf.
[cited by applicant]
Souvik Kundu et al., “A Tunable Robust Pruning Framework Through Dynamic Network Rewiring of DNNS”, arXiv:2011.03083v2 [cs.CV], 8 pages (Nov. 24, 2020).
[cited by applicant]
Namhoon Lee et al., “SNIP: Single-Shot Network Pruning based on Connection Sensitivity”, Int'l Conference on Learning Representations (ICLR) 2019, 15 pages (May 6, 2019).
[cited by applicant]
Tsung-Yi Lin et al., “Feature Pyramid Networks for Object Detection”, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition 2017, pp. 2117-2125 (2017).
[cited by applicant]
Zhuang Liu et al., “Rethinking the Value of Network Pruning”, arXiv:1810.05270v2 [cs.LG], 21 pages (Mar. 5, 2019).
[cited by applicant]
Marc E. McDill, “Forest Resource Management”, Penn State Univ., ch. 11, pp. 203-233 (Jun. 22, 1999), https://faculty.washington.edu/toths/Presentations/Lecture%202/Ch11_LPIntro.pdf.
[cited by applicant]
Hesham Mostafa et al., “Parameter Efficient Training of Deep Convolutional Neural Networks by Dynamic Sparse Reparameterization”, Proceedings of the 36th Int'l Conference on Machine Learning (PMLR), pp. 4646-4655 (May 2…
[cited by applicant]
Junki Park et al., “OPTIMUS: Optimized Matrix Multiplication Structure for Transformer Neural Network Accelerator”, Mar. 15, 2020, pp. 363-378, Proceedings of Machine Learning and Systems, vol. 2.
[cited by applicant]
Prajit Ramachandran et al., “Stand-Alone Self-Attention in Vision Models”, arXiv:1906.05909v1 [cs.CV], 15 pages (Jun. 13, 2019), https://arxiv.org/pdf/1906.05909.pdf.
[cited by applicant]
Christian Szegedy et al., “Going Deeper with Convolutions”, 2015 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 1-9 (2015), doi: 10.1109/CVPR.2015.7298594.
[cited by applicant]
James Bergstra et al., “Algorithms for Hyper-Parameter Optimization”, 2011, 9 pages, Advances in Neural Information Processing Systems 24 (NIPS 2011).
[cited by applicant]
James Bergstra et al., “Random Search for Hyper-Parameter Optimization”, Feb. 1, 2012, 25 pages, J. of Machine Learning Research, vol. 13, No. 2.
[cited by applicant]
Han Cai et al., “Once-For-All: Train One Network and Specialize for Efficient Deployment”, Apr. 29, 2020, 15 pages, arXiv:1908.09791v5 [cs.LG].
[cited by applicant]
Kalyanmoy Deb et al., “A Fast and Elitist Multiobjective Genetic Algorithm: NSGA-II”, Apr. 2002, 16 pages, IEEE Transactions on Evolutionary Computation, vol. 6, No. 2.
[cited by applicant]
Ian Dewancker et al., “Bayesian Optimization Primer”, 2020, 4 pages. Retrieved from the Internet: http://static.sigopt.com/b/20a144d208ef255d3b981ce419667ec25d8412e2/static/pdf/SigOpt_Bayesian_Optimization_Primer.pdf.
[cited by applicant]
Ian Dewancker et al., “Bayesian Optimization for Machine Learning: A Practical Guidebook”, Dec. 14, 2016, 15 pages, arXiv preprint arXiv:1612.04858.
[cited by applicant]
A.E. Eiben et al., “Introduction to Evolutionary Computing”, 2015,295 pages, second edition, Springer Heidelberg New York Dordrecht London.
[cited by applicant]
Wenlan Huang et al., “Survey on Multi-Objective Evolutionary Algorithms”, Aug. 1, 2019, 8 pages, IOP Conf. Series: J. of Physics: Conf. Series, vol. 1288, No. 1, p. 012057.
[cited by applicant]
Christian Igel et al., “Covariance Matrix Adaptation for Multi-objective Optimization”, 2007, pp. 1-28, Evolutionary Computation, vol. 15, No. 1.
[cited by applicant]
Danilo Vasconcellos Vargas et al., “General Subpopulation Framework and Taming the Conflict Inside Populations”, Jan. 2, 2019, 37 pages, arXiv:1901.00266v1 [cs.NE].
[cited by applicant]
Extended European Search Report mailed Jan. 2, 2023 for European Patent Application No. 22186944.9, 14 pages.
[cited by applicant]
Tang et al., “A Semi-Supervised Assessor of Neural Architectures”, arXiv:2005.06821v1 [cs.CV], arxiv.org, Cornell Univ. Library, Ithaca, NY, 10 pages (May 14, 2020).
[cited by applicant]
Kokiopoulou et al., “Fast Task-Aware Architecture Inference”, arXiv:1902.05781v1 [cs.LG], arxiv.org, Cornell Univ. Library, Ithaca, NY, 10 pages (Feb. 15, 2019).
[cited by applicant]
Zichao Lu et al., “NSGANetV2: Evolutionary Multi-objective Surrogate-Assisted Neural Architecture Search”, Computer Vision—ECCV 2020: 16th European Conference, Glasgow, UK, Aug. 23-28, 2020, pp. 35-51 (Aug. 2020).
[cited by applicant]
Extended European Search Report mailed Jan. 10, 2023 for European Patent Application No. 22186932.4, 9 pages.
[cited by applicant]
“Intel Agilex FSeries 027 FPGA R25A Product Specifications” retrieved on Oct. 6, 2021, 3 pages. Retrieved from Internet at: https://www.intel.com/content/www/us/en/products/sku/208599/intel-agilex-fseries-027-fpga-r25a/…
[cited by applicant]
“Intel Atom x6413E Processor Product Specifications”, retrieved on Oct. 6, 2021, 4 pages. Retrieved from Internet at https://ark.intel.com/content/www/us/en/ark/products/207908/intel-atom-x6413e-processor-1-5m-cache-up-…
[cited by applicant]
“Intel® Xeon® Platinum 8362 Processor Product Specifications” retrieved on Oct. 6, 2021, 4 pages. Retrieved from internet at: https://www.intel.com/content/www/us/en/products/sku/217216/intel-xeon-platinum-8362-processo…
[cited by applicant]
Mahrokh Javadi et al., “Combining Manhattan and Crowding distances in Decision Space for Multimodal Multi- objective Optimization Problems”, Sep. 12, 2019, 6 pages, Eurogen 2019.
[cited by applicant]
Will Koehrsen, “A Conceptual Explanation of Bayesian Hyperparameter Optimization for Machine Learning”, Jun. 28, 2018. 17 pages, Toward Data Science, available at: https://towardsdatascience.com/a-conceptual-explanation…
[cited by applicant]
Hanxiao Liu et al., “DARTS: Differentiable Architecture Search”, Apr. 23, 2019, 13 pages, arXiv:1806.09055v2 [cs.LG].
[cited by applicant]
Joseph Mellor et al., “Neural Architecture Search without Training”, Jul. 1, 2021, pp. 7588-7598, Int'l Conference on Machine Learning, PMLR.
[cited by applicant]
Jasper Snoek et al., “Practical Bayesian Optimization of Machine Learning Algorithms”, Aug. 29, 2012, 9 pages, Advances in Neural Information Processing Systems 25 (NIPS 2012).
[cited by applicant]
Eckart Zitzler et al., “SPEA2: Improving the Strength Pareto Evolutionary Algorithm”, May 2001, 21 pages, Computer Engineering and Communication Networks Lab (TIK), Swiss Fed. Inst. of Tech. (ETH), Zurich, Ch, TIK-Repor…
[cited by applicant]
Hanrui Wang et al., “HAT: Hardware-Aware Transformers for Efficient Natural Language Processing”, May 28, 2020, 14 pages, arXiv:2005.14187v1 [cs.CL].
[cited by applicant]
Hanrui Wang et al., “HAT: Hardware-Aware Transformers for Efficient Natural Language Processing”, 2020, 37 pages, ACL 2020 Slides, available at: https://hat.mit.edu/assets/ACL20_HAT_HanruiWang.pdf.
[cited by applicant]
Kalyanmoy Deb, “Multi-Objective Optimization Using Evolutionary Algorithms”, Indian Institute of Technology—Kanpur, Dept. of Mechanical Engineering, Kanpur, India, KanGAL Report No. 2011003, 24 pages (Feb. 10, 2011), ht…
[cited by applicant]
Zhichao Lu et al., “NSGA-Net: Neural Architecture Search using Multi-Objective Genetic Algorithm”, arXiv:1810.03522v2 [cs.CV], 13 pages (Apr. 18, 2019).
[cited by applicant]
Yonglong Tian et al., “Contrastive Representation Distillation”, arXiv:1910.10699v2 [cs.LG], 19 pages (Jan. 18, 2020).
[cited by applicant]
Tianyun Zhang et al., “A Systematic DNN Weight Pruning Framework using Alternating Direction Method of Multipliers”, Proceedings of the European Conference on Computer Vision (ECCV) 2018, in Lecture Notes in Computer Sc…
[cited by applicant]
Ashish Vaswani et al., “Attention Is All You Need”, Advances in Neural Information Processing Systems 30 (NIPS 2017), pp. 5998-6008 (2017), https://proceedings.neurips.cc/paper/2017/file/3f5ee243547dee91fbd053c1c4a845aa…
[cited by applicant]
Chunnan Wang et al., “Multi-Objective Neural Architecture Search Based on Diverse Structures and Adaptive Recommendation”, arXiv:2007.02749v2 [cs.CV], 11 pages (Aug. 13, 2020), https://arxiv.org/pdf/2007.02749.pdf.
[cited by applicant]
Sergey Zagoruyko et al., “Paying More Attention to Attention: Improving the Performance of Convolutional Neural Networks via Attention Transfer”, arXiv:1612.03928v3 [cs.CV], 13 pages (Feb. 12, 2017), https://arxiv.org/p…
[cited by applicant]
Chenzhuo Zhu et al., “Trained Ternary Quantization”, arXiv:1612.01064v3 [cs.LG], 10 pages (Feb. 23, 2017), https://arxiv.org/pdf/1612.01064.pdf.
[cited by applicant]
Barret Zoph et al., “Learning Transferable Architectures for Scalable Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition 2018, pp. 8697-8710 (2018), https://openaccess.thec…
[cited by applicant]
Mohamed S. Abdelfattah et al., “Zero-Cost Proxies for Lightweight NAS”, Mar. 19, 2021, 17 pages, arXiv:2101.08134v2 [cs.LG].
[cited by applicant]
Edward Alibekov et al., “Proxy Functions for Approximate Reinforcement Learning”, 2019, pp. 224-229, IFAC PapersOnLine 52-11.
[cited by applicant]
Bowen Baker et al., “Designing Neural Network Architectures using Reinforcement Learning”, Mar. 22, 2017, 18 pages, arXiv:1611.02167v3.
[cited by applicant]
Han Cai et al., “Proxylessnas: Direct neural architecture search on target task and hardware”, Feb. 23, 2019, 13 pages, arXiv:1812.00332v2 [cs.LG].
[cited by applicant]
Olivier Chapelle et al., “Semi-Supervised Learning”, 2006, 524 pages, MIT Press, Cambridge MA, London England.
[cited by applicant]
Thomas Elsken et al., “Neural architecture search: A survey”, Mar. 19, 2019, 21 pages, J. of Machine Learning Research, vol. 20, No. 1.
[cited by applicant]
Thomas N. Kipf et al., “Semi-Supervised Classification with Graph Convolutional Networks”, Feb. 22, 2017, 14 pages, arXiv:1609.02907v4.
[cited by applicant]
M. Z. Naser et al., “Insights into Performance Fitness and Error Metrics for Machine Learning”, May 17, 2020, 25 pages, arXiv:2006.00887v1.
[cited by applicant]
Kenneth O. Stanley et al., “Evolving Neural Networks through Augmenting Topologies”, Jun. 2002, pp. 99-127, Evolutionary Computation, vol. 10, No. 2.
[cited by applicant]
Mingxing Tan et al., “EfficientNet: Rethinking Model Scaling for Convolutional Neural Networks”, May 24, 2019, pp. 6105-6114 , Int'l Conference on Machine Learning, PMLR.
[cited by applicant]
Jack Turner et al., “BlockSwap: Fisher-guided Block Substitution for Network Compression on a Budget”, Jan. 23, 2020, 15 pages, arXiv:1906.04113v2.
[cited by applicant]
Jesper E. Van Engelen et al., “A survey on semi-supervised learning”, Nov. 15, 2019, pp. 373-440, Machine Learning, vol. 109, No. 2.
[cited by applicant]
Barret Zoph et al., “Neural Architecture Search with Reinforcement Learning”, Feb. 15, 2017, 16 pages, arXiv: 1611.01578v2.
[cited by applicant]
Song Han et al., “Deep Compression: Compressing Deep Neural Networks with Pruning, Trained Quantization and Huffman Coding”, arXiv:1510.00149v5 [cs.CV], 14 pages (Feb. 15, 2016).
[cited by applicant]
Andrew Howard et al., “Searching for MobileNetV3”, Proceedings of the IEEE/Computer Vision Foundation (CVF) Int'l Conference on Computer Vision (ICCV 2019), pp. 1314-1324 (Oct. 2019).
[cited by applicant]
Liam Li et al., “Geometry-Aware Gradient Algorithms for Neural Architecture Search”, arXiv:2004.07802v5 [cs.LG], 25 pages (Mar. 18, 2021).
[cited by applicant]
Mu Li et al., “Scaling Distributed Machine Learning with the Parameter Server”, 11th USENIX Symposium on Operating Systems Design and Implementation (OSDI '14), pp. 583-598 (Oct. 2014).
[cited by applicant]
Misha Khodak et al., “In defense of weight-sharing for neural architecture search: an optimization perspective”, Machine Learning Blog Carnegie Mellon University (ML@CMU), 12 pages (Jul. 17, 2020), https://blog.ml.cmu.e…
[cited by applicant]
Hieu Pham et al. “Efficient Neural Architecture Search via Parameters Sharing”, Proceedings of the 35th Int'l Conference on Machine Learning (PMLR), vol. 80, pp. 4095-4104 (2018).
[cited by applicant]
Aurick Qiao et al., “Litz: Elastic Framework for High-Performance DistributedMachine Learning”, 2018 USENIX Annual Technical Conference (USENIX ATC '18), pp. 631-643 (Jul. 2018).
[cited by applicant]
Mohammad Rastegari et al. “XNOR-Net: Imagenet Classification Using Binary Convolutional Neural Networks”, European Conference on Computer Vision (ECCV), Springer, Cham., pp. 525-542 (Oct. 8, 2016).
[cited by applicant]
Lingxi Xie et al., “Weight-Sharing Neural Architecture Search: A Battle to Shrink the Optimization Gap”, arXiv:2008.01475v2 [cs.CV], 24 pages (Aug. 5, 2020).
[cited by applicant]
Yuge Zhang et al., “Deeper Insights Into Weight Sharing in Neural Architecture Search”, arXiv:2001.01431v1 [cs.LG], 16 pages (Jan. 6, 2020).
[cited by applicant]
Jonathan Frankle et al., “The Lottery Ticket Hypothesis: Finding Sparse, Trainable Neural Networks”, arXiv:1803.03635v5 [cs.LG], 42 pages (Mar. 4, 2019), https://arxiv.org/pdf/1803.03635.pdf.
[cited by applicant]
Jianping Gou et al., “Knowledge distillation: A survey”, Int'l J. of Comp. Vision, vol. 129, No. 6, pp. 1789-1819 (Jun. 2021), https://arxiv.org/pdf/2006.05525.pdf.
[cited by applicant]
Lucas Hansen, “Tiny ImageNet Challenge Submission”, CS231n: Convolutional Neural Networks for Visual Recognition, Stanford University, 6 pages (2015), http://cs231n.stanford.edu/reports/2015/pdfs/lucash_final.pdf.
[cited by applicant]
Kaiming He et al., “Deep Residual Learning for Image Recognition”, Proceedings of the IEEE conference on computer vision and pattern recognition 2016, pp. 770-778 (2016), https:/openaccess.thecvf.com/content_cvpr_2016/p…
[cited by applicant]
Adrián Hernández et al., “Attention Mechanisms and Their Applications to Complex Systems”, Entropy, vol. 23, No. 3, p. 283, 18 pages (Feb. 26, 2021), https://www.mdpi.com/1099-4300/23/3/283/htm.
[cited by applicant]
Geoffrey Hinton et al., “Distilling the Knowledge in a Neural Network”, arXiv preprint arXiv:1503.02531, 9 pages (Mar. 9, 2015), https://arxiv.org/pdf/1503.02531.pdf.
[cited by applicant]
Le Hou et al., “High Resolution Medical Image Analysis with Spatial Partitioning”, arXiv:1909.03108v3 [eess.IV], 5 pages (Sep. 12, 2019), https://arxiv.org/pdf/1909.03108.pdf.
[cited by applicant]
Cheng-Zhi Anna Huang et al., “Music Transformer: Generating Music with Long-Term Structure”, arXiv:1809.04281v3 [cs.LG], 14 pages (Dec. 12, 2018), https://arxiv.org/pdf/1809.04281.pdf.
[cited by applicant]
Judit Ács, “Masking attention weights in PyTorch”, Judit Ács's blog, 4 pages (Dec. 27, 2018; last visited Oct. 2, 2021), http://juditacs.github.io/2018/12/27/masked-attention.html.
[cited by applicant]
Salman Khan et al., “Transformers in Vision: A Survey”, arXiv:2101.01169v2 [cs.CV], 28 pages (Feb. 22, 2021), https://arxiv.org/pdf/2101.01169v2.pdf.
[cited by applicant]
Taehyeon Kim et al., “Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation”, 2021, pp. 2628-2635, (Proceedings of the Thirtieth International Joint Conference on Artificial Intelli…
[cited by applicant]
Taehyeon Kim et al., “Comparing Kullback-Leibler Divergence and Mean Squared Error Loss in Knowledge Distillation”, May 19, 2021, 11 pages, arXiv:2105.08919 [cs.LG].
[cited by applicant]
Salman Khan et al., “Transformers in Vision: A Survey”, ACM Computing Surveys, 38 pages (Accepted Dec. 2021; available online Jan. 6, 2022), https://dl.acm.org/doi/abs/10.1145/3505244.
[cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Sep. 16, 2024, 8 pages.
[cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/504,996, dated Mar. 8, 2024, 19 pages.
[cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Oct. 25, 2024, 8 pages.
[cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Patent Application No. 17/506, 161, dated Jan. 30, 2025, 34 pages.
[cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/504,996, dated Mar. 12, 2025, 9 pages.
[cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/506,161, dated Jan. 30, 2025, 34 pages.
[cited by applicant]
United States Patent and Trademark Office, “Notice of Allowability,” issued in connection with U.S. Appl. No. 17/504,996, dated Apr. 28, 2025, 2 pages.
[cited by applicant]
United States Patent and Trademark Office, “Non-Final Office Action,” issued in connection with U.S. Appl. No. 17/497,736 dated May 7, 2025, 36 pages.
[cited by applicant]
United States Patent and Trademark Office, “Notice of Allowance and Fee(s) Due,” issued in connection with U.S. Appl. No. 17/506,161, dated May 19, 2025, 5 pages.
[cited by applicant]