IP Library Granted Patent US 12,299,561
Granted Patent B2
US 12,299,561 · App. 18/469,980 · Granted May 13, 2025

On-the-fly deep learning in machine learning for autonomous machines

Inventor: Raanan Yonatan Yehezkel Rohekar (Kiryat Ekron, IL)
Assignee: Intel Corporation
G06N3/063G06F9/46G06F18/214G06F18/2148G06F18/2411G06F18/2413G06N3/04G06N3/044G06N3/045G06N3/08G06N3/084G06V10/454G06V10/764G06V10/774G06V10/82G06V10/95G06V10/955G06V20/00G06V40/174G06V2201/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,561
App. No.
18/469,980
Granted
May 13, 2025
Kind
B2
Abstract

A mechanism is described for facilitating on-the-fly deep learning in machine learning for autonomous machines. A method of embodiments, as described herein, includes detecting an output associated with a first deep network serving as a user-independent model associated with learning of one or more neural networks at a computing device having a processor coupled to memory. The method may further include automatically generating training data for a second deep network serving as a user-dependent model, where the training data is generated based on the output. The method may further include merging the user-independent model with the user-dependent model into a single joint model.

Claims (29)

1. A data processing system on a computing device, the data processing system comprising:

one or more processors including a general-purpose graphics processor; and

one or more storage devices comprising a graphics execution environment including a machine learning framework to provide machine learning primitives that are accelerated via the general-purpose graphics processor, the one or more processors to perform operations comprising:

selecting a topology for a first machine learning model for use with the machine learning framework, wherein the first machine learning model comprises a pre-trained machine learning model that was trained on a user-independent dataset to enable user-independent classification of input data; and

training, via the machine learning framework, a second machine learning model based on a dataset including user-dependent data, wherein the user-dependent data includes images and associated labels and the second machine learning model comprises an update of the first machine learning model,

wherein training the second machine learning model includes training the second machine learning model via a primitive provided by the machine learning framework and tuning parameters of the second machine learning model via the machine learning framework, the machine learning framework is a first machine learning framework that is configured to operate with a second machine learning framework, and the second machine learning framework is configured to provide a second primitive to implement a first primitive provided by the first machine learning framework.

2. The data processing system of claim 1 , wherein the first machine learning model comprises a pre-trained machine learning model for computer vision.

3. The data processing system of claim 2 , wherein training the second machine learning model includes training the second machine learning model to perform computer vision operations for image recognition.

4. The data processing system of claim 1 , wherein the first machine learning model comprises a pre-trained machine learning model for speech synthesis.

5. The data processing system of claim 1 , wherein the one or more processors are additionally configured to perform operations including receiving a selection of a topology for the first machine learning model.

6. The data processing system of claim 1 , wherein training the second machine learning model via the primitive includes training the second machine learning model via the first primitive.

7. The data processing system of claim 6 , wherein the second primitive is to cause the general-purpose graphics processor to perform an operation to train a layer of the second machine learning model.

8. The data processing system of claim 1 , wherein the graphics execution environment is a virtualized environment.

9. The data processing system of claim 1 , wherein the general-purpose graphics processor is configurable into partitions and the graphics execution environment is to execute via one or more partitions of general-purpose graphics processor.

10. The data processing system of claim 9 , wherein the general-purpose graphics processor is configured into multiple partitions and the general-purpose graphics processor is to execute multiple graphics execution environments via the multiple partitions.

11. A method comprising:

receiving selection of a topology for a first machine learning model for use with a machine learning framework, the machine learning framework to provide machine learning primitives that are accelerated via a general-purpose graphics processor, wherein the first machine learning model comprises a pre-trained machine learning model that was trained on a user-independent dataset to enable user-independent classification of input data; and

training, via the machine learning framework, a second machine learning model based on a dataset including user-dependent data, wherein the user-dependent data includes images and associated labels and the second machine learning model comprises an update of the first machine learning model, wherein training the second machine learning model includes training the second machine learning model via a primitive provided by the machine learning framework and tuning parameters of the second machine learning model via the machine learning framework, the machine learning framework is a first machine learning framework that is configured to operate with a second machine learning framework, and the second machine learning framework is configured to provide a second primitive to implement a first primitive provided by the first machine learning framework.

12. The method of claim 11 , wherein the first machine learning model comprises a pre-trained machine learning model for computer vision.

13. The method of claim 12 , wherein training the second machine learning model includes training the second machine learning model to perform computer vision operations for image recognition.

14. The method of claim 11 , wherein the first machine learning model comprises a pre-trained machine learning model for speech synthesis.

15. The method of claim 11 , wherein training the second machine learning model via the primitive includes training the second machine learning model via the first primitive.

16. A non-transitory machine-readable medium having instructions stored thereon, the instructions, when executed by one or more processors, cause the one or more processors to perform operations comprising:

selecting or receiving selection of a topology for a first machine learning model for use with a machine learning framework, the machine learning framework to provide machine learning primitives that are accelerated via a general-purpose graphics processor, wherein the first machine learning model comprises a pre-trained machine learning model trained on a user-independent dataset to enable user-independent classification of input data; and

training, via the machine learning framework, a second machine learning model based on a dataset including user-dependent data, wherein the user-dependent data includes images and associated labels and the second machine learning model comprises an update of the first machine learning model, wherein training the second machine learning model includes training the second machine learning model via a primitive provided by the machine learning framework and tuning parameters of the second machine learning model via the machine learning framework, the machine learning framework is a first machine learning framework that is configured to operate with a second machine learning framework, and the second machine learning framework is configured to provide a second primitive to implement a first primitive provided by the first machine learning framework.

17. The non-transitory machine-readable medium of claim 16 , wherein the first machine learning model comprises a pre-trained machine learning model for computer vision.

18. The non-transitory machine-readable medium of claim 17 , wherein training the second machine learning model includes training the second machine learning model to perform computer vision operations for image recognition.

19. The non-transitory machine-readable medium of claim 16 , wherein the first machine learning model comprises a pre-trained machine learning model for speech synthesis.

20. The non-transitory machine-readable medium of claim 16 , wherein training the second machine learning model via the primitive includes training the second machine learning model via the first primitive.

Continuity (7)
Continuation 18322218 · May 23, 2023
Continuation 17400908 · Aug 12, 2021
Continuation 16929976 · Jul 15, 2020
Continuation 16783451 · Feb 6, 2020
Continuation 15659818 · Jul 26, 2017
Provisional Application 62502294 · May 5, 2017
Related Publication 20240005137A1 · Jan 4, 2024
References Cited (59)
US 6751619B1 · Rowstron · 2004 [cited by examiner]
US 7627458B1 · Van Mau · 2009 [cited by examiner]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10572773B2 · Yehezkel Rohekar · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11146564B1 · Ankam et al. · 2021 [cited by applicant]
US 20030093239A1 · Schmit · 2003 [cited by examiner]
US 20040267823A1 · Shapiro · 2004 [cited by examiner]
US 20070198241A1 · Beausoleil · 2007 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160110642A1 · Matsuda et al. · 2016 [cited by applicant]
US 20160342888A1 · Yang et al. · 2016 [cited by applicant]
US 20160379112A1 · He et al. · 2016 [cited by applicant]
US 20170011738A1 · Senior et al. · 2017 [cited by applicant]
US 20170024849A1 · Liu et al. · 2017 [cited by applicant]
US 20170256086A1 · Park · 2017 [cited by examiner]
US 20170300767A1 · Zou et al. · 2017 [cited by applicant]
US 20170308789A1 · Langford et al. · 2017 [cited by applicant]
US 20180012110A1 · Souche · 2018 [cited by examiner]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180293102A1 · Ray · 2018 [cited by examiner]
US 20180293691A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20190114541A1 · Yang et al. · 2019 [cited by applicant]
US 20200081744A1 · Siegl et al. · 2020 [cited by applicant]
US 20210089316A1 · Rash · 2021 [cited by examiner]
CN 108805292A · 2018 [cited by applicant]
CN 111915025A · 2020 [cited by applicant]
Georganas, Evangelos, et al. “Tensor processing primitives: A programming abstraction for efficiency and portability in deep learning workloads.” Proceedings of the International Conference for High Performance Computin… [cited by examiner]
Smith, Micah J., et al. “The machine learning bazaar: Harnessing the ml ecosystem for effective system development.” Proceedings of the 2020 ACM SIGMOD International Conference on Management of Data. 2020. (Year: 2020). [cited by examiner]
Nazir, Zhumakhan, Vladislav Yarovenko, and Jurn-Gyu Park. “Interpretable ML enhanced CNN Performance Analysis of cuBLAS, cuDNN and TensorRT.” Proceedings of the 38th ACM/SIGAPP Symposium on Applied Computing. 2023. (Yea… [cited by examiner]
Radu, Valentin, et al. “Performance aware convolutional neural network channel pruning for embedded GPUs.” 2019 IEEE International Symposium on Workload Characterization (IISWC). IEEE, 2019. (Year: 2019). [cited by examiner]
Yi, Xinyao. “A Study of Performance Programming of CPU, GPU accelerated Computers and SIMD Architecture.” arXiv preprint arXiv:2409.10661 (2024). (Year: 2024). [cited by examiner]
Chetlur, Sharan, et al. “cudnn: Efficient primitives for deep learning.” arXiv preprint arXiv:1410.0759 (2014). (Year: 2014). [cited by examiner]
Final Office Action for U.S. Appl. No. 18/322,218 mailed Mar. 18, 2024, 21 pages. [cited by applicant]
Notice of Publication for CN Application No. 202311277782.9, published as CN117556868A on Feb. 13, 2024. [cited by applicant]
Decision to Grant CN application No. 202010770667.5, mailed Feb. 7, 2024, 10 pages. [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Office Action for U.S. Appl. No. 15/659,818, 17 pages, Mar. 28, 2019. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/659,818, 10 pages, Oct. 17, 2019. [cited by applicant]
Elfwing et al., “Expected energy-based restricted Boltzmann machine for classification”, Science Direct, Neural Networks Journal, Sep. 28, 2014, 10 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/929,976, Aug. 5, 2020, 20 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/929,976, Feb. 12, 2021, 21 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/929,976,May 12, 2021, 13 pages. [cited by applicant]
Office Action for U.S. Appl. No. 16/783,451, Jul. 2, 2021, 23 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/783,451, malled Oct. 18, 2021, 46 pages. [cited by applicant]
De Prado, Miguel, Nuria Pazos, and Luca Benini. “Learning to infer: RL-based search for DNN primitive selection on Heterogeneous Embedded Systems.” 2019 Design, Automation & Test in Europe Conference & Exhibition (DATE)… [cited by applicant]
Chetlur S, Woolley C, Vandermersch P, Cohen J, Tran J, Catanzaro B, Shelhamer E. cudnn: Efficient primitives for deep learning. arXiv preprint arXiv: 1410.0759. Oct. 3, 2014. (Year: 2014). [cited by applicant]
“CUDNN: Efficient Primitives for Deep Learning” Workshop on Intro to Deep Neural Netvvorks Aug. 26 to 27, 2016 Presented by: AmnahNasim Supervised by: Dr. Asifullah Khan DCIS, PIEAS (Year: 2016). [cited by applicant]
Chellapilla, K., Puri, S., & Simard, P. (Oct. 2006). High performance convolutional neural netvvorks for document processing. In Tenth international workshop on frontiers in handwriting recognition. Suvisoft. (Year: 200… [cited by applicant]
Hacker, Christian, Igor Alzenberg, and Jeff Wilson. “GPU simulator of multilayer neural netvvork based on multi-valued neurons.” 2016 International Joint Conference on Neural Netvvorks (IJCN N). IEEE, 2016. (Year: 2016). [cited by applicant]
Khan J, Fultz P, Tamazov A, Lowell D, Liu C, Melesse M, Nandhimandalam M, Nasyrov K, Perminov I, Shah T, Filippov V. MI Open: An open source library for deep learning primitives. arXiv preprint arXiv: 1910.00078. Sep. 3… [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/783,451 mailed Feb. 16, 2022, 19 pages. [cited by applicant]
Office Action for CN Application No. 201810419226.3, Oct. 16, 2024, 10 pages. No translation available. [cited by applicant]