IP Library Granted Patent US 12,367,382
Granted Patent B2
US 12,367,382 · App. 17/323,694 · Granted Jul 22, 2025

Training with adaptive runtime and precision profiling

Inventors: Brian T. Lewis (Palo Alto, CA); Rajkishore Barik (Santa Clara, CA); Murali Sundaresan (Sunnyvale, CA); Leonard Truong (Santa Clara, CA); Feng Chen (Shanghai, CN); Xiaoming Chen (Shanghai, CN); Mike B. Macpherson (Portland, OR)
Assignee: INTEL CORPORATION
G06N3/063G06F7/483G06N3/044G06N3/045G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,367,382
App. No.
17/323,694
Granted
Jul 22, 2025
Kind
B2
Abstract

A mechanism is described for facilitating efficient training of neural networks at computing devices. A method of embodiments, as described herein, includes detecting one or more inputs for training of a neural network, and introducing randomness in floating point (FP) numbers to prevent overtraining of the neural network, where introducing randomness includes replacing less-significant low-order bits of operand and result values with new low-order bits during the training of the neural network.

Claims (42)

1. An apparatus comprising: a graphics hardware processor to:

detect one or more inputs for a first training of a neural network; track, using a hardware counter, precision of floating-point (FP) values during the first training of the neural network by the hardware counter observing exponents of one or more FP registers of hardware to determine a range of values of the exponents stored in the one or more FP registers, wherein the hardware counter is at least one of continuously or periodically reset;

store, in a precision tracking table based on the tracked precision of FP values, a number of bits that are used to represent single-precision FP value used with each layer of the neural network during the first training of the neural network;

provide automated mixed precision for the neural network by enabling selection of a precision for each layer based on the number of bits stored in the precision tracking table, wherein the precisions selected for each layer can differ from other layers; and

perform a second training of the neural network utilizing the precision selected for each layer of the neural network.

2. The apparatus of claim 1 , wherein the hardware counter is to periodically observe the exponents of the one or more FP registers of the hardware to determine the range of values stored in the one or more FP registers.

3. The apparatus of claim 2 , wherein the hardware counter is implemented using one or more bits based on a range of precision values.

4. The apparatus of claim 1 , wherein an adaptive runtime layer of the graphics hardware processor is to determine a best parallelism from a plurality of parallelisms relating to the neural network on a hardware platform.

5. The apparatus of claim 4 , wherein the plurality of parallelisms comprises one or more of a data parallelism, a model parallelism, and a hybrid parallelism.

6. The apparatus of claim 4 , wherein the adaptive runtime layer is to receive the one or more inputs and is communicatively coupled to one or more of the precision tracking table and a profiling table.

7. The apparatus of claim 1 , wherein the neural network comprises one or more of a convolutional neural network (CNN), a deep neural network (DNN), and a recurrent neural network (RNN), wherein the neural network is based on and in communication with one or more processors including the graphics hardware processor, wherein the graphics hardware processor is co-located with an application processor on a common semiconductor package.

8. A method comprising:

detecting, by a graphics hardware processor, one or more inputs for a first training of a neural network;

tracking, using a hardware counter, precision of floating-point (FP) values during the first training of the neural network by the hardware counter observing exponents of one or more FP registers of hardware to determine a range of values of the exponents stored in the one or more FP registers, wherein the hardware counter is at least one of continuously or periodically reset;

storing, in a precision tracking table based on the tracked precision of FP values, a number of bits that are used to represent single-precision FP value used with each layer of the neural network during the first training of the neural network;

providing automated mixed precision for the neural network by enabling selection of a precision for each layer based on the number of bits stored in the precision tracking table, wherein the precisions selected for each layer can differ from other layers; and

performing a second training of the neural network utilizing the precision selected for each layer of the neural network.

9. The method of claim 8 , wherein the hardware counter is to periodically observe the exponents of the one or more FP registers of the hardware to determine the range of values stored in the one or more FP registers.

10. The method of claim 9 , wherein the hardware counter is implemented using one or more bits based on a range of precision values.

11. The method of claim 8 , wherein an adaptive runtime layer of the graphics hardware processor is to determine a best parallelism from a plurality of parallelisms relating to the neural network on a hardware platform.

12. The method of claim 11 , wherein the plurality of parallelisms comprises one or more of a data parallelism, a model parallelism, and a hybrid parallelism.

13. The method of claim 11 , wherein the adaptive runtime layer is to receive the one or more inputs and is communicatively coupled to one or more of the precision tracking table and a profiling table.

14. The method of claim 8 , wherein the neural network comprises one or more of a convolutional neural network (CNN), a deep neural network (DNN), and a recurrent neural network (RNN), wherein the neural network is based on and in communication with one or more processors including the graphics hardware processor, wherein the graphics hardware processor is co-located with an application processor on a common semiconductor package.

15. At least one non-transitory machine-readable medium comprising instructions that when executed by a computing device, cause the computing device to perform operations comprising:

detecting, by a graphics hardware processor of the computing device, one or more inputs for a first training of a neural network;

tracking, using a hardware counter, precision of floating-point (FP) values during the first training of the neural network by the hardware counter observing exponents of one or more FP registers of hardware to determine a range of values of the exponents stored in the one or more FP registers, wherein the hardware counter is at least one of continuously or periodically reset;

storing, in a precision tracking table based on the tracked precision of FP values, a number of bits that are used to represent single-precision FP value used with each layer of the neural network during the first training of the neural network;

providing automated mixed precision for the neural network by enabling selection of a precision for each layer based on the number of bits stored in the precision tracking table, wherein the precisions selected for each layer can differ from other layers; and

performing a second training of the neural network utilizing the precision selected for each layer of the neural network.

16. The at least one non-transitory machine-readable medium of claim 15 , wherein the hardware counter is to periodically observe the exponents of the one or more FP registers of the hardware to determine the range of values stored in the one or more FP registers.

17. The at least one non-transitory machine-readable medium of claim 16 , wherein the hardware counter is implemented using one or more bits based on a range of precision values.

18. The at least one non-transitory machine-readable medium of claim 15 , wherein an adaptive runtime layer of the graphics hardware processor is to determine a best parallelism from a plurality of parallelisms relating to the neural network on a hardware platform.

19. The at least one non-transitory machine-readable medium of claim 18 , wherein the plurality of parallelisms comprises one or more of a data parallelism, a model parallelism, and a hybrid parallelism.

20. The at least one non-transitory machine-readable medium of claim 18 , wherein the adaptive runtime layer is to receive the one or more inputs and is communicatively coupled to one or more of the precision tracking table and a profiling table, wherein the neural network comprises one or more of a convolutional neural network (CNN), a deep neural network (DNN), and a recurrent neural network (RNN), wherein the neural network is based on and in communication with one or more processors including the graphics hardware processor, wherein the graphics hardware processor is co-located with an application processor on a common semiconductor package.

21. A system comprising: a memory; and a graphics processor communicably coupled to the memory, the graphics processor to:

detect one or more inputs for a first training of a neural network;

track, using a hardware counter, precision of floating-point (FP) values during the first training of the neural network by the hardware counter observing exponents of one or more FP registers of hardware to determine a range of values of the exponents stored in the one or more FP registers, wherein the hardware counter is at least one of continuously or periodically reset;

store, in a precision tracking table based on the tracked precision of FP values, a number of bits that are used to represent single-precision FP value used with each layer of the neural network during the first training of the neural network;

provide automated mixed precision for the neural network by enabling selection of a precision for each layer based on the number of bits stored in the precision tracking table, wherein the precisions selected for each layer can differ from other layers; and

perform a second training of the neural network utilizing the precision selected for each layer of the neural network.

22. The system of claim 21 , wherein the hardware counter is to periodically observe the exponents of the one or more FP registers of the hardware to determine the range of values stored in the one or more FP registers.

23. The system of claim 22 , wherein the hardware counter is implemented using one or more bits based on a range of precision values.

Continuity (2)
Continuation 15581031 · Apr 28, 2017
Related Publication 20210350215A1 · Nov 11, 2021
References Cited (33)
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10621486B2 · Yao · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 11017291B2 · Lewis et al. · 2021 [cited by applicant]
US 11676024B2 · Chai · 2023 [cited by examiner]
US 20030101207A1 · Dhong et al. · 2003 [cited by applicant]
US 20090249040A1 · Fujimoto et al. · 2009 [cited by applicant]
US 20120066163A1 · Balls et al. · 2012 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160328647A1 · Lin · 2016 [cited by examiner]
US 20170148433A1 · Catanzaro et al. · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180285733A1 · Mellempudi et al. · 2018 [cited by applicant]
US 20180314935A1 · Lewis et al. · 2018 [cited by applicant]
US 20190057303A1 · Burger · 2019 [cited by examiner]
US 20200026992A1 · Zhang et al. · 2020 [cited by applicant]
EP 3192016B1 · 2019 [cited by examiner]
Han, Song, Huizi Mao, and William J. Dally. “Deep compression: Compressing deep neural networks with pruning, trained quantization and huffman coding.” arXiv preprint arXiv:1510.00149 (2015). (Year: 2015). [cited by examiner]
Shan, Lei, et al. “A dynamic multi-precision fixed-point data quantization strategy for convolutional neural network.” CCF National Conference on Computer Engineering and Technology. Singapore: Springer Singapore, 2016.… [cited by examiner]
Lin, Darryl, Sachin Talathi, and Sreekanth Annapureddy. “Fixed point quantization of deep convolutional networks.” International conference on machine learning. PMLR, 2016. (Year: 2016). [cited by examiner]
Zhang, Hao, et al. “Application-Specific and Reconfigurable AI Accelerator.” Artificial Intelligence and Hardware Accelerators. Cham: Springer International Publishing, 2012. 183-223. (Year: 2012). [cited by examiner]
Goldberg, David. “What every computer scientist should know about floating-point arithmetic.” ACM computing surveys (CSUR) 23.1 (1991): 5-48. (Year: 1991). [cited by examiner]
Nomani, Junaid, and Jakub Szefer. “Predicting program phases and defending against side-channel attacks using hardware performance counters.” Proceedings of the Fourth Workshop on Hardware and Architectural Support for … [cited by examiner]
Lozito, Gabriele-Maria, et al. “FPGA implementations of feed forward neural network by using floating point hardware accelerators.” Advances in Electrical and Electronic Engineering 12.1 (2014): 30-39. (Year: 2014). [cited by examiner]
Huan, Yuxiang, et al. “A multiplication reduction technique with near-zero approximation for embedded learning in IoT devices.” 2016 29th IEEE International System-on-Chip Conference (SOCC). IEEE, 2016. (Year: 2016). [cited by examiner]
Coubariaux et al: “Training Deep Neural Networks with Low Precision Multiplications”, arXiv, Sep. 2015, p. 1-10. (Year: 2015). [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]