IP Library Granted Patent US 12,373,911
Granted Patent B2
US 12,373,911 · App. 18/456,235 · Granted Jul 29, 2025

Compute optimizations for low precision machine learning operations

Inventors: Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Sara S. Baghsorkhi (San Jose, CA); Anbang Yao (Beijing, CN); Kevin Nealis (San Jose, CA); Xiaoming Chen (Shanghai, CN); Altug Koker (El Dorado Hills, CA); Abhishek R. Appu (El Dorado Hills, CA); John C. Weast (Portland, OR); Mike B. Macpherson (Portland, OR); Dukhwan Kim (San Jose, CA); Linda L. Hurd (Cool, CA); Ben J. Ashbaugh (Folsom, CA); Barath Lakshmanan (Chandler, AZ); Liwei Ma (Beijing, CN); Joydeep Ray (Folsom, CA); Ping T. Tang (Edison, NJ); Michael S. Strickland (Sunnyvale, CA)
Assignee: Intel Corporation
G06T1/20G06F7/483G06F9/30014G06F9/30185G06F9/3863G06F9/5044G06N3/044G06N3/045G06N3/063G06N3/084G06N20/00G06F3/14G06T1/60G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,911
App. No.
18/456,235
Granted
Jul 29, 2025
Kind
B2
Abstract

One embodiment provides a general-purpose graphics processing unit comprising a dynamic precision floating-point unit including a control unit having precision tracking hardware logic to track an available number of bits of precision for computed data relative to a target precision, wherein the dynamic precision floating-point unit includes computational logic to output data at multiple precisions.

Claims (50)

1. A graphics processor comprising:

a memory device;

a compressor to compress data to be written to the memory device; and

a streaming multiprocessor coupled with the memory device, wherein the streaming multiprocessor includes a single instruction, multiple thread (SIMT) architecture and the streaming multiprocessor is to concurrently execute multiple threads, including a first thread in parallel with a second thread,

wherein the first thread is configured to process a first instruction to cause a first portion of the streaming multiprocessor to perform a floating-point operation on multiple floating-point input operands,

wherein the second thread is configured to process a second instruction to cause a second portion of the streaming multiprocessor to perform an integer operation on multiple integer operands,

wherein the streaming multiprocessor is to perform operations for a third instruction, the streaming multiprocessor to perform a first operation of the third instruction on 16-bit floating-point input and a second operation of the third instruction on input that includes a 32-bit floating-point input,

wherein the streaming multiprocessor is to perform operations for a fourth instruction, the streaming multiprocessor to perform a third operation on 8-bit integer input and a fourth operation on input that includes a 32-bit integer input, and

wherein the first operation of the fourth instruction includes a multiply and the second operation of the fourth instruction includes an accumulate.

2. The graphics processor as in claim 1 , wherein the 16-bit floating-point input includes a half-precision floating-point input.

3. The graphics processor as in claim 1 , wherein the multiple floating-point input operands include 64-bit data elements.

4. The graphics processor as in claim 1 , wherein the first operation of the third instruction includes a multiply and the second operation of the third instruction includes an accumulate.

5. The graphics processor as in claim 1 , further comprising a level-2 (L2) cache coupled with the compressor.

6. The graphics processor as in claim 5 , wherein the compressor is to losslessly compress the data to be written to the memory device.

7. The graphics processor as in claim 5 , wherein the compressor is to decompress data to be read from the memory device.

8. The graphics processor as in claim 6 , wherein the memory device is a high-bandwidth memory (HBM) device.

9. A method comprising:

decoding a first instruction via an instruction decoder of a graphics processor, the first instruction decoded into a first decoded instruction, wherein the graphics processor includes a streaming multiprocessor coupled to a memory device and a compressor to compress data to be written to the memory device, and the streaming multiprocessor includes a single instruction, multiple thread (SIMT);

executing multiple threads associated with the first decoded instruction via the streaming multiprocessor, wherein the first decoded instruction causes a first portion of the streaming multiprocessor to perform a floating-point operation on multiple floating-point input operands;

decoding a second instruction via the instruction decoder of the graphics processor into a second decoded instruction;

executing multiple threads associated with the second decoded instruction via the streaming multiprocessor, wherein the second decoded instruction causes a second portion of the streaming multiprocessor to perform an integer operation on multiple integer operands, wherein the streaming multiprocessor is to execute a thread of the first decoded instruction in parallel with a thread of the second decoded instruction;

decoding a third instruction via the instruction decoder of the graphics processor into a third decoded instruction;

executing multiple threads associated with the third decoded instruction via the streaming multiprocessor, wherein the streaming multiprocessor performs a first operation of the third decoded instruction on 16 -bit floating-point input and a second operation of the third decoded instruction on input that includes a 32-bit floating-point input;

decoding a fourth instruction via the instruction decoder of the graphics processor into a fourth decoded instruction; and

executing multiple threads associated with the fourth decoded instruction via the streaming multiprocessor, wherein the streaming multiprocessor performs the first operation of the fourth decoded instruction on 8 -bit integer input and a second operation of the fourth decoded instruction on input that includes a 32-bit integer input, wherein the first operation of the fourth decoded instruction includes a multiply and the second operation of the fourth decoded instruction includes an accumulate.

10. The method as in claim 9 , wherein the third instruction is a floating-point instruction, the streaming multiprocessor performs the first operation of the third decoded instruction using a first number of bits associated with a first floating-point precision and the second operation of the third decoded instruction using a second number of bits associated with a second floating-point precision.

11. The method as in claim 9 , wherein the 16-bit floating-point input includes a half-precision floating-point input.

12. The method as in claim 9 , wherein the streaming multiprocessor performs the first operation of the fourth decoded instruction using a first number of bits associated with a first representable range of integer values and the second operation of the fourth decoded instruction using a second number of bits associated with a second representable range of integer values.

13. The method as in claim 12 , wherein the multiple floating-point input operands include 64-bit data elements.

14. The method as in claim 13 , wherein the first operation of the third decoded instruction includes a multiply and the second operation of the third decoded instruction includes an accumulate.

15. The method as in claim 9 , further comprising compressing data associated with the first instruction, second instruction, or third instruction before writing the data to the memory device.

16. The method as in claim 15 , further comprising losslessly compressing the data associated with the first instruction, second instruction, or third instruction before writing the data to the memory device.

17. The method as in claim 9 , further comprising decompressing data associated with the first instruction, second instruction, or third instruction after reading the data from the memory device.

18. A graphics processing system comprising:

a system interface coupled with an interconnect fabric;

a graphics memory device coupled with the interconnect fabric;

a compressor to compress data to be written to the graphics memory device; and

a streaming multiprocessor coupled with the graphics memory device, wherein the streaming multiprocessor includes a single instruction, multiple thread (SIMT) architecture and the streaming multiprocessor is to concurrently execute multiple threads, including a first thread in parallel with a second thread,

wherein the first thread is configured to process a first instruction to cause a first portion of the streaming multiprocessor to perform a floating-point operation on multiple floating-point input operands,

wherein the second thread is configured to process a second instruction to cause a second portion of the streaming multiprocessor to perform an integer operation on multiple integer operands, and

wherein the streaming multiprocessor is to perform operations for a third instruction, the streaming multiprocessor to perform a first operation of the third instruction on 16-bit floating-point input and a second operation of the third instruction on input that includes a 32-bit floating-point input,

wherein the streaming multiprocessor is to perform operations for a fourth instruction, the streaming multiprocessor to perform a third operation on 8-bit integer input and a fourth operation on input that includes a 32-bit integer input, and

wherein the first operation of the fourth instruction includes a multiply and the second operation of the fourth instruction includes an accumulate.

19. The graphics processing system as in claim 18 , wherein the graphics memory device is a high-bandwidth memory (HBM) device.

20. The graphics processing system as in claim 18 , wherein the third instruction is a floating-point instruction, the streaming multiprocessor is to perform the first operation of the third instruction with a first number of bits to enable a first floating-point precision and perform the second operation of the third instruction with a second number of bits to enable a second floating-point precision.

21. The graphics processing system as in claim 18 , wherein the 16-bit floating-point input includes a half-precision floating-point input.

22. The graphics processing system as in claim 18 , further comprising a level-2 (L2) cache coupled with the compressor.

23. The graphics processing system as in claim 22 , wherein the compressor is to losslessly compress the data to be written to the graphics memory device.

24. The graphics processing system as in claim 22 , wherein the compressor is to decompress data to be read from the graphics memory device.

25. The graphics processing system as in claim 18 , wherein the streaming multiprocessor is to perform operations for a fourth instruction and the fourth instruction is an integer instruction, wherein the streaming multiprocessor is to perform a first operation of the fourth instruction with a first number of bits to represent a first range of integer values and perform the second operation with a second number of bits to represent a second range of integer values.

Continuity (9)
Continuation 17978573 · Nov 1, 2022
Continuation 17960611 · Oct 5, 2022
Continuation 17720804 · Apr 14, 2022
Continuation 16983080 · Aug 3, 2020
Continuation 16446265 · Jun 19, 2019
Continuation 16197821 · Nov 21, 2018
Continuation 15789565 · Oct 20, 2017
Continuation 15581167 · Apr 28, 2017
Related Publication 20230401668A1 · Dec 14, 2023
References Cited (152)
US 5974540A · Morikawa · 1999 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 8106914B2 · Oberman et al. · 2012 [cited by applicant]
US 9047171B2 · Fang et al. · 2015 [cited by applicant]
US 9619205B1 · Rowen et al. · 2017 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10726514B2 · Ould-Ahmed-Vall · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20040015533A1 · Hansen et al. · 2004 [cited by applicant]
US 20040205324A1 · Hansen et al. · 2004 [cited by applicant]
US 20050010743A1 · Tremblay et al. · 2005 [cited by applicant]
US 20050066205A1 · Holmer · 2005 [cited by applicant]
US 20060101244A1 · Siu et al. · 2006 [cited by applicant]
US 20090150654A1 · Oberman et al. · 2009 [cited by applicant]
US 20090184869A1 · Talbot et al. · 2009 [cited by applicant]
US 20110078427A1 · Shebanow et al. · 2011 [cited by applicant]
US 20110119446A1 · Blumrich et al. · 2011 [cited by applicant]
US 20110202745A1 · Bordawekar · 2011 [cited by applicant]
US 20110283059A1 · Govindarajan et al. · 2011 [cited by applicant]
US 20120011348A1 · Eichenberger · 2012 [cited by applicant]
US 20120051152A1 · Hollis · 2012 [cited by applicant]
US 20120256653A1 · Cordero et al. · 2012 [cited by applicant]
US 20130031328A1 · Kelleher et al. · 2013 [cited by applicant]
US 20130051502A1 · Khayrallah et al. · 2013 [cited by applicant]
US 20130117541A1 · Choquette et al. · 2013 [cited by applicant]
US 20130211213A1 · DeHennis et al. · 2013 [cited by applicant]
US 20130212139A1 · Gschwind et al. · 2013 [cited by applicant]
US 20140002989A1 · Ahuja et al. · 2014 [cited by applicant]
US 20140143497A1 · Olson et al. · 2014 [cited by applicant]
US 20140173606A1 · Pantaleoni · 2014 [cited by applicant]
US 20140176187A1 · Jayasena et al. · 2014 [cited by applicant]
US 20140181458A1 · Loh et al. · 2014 [cited by applicant]
US 20140201450A1 · Haugen · 2014 [cited by applicant]
US 20140267232A1 · Lum et al. · 2014 [cited by applicant]
US 20140281325A1 · Meaney et al. · 2014 [cited by applicant]
US 20140281366A1 · Felch · 2014 [cited by applicant]
US 20140281653A1 · Gilda et al. · 2014 [cited by applicant]
US 20140281783A1 · Hodges et al. · 2014 [cited by applicant]
US 20150054845A1 · Makarov · 2015 [cited by examiner]
US 20150170021A1 · Lupon et al. · 2015 [cited by applicant]
US 20150234783A1 · Angerer et al. · 2015 [cited by applicant]
US 20150371355A1 · Chen · 2015 [cited by applicant]
US 20150371407A1 · Kim et al. · 2015 [cited by applicant]
US 20150378741A1 · Lukyanov · 2015 [cited by examiner]
US 20160028912A1 · Falcon et al. · 2016 [cited by applicant]
US 20160048464A1 · Nakajima et al. · 2016 [cited by applicant]
US 20160055611A1 · Manevitch · 2016 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160070767A1 · Karras et al. · 2016 [cited by applicant]
US 20160092116A1 · Liu et al. · 2016 [cited by applicant]
US 20160126975A1 · Lutz et al. · 2016 [cited by applicant]
US 20160307482A1 · Huang et al. · 2016 [cited by applicant]
US 20170083827A1 · Robatmili et al. · 2017 [cited by applicant]
US 20170090924A1 · Mishra et al. · 2017 [cited by applicant]
US 20170103319A1 · Henry · 2017 [cited by applicant]
US 20170139677A1 · Lutz et al. · 2017 [cited by applicant]
US 20170168586A1 · Sinha et al. · 2017 [cited by applicant]
US 20170177312A1 · Boehm et al. · 2017 [cited by applicant]
US 20170214930A1 · Loughry · 2017 [cited by examiner]
US 20170263518A1 · Yu · 2017 [cited by examiner]
US 20170323201A1 · Sutskever et al. · 2017 [cited by applicant]
US 20180032796A1 · Kuharenko et al. · 2018 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180067720A1 · Bekas · 2018 [cited by applicant]
US 20180083635A1 · Wu · 2018 [cited by applicant]
US 20180113678A1 · Sadowski et al. · 2018 [cited by applicant]
US 20180246814A1 · Jayasena et al. · 2018 [cited by applicant]
US 20180246816A1 · Smith et al. · 2018 [cited by applicant]
US 20180315157A1 · Ould-Ahmed-Vall · 2018 [cited by applicant]
US 20180365167A1 · Eckert et al. · 2018 [cited by applicant]
US 20190206020A1 · Ould-Ahmed-Vall et al. · 2019 [cited by applicant]
CN 102648449A · 2012 [cited by applicant]
CN 103309702A · 2013 [cited by applicant]
CN 105378651A · 2016 [cited by applicant]
CN 105404889A · 2016 [cited by applicant]
CN 104081341A · 2017 [cited by applicant]
CN 110349075A · 2019 [cited by applicant]
CN 115082283 · 2022 [cited by applicant]
CN 118674603A · 2024 [cited by applicant]
EP 3396547A2 · 2018 [cited by applicant]
EP 3792761A1 · 2021 [cited by applicant]
EP 4099168 · 2022 [cited by applicant]
EP 4160413A1 · 2023 [cited by applicant]
TW 337570 · 1998 [cited by applicant]
TW 200937341A · 2009 [cited by applicant]
TW I514156B · 2015 [cited by applicant]
TW 201616341A · 2016 [cited by applicant]
TW I575477B · 2017 [cited by applicant]
TW 202238509 · 2022 [cited by applicant]
WO 2013039606A1 · 2013 [cited by applicant]
WO 2014190263A2 · 2014 [cited by applicant]
WO 2016028293A1 · 2016 [cited by applicant]
European Search Report for EP22197260, Jan. 30, 2023, 7 pages. [cited by applicant]
Kirk et al., “Programming Massively Parallel Processors—A Hands-On Approach”, Jan. 1, 2010, pp. 1-279, XP055073181. (Abstract included). [cited by applicant]
Office Action and Search Report for TW Application No. 111139972, 4 pages, Mar. 7, 2023. [cited by applicant]
Article 94(3) Communication for EP22197260.7, May 7, 2024, 4 pages. [cited by applicant]
Office Action for CN Application No. 202010848468.1, nailed May 17, 2024, 15 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC for EP Application No. 18164092.1, Sep. 23, 2020, 8 pages. [cited by applicant]
Decision to Grant EP Application No. 19182892.0, Dec. 17, 2020, 1 page. [cited by applicant]
European Search Report for EP19182892 6 pages, Dec. 13, 2019. [cited by applicant]
Extended European Search Report for EP Application No. 20205451.6, Jan. 22, 2021, 12 pages. [cited by applicant]
Extended European Search Report for EP Application No. EP18164092.1, 12 pages, Oct. 29, 2018. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/581,167 mailed Jan. 24, 2019, 29 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/789,565 mailed Apr. 5, 2018, 16 pages. [cited by applicant]
Final Office Action from U.S. Appl. No. 15/581,167, filed Nov. 5, 2019, 25 pages. [cited by applicant]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Junqing Sun et al: High-Performance Mixed-Precision Linear Solver for FPGAs11 , IEEE Transactions on Computers, IEEE, USA, vol. 55, No. 12, Dec. 1, 2008 (Dec. 1, 2008), pp. 1614-1623, XP011227508. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/581,167 mailed Aug. 3, 2018, 17 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/581,167 mailed Jul. 22, 2019, 21 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/789,565 mailed Nov. 16, 2017, 16 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/983,080 mailed Aug. 2, 2021, 28 pages. [cited by applicant]
Notice of Allowance for TW Application No. 108117181 3 pages, Oct. 22, 2019. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/789,565 mailed Oct. 1, 2018, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/983,080 mailed Dec. 29, 2021, 8 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/581,167, 10 pages, filed Mar. 30, 2020. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/197,821, filed Jul. 22, 2020, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/446,265 mailed May 20, 2021, 8 pages. [cited by applicant]
Notice of Publication for CN Application No. 202010848468.1, Publication No. 112330523, mailed Feb. 18, 2021, 102 pages. [cited by applicant]
Notice of Publication of CN Application No. 202110725327.5, mailed on Oct. 19, 2021, 103 pages. [cited by applicant]
Notification of CN Office Action for CN 201910429570.5, Mar. 12, 2020, 6 pages. [cited by applicant]
Notification of CN Publication for Application No 201810392234.3 (Pub No. 108805791), Nov. 13, 2018, 104 pages. [cited by applicant]
Notification of Publication for CN 201910813309.5, Feb. 7, 2020, 5 pages. [cited by applicant]
Notification of TW Publication for Application No. 107105950 (Pub No. 201842478), Dec. 1, 2018, 3 pages. [cited by applicant]
Notification to Grant Patent Right for CN Application No., 201910429570,5 4 pages, Jul. 3, 2020. [cited by applicant]
Office Action for TW Application No., TW107105950, 9 pages, Feb. 23, 2022. [cited by applicant]
Office Action for TW Application No. 108117181 8 pages, Jul. 9, 2019. [cited by applicant]
Office Action for U.S. Appl. No. 16/1446,265, filed Nov. 13, 2020, 19 pages. [cited by applicant]
Office Action received for U.S. Appl. No. 16/197,821, filed Apr. 13, 2020, 15 pages. [cited by applicant]
Pattnaik et al., “Scheduling Techniques for GPU Architectures with Processing-In-Memory Capabilities”, Sep. 11, 2016, pp. 31-44, XP058278466. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Venkataramani Swagath et al: “Quality programmable vector processors for approximate computing”, 2013 46th Annual IEEE/ACM International Symposium on Microarchitecture (Micro), ACM, Dec. 7, 2013 (Dec. 7, 2013), pp. 1-12… [cited by applicant]
Notice of Allowance for TW Patent Application No. 109145734, Sep. 7, 2022, 2 pages. [cited by applicant]
Search Report for TW Application No. 111122428, Oct. 28, 2022, 1 page. [cited by applicant]
Publication Notice for CN 202210661460.3, Sep. 27, 2022, 4 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/978,573 mailed Jan. 26, 2023, 17 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/960,611, mailed Oct. 17, 2023, 28 pages. [cited by applicant]
Notification of Publication for CN202211546793, Jul. 14, 2023, 1 page. [cited by applicant]
Issued patent notification for TW Application No. 111139972, Pat. No. I819861, mailed Oct. 27, 2023, 1 page. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/978,573 mailed Dec. 13, 2023, 20 pages. [cited by applicant]
Office Action for U.S. Appl. No. 17/960,611, mailed Feb. 2, 2024, 31 pages. [cited by applicant]
Office Action for CN201810392234.3, pp. 1-8, Feb. 5, 2024 (no translation). [cited by applicant]
Notification of Publication for TW Application No. 112136546 2 pages. [cited by applicant]
Issue Notification for CN Application No. CN201910813309.5, CN Patent No. ZL201910813309.5, mailed Jun. 19, 2023, 7 pages. [cited by applicant]
Decision to Grant for EP 22197260.7, Nov. 6, 2024, 5 pages. [cited by applicant]
TW Office Action for TW Application No. 112136546 mailed Jan. 22, 2025, 10 pages. [cited by applicant]
Intention to Grant for EP Application No. 20205451.6, Jan. 29, 2025, 5 pages. [cited by applicant]
Notification to Grant Patent Right for CN Application No. 202010848468.1 Mar. 4, 2025, 6 pages. [cited by applicant]
Decision to Grant EP Application No. 22197260.7 Mar. 13, 2025, 2 pages. [cited by applicant]