IP Library › Granted Patent US 12,198,221
Granted Patent B2
US 12,198,221 · App. 18/436,494 · Granted Jan 14, 2025

Compute optimization mechanism for deep neural networks

Inventors: Prasoonkumar Surti (Folsom, CA); Narayan Srinivasa (Portland, OR); Feng Chen (Shanghai, CN); Joydeep Ray (Folsom, CA); Ben J. Ashbaugh (Folsom, CA); Nicolas C. Galoppo Von Borries (Portland, OR); Eriko Nurvitadhi (Hillsboro, OR); Balaji Vembu (Folsom, CA); Tsung-Han Lin (Campbell, CA); Kamal Sinha (Rancho Cordova, CA); Rajkishore Barik (Santa Clara, CA); Sara S. Baghsorkhi (San Jose, CA); Justin E. Gottschlich (Santa Clara, CA); Altug Koker (El Dorado Hills, CA); Nadathur Rajagopalan Satish (Santa Clara, CA); Farshad Akhbari (Chandler, AZ); Dukhwan Kim (San Jose, CA); Wenyin Fu (Folsom, CA); Travis T. Schluessler (Hillsboro, OR); Josh B. Mastronarde (Sacramento, CA); Linda L Hurd (Cool, CA); John H. Feit (Folsom, CA); Jeffery S. Boles (Folsom, CA); Adam T. Lake (Portland, OR); Karthik Vaidyanathan (Berkeley, CA); Devan Burke (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Abhishek R. Appu (El Dorado Hills, CA)
Assignee: Intel Corporation
G06T1/20G06F9/45533G06F9/5061G06F9/5094G06N3/044G06N3/045G06N3/063G06N3/084G06F8/41G06F2009/45583
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,198,221
App. No.
18/436,494
Granted
Jan 14, 2025
Kind
B2
Abstract

Embodiments provide mechanisms to facilitate compute operations for deep neural networks. One embodiment comprises a graphics processing unit comprising one or more multiprocessors, at least one of the one or more multiprocessors including a register file to store a plurality of different types of operands and a plurality of processing cores. The plurality of processing cores includes a first set of processing cores of a first type and a second set of processing cores of a second type. The first set of processing cores are associated with a first memory channel and the second set of processing cores are associated with a second memory channel.

Claims (36)

1. A graphics processing apparatus comprising:

a system interconnect to a host processor; and

a plurality of graphics processing clusters, each of the plurality of graphics processing clusters including a plurality of multiprocessors coupled via a crossbar interconnect, the crossbar interconnect to enable transfer of data from a first multiprocessor of the plurality of multiprocessors to a second multiprocessor of the plurality of multiprocessors, a graphics processing cluster of the plurality of graphics processing clusters including:

a register file to store a plurality of different types of operands;

a first plurality of processing resources of a first type configurable to process a first number of threads having operands stored in a first number of registers of the register file; and

a second plurality of processing resources of a second type configurable to process a second number of threads having operands stored in a second number of registers of the register file, the first number of threads greater than the second number of threads and the second number of registers greater than the first number of registers.

2. The graphics processing apparatus of claim 1 , wherein the first plurality of processing resources is configured to perform multi-dimensional matrix operations on the operands stored in the first number of registers.

3. The graphics processing apparatus of claim 1 , wherein the second plurality of processing resources is configured to perform graphics operations on the operands stored in the second number of registers.

4. The graphics processing apparatus of claim 1 , further comprising compute circuitry to select processing resources from the first plurality of processing resources and the second plurality of processing resources to execute a workload.

5. The graphics processing apparatus of claim 4 , wherein the compute circuitry is to select processing resources of the first type to process a first type of application workload and to select the processing resources of the second type to process a second type of application workload.

6. The graphics processing apparatus of claim 1 , further comprising a memory device coupled with the plurality of graphics processing clusters.

7. The graphics processing apparatus of claim 6 , wherein the memory device includes a high bandwidth memory (HBM) including a plurality of memory channels.

8. The graphics processing apparatus of claim 7 , wherein a first memory channel of the HBM is configured to couple with one or more processing resources of the first plurality of processing resources of a first type and a second memory channel of the HBM is configured to couple with one or more processing resources of the second plurality of processing resources of a second type, the second memory channel distinct from the first memory channel.

9. The graphics processing apparatus of claim 1 , wherein the register file is configured to perform matrix-vector transformations.

10. The graphics processing apparatus of claim 1 , further comprising a shared local memory (SLM) configured to perform matrix-vector transformations.

11. A method comprising:

storing operands for a plurality of different types of operations to a register file of a graphics processor including a plurality of graphics processing clusters, each of the plurality of graphics processing clusters including a plurality of multiprocessors coupled via a crossbar interconnect, the crossbar interconnect to enable transfer of data from a first multiprocessor of the plurality of multiprocessors to a second multiprocessor of the plurality of multiprocessors;

processing a first number of threads having operands stored in a first number of registers of the register file via a first plurality of processing resources of a first type; and

processing second number of threads having operands stored in a second number of registers of the register file via a second plurality of processing resources of a second type, the first number of threads greater than the second number of threads and the second number of registers greater than the first number of registers.

12. The method of claim 11 , comprising:

performing multi-dimensional matrix operations on the operands stored in the first number of registers via the first plurality of processing resources; and

performing graphics operations on the operands stored in the second number of registers via the second plurality of processing resources.

13. The method of claim 11 , further comprising selecting processing resources to execute a workload via compute circuitry of the graphics processor, including selecting processing resources of the first type to process a first type of application workload and selecting the processing resources of the second type to process a second type of application workload.

14. The method of claim 11 , further comprising performing matrix-vector transformations via the register file of the graphics processor.

15. The method of claim 11 , further comprising performing matrix-vector transformations via shared local memory (SLM) of the graphics processor.

16. A graphics processing system comprising:

a system interconnect to a host processor;

a memory device coupled with the system interconnect; and

a graphics processor coupled with the system interconnect and the memory device, the graphics processor including a plurality of graphics processing clusters, each of the plurality of graphics processing clusters including a plurality of multiprocessors coupled via a crossbar interconnect, the crossbar interconnect to enable transfer of data from a first multiprocessor of the plurality of multiprocessors to a second multiprocessor of the plurality of multiprocessors, a graphics processing cluster of the plurality of graphics processing clusters including:

a register file to store a plurality of different types of operands;

a first plurality of processing resources of a first type configurable to process a first number of threads having operands stored in a first number of registers of the register file; and

a second plurality of processing resources of a second type configurable to process a second number of threads having operands stored in a second number of registers of the register file, the first number of threads greater than the second number of threads and the second number of registers from greater than the first number of registers.

17. The graphics processing system of claim 16 , wherein the first plurality of processing resources is configured to perform multi-dimensional matrix operations on the operands stored in the first number of registers and the second plurality of processing resources is configured to perform graphics operations on the operands stored in the second number of registers.

18. The graphics processing system of claim 16 , further comprising compute circuitry to select processing resources to execute a workload, the compute circuitry to select processing resources of the first type to process a first type of application workload and to select the processing resources of the second type to process a second type of application workload.

19. The graphics processing system of claim 16 , wherein the memory device includes a high bandwidth memory (HBM) including a plurality of memory channels, a first memory channel of the HBM is configured to couple with one or more processing resources of the first plurality of processing resources of a first type, and a second memory channel of the HBM is configured to couple with one or more processing resources of the second plurality of processing resources of a second type, the second memory channel distinct from the first memory channel.

20. The graphics processing system of claim 16 , wherein the register file is configured to perform matrix-vector transformations and further comprising a shared local memory (SLM) configured to perform matrix-vector transformations.

Continuity (7)
Continuation 18168207 · Feb 13, 2023
Continuation 17741934 · May 11, 2022
Continuation 17385693 · Jul 26, 2021
Continuation 17145885 · Jan 11, 2021
Continuation 15819093 · Nov 21, 2017
Continuation 15494886 · Apr 24, 2017
Related Publication 20240257294A1 · Aug 1, 2024
References Cited (106)
US 6624818B1 · Mantor et al. · 2003 [cited by applicant]
US 7728841B1 · Nordquist et al. · 2010 [cited by applicant]
US 7747070B2 · Puri · 2010 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 9721190B2 · Vijayanarasimhan et al. · 2017 [cited by applicant]
US 9958932B2 · Williamson et al. · 2018 [cited by applicant]
US 10417731B2 · Surti et al. · 2019 [cited by applicant]
US 10424069B2 · Sun · 2019 [cited by examiner]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10776684B1 · Agarwal et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20060012603A1 · Lindholm et al. · 2006 [cited by applicant]
US 20070273698A1 · Du et al. · 2007 [cited by applicant]
US 20070283356A1 · Du et al. · 2007 [cited by applicant]
US 20080074433A1 · Jiao et al. · 2008 [cited by applicant]
US 20080235316A1 · Du et al. · 2008 [cited by applicant]
US 20080303833A1 · Swift et al. · 2008 [cited by applicant]
US 20090073168A1 · Jiao et al. · 2009 [cited by applicant]
US 20090150654A1 · Oberman et al. · 2009 [cited by applicant]
US 20090160863A1 · Frank · 2009 [cited by applicant]
US 20090265528A1 · Du et al. · 2009 [cited by applicant]
US 20100259536A1 · Toksvig et al. · 2010 [cited by applicant]
US 20110072243A1 · Qiu et al. · 2011 [cited by applicant]
US 20110296428A1 · Rawson, III et al. · 2011 [cited by applicant]
US 20120133654A1 · Redgrave et al. · 2012 [cited by applicant]
US 20120249560A1 · Dickenson · 2012 [cited by applicant]
US 20120317558A1 · Agarwal et al. · 2012 [cited by applicant]
US 20130162661A1 · Bolz et al. · 2013 [cited by applicant]
US 20130226535A1 · Tuan · 2013 [cited by applicant]
US 20130235053A1 · Bourd · 2013 [cited by applicant]
US 20140026146A1 · Jahagirdar et al. · 2014 [cited by applicant]
US 20140071128A1 · Everitt et al. · 2014 [cited by applicant]
US 20140189704A1 · Narvaez et al. · 2014 [cited by applicant]
US 20140281380A1 · Sodhi et al. · 2014 [cited by applicant]
US 20140365548A1 · Mortensen · 2014 [cited by applicant]
US 20140375658A1 · Lichmanov et al. · 2014 [cited by applicant]
US 20150177821A1 · Senthinathan et al. · 2015 [cited by applicant]
US 20150179142A1 · Lehtinen et al. · 2015 [cited by applicant]
US 20150205757A1 · Dally · 2015 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160179574A1 · Merrill, III · 2016 [cited by applicant]
US 20170061569A1 · Sathe · 2017 [cited by applicant]
US 20170097824A1 · Elmer et al. · 2017 [cited by applicant]
US 20170308789A1 · Langford et al. · 2017 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180082399A1 · Martin et al. · 2018 [cited by applicant]
US 20180247190A1 · Chung · 2018 [cited by examiner]
US 20180260220A1 · Lacy et al. · 2018 [cited by applicant]
CN 111539518A · 2020 [cited by applicant]
CN 113705789A · 2021 [cited by applicant]
EP 3396546A1 · 2018 [cited by applicant]
EP 3654185A1 · 2020 [cited by applicant]
EP 3964958A1 · 2022 [cited by applicant]
TW 201214287A · 2012 [cited by applicant]
TW 201629814A · 2016 [cited by applicant]
TW 201839607A · 2018 [cited by applicant]
Communication and European Search Report for EP Application No. 19218493.5, mailed Apr. 14, 2020, 4 pages. [cited by applicant]
Communication Pursuant to Article 94(3) EPC for EP Application No. 19218493.5, mailed Apr. 28, 2020, 6 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC for European Patent Application No. 21205193.2 mailed Feb. 4, 2022, 11 pages. [cited by applicant]
Communication pursuant to Article 94(3) EPC for European Patent Application No. 21205193.2 mailed Aug. 26, 2022, 9 pages. [cited by applicant]
Cormnunication Pursuant to Article 94(3) for EP Application No. 18163807.3, Mar. 30, 2021, 10 pages. [cited by applicant]
Extended European Search Report from EP Application No. 18163807.3, 10 pages, Sep. 26, 2018. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/494,886 mailed Nov. 29, 2018, 20 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/698,217 mailed Nov. 29, 2018, 18 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/819,093 mailed Apr. 5, 2019, 15 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 15/819,093 mailed May 23, 2018, 15 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/531,763 mailed Oct. 28, 2020, 15 pages. [cited by applicant]
Goodfellow et al., “Adaptive Computation and Machine Learning Series,” Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA, USA. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook: A Comprehensive Guide to GPU Programming,” Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA, USA. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/494,886 mailed Apr. 19, 2018, 17 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/698,217 mailed Apr. 19, 2018, 12 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/819,093 mailed Dec. 17, 2019, 17 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/819,093 mailed Feb. 2, 2018, 12 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/819,093 mailed Nov. 19, 2018, 15 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 15/819,093 mailed Oct. 3, 2019, 16 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/531,763 mailed Apr. 16, 2020, 14 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 16/531,763 mailed May 27, 2021, 11 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/145,885 mailed Sep. 16, 2021, 20 pages. [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/385,693 mailed Nov. 12, 2021, 42 pages. [cited by applicant]
Notice of Allowance for TW Application No. 107105953 mailed Dec. 27, 2021, 3 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/494,886 mailed May 8, 2019, 8 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/698,217 mailed May 7, 2019, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/819,093 mailed Jun. 11, 2020, 13 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 15/819,093 mailed Sep. 30, 2020, 13 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/531,763 mailed Sep. 15, 2021, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/145,885 mailed Jan. 26, 2022, 11 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/385,693 mailed Jan. 26, 2022, 9 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/529,862 mailed Sep. 22, 2022, 18 pages. [cited by applicant]
Notification of CN Publication for Application No. 201810368545.6, Nov. 2, 2018, 5 pages. [cited by applicant]
Notification of CN Publication for CN Application No. 202111003293.5, 4 pages, Dec. 22, 2021. [cited by applicant]
Notification of Publication for TW Application No. 110131258, 3 pages, Dec. 7, 2021. [cited by applicant]
Notification of TW OA and Search Report for TW Application No. 107105953), 7 pages, Sep. 29, 2021. [cited by applicant]
Office Action & Search Report for Taiwan Application No. 110131258 mailed Jun. 29, 2022, 14 pages. [cited by applicant]
Office Action and Search Report for Taiwan Patent Application No. 110131258 mailed Jun. 29, 2022, 14 pages. [cited by applicant]
Office Action for TW Application No. 107105953, mailed Oct. 22, 2021, 10 pages. [cited by applicant]
Ross et al., “Intel Processor Graphics: Architecture & Programming,” Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA, USA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Summons to Attend Oral Proceedings for EP Application No. 19218493.5 mailed Feb. 9, 2021, 11 pages. [cited by applicant]
Summons to Attend Oral Proceedings for EP Application No. 19218493.5, mailed Jul. 30, 2020, 8 pages. [cited by applicant]
Notice of Publication for CN202311809249.2, mailed Apr. 10, 2024, 4 pages. [cited by applicant]
NVIDIA Tesla P100 Pascal Whitepaper, 2016, 45 pages. [cited by applicant]
Summons to Attend Oral Proceedings for European Patent Application No. 18 163 807.3 mailed Oct. 6, 2022, 9 pages. [cited by applicant]
Wikipedia, “Floating Point Arithmetic” 28 pages, downloaded on Dec. 7, 2017. [cited by applicant]
Notice of allowance for CN202111003293.5, mailed Oct. 23, 2023, 7 pages. [cited by applicant]
Cited By (1)
US 12,399,743