IP Library Granted Patent US 12,229,867
Granted Patent B2
US 12,229,867 · App. 18/310,015 · Granted Feb 18, 2025

Graphics architecture including a neural network pipeline

Inventors: Hugues Labbe (Granite Bay, CA); Darrel Palke (Portland, OR); Sherine Abdelhak (Beaverton, OR); Jill Boyce (Portland, OR); Varghese George (Folsom, CA); Scott Janus (Loomis, CA); Adam Lake (Portland, OR); Zhijun Lei (Hillsboro, OR); Zhengmin Li (Hillsboro, OR); Mike MacPherson (Portland, OR); Carl Marshall (Portland, OR); Selvakumar Panneer (Portland, OR); Prasoonkumar Surti (Folsom, CA); Karthik Veeramani (Hillsboro, OR); Deepak Vembar (Portland, OR); Vallabhajosyula Srinivasa Somayazulu (Portland, OR)
Assignee: Intel Corporation
G06T15/005G06N3/08G06T1/20G06T1/60G06T15/40G06T17/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,229,867
App. No.
18/310,015
Granted
Feb 18, 2025
Kind
B2
Abstract

One embodiment provides a graphics processor comprising a block of execution resources, a cache memory, a cache memory prefetcher, and circuitry including a programmable neural network unit, the programmable neural network unit comprising a network hardware block including circuitry to perform neural network operations and activation operations for a layer of a neural network, the programmable neural network unit addressable by cores within the block of graphics cores and the neural network hardware block configured to perform operations associated with a neural network configured to determine a prefetch pattern for the cache memory prefetcher.

Claims (34)

1. A graphics processor comprising:

a block of execution resources;

a cache memory;

a cache memory prefetcher having an adjustable prefetch pattern that is adjustable to a learned prefetch pattern, the learned prefetch pattern learned by a neural network; and

circuitry including a programmable neural network unit, the programmable neural network unit comprising a network hardware block including circuitry to perform neural network operations and activation operations for a layer of the neural network, the programmable neural network unit addressable by cores within the block of graphics cores and the neural network hardware block configured to perform operations associated with a neural network configured to determine a prefetch pattern for the cache memory prefetcher, wherein the prefetch pattern determined for the cache memory prefetcher is the learned prefetch pattern, the cache memory prefetcher is to prefetch data according to the learned prefetch pattern, and the learned prefetch pattern is based at least in part on a memory access pattern associated with a workload executed via the block of execution resources, wherein the neural network hardware block is configured, via the neural network, to:

recognize the workload executed via the block of execution resources as the workload associated with the learned prefetch pattern based at least in part on the memory access pattern associated with the workload; and

configure the prefetch pattern for the cache memory prefetcher to the learned prefetch pattern for use with the workload, the learned prefetch pattern one of a plurality of learned prefetch patterns associated with a plurality of workloads.

2. The graphics processor of claim 1 , wherein the cache memory includes multiple levels of cache memory and the cache memory prefetcher is configured to prefetch data into the multiple levels of cache memory.

3. The graphics processor of claim 2 , wherein the cache memory is associated with a unified memory architecture including a local memory of the graphics processor and a system memory coupled with a host processor.

4. The graphics processor of claim 3 , the memory access pattern associated with the workload is to the local memory of the graphics processor.

5. The graphics processor of claim 3 , the memory access pattern associated with the workload is to the system memory coupled with the host processor.

6. The graphics processor as in claim 1 , wherein the neural network hardware block includes a source data buffer, a neural network operations and activation operations block, and an output data buffer.

7. The graphics processor as in claim 6 , wherein the neural network operations and activation operations block is programmably configurable.

8. The graphics processor as in claim 7 , wherein the programmable neural network unit includes a block programming unit to configure layer state information for the neural network hardware block, the layer state information associated with one or more layers of a neural network to be processed by the programmable neural network unit.

9. The graphics processor as in claim 7 , wherein the programmable neural network unit includes a weight cache to cache weights associated with one or more layers of a neural network to be processed by the programmable neural network unit.

10. The graphics processor as in claim 9 , wherein the programmable neural network unit includes multiple neural network hardware blocks, each of the multiple neural network hardware blocks associated with respective layers of the neural network.

11. A graphics processing system comprising:

a memory device; and

a graphics processor coupled with the memory device, the graphics processor including:

a block of execution resources;

a cache memory;

a cache memory prefetcher having an adjustable prefetch pattern that is adjustable to a learned prefetch pattern, the learned prefetch pattern learned by a neural network; and

circuitry including a programmable neural network unit, the programmable neural network unit comprising a network hardware block including circuitry to perform neural network operations and activation operations for a layer of the neural network, the programmable neural network unit addressable by cores within the block of graphics cores and the neural network hardware block configured to perform operations associated with a neural network configured to determine a prefetch pattern for the cache memory prefetcher, wherein the prefetch pattern determined for the cache memory prefetcher is the learned prefetch pattern, the cache memory prefetcher is to prefetch data according to the learned prefetch pattern, and the learned prefetch pattern is based at least in part on a memory access pattern associated with a workload executed via the block of execution resources, wherein the neural network hardware block is configured, via the neural network, to:

recognize the workload executed via the block of execution resources as the workload associated with the learned prefetch pattern based at least in part on the memory access pattern associated with the workload; and

configure the prefetch pattern for the cache memory prefetcher to the learned prefetch pattern for use with the workload, the learned prefetch pattern one of a plurality of learned prefetch patterns associated with a plurality of workloads.

12. The graphics processing system of claim 11 , wherein the cache memory includes multiple levels of cache memory and the cache memory prefetcher is configured to prefetch data into the multiple levels of cache memory.

13. The graphics processing system of claim 12 , wherein the cache memory is associated with a unified memory architecture including a local memory of the graphics processor and a system memory coupled with a host processor.

14. The graphics processing system of claim 13 , the memory access pattern associated with the workload is to the local memory of the graphics processor.

15. The graphics processing system of claim 13 , the memory access pattern associated with the workload is to the system memory coupled with the host processor.

16. The graphics processing system as in claim 11 , wherein the neural network hardware block includes a source data buffer, a neural network operations and activation operations block, and an output data buffer.

17. The graphics processing system as in claim 16 , wherein the neural network operations and activation operations block is programmably configurable.

18. The graphics processing system as in claim 17 , wherein the programmable neural network unit includes a block programming unit to configure layer state information for the neural network hardware block, the layer state information associated with one or more layers of a neural network to be processed by the programmable neural network unit.

19. The graphics processing system as in claim 17 , wherein the programmable neural network unit includes a weight cache to cache weights associated with one or more layers of a neural network to be processed by the programmable neural network unit.

20. The graphics processing system as in claim 19 , wherein the programmable neural network unit includes multiple neural network hardware blocks, each of the multiple neural network hardware blocks associated with respective layers of the neural network.

Continuity (6)
Continuation 17500631 · Oct 13, 2021
Continuation 16537140 · Aug 9, 2019
Provisional Application 62717603 · Aug 10, 2018
Provisional Application 62717685 · Aug 10, 2018
Provisional Application 62717593 · Aug 10, 2018
Related Publication 20230360307A1 · Nov 9, 2023
References Cited (24)
US 6459434B1 · Cohen et al. · 2002 [cited by applicant]
US 7873812B1 · Mimar · 2011 [cited by applicant]
US 10528864B2 · Dally et al. · 2020 [cited by applicant]
US 10860922B2 · Dally et al. · 2020 [cited by applicant]
US 10891538B2 · Dally et al. · 2021 [cited by applicant]
US 20140108740A1 · Rafacz · 2014 [cited by examiner]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20180039886A1 · Umuroglu · 2018 [cited by examiner]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180314932A1 · Schwartz et al. · 2018 [cited by applicant]
US 20190114533A1 · Ng · 2019 [cited by examiner]
US 20190114548A1 · Wu · 2019 [cited by applicant]
US 20190236829A1 · Hakura · 2019 [cited by applicant]
US 20190294959A1 · Vantrease · 2019 [cited by examiner]
Liao, “Machine Learning-Based Prefetch Optimization for Data Center Applications” (Year 2009), ACM 978-60558-744-8, SC09 Nov. 14-20, 2009, Portland, Oregon, USA, pp. 1-10; (Year 2009). [cited by examiner]
Goodfellow, et al. “Adaptive Computation and Machine Learning Series”, Book, Nov. 18, 2016, pp. 98-165, Chapter 5, The MIT Press, Cambridge, MA. [cited by applicant]
Ross, et al. “Intel Processor Graphics: Architecture & Programming”, Power Point Presentation, Aug. 2015, 78 pages, Intel Corporation, Santa Clara, CA. [cited by applicant]
Shane Cook, “CUDA Programming”, Book, 2013, pp. 37-52, Chapter 3, Elsevier Inc., Amsterdam Netherlands. [cited by applicant]
Nicholas Wilt, “The CUDA Handbook; A Comprehensive Guide to GPU Programming”, Book, Jun. 22, 2013, pp. 41-57, Addison-Wesley Professional, Boston, MA. [cited by applicant]
Stephen Junkins, “The Compute Architecture of Intel Processor Graphics Gen9”, paper, Aug. 14, 2015, 22 pages, Version 1.0, Intel Corporation, Santa Clara, CA. [cited by applicant]
Office Action for U.S. Appl. No. 16/537,140, filed Jul. 16, 2020, 22 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 16/537,140, filed Jun. 11, 2021, 15 pages. [cited by applicant]
Final Office Action for U.S. Appl. No. 16/537,140, filed Oct. 30, 2020, 22 pages. [cited by applicant]
Notice of Allowance for U.S. Appl. No. 17/500,631, filed Feb. 14, 2023, 9 pages. [cited by applicant]