IP Library Granted Patent US 11,620,256
Granted Patent B2
US 11,620,256 · App. 17/732,308 · Granted Apr 4, 2023

Systems and methods for improving cache efficiency and utilization

Inventors: Altug Koker (El Dorado Hills, CA); Joydeep Ray (Folsom, CA); Ben Ashbaugh (Folsom, CA); Jonathan Pearce (Hillsboro, OR); Abhishek Appu (El Dorado Hills, CA); Vasanth Ranganathan (El Dorado Hills, CA); Lakshminarayanan Striramassarma (Folsom, CA); Elmoustapha Ould-Ahmed-Vall (Chandler, AZ); Aravindh Anantaraman (Folsom, CA); Valentin Andrei (San Jose, CA); Nicolas Galoppo Von Borries (Portland, OR); Varghese George (Folsom, CA); Yoav Harel (Carmichael, CA); Arthur Hunter, Jr. (Cameron Park, CA); Brent Insko (Portland, OR); Scott Janus (Loomis, CA); Pattabhiraman K (Bangalore, IN); Mike Macpherson (Portland, OR); Subramaniam Maiyuran (Gold River, CA); Marian Alin Petre (San Mateo, CA); Murali Ramadoss (Folsom, CA); Shailesh Shah (Folsom, CA); Kamal Sinha (Folsom, CA); Prasoonkumar Surti (Folsom, CA); Vikranth Vemulapalli (Folsom, CA)
Assignee: Intel Corporation
G06F15/7839G06F7/5443G06F7/575G06F7/588G06F9/3001G06F9/3004G06F9/30014G06F9/30036G06F9/30043G06F9/30047G06F9/30065G06F9/30079G06F9/3887G06F9/5011G06F9/5077G06F12/0215G06F12/0238G06F12/0246G06F12/0607G06F12/0802G06F12/0804G06F12/0811G06F12/0862G06F12/0866G06F12/0871G06F12/0875G06F12/0882G06F12/0891G06F12/0893G06F12/0895G06F12/0897G06F12/1009G06F12/128G06F15/8046G06F17/16G06F17/18G06T1/20G06T1/60H03M7/46G06F9/3802G06F9/3818G06F9/3867G06F2212/1021G06F2212/1044G06F2212/302G06F2212/401G06F2212/455G06F2212/60G06N3/08G06T15/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,620,256
App. No.
17/732,308
Granted
Apr 4, 2023
Kind
B2
Abstract

Systems and methods for improving cache efficiency and utilization are disclosed. In one embodiment, a graphics processor includes processing resources to perform graphics operations and a cache controller of a cache coupled to the processing resources. The cache controller is configured to control cache priority by determining whether default settings or an instruction will control cache operations for the cache.

Claims (55)

1. A graphics processing unit (GPU) comprising:

a plurality of groups of cores, each group of cores including:

a plurality of cores of a first type; and

a plurality of cores of a second type, wherein the plurality of cores of the second type are tensor cores;

a plurality of combined level 1 (L1) cache and shared memory units, each corresponding to a different group of cores of the plurality of groups of cores;

a level 2 (L2) cache to be shared by the plurality of groups of cores;

a plurality of memory controllers to couple the GPU to a memory; and

cache controller circuitry associated with the L2 cache in response to a load instruction from a first core of the plurality of groups of cores, to:

select a cache control, based on the load instruction, out of a plurality of stored cache controls, wherein at least some of the cache controls have different cache eviction priorities; and

apply the selected cache control to data allocated into the L2 cache.

2. The GPU of claim 1 , wherein the plurality of cache controls are to be stored in a data structure.

3. The GPU of claim 1 , wherein the selected cache control has a streaming cache eviction priority.

4. The GPU of claim 1 , wherein the selected cache control is a streaming cache control.

5. The GPU of claim 1 , wherein the load instruction is to indicate the data is for a global address space.

6. The GPU of claim 1 , further comprising:

scheduler/dispatcher circuitry to schedule and dispatch graphics threads for execution on the plurality of groups of cores; and

a plurality of groups of texture units, each corresponding to a different group of cores of the plurality of groups of cores.

7. The GPU of claim 6 , further comprising input/output (I/O) circuitry to couple the GPU to one or more I/O devices.

8. The GPU of claim 1 , wherein each group of cores of the plurality of groups of cores includes a ray tracing core.

9. A method, performed by a graphics processing unit (GPU), the method comprising:

processing data with a plurality of groups of cores, including:

processing graphics data with a plurality of cores of a first type in each of the groups;

performing matrix operations with a plurality of cores of a second type in each of the groups, wherein the plurality of cores of the second type are tensor cores; and

storing data in a plurality of combined level 1 (L1) cache and shared memory units, each corresponding to a different group of cores of the plurality of groups of cores;

sharing a level 2 (L2) cache by the plurality of groups of cores;

accessing data from a memory by a plurality of memory controllers of the GPU; and

performing a load instruction received from a first core of the plurality of groups of cores, including:

selecting a cache control, based on the load instruction, out of a plurality of stored cache controls, wherein at least some of the cache controls have different cache eviction priorities; and

applying the selected cache control to data allocated into the L2 cache.

10. The method of claim 9 , wherein the plurality of cache controls are stored in a data structure.

11. The method of claim 9 , wherein selecting the cache control comprises selecting a cache control having a streaming cache eviction priority.

12. The method of claim 9 , wherein selecting the cache control comprises selecting a streaming cache control.

13. The method of claim 9 , further comprising scheduling and dispatching graphics threads for execution on the plurality of groups of cores.

14. The method of claim 9 , further comprising performing ray tracing with a ray tracing core in each of the groups of cores.

15. A system comprising:

a memory; and

a graphics processing unit (GPU) coupled with the memory, the GPU comprising:

a plurality of groups of cores, each group of cores including:

a plurality of cores of a first type; and

a plurality of cores of a second type, wherein the plurality of cores of the second type are tensor cores;

a plurality of combined level 1 (L1) cache and shared memory units, each corresponding to a different group of cores of the plurality of groups of cores;

a level 2 (L2) cache to be shared by the plurality of groups of cores;

a plurality of memory controllers to couple the GPU to a memory; and

cache controller circuitry associated with the L2 cache in response to a load instruction from a first core of the plurality of groups of cores, to:

select a cache control, based on the load instruction, out of a plurality of stored cache controls, wherein at least some of the cache controls have different cache eviction priorities; and

apply the selected cache control to data allocated into the L2 cache.

16. The system of claim 15 , further comprising a data storage device coupled with the GPU and the memory, and wherein the plurality of cache controls are to be stored in a data structure.

17. The system of claim 15 , further comprising a network controller coupled with the memory, and wherein the selected cache control has a streaming cache eviction priority.

18. The system of claim 15 , further comprising a touch sensor coupled with the GPU, and wherein the selected cache control is a streaming cache control.

19. The system of claim 15 , further comprising a data storage device coupled with the GPU and the memory, and wherein the load instruction is to indicate the data is for a global address space.

20. The system of claim 15 , further comprising a network controller coupled with the memory, and wherein the GPU further comprises:

scheduler/dispatcher circuitry to schedule and dispatch graphics threads for execution on the plurality of groups of cores; and

a plurality of groups of texture units, each corresponding to a different group of cores of the plurality of groups of cores.

21. The system of claim 20 , further comprising at least one I/O device, and wherein the GPU further comprises input/output (I/O) circuitry to couple the GPU to the at least one I/O device.

22. The system of claim 15 , wherein each group of cores of the plurality of groups of cores includes a ray tracing core.

Continuity (5)
Continuation 17428530
Provisional Application 62819337 · Mar 15, 2019
Provisional Application 62819435 · Mar 15, 2019
Provisional Application 62819361 · Mar 15, 2019
Related Publication 20220261347A1 · Aug 18, 2022
Cited By (16)
US 12,198,222 US 12,204,487 US 12,210,477 US 12,217,053 US 12,242,414 US 12,293,431 US 12,321,310 US 12,361,600 US 12,386,779 US 12,411,695 US 12,493,922 US 12,554,674 US 12,561,276 US 12,561,277 US 12,670,121 US 12,688,146