IP Library Granted Patent US 9,952,977
Granted Patent B2
US 9,952,977 · App. 12/890,476 · Granted Apr 24, 2018

Cache operations and policies for a multi-threaded client

Inventors: Steven James Heinrich (Madison, AL); Alexander L. Minkin (Los Altos, CA); Brett W. Coon (San Jose, CA); Rajeshwaran Selvanesan (Milpitas, CA); Robert Steven Glanville (Cupertino, CA); Charles McCarver (Madison, AL); Anjana Rajendran (San Jose, CA); Stewart Glenn Carlton (Madison, AL); John R. Nickolls (Los Altos, CA); Brian Fahs (Los Altos, CA)
Assignee: NVIDIA CORPORATION
G06F12/0842G06F12/0897
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,952,977
App. No.
12/890,476
Granted
Apr 24, 2018
Kind
B2
Abstract

A method for managing a parallel cache hierarchy in a processing unit. The method including receiving an instruction that includes a cache operations modifier that identifies a level of the parallel cache hierarchy in which to cache data associated with the instruction; and implementing a cache replacement policy based on the cache operations modifier.

Claims (32)

1. A method implemented on a computer for managing a parallel cache hierarchy in a processor, the method comprising:

receiving an instruction that includes a cache operations modifier that identifies a first level of the parallel cache hierarchy in which to cache data associated with the instruction;

selecting a first cache replacement policy from among a plurality of cache replacement policies, wherein the first cache replacement policy is selected based on whether the instruction comprises a load instruction or a store instruction; and

executing the instruction and performing at least one related cache operation within the first level of the parallel cache hierarchy according to the first cache replacement policy.

2. The method of claim 1 , wherein a first portion of the first cache memory in the parallel cache hierarchy is allocated for storing streaming data, and a second portion of the first cache memory is allocated for storing non-streaming data.

3. The method of claim 2 , wherein the instruction is associated with streaming data, and the streaming data is stored in the first portion of the first cache memory when at least one cache line in the first portion of the first cache memory is marked as invalid.

4. The method of claim 2 , wherein the instruction is associated with non-stream ing data.

5. The method of claim 4 , wherein the non-streaming data is stored in the second portion of the first cache memory when at least one cache line in the second portion of the first cache memory is marked as invalid.

6. The method of claim 4 , wherein the non-streaming data is stored in the first portion of the first cache memory when none of the cache lines in the second portion of the first cache memory is marked as invalid, and at least one cache line in the first portion of the first cache memory is marked as invalid.

7. The method of claim 4 , wherein the non-streaming data is stored in the first portion of the first cache memory when none of the cache lines in either the first portion of the first cache memory or the second portion of the first cache memory is marked as invalid.

8. The method of claim 1 , wherein cache tags associated with the first cache memory are either global or local, and the first cache replacement policy results in all valid cache lines having global cache tags being invalidated in one clock cycle and/or all valid cache lines having local cache tags being invalidated in one clock cycle.

9. The method of claim 1 , wherein the cache operations modifier is associated with a fully covered load instruction, and the first cache replacement policy results in a cache line being invalidated after a last use of the fully covered load instruction.

10. The method of claim 1 , wherein the processor comprises a graphics processing unit.

11. The method of claim 1 , wherein the first cache replacement policy specifies a first policy with respect to a first cache memory and a second policy with respect to a second cache memory.

12. The method of claim 11 , wherein the first policy is different from the second policy.

13. The method of claim 11 , wherein the first cache memory comprises a level one cache memory and the second cache memory comprises a level two cache memory.

14. The method of claim 1 , wherein the instruction comprises a load instruction and wherein selecting the first cache replacement policy comprises selecting from among an evict normal, an evict first, a non-cached, a fetch volatile, and a last use cache replacement policy.

15. The method of claim 1 , wherein the instruction comprises a store instruction and wherein selecting the first cache replacement policy comprises selecting from among an evict normal, an evict first, a non-cached, and a write-through cache replacement policy.

16. A system for managing a parallel cache hierarchy, the system comprising:

a processor, configured to:

receive an instruction that includes a cache operations modifier that identifies a first level of the parallel cache hierarchy in which to cache data associated with the instruction,

select a first cache replacement policy from among a plurality of cache replacement policies, wherein the first cache replacement policy is selected based on the cache operations modifier and based on whether the instruction comprises a load instruction or a store instruction; and

execute the instruction and performing at least one related cache operation within the first level of the parallel cache hierarchy according to the first cache replacement policy.

17. The system of claim 16 , wherein a first portion of the first cache memory in the parallel cache hierarchy is allocated for storing streaming data, and a second portion of the first cache memory is allocated for storing non-streaming data.

18. The system of claim 17 , wherein the instruction is associated with streaming data, and the streaming data is stored in the first portion of the first cache memory when at least one cache line in the first portion of the first cache memory is marked as invalid.

19. The system of claim 17 , wherein the instruction is associated with non-stream ing data.

20. The system of claim 19 , wherein the non-streaming data is stored in the second portion of the first cache memory when at least one cache line in the second portion of the first cache memory is marked as invalid.

21. The system of claim 19 , wherein the non-streaming data is stored in the first portion of the first cache memory when none of the cache lines in the second portion of the first cache memory is marked as invalid, and at least one cache line in the first portion of the first cache memory is marked as invalid.

22. The system of claim 19 , wherein the non-streaming data is stored in the first portion of the first cache memory when none of the cache lines in either the first portion of the first cache memory or the second portion of the first cache memory is marked as invalid.

23. The system of claim 16 , wherein cache tags associated with the first cache memory are either global or local, and the first cache replacement policy results in all valid cache lines having global cache tags being invalidated in one clock cycle and/or all valid cache lines having local cache tags being invalidated in one clock cycle.

24. The system of claim 16 , wherein the cache operations modifier is associated with a fully covered load instruction, and the first cache replacement policy results in a cache line being invalidated after a last use of the fully covered load instruction.

25. The system of claim 16 , wherein the processor comprises a graphics processing unit.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2010
From: HEINRICH, STEVEN JAMES; MINKIN, ALEXANDER L.; COON, BRETT W.; SELVANESAN, RAJESHWARAN; GLANVILLE, ROBERT STEVEN; MCCARVER, CHARLES; RAJENDRAN, ANJANA; CARLTON, STEWART GLENN; NICKOLLS, JOHN R.; FAHS, BRIAN
To: NVIDIA CORPORATION
Reel/Frame 025405/0593 →
Continuity (1)
Related Publication 20110078381A1 · Mar 31, 2011