IP Library › Granted Patent US 9,477,526
Granted Patent B2
US 9,477,526 · App. 14/147,395 · Granted Oct 25, 2016

Cache utilization and eviction based on allocated priority tokens

Inventors: Daniel Robert Johnson (Austin, TX); Minsoo Rhu (Austin, TX); James M. O'Connor (Austin, TX); Stephen William Keckler (Austin, TX)
Assignee: NVIDIA Corporation
G06F9/5027G06F9/5016G06F9/48G06F2209/5021Y02B60/142
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,477,526
App. No.
14/147,395
Granted
Oct 25, 2016
Kind
B2
Abstract

A system, method, and computer program product are provided for providing prioritized access for multithreaded processing. The method includes the steps of allocating threads to process a workload and assigning a set of priority tokens to at least a portion of the threads. Access to a resource, by each one of the threads, is based on the priority token assigned to the thread and the threads are executed by a multithreaded processor to process the workload.

Claims (44)

1. A method comprising:

allocating threads to process a workload;

assigning a set of reallocatable priority tokens with different priorities to at least a portion of the threads, wherein access to a resource, by each one of the threads, is based on the priority token assigned to the thread; and

executing, by a multithreaded processor, the threads to process the workload, the executing comprising:

determining that no cache entries are available to store data to complete a memory operation associated with a thread;

determining that a priority token assigned to the thread is lower priority compared with priority tokens associated with the cache entries; and

transmitting the memory operation to a memory system to complete the memory operation for the thread.

2. The method of claim 1 , wherein the resource is at least one of a storage resource or a communication resource.

3. The method of claim 1 , wherein the resource is a cache memory.

4. The method of claim 3 , wherein an eviction policy for the cache memory is applied based on the priority token.

5. The method of claim 3 , wherein an allocation policy for the cache memory is applied based on the priority token.

6. The method of claim 1 , further comprising determining a maximum number of priority tokens in the set.

7. The method of claim 6 , further comprising increasing or decreasing the maximum number of priority tokens in the set.

8. The method of claim 1 , further comprising acquiring, by a first thread, a priority token when a first instruction is reached during execution of a sequence of instructions.

9. The method of claim 1 , further comprising releasing, by a first thread, a priority token when a particular instruction is reached during execution of a sequence of instructions.

10. The method of claim 1 , further comprising releasing, by a first thread, a priority token when the first thread completes execution of a sequence of instructions for the workload and exits.

11. The method of claim 1 , further comprising storing the priority token assigned to a second thread when a first cache entry is allocated for storing data associated with the second thread.

12. The method of claim 1 , further comprising assigning a second set of second priority tokens to at least a second portion of the threads, wherein access to a second resource, by each one of the threads, is based on the second priority token assigned to the thread.

13. The method of claim 12 , wherein the resource is a first cache memory and the second resource is a second cache memory.

14. The method of claim 1 , further comprising giving scheduling priority for execution to the portion of the threads to which the priority tokens are assigned.

15. A method comprising:

allocating threads to process a workload;

assigning a set of priority tokens with different priorities to at least a portion of the threads, wherein access to a resource, by each one of the threads, is based on the priority token assigned to the thread; and

executing, by a multithreaded processor, the threads to process the workload, the executing comprising:

determining that no cache entries are available to store data to complete a memory operation associated with a thread;

determining that a first cache entry is allocated to store data associated with a second thread that has not been assigned a priority token;

evicting data from the first cache entry when the thread has been assigned a priority token; and

transmitting the memory operation to a memory system to complete the memory operation when the thread has been assigned the priority token.

16. A system comprising:

a multithreaded processor that is configured to:

allocate threads to process a workload;

assign a set of reallocatable priority tokens with different priorities to at least a portion of the threads, wherein access to a resource, by each one of the threads, is based on the priority token assigned to the thread; and

execute, by the multithreaded processor, the threads to process the workload, comprising:

determining that no cache entries are available to store data to complete a memory operation associated with a thread;

determining that a priority token assigned to the thread is lower priority compared with priority tokens associated with the cache entries; and

transmitting the memory operation to a memory system to complete the memory operation for the thread.

17. The system of claim 16 , wherein the resource is a cache memory.

18. A non-transitory computer-readable storage medium storing instructions that, when executed by a multithreaded processor, causes the multithreaded processor to perform steps comprising:

allocating threads to process a workload;

assigning a set of reallocatable priority tokens with different priorities to at least a portion of the threads, wherein access to a resource, by each one of the threads, is based on the priority token assigned to the thread; and

executing the threads to process the workload, the executing comprising:

determining that no cache entries are available to store data to complete a memory operation associated with a thread;

determining that a priority token assigned to the thread is lower priority compared with priority tokens associated with the cache entries; and

transmitting the memory operation to a memory system to complete the memory operation for the thread.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 24, 2014
From: JOHNSON, DANIEL ROBERT; RHU, MINSOO; O'CONNOR, JAMES M.; KECKLER, STEPHEN WILLIAM
To: NVIDIA CORPORATION
Reel/Frame 034584/0584 →
Continuity (2)
Provisional Application 61873778 · Sep 4, 2013
Related Publication 20150067691A1 · Mar 5, 2015