IP Library › Granted Patent US 10,121,220
Granted Patent B2
US 10,121,220 · App. 14/698,024 · Granted Nov 6, 2018

System and method for creating aliased mappings to minimize impact of cache invalidation

Inventor: Jeffrey Bolz (Austin, TX)
Assignee: Nvidia Corporation
G06T1/20G06F9/3885G06F12/0833G06T15/005G06T15/04G09G5/395G06F2212/621
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,121,220
App. No.
14/698,024
Granted
Nov 6, 2018
Kind
B2
Abstract

A parallel processor and a method of reducing texture cache invalidation are disclosed. In one embodiment, the parallel processor includes a cache configured to receive lines of data; and a parallel execution unit associated with the cache and configured to execute parallel counterparts of an operation. The parallel counterparts, when executed, are configured to create, in the cache, corresponding aliases of a line of data pertaining to the operation such that the parallel counterparts are operable to invalidate only the corresponding aliases.

Claims (29)

1. A parallel processor, comprising:

a cache configured to receive data pertaining to an operation; and

a parallel execution unit associated with said cache and configured to execute parallel counterparts of said operation such that, when executed in parallel, said parallel counterparts is configured to create in said cache multiple aliases for said data, each of said multiple aliases corresponding to a different copy of said data,

wherein each of said parallel counterparts is operable to invalidate only one of said different copies of said data.

2. The parallel processor as recited in claim 1 wherein said multiple aliases are created in said cache by executing a texture load.

3. The parallel processor as recited in claim 1 wherein said parallel processor is a graphics processing unit.

4. The parallel processor as recited in claim 1 wherein said operation includes blending color values to accomplish transparency.

5. The parallel processor as recited in claim 1 wherein said parallel counterparts are configured to store said multiple aliases by executing a surface store.

6. The parallel processor as recited in claim 1 wherein said operation prevents said parallel counterparts from working on a same portion of said data.

7. The parallel processor as recited in claim 1 wherein said multiple aliases correspond to said parallel counterparts based on a number of said parallel counterparts using said multiple aliases.

8. A method of reducing an impact of cache invalidation, comprising:

executing, using a parallel processor, parallel counterparts of a surface operation in parallel, said operation pertaining to data in a cache;

causing, with said parallel counterparts, multiple aliases for said data to be loaded into said cache, each of said multiple aliases corresponding to a different copy of said data; and

allowing each of said parallel counterparts of said operation to invalidate only one of said different copies of said data.

9. The method as recited in claim 8 wherein said cache is associated with a parallel processor and said multiple aliases are loaded by executing a texture load.

10. The method as recited in claim 8 wherein said parallel processor is a graphics processing unit.

11. The method as recited in claim 8 wherein said cache is a texture cache.

12. The method as recited in claim 11 wherein said multiple aliases are stored by executing a surface store.

13. The method as recited in claim 8 wherein said operation prevents said parallel counterparts from working on a same portion of said data.

14. The method as recited in claim 8 wherein said multiple aliases correspond to said parallel counterparts based on a number of said parallel counterparts using said multiple aliases.

15. A non-transitory computer readable medium storing a surface operation that is executable by a processor, said surface operation comprising:

executing parallel counterparts of said surface operation in parallel, said operation pertaining to data in a cache;

causing, with said parallel counterparts, multiple aliases for said data to be loaded into said cache, each of said multiple aliases corresponding to a different copy of said data; and

allowing each of said parallel counterparts of said operation to invalidate only one of said different copies of said data.

16. The non-transitory computer readable medium as recited in claim 15 wherein said multiple aliases are loaded by executing a texture load.

17. The non-transitory computer readable medium as recited in claim 15 wherein said cache is a texture cache.

18. The non-transitory computer readable medium as recited in claim 15 wherein said multiple aliases are stored by executing a surface store.

19. The non-transitory computer readable medium as recited in claim 15 wherein said operation prevents said parallel counterparts from working on a same portion of said data.

20. The non-transitory computer readable medium as recited in claim 15 wherein said multiple aliases correspond to said parallel counterparts based on a number of said parallel counterparts using said multiple aliases.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2015
From: BOLZ, JEFFREY
To: NVIDIA CORPORATION
Reel/Frame 035513/0449 →
Continuity (1)
Related Publication 20160321773A1 · Nov 3, 2016