IP Library › Granted Patent US 8,935,475
Granted Patent B2
US 8,935,475 · App. 13/436,767 · Granted Jan 13, 2015

Cache management for memory operations

Inventors: Anthony Asaro (Toronto, CA); Kevin Normoyle (Los Gatos, CA); Mark Hummel (Franklin, MA); Norman Rubin (Cambridge, MA); Mark Fowler (Hopkinton, MA)
Assignees: ATI Technologies ULC; Advanced Micro Devices, Inc.
G06F12/0891G06F12/0837G06F12/128G06F12/123G06F12/0815G06F12/126
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,935,475
App. No.
13/436,767
Granted
Jan 13, 2015
Kind
B2
Abstract

Embodiments of the present invention provides for the execution of threads and/or workitems on multiple processors of a heterogeneous computing system in a manner that they can share data correctly and efficiently. Disclosed method, system, and article of manufacture embodiments include, responsive to an instruction from a sequence of instructions of a work-item, determining an ordering of visibility to other work-items of one or more other data items in relation to a particular data item, and performing at least one cache operation upon at least one of the particular data item or the other data items present in any one or more cache memories in accordance with the determined ordering. The semantics of the instruction includes a memory operation upon the particular data item.

Claims (52)

1. A method, comprising:

responsive to an instruction from a sequence of instructions of a work-item performed by a processor in a heterogeneous computing system, determining an ordering of visibility to other work-items of one or more other data items in relation to a particular data item, wherein semantics of the instruction includes a memory operation upon the particular data item; and

performing at least one cache operation upon at least one of the particular data item or the other data items present in any one or more cache memories in accordance with the determined ordering,

wherein the heterogeneous computing system includes one or more central processing units (CPUs) and one or more Advanced Processing Devices (APDs).

2. The method of claim 1 , wherein the cache operation includes at least one of a cache flush operation or a cache invalidate operation.

3. The method of claim 1 , wherein the determining an ordering includes: determining the ordering in accordance with a set of visibility rules.

4. The method of claim 3 , wherein the determining the ordering in accordance with a set of visibility rules comprises:

identifying a relative ordering of the instruction and respective instructions corresponding to the other data items, wherein the relative ordering is based, at least in part, upon positions of the instruction and the respective instructions in the sequence of instructions.

5. The method of claim 4 , wherein the relative ordering is further based upon memory addresses associated with the particular data item and the one or more other data items.

6. The method of claim 4 , wherein the relative ordering is further based upon whether there is a synchronization operation between the particular data item and the one or more other data items.

7. The method of claim 1 , wherein the performing at least one cache operation includes:

identifying one or more caches having a subset of the other data items, wherein the subset includes data items sequenced before the particular data item;

writing the subset to a common memory from the identified one or more caches; and

writing the particular data item in accordance with the instruction to the common memory, wherein the memory operation is a store operation, and wherein the writing of the particular data item is executed after the writing of the subset.

8. The method of claim 7 , wherein the performing at least one cache operation further includes:

invalidating entries in respective ones of the one or more caches, wherein the invalidated entries correspond to the particular data item.

9. The method of claim 1 , wherein the performing at least one cache operation includes:

identifying one or more caches having a subset of the other data items, wherein the subset includes data items sequenced before the particular data item;

writing the subset to a common memory from the identified one or more caches; and

reading the particular data item in accordance with the instruction, wherein the memory operation is a load operation, and wherein the reading of the particular data item is executed after the writing of the subset.

10. The method of claim 1 , wherein the performing at least one cache operation includes:

selectively flushing data items from one or more caches, in accordance with the determined ordering.

11. The method of claim 1 , wherein the performing at least one cache operation includes:

selectively invalidating data items from one or more caches, in accordance with the determined ordering.

12. A system comprising:

a central processing unit (CPU);

an advanced processing device (APD);

a common memory accessible to the CPU and the APD;

one or more cache memories, wherein each cache memory is associated with the CPU or the APD;

a memory order determiner configured to execute on one or more of the CPU or the APD, and further configured to:

responsive to an instruction from a sequence of instructions of a work-item, determine an ordering of visibility to other work-items of one or more other data items in relation to a particular data item, wherein semantics of the instruction includes a memory operation upon the particular data item; and

a cache updater configured to:

perform at least one cache operation upon at least one of the particular data item or the other data items present in any one or more cache memories in accordance with the determined ordering.

13. The system of claim 12 , wherein the memory order determiner is further configured to:

determining the ordering in accordance with a set of visibility rules.

14. The system of claim 13 , wherein the memory order determiner is further configured to:

identify a relative ordering of the instruction and respective instructions corresponding to the other data items, wherein the relative ordering is based, at least in part, upon positions of the instruction and the respective instructions in the sequence of instructions.

15. An article of manufacture comprising a non-transitory computer readable storage medium having instructions encoded thereon that, in response to execution by a computing device in a heterogeneous computing system, cause the computing device to perform operations comprising:

responsive to an instruction from a sequence of instructions of a work-item, determining an ordering of visibility to other work-items of one or more other data items in relation to a particular data item, wherein semantics of the instruction includes a memory operation upon the particular data item; and

performing at least one cache operation upon at least one of the particular data item or the other data items present in any one or more cache memories in accordance with the determined ordering,

wherein the heterogeneous computing system includes one or more central processing units (CPUs) and one or more Advanced Processing Devices (APDs).

16. The article of manufacture of claim 15 , wherein the determining an ordering includes:

determining the ordering in accordance with a set of visibility rules.

17. The article of manufacture of claim 16 , wherein the determining the ordering in accordance with a set of visibility rules comprises:

identifying a relative ordering of the instruction and respective instructions corresponding to the other data items, wherein the relative ordering is based, at least in part, upon positions of the instruction and the respective instructions in the sequence of instructions.

18. An apparatus for sharing data between work-items the apparatus includes one or more central processing units (CPUs) and one or more Advanced Processing Devices (APDs) being configured to:

responsive to an instruction from a sequence of instructions of a work-item, determine an ordering of visibility to other work-items of one or more other data items in relation to a particular data item, wherein semantics of the instruction includes a memory operation upon the particular data item; and

perform at least one cache operation upon at least one of the particular data item or the other data items present in any one or more cache memories in accordance with the determined ordering.

19. The apparatus of claim 18 , further configured to:

determine the ordering in accordance with a set of visibility rules.

20. The apparatus of claim 19 , further configured to:

identify a relative ordering of the instruction and respective instructions corresponding to the other data items, wherein the relative ordering is based, at least in part, upon positions of the instruction and the respective instructions in the sequence of instructions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2012
From: ASARO, ANTHONY; NORMOYLE, KEVIN; HUMMEL, MARK; RUBIN, NORMAN; FOWLER, MARK
To: ATI TECHNOLOGIES ULC; ADVANCED MICRO DEVICES, INC.
Reel/Frame 028377/0144 →
Continuity (1)
Related Publication 20130262775A1 · Oct 3, 2013