IP Library › Granted Patent US 11,068,410
Granted Patent B2
US 11,068,410 · App. 16/291,154 · Granted Jul 20, 2021

Multi-core computer systems with private/shared cache line indicators

Inventors: Alberto Ros (Cartagena, ES); Stefanos Kaxiras (Uppsala, SE)
Assignee: ETA SCALE AB
G06F12/1045G06F12/084G06F12/0808G06F12/0811G06F12/0815G06F12/0891G06F12/0897G06F2212/1021G06F2212/6042G06F2212/684
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,068,410
App. No.
16/291,154
Granted
Jul 20, 2021
Kind
B2
Abstract

According to embodiments described herein, the hierarchical complexity for coherence protocols associated with clustered cache architectures can be encapsulated in a simple function, i.e., that of determining when a data block is shared entirely within a cluster (i.e., a sub-tree of the hierarchy) and is private from the outside. This allows embodiments to eliminate complex recursive coherence operations that span the hierarchy and instead employ simple coherence mechanisms such as self-invalidation and write-through but which are restricted to operate where a data block is shared. Thus embodiments recognize that, in the context of clustered cache hierarchies, data can be shared entirely within one cluster but can be private (unshared) to this cluster when viewed from the perspective of other clusters. This characteristic of the data can be determined and then used to locally simplify coherence protocols.

Claims (35)

1. A computer system comprising:

multiple processor cores;

at least one local cache memory associated with, and operatively coupled to, a respective one of the multiple processor cores for storing one or more cache lines of data accessible only by the associated core;

at least one intermediary cache memory which is coupled to a subset of the multiple processor cores and which stores one or more cache lines of data; and

at least one shared memory, the shared memory being operatively coupled to all of the cores and which stores multiple data blocks,

wherein each cache line has a bit that signifies whether this cache line is private or shared in said shared memory,

wherein a common shared level is identified among the at least one intermediary cache memory and the at least one shared memory, the identifying being based on which of the at least one intermediary cache memory and the at least one shared memory is shared between two of the multiple processor cores,

wherein the common shared level is a level within computer system memory where a bit's value for a memory block becomes shared from being private in levels closer to local cache memory,

wherein a cache coherence operation is selected among a plurality of cache coherence operations based on said common shared level being identified, and

wherein said cache coherence operation is performed upon the occurrence of a coherence event.

2. The computer system of claim 1 , wherein if a cache line's bit is set to private, then that cache line is not self-invalidated and is written back when the coherence event associated with that cache line occurs.

3. The computer system of claim 1 , wherein if a cache line's bit is set to shared, then that cache line is self-invalidated and written through when the coherence event associated with that cache line occurs.

4. A computer system comprising:

multiple processor cores;

a clustered cache memory hierarchy including:

at least one local cache memory associated with and operatively coupled to each core for storing one or more cache lines accessible only by the associated core; and

a shared memory, the shared memory being operatively coupled to other shared memories or the local cache memories and accessible by a subset of cores that are transitively coupled to said shared memory via any number of local memories and intermediate shared memories, the shared memory being capable of storing a plurality of cache lines,

wherein each cache line has a private/shared bit that signifies whether this cache line is private or shared in said shared memory,

wherein a common shared level is identified for a memory block stored in the clustered cache hierarchy, the identifying being based on which of the clustered cache memory hierarchy is shared between two of the multiple processor cores,

wherein the common shared level is a level within the clustered cache memory hierarchy where a private/shared bit's value for a memory block becomes shared from being private in levels closer to local cache memory,

wherein a cache coherence operation is selected among a plurality of cache coherence operations based on said common shared level being identified, and wherein said cache coherence operation is performed upon the occurrence of a coherence event.

5. The computer system of claim 4 , wherein if a cache line's private/shared bit is set to private, then that cache line is not self-invalidated and is written back when the coherence event associated with that cache line occurs.

6. The computer system of claim 4 , wherein if a cache line's private/shared bit is set to shared, then that cache line is self-invalidated and written through when the coherence event associated with that cache line occurs.

7. The computer system of claim 4 , wherein the common shared level is also a level of a root cache of a cache cluster.

8. A method of a computer system including multiple processor cores and a clustered cache memory hierarchy, the method comprising:

storing, in at least one local cache memory associated with and operatively coupled to each of the multiple processor cores, cache lines accessible only by an associated core;

storing, in a shared memory being operatively coupled to other shared memories or the at least one local cache memory and accessible by a subset of cores that are transitively coupled to said shared memory via any number of local memories and intermediate shared memories, a plurality of cache lines, wherein each cache line has a private/shared bit that signifies whether this cache line is private or shared in said shared memory;

identifying a common shared level for a memory block stored in the clustered cache hierarchy, the identifying being based on which of the clustered cache memory hierarchy is shared between two of the multiple processor cores, wherein the common shared level is a level within the clustered cache memory hierarchy where a private/shared bit's value for a memory block becomes shared from being private in levels closer to local cache memory;

selecting a cache coherence operation among a plurality of cache coherence operations based on said common shared level being identified; and

performing said cache coherence operation upon the occurrence of a coherence event.

9. The method of claim 8 , comprising:

when a cache line's private/shared bit is set to private, not self-invalidating and writing back the cache line when the coherence event associated with that cache line occurs.

10. The method of claim 8 , comprising:

when a cache line's private/shared bit is set to shared, self-invalidating and writing through the cache line when the coherence event associated with that cache line occurs.

11. The method of claim 8 , wherein the common shared level is also a level of a root cache of a cache cluster.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2025
From: ETA SCALE AB
To: ARRAY CACHE TECHNOLOGIES LLC
Reel/Frame 071331/0090 →
Continuity (3)
Division 15015274 · Feb 4, 2016
Provisional Application 62112347 · Feb 5, 2015
Related Publication 20190205262A1 · Jul 4, 2019
Cited By (3)
US 12,566,704 US 12,675,408 US 12,724,712