IP Library › Granted Patent US 11,188,464
Granted Patent B2
US 11,188,464 · App. 16/710,203 · Granted Nov 30, 2021

System and method for self-invalidation, self-downgrade cachecoherence protocols

Inventors: Alberto Ros (Cartagena, ES); Stefanos Kaxiras (Uppsala, SE)
Assignee: ETA SCALE AB
G06F12/0831G06F9/52G06F12/084G06F12/0811G06F12/0815G06F12/1018G06F12/128G06F2212/1016G06F2212/283G06F2212/621
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,188,464
App. No.
16/710,203
Granted
Nov 30, 2021
Kind
B2
Abstract

Methods and systems for self-invalidating cachelines in a computer system having a plurality of cores are described. A first one of the plurality of cores, requests to load a memory block from a cache memory local to the first one of the plurality of cores, which request results in a cache miss. This results in checking a read-after-write detection structure to determine if a race condition exists for the memory block. If a race condition exists for the memory block, program order is enforced by the first one of the plurality of cores at least between any older loads and any younger loads with respect to the load that detects the prior store in the first one of the plurality of cores that issued the load of the memory block and causing one or more cache lines in the local cache memory to be self-invalidated.

Claims (32)

1. A method for self-invalidating cachelines in a computer system having a plurality of cores, the method comprising:

requesting, by a first one of the plurality of cores, to load a memory block from a cache memory local to the first one of the plurality of cores, which request results in a cache miss;

checking a read-after-write detection structure to determine if a race condition exists for the memory block; and

if a race condition exists for the memory block, enforcing program order at least between any older loads and any younger loads with respect to the load that detects a prior store in the first one of the plurality of cores that issued the load of the memory block and causing one or more cache lines in the local cache memory to be self-invalidated.

2. The method of claim 1 , wherein the read-after-write detection structure contains address information associated with stores made by the plurality of cores and wherein the step of checking further comprises:

comparing a target address of the load with the address information in the read-after-write detection structure to determine if the race condition exists.

3. The method of claim 1 , further comprising:

adding, to the read-after-write detection structure, an address associated with an executed store to address information stored in the read-after-write detection structure associated with each of the plurality of cores except for a core which executed the store.

4. The method of claim 1 , wherein checking the read-after-write detection structure comprises accessing a signature table at a level of a shared lower level cache.

5. The method of claim 4 , wherein checking the read-after-write detection structure comprises matching in the signature table a target address of the memory block accessed by the load with a target address of the memory block previously written by one or more stores from one or more other cores.

6. The method of claim 1 , further comprising:

sending, by the first one of the plurality of cores, a request to a shared lower level cache when the cache miss results.

7. The method of claim 1 , wherein the prior store is ordered in the read-after-write detection structure, and wherein the enforcing program order comprises performing the load of the memory block in program order relative to the prior store.

8. The method of claim 1 , wherein the prior store is not ordered, and wherein the enforcing program order comprises delaying performing the load of the memory block.

9. The method of claim 1 , wherein the causing one or more cachelines in the local cache memory to be self-invalidated comprises including an indication that the race condition exists in a response to the load that detects the prior store.

10. The method of claim 9 , wherein the causing one or more cachelines in the local cache memory to be self-invalidated comprises including an identity of a core that executed the prior store.

11. A computer system comprising:

a plurality of cores; and

a read-after-write detection structure,

wherein a first one of the plurality of cores is configured to request to load a memory block from a cache memory local to the first one of the plurality of cores, which request results in a cache miss;

wherein the computer system is configured to check the read-after-write detection structure to determine if a race condition exists for the memory block; and

wherein, if a race condition exists for the memory block, the computer system is configured to enforce program order at least between any older loads and any younger loads with respect to the load that detects a prior store in the first one of the plurality of cores that issued the load of the memory block and cause one or more cache lines in the local cache memory to be self-invalidated.

12. The computer system of claim 11 , wherein the read-after-write detection structure contains address information associated with stores made by the plurality of cores and wherein the computer system is configured to check the read-after-write detection structure to determine if the race condition exists for the memory block by:

comparing a target address of the load with the address information in the read-after-write detection structure to determine if the race condition exists.

13. The computer system of claim 11 , wherein the computer system is configured to add, to the read-after-write detection structure, an address associated with an executed store to address information stored in the read-after-write detection structure associated with each of the plurality of cores except for a core which executed the store.

14. The computer system of claim 11 , wherein the computer system is configured to check the read-after-write detection structure by accessing a signature table at a level of a shared lower level cache.

15. The computer system of claim 14 , wherein the computer system is configured to check the read-after-write detection structure by matching in the signature table a target address of the memory block accessed by the load with a target address of the memory block previously written by one or more stores from one or more other cores.

16. The computer system of claim 11 , wherein the first one of the plurality of cores is configured to send a request to a shared lower level cache when the cache miss results.

17. The computer system of claim 11 , wherein the prior store is ordered in the read-after-write detection structure, and wherein the computer system is configured to enforce program order by performing the load of the memory block in program order relative to the prior store.

18. The computer system of claim 11 , wherein the prior store is not ordered, and wherein the computer system is configured to enforce program order by delaying performing the load of the memory block.

19. The computer system of claim 11 , wherein the computer system is configured to cause one or more cachelines in the local cache memory to be self-invalidated by including an indication that the race condition exists in a response to the load that detects the prior store.

20. The computer system of claim 19 , wherein the computer system is configured to cause one or more cachelines in the local cache memory to be self-invalidated by including an identity of a core that executed the prior store.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2025
From: ETA SCALE AB
To: ARRAY CACHE TECHNOLOGIES LLC
Reel/Frame 071331/0090 →
Continuity (3)
Division 15855378 · Dec 27, 2017
Provisional Application 62439189 · Dec 27, 2016
Related Publication 20200110703A1 · Apr 9, 2020