IP Library Granted Patent US 12,596,647
Granted Patent B2
US 12,596,647 · App. 18/413,211 · Granted Apr 7, 2026

Cache management using shared cache line storage

Inventors: Sanjay Patel (San Ramon, CA); Yogesh Shamkant Thombre (Pleasanton, CA)
Assignee: Akeana, Inc.
G06F12/0828G06F12/0868G06F12/0873
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,596,647
App. No.
18/413,211
Granted
Apr 7, 2026
Kind
B2
Abstract

Techniques for cache management based on cache management using memory queues are disclosed. A plurality of processor cores is accessed. The plurality of processor cores comprises a coherency domain. Two or more processor cores within the plurality of processor cores generate read operations for a common memory structure coupled to the plurality of processor cores. Coherency for the coherency domain is managed using a compute coherency block (CCB). The CCB includes a memory queue for controlling transfer of cache lines determined by the CCB. The memory queue includes an evict queue and a miss queue. Snoop requests are generated by the CCB. The snoop requests correspond to entries in the memory queue. Cache lines are transferred between the CCB and a bus interface unit. The transferring is controlled by the memory queue. The bus interface unit controls memory accesses.

Claims (36)

1 . A processor-implemented method for cache management comprising:

accessing a plurality of processor cores, wherein the plurality of processor cores comprises a coherency domain, and wherein two or more processor cores within the plurality of processor cores generate read operations for a common memory structure coupled to the plurality of processor cores, wherein the memory accesses are performed over a network-on-chip, wherein non-cacheable memory accesses are targeted to memory mapped input and output (I/O) locations within the common memory structure, wherein the non-cacheable memory accesses are aligned to a cache line boundary;

managing coherency for the coherency domain using a compute coherency block (CCB), wherein the CCB includes a memory queue for managing cache line transfers determined by the CCB;

generating snoop requests, by the CCB, wherein the snoop requests correspond to entries in the memory queue; and

transferring cache lines between the compute coherency block and a bus interface unit, wherein the transferring is controlled by the memory queue, wherein the transferring occurs from the CCB to the bus interface unit when the cache line is an evicted cache line; and

committing the evicted cache line to the common memory structure.

2 . The method of claim 1 wherein the transferring is initiated based on a response to the snoop requests.

3 . The method of claim 2 wherein the transferring occurs from the bus interface unit to the CCB when the cache line is a pending cache line fill.

4 . The method of claim 1 wherein the memory queue comprises an evict queue.

5 . The method of claim 4 wherein the evict queue controls transferring cache lines evicted from the CCB.

6 . The method of claim 5 wherein the cache lines are transferred from the CCB to the bus interface unit, based on completion of all required, pending snoop responses.

7 . The method of claim 1 wherein the memory queue comprises a miss queue.

8 . The method of claim 7 wherein the miss queue controls transfer of cache lines read from the common memory structure to the CCB.

9 . The method of claim 8 wherein cache lines in the bus interface unit are scheduled for transfer to the CCB by the miss queue, based on completion of all required, pending snoop responses.

10 . The method of claim 1 wherein cache lines are stored in a bus interface unit cache prior to commitment to the common memory structure.

11 . The method of claim 1 wherein cache lines are stored in a bus interface unit cache pending completion of a cache line fill from the common memory structure.

12 . The method of claim 11 wherein cacheable memory accesses are targeted to storage locations within the common memory structure.

13 . The method of claim 12 wherein the cache line comprises 512 bits.

14 . The method of claim 1 wherein the network-on-chip is included in the coherency domain.

15 . The method of claim 1 wherein the non-cacheable memory accesses comprise 8, 16, 32, or 64 bits.

16 . The method of claim 1 wherein each processor within the plurality of processor cores is coupled to a dedicated local cache.

17 . The method of claim 16 wherein the dedicated local cache is included in the coherency domain.

18 . The method of claim 1 wherein the CCB comprises a common ordering point for coherency management.

19 . A computer program product embodied in a non-transitory computer readable medium for cache management, the computer program product comprising code which causes one or more processors to generate semiconductor logic for:

accessing a plurality of processor cores, wherein the plurality of processor cores comprises a coherency domain, and wherein two or more processor cores within the plurality of processor cores generate read operations for a common memory structure coupled to the plurality of processor cores, wherein the memory accesses are performed over a network-on-chip, wherein non-cacheable memory accesses are targeted to memory mapped input and output (I/O) locations within the common memory structure, wherein the non-cacheable memory accesses are aligned to a cache line boundary;

managing coherency for the coherency domain using a compute coherency block (CCB), wherein the CCB includes a memory queue for storing cache lines determined by the CCB;

generating snoop requests, by the CCB, wherein the snoop requests correspond to entries in the memory queue; and

transferring cache lines between the compute coherency block and a bus interface unit, wherein the transferring is controlled by the memory queue, wherein the transferring occurs from the CCB to the bus interface unit when the cache line is an evicted cache line; and

committing the evicted cache line to the common memory structure.

20 . An apparatus for cache management comprising:

a plurality of processor cores, wherein the plurality of processor cores comprises a coherency domain, and wherein two or more processor cores within the plurality of processor cores generate memory accesses for a common memory structure coupled to the plurality of processor cores, wherein the memory accesses are performed over a network-on-chip, wherein non-cacheable memory accesses are targeted to memory mapped input and output (I/O) locations within the common memory structure, wherein the non-cacheable memory accesses are aligned to a cache line boundary;

a compute coherency block (CCB) for managing coherency within the coherency domain, wherein the CCB is coupled to the plurality of processor cores and includes a memory queue for storing cache lines determined by the CCB;

a bus interface unit, wherein the bus interface unit is coupled to the CCB and implements memory accesses; and wherein

the CCB generates snoop requests, wherein the snoop requests correspond to entries in the memory queue; and

cache lines are transferred between the compute coherency block and the bus interface unit, under control of the memory queue, wherein transferring occurs from the CCB to the bus interface unit when the cache line is an evicted cache line; and

the evicted cache line is committed to the common memory structure.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2025
From: PATEL, SANJAY; THOMBRE, YOGESH SHAMKANT
To: AKEANA, INC.
Reel/Frame 070933/0318 →
Continuity (16)
Provisional Application 63605620 · Dec 4, 2023
Provisional Application 63602514 · Nov 24, 2023
Provisional Application 63547574 · Nov 7, 2023
Provisional Application 63547404 · Nov 6, 2023
Provisional Application 63546769 · Nov 1, 2023
Provisional Application 63545961 · Oct 27, 2023
Provisional Application 63542797 · Oct 6, 2023
Provisional Application 63526009 · Jul 11, 2023
Provisional Application 63521365 · Jun 16, 2023
Provisional Application 63471283 · Jun 6, 2023
Provisional Application 63467335 · May 18, 2023
Provisional Application 63463371 · May 2, 2023
Provisional Application 63462542 · Apr 28, 2023
Provisional Application 63444619 · Feb 10, 2023
Provisional Application 63439761 · Jan 18, 2023
Related Publication 20240241830A1 · Jul 18, 2024
References Cited (24)
US 5692202A · Kardach · 1997 [cited by examiner]
US 6934809B2 · Tremblay et al. · 2005 [cited by applicant]
US 7065632B1 · Col · 2006 [cited by examiner]
US 7506105B2 · Al-Sukhni et al. · 2009 [cited by applicant]
US 10013356B2 · Chou · 2018 [cited by applicant]
US 10671394B2 · Britto et al. · 2020 [cited by applicant]
US 10929948B2 · Benthin et al. · 2021 [cited by applicant]
US 11288405B2 · Belgarric et al. · 2022 [cited by applicant]
US 11403099B2 · Cerny et al. · 2022 [cited by applicant]
US 11403225B2 · Zheng et al. · 2022 [cited by applicant]
US 11429529B2 · Hornung et al. · 2022 [cited by applicant]
US 11442863B2 · Shulyak et al. · 2022 [cited by applicant]
US 11474130B2 · Lentz et al. · 2022 [cited by applicant]
US 11486911B2 · Tuncer et al. · 2022 [cited by applicant]
US 20040003336A1 · Cypher · 2004 [cited by examiner]
US 20080109606A1 · Lataille · 2008 [cited by examiner]
US 20080162661A1 · Mannava · 2008 [cited by examiner]
US 20120060169A1 · Steffens · 2012 [cited by examiner]
US 20160062890A1 · Salisbury · 2016 [cited by examiner]
US 20220004639A1 · Yardi et al. · 2022 [cited by applicant]
US 20220029780A1 · Dafali · 2022 [cited by applicant]
US 20220197657A1 · Soundararajan et al. · 2022 [cited by applicant]
US 20230315449A1 · Cunningham · 2023 [cited by examiner]
WO 2022117687A1 · 2022 [cited by applicant]