IP Library Granted Patent US 12,386,751
Granted Patent B2
US 12,386,751 · App. 17/809,484 · Granted Aug 12, 2025

Composable infrastructure enabled by heterogeneous architecture, delivered by CXL based cached switch SOC and extensible via cxloverethernet (COE) protocols

Inventors: Shreyas Shah (San Jose, CA); George Apostol, Jr. (Los Gatos, CA); Nagarajan Subramaniyan (San Jose, CA); Jack Regula (Durham, NC); Jeffrey S. Earl (San Jose, CA)
Assignee: Avago Technologies International Sales Pte. Limited
G06F12/0868G06F12/0646G06F12/0815G06F12/0837G06F12/0862G06F12/1466G06F13/1642G06F13/1668G06F13/1673G06F13/4022G06F13/4221G06F2213/0026G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,386,751
App. No.
17/809,484
Granted
Aug 12, 2025
Kind
B2
Abstract

Described herein are systems, methods, and products utilizing a cache coherent switch on chip. The cache coherent switch on chip may utilize Compute Express Link (CXL) interconnect open standard and allow for multi-host access and the sharing of resources. The cache coherent switch on chip provides for resource sharing between components while independent of a system processor, removing the system processor as a bottleneck. Cache coherent switch on chip may further allow for cache coherency between various different components. Thus, for example, memories, accelerators, and/or other components within the disclose systems may each maintain caches, and the systems and techniques described herein allow for cache coherency between the different components of the system with minimal latency.

Claims (38)

1. A system comprising:

a first server device comprising:

a processor;

a first accelerator, separate from the processor, and configured to accelerate one or more types of workloads;

a second accelerator, separate from the processor, and configured to accelerate the one or more types of workloads;

a first cache coherent switch on chip, communicatively coupled to the first accelerator and the second accelerator via a Compute Express Link (CXL) protocol, wherein the first cache coherent switch on chip is configured to bypass the processor to provide cache coherency between the first accelerator and the second accelerator; and

wherein the first cache coherent switch comprises one or more cache hierarchies, the one or more cache hierarchies indicating a priority for the one or more caches coupled to the first cache coherent switch.

2. The system of claim 1 , further comprising:

a third accelerator, wherein the first cache coherent switch on chip is further configured to provide cache coherency between first accelerator, the second accelerator, and the third accelerator.

3. The system of claim 1 , further comprising:

a first network interface card, communicatively coupled to the first cache coherent switch on chip and to a network.

4. The system of claim 3 , wherein the first network interface card is configured to:

receive, from the network, cache coherent data; and

provide, to the first cache coherent switch on chip, the cache coherent data.

5. The system of claim 3 , further comprising:

a second server device comprising:

a third accelerator;

a second cache coherent switch on chip, communicatively coupled to the third accelerator and configured to:

receive, from a second network interface card, the cache coherent data; and

provide, to the third accelerator, the cache coherent data; and

the second network interface card, communicatively coupled to the second cache coherent switch on chip and to the first network interface card via the network, wherein the first cache coherent switch on chip and the second cache coherent switch on chip are configured to provide cache coherency between the first accelerator, the second accelerator, and the third accelerator by:

receiving cache coherent data from the first accelerator;

providing the cache coherent data to the second accelerator; and

providing the cache coherent data to the first network interface card for communication to the second cache coherent switch on chip via the network and the second network interface card.

6. The system of claim 1 , further comprising a memory, wherein the first cache coherent switch on chip is further configured to provide cache coherency to the memory.

7. The system of claim 1 , wherein the first cache coherent switch on chip comprises a microprocessor, and wherein the microprocessor is configured to direct cache coherent data to the first accelerator and/or the second accelerator to provide the cache coherency.

8. A method comprising:

receiving, with a cache coherent switch on chip from a network interface card, cache coherent data addressed to a first accelerator, the cache coherent switch comprising one or more cache hierarchies, the one or more cache hierarchies indicating a priority for the one or more caches coupled to the first cache coherent switch;

providing, by the cache coherent switch on chip to the first accelerator while bypassing the processor, the cache coherent data;

receiving, with the cache coherent switch on chip from the first accelerator, a bias change;

providing, by the cache coherent switch on chip to a processor, the bias change;

receiving, with the cache coherent switch on chip from the processor, line resolved data; and

providing, by the cache coherent switch on chip to the first accelerator, the line resolved data to cause the first accelerator to write the cache coherent data into a cache coherent memory of a second accelerator;

wherein the first accelerator and the second accelerator, separate from the processor, are configured to accelerate one or more types of workloads.

9. The method of claim 8 , further comprising:

receiving, with the cache coherent switch on chip from the processor, snoop data; and

providing, by the cache coherent switch on chip to a first snooping component, the snoop data, wherein the snoop data indicates the first snooping component.

10. The method of claim 9 , wherein the first snooping component is the second accelerator.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2023
From: ELASTICS.CLOUD, INC.
To: AVAGO TECHNOLOGIES INTERNATIONAL SALES PTE. LIMITED
Reel/Frame 065104/0547 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2022
From: SHAH, SHREYAS; APOSTOL, GEORGE, JR.; SUBRAMANIYAN, NAGARAJAN; REGULA, JACK; EARL, JEFFREY S.
To: ELASTICS.CLOUD, INC.
Reel/Frame 060424/0203 →
Continuity (2)
Provisional Application 63223045 · Jul 18, 2021
Related Publication 20230027178A1 · Jan 26, 2023
References Cited (48)
US 11074208B1 · Dastidar et al. · 2021 [cited by applicant]
US 11388268B1 · Siva et al. · 2022 [cited by applicant]
US 11573898B2 · Passint · 2023 [cited by examiner]
US 20150012679A1 · Davis et al. · 2015 [cited by applicant]
US 20150169452A1 · Persson et al. · 2015 [cited by applicant]
US 20160299860A1 · Harriman · 2016 [cited by applicant]
US 20160381176A1 · Cherubini et al. · 2016 [cited by applicant]
US 20190042518A1 · Marolia · 2019 [cited by examiner]
US 20190227936A1 · Jang · 2019 [cited by applicant]
US 20200192798A1 · Natu · 2020 [cited by applicant]
US 20200322287A1 · Connor et al. · 2020 [cited by applicant]
US 20200341930A1 · Cannata et al. · 2020 [cited by applicant]
US 20210011755A1 · Shah · 2021 [cited by applicant]
US 20210075633A1 · Sen et al. · 2021 [cited by applicant]
US 20210117244A1 · Herdrich et al. · 2021 [cited by applicant]
US 20210132999A1 · Haywood · 2021 [cited by examiner]
US 20210240655A1 · Das Sharma · 2021 [cited by applicant]
US 20210311643A1 · Shanbhouge et al. · 2021 [cited by applicant]
US 20210311646A1 · Malladi et al. · 2021 [cited by applicant]
US 20210311739A1 · Malladi et al. · 2021 [cited by applicant]
US 20210311900A1 · Malladi et al. · 2021 [cited by applicant]
US 20210318976A1 · Zhang et al. · 2021 [cited by applicant]
US 20210320866A1 · Le et al. · 2021 [cited by applicant]
US 20210374056A1 · Malladi et al. · 2021 [cited by applicant]
US 20210382838A1 · Mittal et al. · 2021 [cited by applicant]
US 20220124038A1 · Leguay et al. · 2022 [cited by applicant]
US 20220147476A1 · Nam et al. · 2022 [cited by applicant]
US 20220164288A1 · Ramagiri · 2022 [cited by examiner]
US 20220292026A1 · Hornung et al. · 2022 [cited by applicant]
US 20220326874A1 · Del Gatto et al. · 2022 [cited by applicant]
US 20220350767A1 · McGraw et al. · 2022 [cited by applicant]
US 20220383961A1 · Lien et al. · 2022 [cited by applicant]
US 20220398207A1 · Norrie et al. · 2022 [cited by applicant]
US 20220405212A1 · Kakaiya et al. · 2022 [cited by applicant]
US 20230012822A1 · Shah et al. · 2023 [cited by applicant]
US 20230017583A1 · Shah et al. · 2023 [cited by applicant]
US 20230017643A1 · Shah et al. · 2023 [cited by applicant]
US 20230409302A1 · Kodama et al. · 2023 [cited by applicant]
“FlexPod Datacenter with Citrix VDI and VMware vSphere 7 for up to 2500 Seats”, Cisco, Published Apr. 2022, http://www.cisco.com/go/designzone, 497 pages. [cited by applicant]
Amir Roozbeh, “Realizing Next-Generation Data Centers via Software-Defined “Hardware” Infrastructures and Resource Disaggregation”, Doctoral Thesis KTH Royal Institute of Technology, 227 pages. [cited by applicant]
Davide Giri et al, “NoC-Based Support of Heterogeneous Cache-Coherence Models for Accelerators”, 2018 Twelfth IEEE/ACM International Symposium on Networks-on-Chip (NOCS), IEEE Oct. 4, 18, pp. 1-8, Section III and figure… [cited by applicant]
Int'l Application Serial No. PCT/US22/73233, ISR/WO mailed Oct. 14, 2022 9 pgs. [cited by applicant]
Kshitij Bhardwaj et al, “Determining Optimal Coherence Interface for Many-Accelerator SoC's Using Bayesian Optimization”, IEEE Computer Architecture Letters, IEEE Sep. 16, 2019, pp. 119-123 Section 3.1; and figure 2. [cited by applicant]
Kshitij Bhardwaj, et al, “A Comprehensive Methodology to Determine Optimal Coherence Interfaces for Many-Accelerator SoC's”, ISLPED '20 Proceedings of the ACM/IEEE International Symposium on Low Power Electronics and De… [cited by applicant]
Prateek Shantharama, et al., “Hardware Accelerated Platforms and Infrastructures for Network Functions: A Survey of Enabling Technologies and Research Studies”.IEEE Jul. 9, 2020, Digital Object Identifier 10.1109/Access… [cited by applicant]
Yakun Sophia Shao, et al. “Co-Designing Accelerators and SoC Interfaces using gem5-Aladdin”, 2016 49th Annual IEEE/ACM International Symposium on Microarchitecture (Micro). IFEE, Oct. 15, 2016, pp. 1-12, pp. 3-5 and fig… [cited by applicant]
Zuckerman, et al. “Cohmeleon: Learning-Based Orchestration ofAccelerator Coherence in Heterogeneous SoCs”, Columbia University, New York, New York, arXiv:2109.06382v1 [cs.AR] Sep. 14, 2021, 14 pages. [cited by applicant]
Non-Final Office Action on U.S. Appl. No. 18/605,301 DTD May 7, 2025. [cited by applicant]