IP Library › Granted Patent US 12,430,031
Granted Patent B2
US 12,430,031 · App. 18/641,779 · Granted Sep 30, 2025

Computing device with independently coherent nodes

Inventors: Siamak Tavallaei (Spring, TX); Ishwar Agarwal (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06F3/061G06F3/0655G06F3/0673G06F13/1668
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,430,031
App. No.
18/641,779
Granted
Sep 30, 2025
Kind
B2
Abstract

A computing device includes a system-on-a-chip. The computing device comprises a network interface controller (NIC) that hosts a plurality of virtual functions and physical functions. Two or more compute nodes are coupled to the NIC. Each compute node is configured to operate a plurality of Virtual Machines (VMs). Each VM is configured to operate in conjunction with a virtual function via a virtual function driver. A dedicated VM operates in conjunction with a virtual NIC using a physical function hosted by the NIC via a physical function driver hosted by the compute node. The computing device further comprises a fabric manager configured to own a physical function of the NIC, to bind virtual functions hosted by the NIC to individual compute nodes, and to pool I/O devices across the two or more compute nodes.

Claims (55)

1. A computing device, comprising:

two or more compute nodes, each compute node comprising an independently coherent domain that is not coherent with other compute nodes; and

a central IO die communicatively coupled to each of the two or more compute nodes, the central IO die comprising:

a compute die port coupled to each compute node;

a memory port communicatively coupled to disaggregated memory at an external device via an IO port; and

a plurality of distributed home agents communicatively coupled to each compute die port, each distributed home agent communicatively coupled via a target address decoder (TAD) to at least one memory port, each distributed home agent configured to:

receive a request for disaggregated memory from a requesting compute node via a respective compute die port;

pass the received request to the TAD; and

at the TAD, map the received request to a target memory port.

2. The computing device of claim 1 , wherein the compute die port decodes an address within the received request and maps the decoded address to a respective distributed home agent.

3. The computing device of claim 1 , wherein the target memory port is coupled to the external device comprising the disaggregated memory.

4. The computing device of claim 1 , wherein the distributed home agent further comprises a cache currency mechanism comprising a snoop filter, the cache currency mechanism coupled to the TAD via content addressable memory.

5. The computing device of claim 4 , wherein the distributed home agent is further configured to:

receive a request for memory access;

using the snoop filter, determine whether an updated copy of the request for memory access exists in the content addressable memory; and

present the updated copy of the request for memory access.

6. The computing device of claim 4 , wherein the IO port further includes a serial bus interconnect (SBI) currency port configured to allow parallel lookups to external cache and to disaggregated memory.

7. The computing device of claim 6 , wherein the distributed home agent is coupled to one or more additional external devices via the IO port.

8. The computing device of claim 7 , wherein the central IO die is configured to:

receive a request for an IO virtual address from the external device via an SBI root port;

at the SBI root port, translate the IO virtual address into a known physical address;

decode the known physical address such that any request for a given physical address targets the same distributed home agent;

send the decoded known physical address to the respective distributed home agent; and

at the distributed home agent, translate the decoded known physical address for the external device.

9. The computing device of claim 1 , wherein the computing device is a multi-chiplet device.

10. A computing device, comprising:

two or more compute nodes, each compute node comprising an independently coherent domain that is not coherent with other compute nodes; and

a central IO die communicatively coupled to each of the two or more compute nodes, the central IO die comprising:

a compute die port coupled to each compute node;

a memory port communicatively coupled to disaggregated memory at an external device via an IO port; and

a plurality of distributed memory access request managers communicatively coupled to each compute die port, each distributed memory access request manager communicatively coupled via a memory access routing agent to at least one memory port, each distributed memory access request manager configured to:

receive a request for disaggregated memory from a requesting compute node via a respective compute die port;

pass the received request to the memory access routing agent;

at the memory access routing agent, map the received request to a target memory port.

11. The computing device of claim 10 , wherein the compute die port decodes an address within the received request and maps the decoded address to a respective distributed memory access request manager.

12. The computing device of claim 10 , wherein the target memory port is coupled to the external device comprising the disaggregated memory.

13. The computing device of claim 10 , wherein the distributed memory access request manager further comprises a cache currency mechanism comprising a snoop filter, the cache currency mechanism coupled to the memory access routing agent via content addressable memory.

14. The computing device of claim 13 , wherein the distributed memory access request manager is further configured to:

receive a request for memory access;

using the snoop filter, determine whether an updated copy of the request for memory access exists in the content addressable memory; and

present the updated copy of the request for memory access.

15. The computing device of claim 13 , wherein the IO port further includes a serial bus interconnect (SBI) currency port configured to allow parallel lookups to external cache and to disaggregated memory.

16. The computing device of claim 15 , wherein the distributed memory access request manager is coupled to one or more additional external devices via the IO port.

17. The computing device of claim 16 , wherein the central IO die is configured to:

receive a request for an IO virtual address from the external device via an SBI root port;

at the SBI root port, translate the IO virtual address into a known physical address;

decode the known physical address such that any request for a given physical address targets the same distributed memory access request manager;

send the decoded known physical address to the respective distributed memory access request manager; and

at the distributed memory access request manager, translate the decoded known physical address for the external device.

18. The computing device of claim 10 , wherein the computing device is a multi-chiplet device.

19. A method for a computing device comprising a central IO die communicatively coupled to two or more compute nodes, each compute node comprising an independently coherent domain that is not coherent with other compute nodes, the method comprising:

at a home agent communicatively coupled to one of the compute nodes via a compute die port, receiving a request for disaggregated memory at an external device from the communicatively coupled compute node, the disaggregated memory communicatively coupled to a first memory port of the central IO die via an IO port;

passing the received request to a target address decoder (TAD) communicatively coupled to at least the first memory port via the home agent; and

at the TAD, mapping the received request to the first memory port.

20. The method of claim 19 , wherein the computing device is a multi-chiplet device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 5, 2024
From: AGARWAL, ISHWAR; TAVALLAEI, SIAMAK
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 067915/0913 →
Continuity (3)
Continuation 18049224 · Oct 24, 2022
Continuation 17016156 · Sep 9, 2020
Related Publication 20240377953A1 · Nov 14, 2024
References Cited (9)
US 8086765B2 · Turner · 2011 [cited by examiner]
US 8190839B2 · Bouvier · 2012 [cited by examiner]
US 10649915B1 · Bielski · 2020 [cited by examiner]
US 20030229721A1 · Bonola · 2003 [cited by examiner]
US 20100235598A1 · Bouvier · 2010 [cited by examiner]
US 20200371692A1 · Van Doorn · 2020 [cited by examiner]
US 20220004488A1 · Paul · 2022 [cited by examiner]
US 20220342835A1 · Kim · 2022 [cited by examiner]
US 20230325333A1 · Dastidar · 2023 [cited by examiner]