IP Library › Granted Patent US 11,966,335
Granted Patent B2
US 11,966,335 · App. 17/876,110 · Granted Apr 23, 2024

Hardware interconnect with memory coherence

Inventors: Kiran Suresh Puranik (Fremont, CA); Prakash Chauhan (Los Gatos, CA)
Assignee: Google LLC
G06F12/0815G06F2212/301
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,966,335
App. No.
17/876,110
Granted
Apr 23, 2024
Kind
B2
Abstract

Aspects of the disclosure are directed to hardware interconnects and corresponding devices and systems for non-coherently accessing data in shared memory devices. Data produced and consumed by devices implementing the hardware interconnect can read and write directly to a memory device shared by multiple devices, and limit coherent memory transactions to relatively smaller flags and descriptors used to facilitate data transmission as described herein. Devices can communicate less data on input/output channels, and more data on memory and cache channels that are more efficient for data transmission. Aspects of the disclosure are directed to devices configured to process data that is read from the shared memory device. Devices, such as hardware accelerators, can receive data indicating addresses for different data buffers with data for processing, and non-coherently read or write the contents of the data buffers on a memory device shared between the accelerators and a host device.

Claims (49)

1. A system comprising:

a host device and an accelerator communicatively coupled over a hardware interconnect supporting memory-coherent data transmission between the host device and the accelerator, the host device comprising a host cache and the accelerator comprising an accelerator cache;

wherein the host device is configured to:

read or write data to one or more data buffers to a first memory device shared between the host device and the accelerator,

write control information to a second computing device, the control information comprising one or more flags indicating the status of one or more data buffers in the first memory device,

wherein the accelerator is configured to:

non-coherently read or write data from or to the one or more data buffers of the first memory device based on the control information, and after non-coherently reading or writing the data, coherently write updated control information to the second memory device, wherein coherently writing the updated control information causes the control information in the host cache to be updated,

wherein the accelerator is further configured to receive, at an enqueue register, one or more command descriptors, each command descriptor specifying a respective data descriptor in the accelerator cache, and

wherein to non-coherently read or write data from or to the one or more data buffers, the accelerator is configured to read addresses from the one or more command descriptors corresponding to the one or more data descriptors.

2. The system of claim 1 , wherein the accelerator is configured to receive the one or more command descriptors as a deferred memory write (DMWr) transaction.

3. The system of claim 1 , wherein the one or more command descriptors are received from an application executed on a virtual machine hosted by the host device.

4. The system of claim 1 ,

wherein hardware interconnect comprises a plurality of channels; and

wherein the accelerator is configured to:

receive the one or more command descriptors over a first channel of the plurality of channels dedicated to input/output data communication; and

non-coherently read or write data from or to the one or more data buffers over a second channel of the plurality of channels dedicated to communication between memory devices connected to the first or second computing devices.

5. The system of claim 4 , wherein to coherently write the updated control information to the second memory device, the accelerator is further configured to cause the updated control information to be sent to the host cache of the host device over a third channel of the plurality of channels dedicated to updating contents of the accelerator or host cache.

6. A first computing device comprising:

a first cache; and

one or more processors coupled to a first memory device shared between the first computing device and a second computing device, the one or more processors are configured to:

cache control information in the first cache, the control information comprising one or more flags indicating the status of one or more data buffers in the memory device and accessed from a second memory device connected to the second computing device; and

non-coherently read or write contents of the one or more data buffers based on the control information, and after non-coherently reading or writing the contents of the one or more buffers, coherently write updated control information to the second memory device, wherein coherently writing the updated control information causes the control information in the first or second cache to also be updated,

wherein the first computing device is configured to communicate with the second computing device over a hardware interconnect comprising a plurality of channels and configured for memory-coherent data transmission,

wherein the one or more processors are further configured to:

coherently read or write the control information over a first channel dedicated to input/output (I/O) data communication; and

non-coherently read or write the contents of the one or more data buffers over a second channel dedicated to communication between memory devices connected to the first or second computing device,

wherein to coherently write the updated control information to the second memory device, the one or more processors are configured to cause the updated control information to be sent to the second cache of the second memory device over a third channel dedicated to updating contents of the first or second cache, and

wherein the one or more processors are further configured to:

receive, over the I/O channel, a command descriptor, the command descriptor comprising respective addresses for a source data buffer and a destination data buffer in the first memory device;

cache the respective addresses for the source and destination data buffers to the first cache; and

non-coherently read or write the contents of the source and destination data buffer using the respective cached addresses.

7. The first computing device of claim 6 , wherein the first computing device is a hardware accelerator device comprising one or more accelerator cores and the first cache is an accelerator cache for the hardware accelerator device.

8. The first computing device of claim 7 , wherein the one or more processors are configured to non-coherently read the contents of the one or more data buffers based on the value of one or more of the plurality of flags indicating that the contents of the one or more data buffers are ready for consumption.

9. The first computing device of claim 8 , wherein the one or more of the plurality of flags are set by the second computing device configured to write the contents to the one or more data buffers.

10. The first computing device of claim 6 , where the control information further comprises data descriptors, each data descriptor identifying an address for a respective source data buffer for the first computing device to read from, or for a respective destination data buffer for the first computing to write to.

11. One or more non-transitory computer-readable storage media encoded with instructions that, when executed by one or more processors of a first computing device comprising a first cache and coupled to a first memory device shared between the first computing device and a second computing device, causes the one or more processors to perform operations comprising:

caching control information in the first cache, the control information comprising one or more flags indicating the status of one or more data buffers in the memory device and accessed from a second memory device connected to the second computing device; and

non-coherently reading or writing contents of the one or more data buffers based on the control information, and after non-coherently reading or writing the contents of the one or more buffers, coherently write updated control information to the second memory device, wherein coherently writing the updated control information causes the control information in the first or second cache to also be updated,

wherein the first computing device is configured to communicate with the second computing device over a hardware interconnect comprising a plurality of channels and configured for memory-coherent data transmission,

wherein the operations further comprise:

coherently reading or writing the control information over a first channel dedicated to input/output (I/O) data communication, and

non-coherently reading or writing the contents of the one or more data buffers over a second channel dedicated to communication between memory devices connected to the first or second computing device,

wherein coherently writing the updated control information to the second memory device comprises causing the updated control information to be sent to the second cache of the second memory device over a third channel dedicated to updating contents of the first or second cache, and

wherein the operations further comprise:

receiving, over the I/O channel, a command descriptor, the command descriptor comprising respective addresses for a source data buffer and a destination data buffer in the first memory device;

caching the respective addresses for the source and destination data buffers to the first cache; and

non-coherently reading or writing the contents of the source and destination data buffer using the respective cached addresses.

12. The non-transitory computer-readable storage media of claim 11 , wherein the first computing device is a hardware accelerator device comprising one or more accelerator cores and the first cache is an accelerator cache for the hardware accelerator device.

13. The non-transitory computer-readable storage media of claim 12 , wherein the operations further comprise non-coherently reading the contents of the one or more data buffers based on the value of one or more of the plurality of flags indicating that the contents of the one or more data buffers are ready for consumption.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2022
From: PURANIK, KIRAN SURESH; CHAUHAN, PRAKASH
To: GOOGLE LLC
Reel/Frame 060696/0783 →
Continuity (2)
Provisional Application 63231397 · Aug 10, 2021
Related Publication 20230052808A1 · Feb 16, 2023