IP Library Granted Patent US 11,586,578
Granted Patent B1
US 11,586,578 · App. 17/080,642 · Granted Feb 21, 2023

Machine learning model updates to ML accelerators

Inventor: Jaideep Dastidar (San Jose, CA)
Assignee: XILINX, INC.
G06F15/7825G06F3/067G06F9/544G06F9/546G06F13/4282G06N20/00H04L12/66G06F2213/0026
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,586,578
App. No.
17/080,642
Granted
Feb 21, 2023
Kind
B1
Abstract

Examples herein describe a peripheral I/O device with a hybrid gateway that permits the device to have both I/O and coherent domains. As a result, the compute resources in the coherent domain of the peripheral I/O device can communicate with the host in a similar manner as CPU-to-CPU communication in the host. The dual domains in the peripheral I/O device can be leveraged for machine learning (ML) applications. While an I/O device can be used as an ML accelerator, these accelerators previously only used an I/O domain. In the embodiments herein, compute resources can be split between the I/O domain and the coherent domain where a ML engine is in the I/O domain and a ML model is in the coherent domain. An advantage of doing so is that the ML model can be coherently updated using a reference ML model stored in the host.

Claims (35)

1. An accelerator device, comprising:

an interface configured to communicatively couple the accelerator device to a host;

I/O logic comprising a machine learning (ML) engine assigned to an I/O domain; and

coherent logic comprising parameters corresponding to a ML algorithm, wherein the parameters are assigned to a coherent domain, wherein the parameters share the coherent domain with compute resources in the host.

2. The accelerator device of claim 1 , wherein the interface is configured to use a coherent interconnect protocol to extend the coherent domain of the host into the accelerator device.

3. The accelerator device of claim 2 , wherein the interface comprises an update agent configured to use a cache-coherent shared-memory multiprocessor paradigm to update the parameters in response to changes made in reference parameters stored in memory associated with the host.

4. The accelerator device of claim 1 , further comprising:

a NoC coupled to the I/O logic and the coherent logic, wherein at least one of the NoC, programmable logic (PL)-to-PL messages, and wire signaling is configured to permit the parameters to be transferred from the coherent logic to the I/O logic.

5. The accelerator device of claim 1 , wherein the ML engine is configured to process a ML data set received from the host using the parameters.

6. The accelerator device of claim 1 , further comprising:

a programmable logic (PL) array, wherein a first plurality of PL blocks in the PL array are part of the I/O logic and are assigned to the I/O domain and a second plurality of PL blocks in the PL array are part of the coherent logic and are assigned to the coherent domain.

7. The accelerator device of claim 6 , further comprising:

a plurality of memory blocks, wherein a first subset of the plurality of memory blocks are part of the I/O logic and are assigned to the I/O domain and a second subset of the plurality of memory blocks are part of the coherent logic and are assigned to the coherent domain, wherein the first subset of the plurality of memory blocks can communicate with the first plurality of PL blocks but not directly communicate with the second plurality of PL blocks and the second subset of the plurality of memory blocks can communicate with the second plurality of PL blocks but not directly communicate with the first plurality of PL blocks.

8. An accelerator device, comprising:

an interface configured to communicatively couple the accelerator device to a host;

I/O logic comprising a machine learning (ML) engine assigned to an I/O domain; and

coherent logic configured to store updates for changing how the ML engine processes data, wherein the updates are assigned to a coherent domain, wherein the updates share the coherent domain with compute resources in the host.

9. The accelerator device of claim 8 , wherein the interface is configured to use a coherent interconnect protocol to extend the coherent domain of the host into the accelerator device.

10. The accelerator device of claim 9 , wherein the interface comprises an update agent configured to use a cache-coherent shared-memory multiprocessor paradigm to store the updates in response to changes made in reference parameters stored in memory associated with the host.

11. The accelerator device of claim 8 , further comprising:

a NoC coupled to the I/O logic and the coherent logic, wherein at least one of the NoC, programmable logic (PL)-to-PL messages, and wire signaling is configured to permit the updates to be transferred from the coherent logic to the I/O logic.

12. The accelerator device of claim 11 , wherein the ML engine is configured to process a ML data set received from the host based on the updates.

13. The accelerator device of claim 8 , further comprising:

a programmable logic (PL) array, wherein a first plurality of PL blocks in the PL array are part of the I/O logic and are assigned to the I/O domain and a second plurality of PL blocks in the PL array are part of the coherent logic and are assigned to the coherent domain.

14. The accelerator device of claim 13 , further comprising:

a plurality of memory blocks, wherein a first subset of the plurality of memory blocks are part of the I/O logic and are assigned to the I/O domain and a second subset of the plurality of memory blocks are part of the coherent logic and are assigned to the coherent domain, wherein the first subset of the plurality of memory blocks can communicate with the first plurality of PL blocks but not directly communicate with the second plurality of PL blocks and the second subset of the plurality of memory blocks can communicate with the second plurality of PL blocks but not directly communicate with the first plurality of PL blocks.

15. An accelerator device, comprising:

an interface configured to communicatively couple the accelerator device to a host, the host storing a reference ML model;

I/O logic comprising a machine learning (ML) engine assigned to an I/O domain; and

coherent logic configured to store a sub-portion of the reference ML model, wherein the sub-portion of the reference ML model shares a coherent domain with compute resources in the host.

16. The accelerator device of claim 15 , wherein the coherent logic is configured to store parameters corresponding to the sub-portion of the reference ML model, wherein the parameters are unique to the sub-portion of the reference ML model.

17. The accelerator device of claim 15 , wherein the reference ML model in the host stores a plurality of ML models.

18. The accelerator device of claim 15 , wherein the accelerator device is disposed in a container with a plurality of accelerator devices, wherein each of the plurality of accelerator devices comprises a sub-portion of the reference ML model.

19. The accelerator device of claim 18 , wherein the plurality of accelerator devices are transmitted different data sets to process using the sub-portion of the reference ML model.

20. The accelerator device of claim 18 , wherein the plurality of accelerator devices are transmitted the same data set to process using the sub-portion of the reference ML model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 26, 2021
From: DASTIDAR, JAIDEEP; MITTAL, MILLIND
To: XILINX, INC.
Reel/Frame 055039/0407 →
Continuity (1)
Division 16396540 · Apr 26, 2019
Cited By (2)
US 12,253,940 US 12,493,576