IP Library Granted Patent US 11,948,073
Granted Patent B2
US 11,948,073 · App. 16/117,302 · Granted Apr 2, 2024

Machine learning inference engine scalability

Inventors: Lei Zhang (Richmond Hill, CA); Sateesh Lagudu (Hyderabad, IN); Allen Rush (Danville, CA)
Assignees: Advanced Micro Devices, Inc.; ATI Technologies ULC
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,948,073
App. No.
16/117,302
Granted
Apr 2, 2024
Kind
B2
Abstract

Systems, apparatuses, and methods for adaptively mapping a machine learning model to a multi-core inference accelerator engine are disclosed. A computing system includes a multi-core inference accelerator engine with multiple inference cores coupled to a memory subsystem. The system also includes a control unit which determines how to adaptively map a machine learning model to the multi-core inference accelerator engine. In one implementation, the control unit selects a mapping scheme which minimizes the memory bandwidth utilization of the multi-core inference accelerator engine. In one implementation, this mapping scheme involves having one inference core of the multi-core inference accelerator engine fetch given data and broadcast the given data to other inference cores of the inference accelerator engine. Each inference core fetches second data unique to the respective inference core. The inference cores then perform computations on the first and second data in order to implement the machine learning model.

Claims (49)

1. A system comprising:

a plurality of computing cores; and

a control unit;

wherein responsive to receipt of data indicative of a type of machine learning model, the control unit comprises circuitry configured to:

retrieve an identification of a preferred bandwidth reduction scheme that is mapped to the machine learning model from stored data that maps a plurality of machine learning models to different bandwidth reduction schemes; and

program a first computing core of the plurality of computing cores to implement the preferred bandwidth reduction scheme, wherein the bandwidth reduction scheme causes:

at least one of the plurality of computing cores to fetch and broadcast first data; and

two or more of the plurality of computing cores to receive the first data broadcast by the at least one of the plurality of computing cores;

wherein each of the two or more computing cores of the plurality of computing cores is configured to:

fetch second data for processing with the first data; and

perform one or more computations using the first data and the second data in order to perform a computing operation.

2. The system as recited in claim 1 , wherein the second data comprises a set of coefficients associated with a corresponding set of filters.

3. The system as recited in claim 1 , wherein the first data comprises input channel data.

4. The system as recited in claim 1 , wherein the first computing core is selected based on an identified machine learning model type from a plurality of identified machine learning model types.

5. The system as recited in claim 1 , wherein the bandwidth reduction scheme comprises an identification of one or more layers of a machine learning model type.

6. The system as recited in claim 1 , wherein the bandwidth reduction scheme comprises an identification of computing cores to which the first computing core broadcasts the first data.

7. The system as recited in claim 5 , wherein the control unit is further configured to determine which portions of the machine learning model to map to the plurality of computing cores based on a number and size of filters of the machine learning model.

8. A method comprising:

responsive to receiving data indicative of a type of machine learning model:

retrieving an identification of a preferred bandwidth reduction scheme that is mapped to the machine learning model from stored data that maps a plurality of machine learning models to different bandwidth reduction schemes;

programming a first computing core of a plurality of computing cores to implement the preferred bandwidth reduction scheme, wherein the bandwidth reduction scheme causes:

at least one of the plurality of computing cores to fetch and broadcast first data; and

two or more of the plurality of computing cores to receive the first data broadcast by the at least one of the plurality of computing cores;

each of the two or more computing cores of the plurality of computing cores:

fetching second data for processing with the first data; and

performing one or more computations using the first data and the second data in order to perform a computing operation.

9. The method as recited in claim 8 , wherein the second data comprises a set of coefficients associated with a corresponding set of filters.

10. The method as recited in claim 8 , wherein the first data comprises input channel data.

11. The method as recited in claim 8 , wherein the first computing core is selected based on an identified memory bandwidth reduction scheme.

12. The method as recited in claim 8 , wherein the bandwidth reduction scheme comprises an identification of one or more layers of a machine learning model type.

13. The method as recited in claim 12 , further comprising determining which portions of the machine learning model to map to the plurality of computing cores based on a size of an input dataset.

14. The method as recited in claim 12 , further comprising determining which portions of the machine learning model to map to the plurality of computing cores based on a number and size of filters of the machine learning model.

15. An apparatus comprising:

a plurality of computing cores;

a control unit comprising circuitry; and

a table for mapping memory bandwidth reduction schemes to a processor core designated to broadcast data to one or more other processor cores;

wherein responsive to receipt of data indicative of a type of machine learning model, the control unit is configured to:

access the table to determine a given memory bandwidth reduction scheme to use in performing an indicated computing operation;

program a first computing core of the plurality of computing cores to implement the given bandwidth reduction scheme, wherein the bandwidth reduction scheme causes:

at least one of the plurality of computing cores to fetch and broadcast first data; and

two or more of the plurality of computing cores to receive the first data broadcast by the at least one of the plurality of computing cores;

wherein each of the two or more computing cores of the plurality of computing cores is configured to:

fetch second data for processing with the first data; and

perform one or more computations using the first data and the second data in order to perform the computing operation.

16. The apparatus as recited in claim 15 , wherein the second data comprises a set of coefficients associated with a corresponding set of filters.

17. The apparatus as recited in claim 15 , wherein the first data comprises input channel data.

18. The apparatus as recited in claim 15 , wherein the one or more computations are performed as part of a machine learning model.

19. The apparatus as recited in claim 18 , wherein the indicated computing operation specifies a type of machine learning model to implement.

20. The apparatus as recited in claim 19 , wherein the control unit is further configured to determine which portions of the machine learning model to map to the plurality of computing cores based on a size of an input dataset.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2018
From: ZHANG, LEI
To: ATI TECHNOLOGIES ULC
Reel/Frame 046753/0780 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 30, 2018
From: LAGUDU, SATEESH; RUSH, ALLEN
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 046753/0885 →
Continuity (2)
Provisional Application 62660817 · Apr 20, 2018
Related Publication 20190325305A1 · Oct 24, 2019
Cited By (5)
US 12,541,525 US 12,602,498 US 12,675,643 US 12,699,847 US 12,705,031