IP Library Granted Patent US 10,891,136
Granted Patent B1
US 10,891,136 · App. 16/420,103 · Granted Jan 12, 2021

Data transmission between memory and on chip memory of inference engine for machine learning via a single data gathering instruction

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,891,136
App. No.
16/420,103
Granted
Jan 12, 2021
Kind
B1
Abstract

A system to support data gathering for a machine learning (ML) operation comprises a memory unit configured to maintain data for the ML operation in a plurality of memory blocks each accessible via a memory address. The system further comprises an inference engine comprising a plurality of processing tiles each comprising one or more of an on-chip memory (OCM) configured to load and maintain data for local access by components in the processing tile. The system also comprises a core configured to program components of the processing tiles of the inference engine according to an instruction set architecture (ISA) and a data streaming engine configured to stream data between the memory unit and the OCMs of the processing tiles of the inference engine wherein data streaming engine is configured to perform a data gathering operation via a single data gathering instruction of the ISA at the same time.

Claims (48)

1. A system to support data gathering for a machine learning (ML) operation, comprising:

a memory unit configured to maintain data for the ML operation, wherein the memory unit includes a plurality of memory blocks each accessible via a memory address;

an inference engine comprising a plurality of processing tiles, wherein each processing tile comprises at least:

an on-chip memory (OCM) configured to load and maintain data for local access by components in the processing tile; and

one or more processing units configured to perform one or more computation tasks of the ML operation on the data in the OCM;

a core configured to

program components of the plurality of processing tiles of the inference engine by translating one or more commands from a host into a set of programming instructions for the ML operation according to an instruction set architecture (ISA) designed for data processing in a data-path; and

specify one or more processing tiles via a programming instruction, wherein the programming instruction identifies the one or more OCMs of the one or more processing tiles to have data written into;

a programmable processor configured to stream data between the memory unit and the OCMs of the plurality of processing tiles of the inference engine wherein the programmable processor is configured to perform a data gathering operation via a single data gathering instruction of the ISA to

gather data from one or more memory blocks of the plurality of memory blocks of the memory unit for the ML operation at the same time; and

write the gathered data into the OCM of each of the specified one or more processing tiles for the one or more processing units of the specified one or more processing tiles to perform the one or more computation tasks of the ML operation.

2. The system of claim 1 , wherein:

the memory unit is a double data rate (DDR) memory.

3. The system of claim 2 , wherein:

the programmable processor is configured to write the gathered data into the OCM of each of the specified one or more processing tiles via DDR to OCM direct memory access (DMA).

4. The system of claim 1 , wherein:

the core is configured to maintain the memory addresses to the one or more memory blocks of the plurality of memory blocks of the memory unit in an array to create a two-level indirect reference to the entire list of the one or more memory blocks from which the data is to be gathered in the single data gathering instruction irrespective of a number of the one or more memory blocks from which the data is to be gathered.

5. The system of claim 4 , wherein:

the array of the memory addresses to the one or more memory blocks from which data is to be gathered is contained in the memory unit.

6. The system of claim 5 , wherein:

the core is configured to provide address of the array of the memory addresses to the one or more memory blocks in the memory unit from which data is to be gathered as an input parameter to the single data gathering instruction.

7. The system of claim 6 , wherein:

the single data gathering instruction further takes as its input parameters, in addition to the address of the array of addresses number of pointers to the one or more memory blocks from which data is to be gathered.

8. The system of claim 6 , wherein:

the single data gathering instruction further takes as its input parameters, length of a line in each memory block to be loaded by each pointer to the memory block.

9. The system of claim 6 , wherein:

the single data gathering instruction further takes, as its input parameters, address of the OCM in each of the specified one or more processing tiles into which the data gathered from the memory unit is to be written into sequentially.

10. The system of claim 6 , wherein:

the single data gathering instruction further takes, as its input parameters, a Boolean indicator, which signifies if the data being transferred from the memory unit to the OCMs of the specified one or more processing tiles is either signed or unsigned.

11. A method to support data gathering for a machine learning (ML) operation, comprising:

maintaining data in a memory unit for the ML operation, wherein the memory unit includes a plurality of memory blocks each accessible via a memory address;

programming components of a plurality of processing tiles of an inference engine by translating one or more commands from a host into a set of programming instructions for the ML operation according to an instruction set architecture (ISA) designed for data processing in a data-path, wherein each processing tile comprises at least an on-chip memory (OCM) configured to load and maintain data for local access by components in the processing tile and one or more processing units configured to perform one or more computation tasks of the ML operation on the data in the OCM;

specifying one or more processing tiles via a programming instruction, wherein the programming instruction identifies the one or more OCMs of the one or more processing tiles to have data written into; and

streaming data between the memory unit and the OCMs of the plurality of processing tiles of the inference engine by performing a data gathering operation via a single data gathering instruction of the ISA to

gather data from one or more memory blocks of the plurality of memory blocks of the memory unit for the ML operation at the same time; and

write the gathered data into the OCM of each of the specified one or more processing tiles for the one or more processing units of the specified one or more processing tiles to perform the one or more computation tasks of the ML operation.

12. The method of claim 11 , wherein:

the memory unit is a double data rate (DDR) memory.

13. The method of claim 12 , wherein:

the writing the gathered data into the OCM of each of the specified one or more processing tiles is via DDR to OCM direct memory access (DMA).

14. The method of claim 11 , further comprising:

maintaining the memory addresses to the one or more memory blocks of the plurality of memory blocks of the memory unit in an array to create a two-level indirect reference to the entire list of the one or more memory blocks from which the data is to be gathered in the single data gathering instruction irrespective of the number of the one or more memory blocks from which the data is to be gathered.

15. The method of claim 14 , further comprising:

containing the array of the memory addresses to the one or more memory blocks from which data is to be gathered in the memory unit.

16. The method of claim 15 , further comprising:

providing address of the array of the memory addresses to the one or more memory blocks from which data is to be gathered in the memory unit as an input parameter to the single data gathering instruction.

17. The method of claim 16 , further comprising:

including, in addition to the address of the array of addresses, as input parameters to the single data gathering instruction one or more of: number of pointers to the one or more memory blocks from which data is to be gathered, length of a line in each memory block to be loaded by each pointer to the memory block, address of the OCM in each of the specified one or more processing tiles into which the data gathered from the memory unit is to be written into sequentially, and a Boolean indicator, which signifies if the data being transferred from the memory unit to the OCMs of the specified one or more processing tiles is either signed or unsigned.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2021
From: CAVIUM INTERNATIONAL
To: MARVELL ASIA PTE, LTD.
Reel/Frame 055334/0579 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 18, 2021
From: MARVELL INTERNATIONAL LTD.
To: CAVIUM INTERNATIONAL
Reel/Frame 055334/0589 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: SODANI, AVINASH
To: MARVELL SEMICONDUCTOR, INC.
Reel/Frame 054540/0302 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 4, 2020
From: MARVELL SEMICONDUCTOR, INC.
To: MARVELL INTERNATIONAL LTD.
Reel/Frame 054540/0332 →