IP Library Granted Patent US 11,403,519
Granted Patent B2
US 11,403,519 · App. 16/841,601 · Granted Aug 2, 2022

Machine learning network implemented by statically scheduled instructions, with system-on-chip

Inventors: Nishit Shah (Sunnyvale, CA); Reed Kotler (San Jose, CA); Srivathsa Dhruvanarayan (Saratoga, CA); Moenes Zaher Iskarous (San Jose, CA); Kavitha Prasad (San Jose, CA); Yogesh Laxmikant Chobe (Santa Clara, CA); Sedny S. J Attia (Santa Cruz, CA); Spenser Don Gilliland (San Jose, CA)
Assignee: SiMa Technologies, Inc.
G06N3/063G06F8/41G06F9/30087G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,403,519
App. No.
16/841,601
Granted
Aug 2, 2022
Kind
B2
Abstract

A compiler receives a description of a machine learning network and generates a computer program that implements the machine learning network. The computer program includes statically scheduled instructions that are executed by a mesh of processing elements (Tiles). The instructions executed by the Tiles are statically scheduled because the compiler can determine which instructions are executed by which Tiles at what times. For example, for the statically scheduled instructions, there are no conditions, branching or data dependencies that can be resolved only at run-time, and which would affect the timing and order of the execution of the instructions.

Claims (29)

1. A device that performs a task using a machine learning network, the device comprising:

a machine learning accelerator (MLA) comprising a plurality of Tiles and associated on-chip memory on a semiconductor die, wherein the Tiles are organized into one or more meshes of interconnected Tiles; and

non-transitory data storage that is not on the same semiconductor die as the Tiles and on-chip memory but is coupled to the MLA, the data storage storing an executable computer program that implements the machine learning network on the MLA, the computer program comprising deterministic phases of Tile instructions that implement computations in the machine learning network and non-deterministic phases that include instructions for data transfer between the on-chip memory and the non-transitory data storage; wherein the Tile instructions comprise (a) Tile compute instructions for performing computations in the machine learning network using data stored in the on-chip memory, and (b) Tile data transfer instructions for data transfer along data transfer paths between different Tiles and between Tiles and the on-chip memory;

wherein each deterministic phase utilizes multiple Tiles with concurrent execution of Tile compute instructions and Tile data transfer instructions by different Tiles and the Tiles are configured to execute the Tile instructions within each deterministic phase according to a static schedule relative to the other Tile instructions in the same deterministic phase with unconditional start times for every Tile instruction within the same deterministic phase so that the Tiles execute the Tile instructions without any run-time determination of whether the data, Tiles or data transfer paths required for the Tile instructions are available; and

a controller external to the mesh of Tiles, the controller configured to execute the instructions in the non-deterministic phases.

2. The device of claim 1 wherein the memory in the MLA comprises SRAM, and the non-transitory data storage comprises DRAM.

3. The device of claim 1 further comprising:

a pipeline of one or more additional programmable processors that perform the task by executing instructions from the computer program, wherein the pipeline includes the MLA.

4. The device of claim 3 wherein the computer program configures the programmable processors into the pipeline.

5. The device of claim 3 wherein the pipeline also includes a general purpose CPU.

6. The device of claim 3 wherein the pipeline also includes an application-specific processor.

7. The device of claim 3 wherein all of the programmable processors in the pipeline are on the same semiconductor die as the MLA.

8. The device of claim 3 further comprising a master controller that coordinates operation of the MLA and the programmable processors in the pipeline.

9. The device of claim 3 further comprising one or more sensors that provide input samples to the pipeline.

10. The device of claim 9 wherein the input samples comprises images captured by the one or more sensors.

11. The device of claim 3 wherein the pipeline of programmable processors performs the task without utilizing remote compute resources.

12. The device of claim 1 wherein:

the non-transitory data storage stores multiple executable computer programs for performing different tasks, each executable computer program comprising Tile instructions that implement computations in one of a plurality of corresponding machine learning networks used to perform the different tasks; and

the Tiles are partitioned to simultaneously implement the corresponding machine learning networks for the different tasks.

13. The device of claim 1 further comprising:

additional programmable processors;

wherein the data storage stores multiple different executable computer programs, each computer program configuring some or all of the programmable processors and the MLA as different pipelines to perform different tasks.

14. The device of claim 1 wherein the task is a real-time task.

15. The device of claim 1 wherein the task is one of an agricultural task, an industrial task, a surveillance task and a robotics task.

16. The device of claim 1 wherein the task is one of a computer vision task, an image analysis task, an image understanding task, a speech recognition task, an audio analysis task, an audio understanding task, and a natural language processing task.

17. The device of claim 1 wherein the task is one of a health monitoring task, and a personalized health task.

18. The device of claim 1 wherein the device is an edge device.

19. The device of claim 1 wherein the device is a portable device.

20. The device of claim 1 wherein the MLA implements computations in the machine learning network at a speed of at least 50 trillion operations per second at a power consumption of not more than 5 watts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 7, 2020
From: SHAH, NISHIT; KOTLER, REED; DHRUVANARAYAN, SRIVATHSA; ISKAROUS, MOENES ZAHER; PRASAD, KAVITHA; CHOBE, YOGESH LAXMIKANT; ATTIA, SEDNY S.J; GILLILAND, SPENSER DON
To: SIMA TECHNOLOGIES, INC.
Reel/Frame 052337/0001 →
Continuity (2)
Continuation 16840216 · Apr 3, 2020
Related Publication 20210312322A1 · Oct 7, 2021
Cited By (2)
US 12,699,552 US 12,717,375