IP Library › Granted Patent US 10,705,967
Granted Patent B2
US 10,705,967 · App. 16/160,270 · Granted Jul 7, 2020

Programmable interface to in-memory cache processor

Inventors: Amrita Mathuriya (Portland, OR); Sasikanth Manipatruni (Portland, OR); Victor Lee (Santa Clara, CA); Huseyin Sumbul (Portland, OR); Gregory Chen (Portland, OR); Raghavan Kumar (Hillsboro, OR); Phil Knag (Hillsboro, OR); Ram Krishnamurthy (Portland, OR); Ian Young (Portland, OR); Abhishek Sharma (Hillsboro, OR)
Assignee: Intel Corporation
G06F12/0875G06F3/0604G06F3/0629G06F3/0673G06F8/41G06F9/45508G06N3/04G06N3/0445G06N3/063G06F2212/251
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,705,967
App. No.
16/160,270
Filed
Oct 15, 2018
Granted
Jul 7, 2020
Kind
B2
Art Unit
2191
USPC
717/140
Abstract

The present disclosure is directed to systems and methods of implementing a neural network using in-memory mathematical operations performed by pipelined SRAM architecture (PISA) circuitry disposed in on-chip processor memory circuitry. A high-level compiler may be provided to compile data representative of a multi-layer neural network model and one or more neural network data inputs from a first high-level programming language to an intermediate domain-specific language (DSL). A low-level compiler may be provided to compile the representative data from the intermediate DSL to multiple instruction sets in accordance with an instruction set architecture (ISA), such that each of the multiple instruction sets corresponds to a single respective layer of the multi-layer neural network model. Each of the multiple instruction sets may be assigned to a respective SRAM array of the PISA circuitry for in-memory execution. Thus, the systems and methods described herein beneficially leverage the on-chip processor memory circuitry to perform a relatively large number of in-memory vector/tensor calculations in furtherance of neural network processing without burdening the processor circuitry.

Claims (41)

1. A processor, comprising:

memory circuitry that includes a plurality of static random access memory (SRAM) arrays;

processor circuitry to:

compile data representative of a multi-layer neural network model and one or more neural network data inputs from a first programming language to an intermediate domain-specific language (DSL); and

compile the representative data from the intermediate DSL to multiple instruction sets in accordance with an instruction set architecture (ISA), wherein each of the multiple instruction sets corresponds to a respective layer of the multi-layer neural network model; and

memory controller circuitry to:

configure multiple of the SRAM arrays within the plurality of SRAM arrays to operate as serially coupled, pipelined SRAM architecture (PISA) circuitry; and

assign each of the multiple instruction sets to a respective one of the multiple SRAM arrays for execution by the respective one SRAM array.

2. The processor of claim 1 , wherein the first programming language is a high-level programming language of a group that includes one or more of Python, C, C++, R, Lisp, Prolog, or Java.

3. The processor of claim 1 , wherein the intermediate DSL includes a data-flow graph of layer descriptors.

4. The processor of claim 1 wherein each of the multiple SRAM arrays comprises a SRAM array having in-memory integer compute capability (C-SRAM).

5. The processor of claim 1 wherein the system comprises a multi-chip module that includes processor circuitry, the memory circuitry, and the neural network control circuitry.

6. The processor of claim 1 , wherein to compile the representative data from the first programming language to the intermediate DSL includes to generate a data-flow graph of layer descriptors representing the multi-layer neural network model.

7. The processor of claim 6 , wherein to compile the representative data from the intermediate DSL to the multiple instruction sets includes to optimize each layer descriptor within the generated data-flow graph.

8. A non-transitory machine-readable storage device having instructions that, when executed by processing circuitry, cause the processing circuitry to perform neural network processing operations comprising to:

receive, in a first programming language, data representative of a multi-layer neural network model and one or more neural network data inputs;

compile, via the processor circuitry, the data representative of the multi-layer neural network model and the one or more neural network data inputs from the first programming language to an intermediate domain-specific language (DSL);

compile, via the processor circuitry, the data representative of the multi-layer neural network model and the one or more neural network data inputs from the intermediate DSL to multiple instruction sets, each of the multiple instruction sets being in accordance with an instruction set architecture (ISA) and corresponding to a respective layer of the multi-layer neural network model;

configure, via the processor circuitry and memory circuitry that is coupled to the processor circuitry, multiple static random access memory (SRAM) arrays of a plurality of SRAM arrays within the memory circuitry to operate as serially coupled pipelined SRAM architecture (PISA) circuitry; and

assign, via the processor circuitry, each of the multiple instruction sets to a respective one of the multiple SRAM arrays.

9. The non-transitory machine-readable storage device of claim 8 , wherein to compile the representative data from the intermediate DSL to the multiple instruction sets includes using a software stack specific to an in-memory compute accelerator.

10. The non-transitory machine-readable storage device of claim 8 further comprising initiating execution of each of the multiple instruction sets by the respective one of the multiple SRAM arrays.

11. The non-transitory machine-readable storage device of claim 8 wherein each of the multiple SRAM arrays comprises an SRAM array having in-memory integer compute capability (C-SRAM).

12. The non-transitory machine-readable storage device of claim 8 wherein the method is performed via a multi-chip module that includes the processor circuitry and the memory circuitry.

13. The non-transitory machine-readable storage device of claim 8 wherein to compile the representative data from the first programming language to the intermediate DSL includes to generate a data-flow graph of layer descriptors.

14. The non-transitory machine-readable storage device of claim 13 wherein to compile the representative data from the intermediate DSL to the multiple instruction sets includes to optimize each layer descriptor within the generated data-flow graph.

15. A system, comprising:

system memory to store data representative of a multi-layer neural network model and one or more neural network data inputs;

processor circuitry coupled to the system memory, the processor circuitry comprising processor memory circuitry that includes a plurality of static random access memory (SRAM) arrays;

wherein the processor circuitry is to:

compile the data representative of the multi-layer neural network model and the one or more neural network data inputs from a first programming language to an intermediate domain-specific language (DSL);

compile the representative data from the intermediate DSL to multiple instruction sets in accordance with an instruction set architecture (ISA), wherein each of the multiple instruction sets corresponds to a respective layer of the multi-layer neural network model;

configure multiple of the SRAM arrays within the plurality of SRAM arrays to operate as serially coupled, pipelined SRAM architecture (PISA) circuitry; and

assign each of the multiple instruction sets to a respective one of the multiple SRAM arrays for execution by the respective one SRAM array; and

PISA memory circuitry coupled to the system memory and coupled to the processor via the processor memory circuitry.

16. The system of claim 15 , wherein the first programming language is a high-level programming language of a group that includes one or more of Python, C, C++, R, Lisp, Prolog, or Java.

17. The system of claim 15 , wherein the intermediate DSL includes a data-flow graph of layer descriptors.

18. The system of claim 15 wherein each of the multiple SRAM arrays comprises a SRAM array having in-memory integer compute capability (C-SRAM).

19. The system of claim 15 wherein the system comprises a multi-chip module that includes the processor circuitry and the processor memory circuitry.

20. The system of claim 15 , wherein to compile the representative data from the first programming language to the intermediate DSL includes to generate a data-flow graph of layer descriptors representing the multi-layer neural network model.

21. The system of claim 20 , wherein to compile the representative data from the intermediate DSL to the multiple instruction sets includes to optimize each layer descriptor within the generated data-flow graph.

Assignments (2)
CORRECTIVE ASSIGNMENT TO CORRECT THE CORRECT TENTH INVENTORS NAME PREVIOUSLY RECORDED AT REEL: 048576 FRAME: 0815. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Jun 15, 2020
From: MATHURIYA, AMRITA; MANIPATRUNI, SASIKANTH; LEE, VICTOR; SUMBUL, HUSEYIN; CHEN, GREGORY; KUMAR, RAGHAVAN; KNAG, PHIL; KRISHNAMURTHY, RAM; YOUNG, IAN; SHARMA, ABHISHEK
To: INTEL CORPORATION
Reel/Frame 052945/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2019
From: MATHURIYA, AMRITA; MANIPATRUNI, SASIKANTH; LEE, VICTOR; SUMBUL, HUSEYIN; CHEN, GREGORY; KUMAR, RAGHAVAN; KNAG, PHIL; KRISHNAMURTHY, RAM; YOUNG, IAN; ABIDSHEK, SHARMA
To: INTEL CORPORATION
Reel/Frame 048576/0815 →
Continuity (1)
Related Publication 20190057036A1 · Feb 21, 2019