IP Library Granted Patent US 10,282,659
Granted Patent B2
US 10,282,659 · App. 15/600,807 · Granted May 7, 2019

Device for implementing artificial neural network with multiple instruction units

Inventors: Shaoxia Fang (Beijing, CN); Lingzhi Sui (Beijing, CN); Qian Yu (Beijing, CN); Junbin Wang (Beijing, CN); Yi Shan (Beijing, CN)
Assignee: BEIJING DEEPHI INTELLIGENT TECHNOLOGY CO., LTD.
G06N3/063G06F7/5443G06F17/153G06N3/049G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,282,659
App. No.
15/600,807
Granted
May 7, 2019
Kind
B2
Abstract

The present disclosure relates to a processor for implementing artificial neural networks, for example, convolutional neural networks. The processor includes a memory controller group, an on-chip bus and a processor core, wherein the processor core further includes a register map, a first instruction unit, a second instruction unit, an instruction distributing unit, a data transferring controller, a buffer module and a computation module. The processor of the present disclosure may be used for implementing various neural networks with increased computation efficiency.

Claims (38)

1. A processor for implementing an artificial neural network, comprising:

a memory controller group, which includes one or more memory controller, wherein each memory controller is configured for accessing a corresponding external storage chip, said external storage chip being configured for storing neural network data and instructions;

an on-chip bus, configured for communicating between the memory controller group and a processor core array; and

a processor core array, which includes one or more processor core, wherein each processor core further comprises:

a register map, configured for configuring operation parameters of the processor core and obtaining operation status of the processor core;

a first instruction unit, configured for obtaining and decoding instructions stored in the external storage chip;

a second instruction unit, configured for obtaining and decoding instructions stored in the external storage chip;

an instruction distributing unit, configured for selectively launching one of the first instruction unit and the second instruction unit, and obtaining the decoded result of said one of the first instruction unit and the second instruction unit;

a data transferring controller, configured for writing the neural network data received from the external storage chip into a data writing scheduling unit based on the decoded result, and for writing computation results of a computation module back to the external storage chip;

a buffer module, configured for storing the neural network data and the computation results of the computation module, said computation result including intermediate computation result and final computation result; and

a computation module, which includes one or more computation units, each configured for performing operations of the neural network.

2. The processor according to claim 1 , wherein the first instruction unit further comprises:

a first instruction obtaining unit, configured for obtaining instructions stored in the external storage chip; and

a first instruction decoding unit, configured for decoding the instructions obtained by the first instruction obtaining unit.

3. The processor according to claim 1 , wherein the second instruction unit further comprises:

a second instruction obtaining unit, configured for obtaining instructions stored in the external storage chip; and

a second instruction decoding unit, configured for decoding the instructions obtained by the second instruction obtaining unit.

4. The processor according to claim 1 , wherein the instruction distributing unit further parses the decoded result of the first instruction unit or the second instruction unit.

5. The processor according to claim 1 , wherein when the executing process of one of the first instruction unit and the second instruction unit has been delayed, the instruction distributing unit stops said one of the first instruction unit and the second instruction unit and initiates the other of the first instruction unit and the second instruction unit.

6. The processor according to claim 1 , wherein the computation module further comprises:

a convolution operation unit array, which comprises a plurality of convolution operation units, each of which being configured for performing convolution operation and obtaining convolution operation results; and

a hybrid computation unit, configured for performing hybrid computation and obtaining hybrid computation results.

7. The processor according to claim 6 , wherein the convolution operation unit further comprises:

a multiplier array, configured for performing multiplication operations and obtaining multiplication operation results;

an adder tree, coupled to the multiplier array and configured for summing the multiplication operation results; and

a non-linear operation array, coupled to the adder tree and configured for applying a non-linear function to the output of adder tree.

8. The processor according to claim 6 , wherein the hybrid computation unit further comprises computation units for performing pooling operation, element-wise operation, resizing operation, full connected operation.

9. The processor according to claim 1 , wherein the buffer module further comprises:

a buffer pool, configured for storing the neural network data and the computation results of the computation module, said computation result including intermediate computation result and final computation result;

a data writing scheduling unit, configured for writing the neural network data and the computation results of the computation module into the buffer pool;

a data reading scheduling unit, configured for reading data needed for computation and the computation results of the computation module from the buffer pool.

10. The processor according to claim 9 , wherein the buffer pool further comprises one or more blocks of buffer.

11. The processor according to claim 9 , wherein the data writing scheduling unit further comprises:

one or more writing scheduling channel, each writing scheduling channel communicating with the output of a corresponding computation unit of said one or more computation units;

a writing arbitration unit, configured for ranking the priority among said one or more writing scheduling channels, so as to schedule writing operations of computation results from said one or more computation units.

12. The processor according to claim 9 , wherein the data reading scheduling unit further comprises:

one or more reading scheduling channel, each reading scheduling channel communicating with the input of a corresponding computation unit of said one or more computation units;

a reading arbitration unit, configured for ranking the priority among said one or more reading scheduling channel, so as to schedule reading operations into the input of said one or more computation units.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2019
From: BEIJING DEEPHI INTELLIGENT TECHNOLOGY CO., LTD.
To: XILINX, INC.
Reel/Frame 050377/0436 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S NAME PREVIOUSLY RECORDED AT REEL: 044346 FRAME: 0217. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 8, 2018
From: FANG, SHAOXIA; SUI, LINGZHI; YU, QIAN; WANG, JUNBIN; SHAN, YI
To: BEIJING DEEPHI INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 045529/0157 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2017
From: FANG, SHAOXIA; SUI, LINGZHI; YU, QIAN; WANG, JUNBIN; SHAN, YI
To: BEIJING DEEPHI INTELLIGENCE TECHNOLOGY CO., LTD.
Reel/Frame 044346/0217 →
Priority Claims (1)
CN 2017 1 0258566 · Apr 19, 2017 · national
Continuity (1)
Related Publication 20180307974A1 · Oct 25, 2018
Cited By (2)
US 12,442,074 US 12,443,832