IP Library Granted Patent US 10,902,315
Granted Patent B2
US 10,902,315 · App. 15/600,806 · Granted Jan 26, 2021

Device for implementing artificial neural network with separate computation units

Inventors: Shaoxia Fang (Beijing, CN); Lingzhi Sui (Beijing, CN); Qian Yu (Beijing, CN); Junbin Wang (Beijing, CN); Yi Shan (Beijing, CN)
Assignee: XILINX, INC.
G06N3/063G06F7/5443G06N3/0454G06N3/08G06F2207/4824
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,902,315
App. No.
15/600,806
Granted
Jan 26, 2021
Kind
B2
Abstract

The present disclosure relates to a processor for implementing artificial neural networks, for example, convolutional neural networks. The processor includes a memory controller group, an on-chip bus and a processor core, wherein the processor core further includes a register map, an instruction module, a data transferring controller, a data writing scheduling unit, a buffer module, a convolution operation unit and a hybrid computation unit. The processor of the present disclosure may be used for implementing various neural networks with increased computation efficiency.

Claims (37)

1. A processor for implementing an artificial neural network, comprising:

a memory controller group, which includes one or more memory controller, wherein each memory controller is configured for accessing a corresponding external storage chip, said external storage chip being configured for storing neural network data and instructions;

an on-chip bus, configured for communicating between the memory controller group and a processor core array; and

the processor core array, which includes one or more processor core, wherein each processor core further comprises:

a register map, configured for configuring operation parameters of the processor core and obtaining operation status of the processor core;

an instruction module, configured for obtaining and decoding instructions stored in the external storage chip;

a data transferring controller, configured for writing neural network data received from the external storage chip into a data writing scheduling unit based on the decoded result of the instruction module, and for writing computation results of one or more convolution operation units and a hybrid computation unit back to the external storage chip;

a buffer module, configured for storing the neural network data and the computation results, said computation result including intermediate computation result and final computation result;

one or more convolution operation units, each of which being configured for performing convolution operation and obtaining convolution operation results; and

a hybrid computation unit, configured for performing hybrid computation and obtaining hybrid computation results.

2. The processor according to claim 1 , wherein the convolution operation unit further comprises:

a multiplier array, configured for performing multiplication operations and obtaining multiplication operation results;

an adder tree, which is coupled to the multiplier array and is configured for summing the multiplication operation results; and

a non-linear operation array, which is coupled to the adder tree and is configured for applying a non-linear function to the output of the adder tree.

3. The processor according to claim 1 , wherein the hybrid computation unit further comprises computation units for performing pooling operation, element-wise operation, resizing operation, or full connected operation.

4. The processor according to claim 1 , wherein the buffer module further comprises:

a buffer pool, configured for storing the neural network data and the computation results of the convolution operation unit and the hybrid computation unit, said computation result including intermediate computation result and final computation result;

the data writing scheduling unit, configured for writing the neural network data and the computation results of the convolution operation unit and the hybrid computation unit into the buffer pool;

a data reading scheduling unit, configured for reading data needed for computation and the computation results of the convolution operation unit and the hybrid computation unit from the buffer pool.

5. The processor according to claim 4 , wherein the buffer pool further comprises one or more blocks of buffer.

6. The processor according to claim 4 , wherein the data writing scheduling unit further comprises:

one or more writing scheduling channel, each writing scheduling channel communicating with the output of a corresponding computation unit of said one or more convolution operation units and the hybrid computation unit;

a writing arbitration unit, configured for ranking the priority among said one or more writing scheduling channels, so as to schedule writing operations of computation results from said one or more convolution operation units and the hybrid computation unit.

7. The processor according to claim 4 , wherein the data reading scheduling unit further comprises:

one or more reading scheduling channel, each reading scheduling channel communicating with the input of a corresponding computation unit of said one or more convolution operation units and the hybrid computation unit;

a reading arbitration unit, configured for ranking the priority among said one or more reading scheduling channel, so as to schedule reading operations into input of said one or more convolution operation units and the hybrid computation unit.

8. The processor according to claim 1 , wherein the instruction module further comprises:

a first instruction unit, configured for obtaining and decoding instructions stored in the external storage chip;

a second instruction unit, configured for obtaining and decoding instructions stored in the external storage chip; and

an instruction distributing unit, configured for selectively launching one of the first instruction unit and the second instruction unit, and obtaining the decoded result of said one of the first instruction unit and the second instruction unit.

9. The processor according to claim 8 , wherein the first instruction unit further comprises:

a first instruction obtaining unit, configured for obtaining instructions stored in the external storage chip; and

a first instruction decoding unit, configured for decoding the instructions obtained by the first instruction obtaining unit.

10. The processor according to claim 8 , wherein the second instruction unit further comprises:

a second instruction obtaining unit, configured for obtaining instructions stored in the external storage chip; and

a second instruction decoding unit, configured for decoding the instructions obtained by the second instruction obtaining unit.

11. The processor according to claim 8 , wherein the instruction distributing unit further parses the decoded result of the first instruction unit or the second instruction unit.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 16, 2019
From: BEIJING DEEPHI INTELLIGENT TECHNOLOGY CO., LTD.
To: XILINX, INC.
Reel/Frame 050377/0436 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE'S NAME PREVIOUSLY RECORDED AT REEL: 044346 FRAME: 0121. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Mar 8, 2018
From: SUI, LINGZHI; FANG, SHAOXIA; YU, QIAN; WANG, JUNBIN; SHAN, YI
To: BEIJING DEEPHI INTELLIGENT TECHNOLOGY CO., LTD.
Reel/Frame 045529/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 11, 2017
From: FANG, SHAOXIA; SUI, LINGZHI; YU, QIAN; WANG, JUNBIN; SHAN, YI
To: BEIJING DEEPHI INTELLIGENCE TECHNOLOGY CO., LTD
Reel/Frame 044346/0121 →
Priority Claims (1)
CN 2017 1 0258133 · Apr 19, 2017 · national
Continuity (1)
Related Publication 20180307976A1 · Oct 25, 2018
Cited By (3)
US 12,442,074 US 12,443,832 US 12,718,062