IP Library Granted Patent US 8,442,927
Granted Patent B2
US 8,442,927 · App. 12/697,823 · Granted May 14, 2013

Dynamically configurable, multi-ported co-processor for convolutional neural networks

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,442,927
App. No.
12/697,823
Granted
May 14, 2013
Kind
B2
Abstract

A coprocessor and method for processing convolutional neural networks includes a configurable input switch coupled to an input. A plurality of convolver elements are enabled in accordance with the input switch. An output switch is configured to receive outputs from the set of convolver elements to provide data to output branches. A controller is configured to provide control signals to the input switch and the output switch such that the set of convolver elements are rendered active and a number of output branches are selected for a given cycle in accordance with the control signals.

Claims (10)

1. A method for processing convolutional neural networks (CNN), comprising:

determining a workload for a convolutional neural networks;

generating a control signal based upon the workload to permit a selection of a type of parallelism to be employed in processing a layer of the CNN;

configuring an input switch to enable a number of convolvers which convolve an input in accordance with the control signal;

configuring an output switch to enable a number of output branches for a given cycle in accordance with the control signal; and

processing outputs from the output branches, said processing including reconfiguring the input switch and the output switch in accordance with a next layer of the CNN to be processed.

2. The method as recited in claim 1 , further comprising selecting the number of convolvers and the number of output branches based upon a type of parallelism needed for a given neural network layer.

3. The method as recited in claim 2 , wherein the type of parallelism includes one of parallelism within a convolution operation, inter-output parallelism and intra-output parallelism.

4. The method as recited in claim 1 , further comprising providing a memory subsystem having at least two banks, wherein a first bank provides input storage and a second bank provides output storage to enable a stateless, streaming coprocessor architecture.

5. The method as recited in claim 4 , further comprising providing a third memory bank for storing intermediate results.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE REMOVE 8538896 AND ADD 8583896 PREVIOUSLY RECORDED ON REEL 031998 FRAME 0667. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 30, 2017
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 042754/0703 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2014
From: NEC LABORATORIES AMERICA, INC.
To: NEC CORPORATION
Reel/Frame 031998/0667 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 1, 2010
From: CHAKRADHAR, SRIMAT; SANKARADAS, MURUGAN; JAKKULA, VENKATA S.; CADAMBI, SRIHARI
To: NEC LABORATORIES AMERICA, INC.
Reel/Frame 023879/0823 →