IP Library Granted Patent US 9,858,220
Granted Patent B2
US 9,858,220 · App. 14/660,589 · Granted Jan 2, 2018

Computing architecture with concurrent programmable data co-processor

Inventors: Eugenio Culurciello (Lafayette, IN); Berin Eduard Martini (Lafayette, IN); Vinayak Anand Gokhale (West Lafayette, IN); Jonghoon Jin (West Lafayette, IN); Aysegul Dundar (Lafayette, IN)
Assignee: Purdue Research Foundation
G06F13/28G06F9/4411G06F9/46G06F12/023G06N3/063G06F2212/261Y02B60/1228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,858,220
App. No.
14/660,589
Granted
Jan 2, 2018
Kind
B2
Abstract

A coprocessor (PL) is disclosed. The PL includes a memory router, at least one collection block that is configured to transfer data to/from the memory router, each collection block includes a collection router that is configured to i) transfer data to/from the memory router, ii) transfer data to/from at least one collection router of a neighboring collection block, and iii) transfer data to/from blocks within the collection block, and at least one programmable operator that is configured to i) transfer data to/from the collection router, and ii) perform a programmable operation on data received from the collection router.

Claims (26)

1. A coprocessor (PL) unit, comprising:

a memory router configured to i) transfer data to/from an external memory device, the transfer of data being initiated by an external processing system (PS) and ii) distribute the data to a plurality of blocks within the PL unit; and

at least one collection block configured to transfer data to/from the memory router, each collection block including

a collection router configured to i) transfer data to/from the memory router, ii) transfer data to/from at least one collection router of a neighboring collection block, and iii) transfer data to/from blocks within the collection block,

at least one programmable operator configured to i) transfer data to/from the collection router, and ii) perform a programmable operation on data received from the collection router,

the PL unit configured to perform programmable operations on data transferred from the external memory and provide the operated-on data to the external memory with substantially zero overhead to the PS.

2. The PL unit of claim 1 , the at least one collection block further comprising:

at least one multiply-accumulator (MAC) block configured to i) transfer data to/from the collection router, ii) transfer data to/from the at least one programmable operator, and iii) perform multiply and accumulate operations on data received from the collection router and/or the at least one programmable operator.

3. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes max-pooling operations.

4. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes pixel-wise subtraction operations.

5. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes pixel-wise addition operations.

6. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes pixel-wise multiplication operations.

7. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes pixel-wise division operations.

8. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes non-linear operations.

9. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes MAC operations.

10. The PL unit of claim 1 , the programmable operation performed by the programmable operator includes non-linear operations.

11. The PL unit of claim 1 , the at least one collection unit is implemented based on field programmable gate array technology.

12. The PL unit of claim 1 , the at least one collection unit is implemented based on field programmable gate array (FPGA) technology.

13. The PL unit of claim 1 , the at least one collection unit is implemented based on application specific integrated circuit (ASIC) technology.

14. The PL unit of claim 1 , the PL unit includes a plurality of collection units.

15. The PL unit of claim 1 , the PL unit includes at least 50 collections units.

16. The PL unit of claim 1 , the PL unit includes at least 5 collections units.

17. The PL unit of claim 1 , the PL unit includes at least 5 collections units.

18. The PL unit of claim 2 , the collection router is further configured to transfer data to/from the at least one MAC.

19. The PL unit of claim 10 , the non-linear operations include applying a sigmoid function to a series.

20. The PL unit of claim 18 , the at least one programmable operator further configured to perform a programmable operation on data received from the at least one MAC.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2017
From: CULURCIELLO, EUGENIO; MARTINI, BERIN EDUARD; GOKHALE, VINAYAK ANAND; JIN, JONGHOON; DUNDAR, AYSEGUL
To: PURDUE RESEARCH FOUNDATION
Reel/Frame 044193/0515 →
Continuity (2)
Provisional Application 61954544 · Mar 17, 2014
Related Publication 20150261702A1 · Sep 17, 2015