IP Library › Granted Patent US 12,190,224
Granted Patent B2
US 12,190,224 · App. 17/136,744 · Granted Jan 7, 2025

Processing elements array that includes delay queues between processing elements to hold shared data

Inventors: Yao-Hua Chen (Changhua County, TW); Yu-Xiang Yen (Chiayi County, TW); Wan-Shan Hsieh (Taoyuan, TW); Chih-Tsun Huang (Hsinchu, TW); Juin-Ming Lu (Hsinchu, TW); Jing-Jia Liou (Hsinchu, TW)
Assignee: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
G06N3/049G06F9/4881G06F15/8023G06F15/8046G06N3/063G06F9/5027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,190,224
App. No.
17/136,744
Granted
Jan 7, 2025
Kind
B2
Abstract

A processing element architecture adapted to a convolution comprises a plurality of processing elements and a delayed queue circuit. The plurality of processing elements includes a first processing element and a second processing element, wherein the first processing element and the second processing element perform the convolution according to a shared datum at least. The delayed queue circuit connects to the first processing element and connects to the second processing element. The delayed queue circuit receives the shared datum sent by the first processing element, and sends the shared datum to the second processing element after receiving the shared datum and waiting for a time interval.

Claims (13)

1. A processing element cluster adapted to a convolution comprising:

a first processing element set comprising a plurality of first processing elements;

a second processing element set comprising a plurality of second processing elements;

a bus connecting to the first processing element set and the second processing element set, the bus provides a plurality of shared data to each of the plurality of first processing elements; and

a plurality of delayed queue circuits, wherein one of the plurality of delayed queue circuits connects to one of the plurality of first processing elements and connects to one of the plurality of second processing elements; another one of the plurality of delayed queue circuits connects to two of the plurality of second processing elements, and each of the plurality of delayed queue circuits sends one of the plurality of shared data; wherein

each of the plurality of first processing elements of the first processing element set comprises a storage device storing said one of the plurality of shared data; and

each of the plurality of second processing elements of the second processing element set does not comprises the storage device storing said one of the plurality of shared data.

2. The processing element cluster of claim 1 , wherein the storage device is a first storage device, and each of the plurality of first processing elements and the plurality of second processing elements further comprises:

a second storage device storing a private datum; and

a computing circuit electrically connecting to the first storage device and the second storage device, wherein the computing circuit performs the convolution according to said one of the plurality of shared data and the private datum.

3. The processing element cluster of claim 1 , wherein

the first processing element set and the second processing element set form a two-dimensional array with M rows and N columns, each of the M rows has one of the plurality of first processing elements and (N−1) of the plurality of second processing elements; and

the plurality of delayed queue circuits is divided into M sets and each of the M sets has (N−1) delayed queue circuits.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 10, 2021
From: CHEN, YAO-HUA; YEN, YU-XIANG; HSIEH, WAN-SHAN; HUANG, CHIH-TSUN; LU, JUIN-MING; LIOU, JING-JIA
To: INDUSTRIAL TECHNOLOGY RESEARCH INSTITUTE
Reel/Frame 055552/0913 →
Continuity (1)
Related Publication 20220207323A1 · Jun 30, 2022
References Cited (28)
US 9710748B2 · Ross et al. · 2017 [cited by applicant]
US 9747546B2 · Ross et al. · 2017 [cited by applicant]
US 10467501B2 · N et al. · 2019 [cited by applicant]
US 10586148B2 · Henry et al. · 2020 [cited by applicant]
US 11194490B1 · Sunkavalli · 2021 [cited by examiner]
US 11347916B1 · Desai · 2022 [cited by examiner]
US 20180032859A1 · Park et al. · 2018 [cited by applicant]
US 20180046900A1 · Dally et al. · 2018 [cited by applicant]
US 20180046906A1 · Dally et al. · 2018 [cited by applicant]
US 20180046916A1 · Dally et al. · 2018 [cited by applicant]
US 20180121796A1 · Deisher · 2018 [cited by examiner]
US 20180189641A1 · Boesch et al. · 2018 [cited by applicant]
US 20190114536A1 · Tsung · 2019 [cited by examiner]
US 20190311243A1 · Whatmough · 2019 [cited by examiner]
US 20210334142A1 · Wang · 2021 [cited by examiner]
CN 107657581A · 2018 [cited by applicant]
TW I645301B · 2018 [cited by applicant]
TW 662485B · 2019 [cited by applicant]
TW 201945988A · 2019 [cited by applicant]
S. Subathradevi, C. Vennila. “Systolic array multiplier for augmenting data center networks communication link” Cluster Computing ( Year: 2019). [cited by examiner]
Liqiang Lu, Yun Liang. “SpWA: An Efficient Sparse Winograd Convolutional Neural Networks Accelerator on FPGAs” 55th ACM/ESDA/IEEE Design Automation Conference (DAC) (Year: 2018). [cited by examiner]
Kim et al. “A Novel Zero Weight/Activation-Aware Hardware Architecture of Convolutional Neural Network” Design, Automation & Test in Europe Conference & Exhibition (DATE), Mar. 2017. [cited by applicant]
Shin et al. “DNPU: An 8.1TOPS/W Reconfigurable CNN-RNN Processor for General-Purpose Deep Neural Networks” 2017 IEEE International Solid-State Circuits Conference (ISSCC) Feb. 2017. [cited by applicant]
Han et al. “EIE: Efficient Inference Engine on Compressed Deep Neural Network” ISCA 2016; Feb. 4, 2016. [cited by applicant]
Jouppi et al. “In-Datacenter Performance Analysis of a Tensor Processing Unit” ISCA '17: Proceedings of the 44th Annual International Symposium on Computer Architecture; Jun. 2017 pp. 1-12. [cited by applicant]
Parashar et al. “SCNN: An Accelerator for Compressed-sparse Convolutional Neural Networks” 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA) Jun. 2017. [cited by applicant]
Lee et al. “UNPU: a 50.6TOPS/W Unified Deep Neural Network Accelerator with 1b-to-16b Fully-Variable Weight Bit-Precision” 2018 IEEE International Solid-State Circuits Conference—(ISSCC) Feb. 2018. [cited by applicant]
Chen et al. “Eyeriss: An Energy-Efficient Reconfigurable Accelerator for Deep Convolutional Neural Networks” IEEE Journal of Solid-State Circuits ( vol. 52, Issue: 1, Jan. 2017). [cited by applicant]