IP Library › Granted Patent US 12,625,675
Granted Patent B2
US 12,625,675 · App. 17/480,925 · Granted May 12, 2026

Convolutional computation device

Inventor: Teppei Hirotsu (Tokyo, JP)
Assignee: DENSO CORPORATION
G06F7/5443G06F17/153G11C19/38
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,625,675
App. No.
17/480,925
Granted
May 12, 2026
Kind
B2
Abstract

A convolutional computation device includes a two-dimensional circulation shift register unit and one or more multiplier-accumulators. The two-dimensional circulation shift register has storage elements, cyclically shifts the data among the storage elements, provides one or more input window in a predetermined area, and selects the data stored in one of the storage elements disposed in the input window as input data. The one or more multiplier-accumulators generate output data by performing a multiply-accumulate operation on the input data input from the two-dimensional circulation shift register unit and weight data from a predetermined filter.

Claims (41)

1 . A convolutional computation device comprising:

a two-dimensional circulation shift register unit that has a plurality of storage elements arranged two-dimensionally and respectively storing data, cyclically shifts the data among the plurality of storage elements, provides at least one input window in a predetermined area, and selects the data stored in one of the plurality of storage elements disposed in the at least one input window as input data; and

at least one multiplier-accumulator that generates output data by performing a multiply-accumulate operation on the input data input from the two-dimensional circulation shift register unit and weight data from a predetermined filter,

wherein:

each storage element of the plurality of storage elements includes a multiplexer and a flip-flop, the flip-flop in each storage element is configured to store data output by the multiplexer in each storage element, the multiplexer in each storage element is directly connected by signal lines to the flip-flops of directly adjacent storage elements of the plurality of storage elements, the directly adjacent storage elements being directly adjacent via signal lines in each of an up direction and a down direction from each storage element and directly adjacent via signal lines in each of a left direction and a right direction from each storage element, and the multiplexer in each storage element is configured to select one data element of the directly adjacent storage elements and output data from the selected one data element.

2 . The convolutional computation device according to claim 1 , wherein:

the at least one multiplier-accumulator performs the multiply-accumulate operation based on a Winograd algorithm;

the predetermined filter is a 3 rows and 3 columns matrix filter;

the input data is a 5 rows and 5 columns matrix input data; and

the output data is a 3 rows and 3 columns matrix output data.

3 . The convolutional computation device according to claim 2 ,

wherein the Winograd algorithm is defined by an equation:

Y=A T [[GgG T ]⊙[B T dB]] A   (2)

where

Y is an output data;

A, B and G are constant matrixes;

g is a weight data;

d is an input data;

a weight term of GgG T is calculated in advance; and

the constant matrixes B and A have one of the elements 0, ±1, ±2, ±3, and ±4.

4 . The convolutional computation device according to claim 1 , wherein:

the two-dimensional circulation shift register unit has a plurality of input windows including the at least one input window to select a plurality of the input data,

the at least one multiplier-accumulator further comprises a plurality of multiplier-accumulators into which the plurality of the input data are input from the two-dimensional circulation shift register unit, respectively.

5 . The convolutional computation device according to claim 1 , wherein:

the two-dimensional circulation shift register unit changes an input window area of the at least one input window or a shift amount of the shift that cyclically shifts the data.

6 . The convolutional computation device according to claim 1 , wherein:

each storage element is connected to only four storage elements arranged vertically and horizontally of the each storage element.

7 . The convolutional computation device according to claim 1 , further comprising a memory interface connected to the two-dimensional circulation shift register unit, wherein

data elements are sequentially input from the memory interface to each storage element in a bottom row of the two-dimensional circulation shift register unit.

8 . The convolutional computation device according to claim 1 , wherein

the at least one multiplier-accumulator includes an input register, a weight register, at least one multiplier, and an adder tree,

the input register is configured to hold the input data in a plurality of input data element storage areas,

the weight register is configured to hold the weight data in a plurality of weight data element storage areas,

the at least one multiplier is configured to multiply each input data element of the plurality of input data element storage areas with respectively each weight data element of the plurality of weight data element storage areas to generate multiplier results, and

the adder tree is configured to calculate a total multiplication result from the multiplier results generated by the at least one multiplier.

9 . The convolutional computation device according to claim 1 , wherein

the at least one input window comprises a plurality of input window areas which are configured to be switched with each other in the two-dimensional circulation shift register unit.

10 . The convolutional computation device according to claim 1 , wherein the at least one input window comprises a plurality of input window areas which are configured to be sequentially switched to select the input data, and the output data is sequentially generated from the input data.

11 . The convolutional computation device according to claim 1 , wherein the multiply-accumulate operation is performed only by a bit shift operation and an addition operation.

12 . The convolutional computation device according to claim 1 , wherein

the plurality of storage elements which are arranged two-dimensionally are arranged in rows and columns.

Assignments (2)
MERGER Recorded Mar 21, 2024
From: NSITEXE, INC.
To: DENSO CORPORATION
Reel/Frame 066861/0661 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: HIROTSU, TEPPEI
To: DENSO CORPORATION; NSITEXE,INC.
Reel/Frame 057549/0599 →
Priority Claims (1)
JP 2019-062744 · Mar 28, 2019 · national
Continuity (2)
Continuation PCTJP2020012728 · Mar 23, 2020
Related Publication 20220004364A1 · Jan 6, 2022
References Cited (14)
US 5764557A · Hara et al. · 1998 [cited by applicant]
US 9747548B2 · Ross et al. · 2017 [cited by applicant]
US 9986187B2 · Meixner · 2018 [cited by examiner]
US 11106972B1 · Verheyen · 2021 [cited by examiner]
US 20180005074A1 · Shacham et al. · 2018 [cited by applicant]
US 20180032312A1 · Hansen et al. · 2018 [cited by applicant]
US 20180329479A1 · Meixner · 2018 [cited by examiner]
US 20190205780A1 · Sakaguchi · 2019 [cited by examiner]
JP H06318194A · 1994 [cited by applicant]
JP 2019179410A · 2019 [cited by applicant]
Stein, Jonathan Y. Digital Signal Processing: A computer Science Perspective [textbook]. John Wiley & Sons, Inc. Print ISBN 0-471-29546-9, Online ISBN 0-471-20059-X. pp. 570-644. Retried from the Internet: <URL:https://… [cited by examiner]
Maji, et al. Efficient Winograd or Cook-Toom Convolution Kernel Implementation on Widely Used Mobile CPUs [online]. Retrieved from the Internet: <URL: https://www.emc2-ai.org/hpca-19> (Year: 2019). [cited by examiner]
EMC^2. The 2nd EMC2—Energy Efficient Machine Learning and Cognitive Computing Co-located with the 25th IEEE International Symposium on High-Performance Computer Architecture HPCA 2019 [online]. Retrieved from the Intern… [cited by examiner]
Aydonat, Utku et al: “An OpenCL(TM) Deep Learning Accelerator on Arria 10”, Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems , CHI '17, ACM Press, New York, New York, USA, Feb. 22, 2017 (Feb.… [cited by applicant]