IP Library › Granted Patent US 12,541,674
Granted Patent B2
US 12,541,674 · App. 17/743,476 · Granted Feb 3, 2026

Method and non-transitory computer readable medium for compute-in-memory macro arrangement, and electronic device applying the same

Inventors: Chien Te Tung (Taipei, TW); Chih Feng Juan (New Taipei, TW); Jen-Wei Liang (Taipei, TW)
Assignee: Novatek Microelectronics Corp.
G06N3/0464G06F7/49942G11C7/1006G11C7/1012G11C7/1063G11C7/109G11C11/54
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,674
App. No.
17/743,476
Filed
May 13, 2022
Granted
Feb 3, 2026
Kind
B2
Art Unit
2825
USPC
365/148
Abstract

A method and a non-transitory computer readable medium for CIM arrangement, and an electronic device applying the same are proposed. The method for CIM arrangement includes to obtain information of the number of CIM macros and information of the dimension of each of the CIM micros, to obtain information of the number of input channels and the number of output channels of a designated convolutional layer of a designate neural network, and to determine a CIM macro arrangement for arranging the CIM macros according to the number of the CIM macros, the dimension of each of the CIM macros, the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network, for applying convolution operation to the input channels to generate the output channels.

Claims (43)

1 . A method for compute-in-memory (CIM) macro arrangement comprising:

obtaining information of the number of a plurality of CIM macros and information of a dimension of each of the CIM macros;

obtaining information of the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network; and

determining a CIM macro arrangement for arranging the CIM macros according to the number of the CIM macros, the dimension of each of the CIM macros, the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network, for applying convolution operation to the input channels to generate the output channels,

wherein the determined CIM macro arrangement provides a summation of a vertical dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the input channels.

2 . The method according to claim 1 , wherein the step of determining the CIM macro arrangement according to the number of the CIM macros, the dimensions of each of the CIM macros, and the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network comprises:

determining the CIM macro arrangement capable of performing the convolution of the filters and the input channels according to latency, energy consumption, and utilization.

3 . The method according to claim 2 ,

wherein the latency is associated with at least one of a DRAM latency, a latency for loading weights into the CIM macros, and a processing time of the CIM macros,

wherein the energy consumption is associated with energy cost for accessing at least one memory including an on-chip SRAM which is in a same chip as the CIM macros and a DRAM outside the chip, and

wherein the utilization is a ratio of used part of the CIM macros to all of the CIM macros.

4 . The method according to claim 1 , wherein the determined CIM macro arrangement further provides a summation of a horizontal dimension of the CIM macros adapted for performing the convolution of the filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the filters.

5 . An electronic apparatus comprising:

a plurality of compute-in-memory (CIM) macros, wherein the CIM macros are arranged in a predetermined CIM macro arrangement based on the number of the CIM macros, the dimensions of each of the CIM macros, and the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network, wherein the predetermined CIM macro arrangement provides a summation of a vertical dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the input channels; and

a processing circuit, configured to:

load weights in the arranged CIM macros; and

input a plurality of input channels of one input feature map into the arranged CIM macros with the loaded weights for a convolutional operation to generate an output activation of one of a plurality of output feature maps.

6 . The electronic apparatus according to claim 5 ,

wherein the processing circuit loads the weights of the filters in the arranged CIM macros based on the predetermined CIM macro arrangement, the number of the filters, height and width of each kernel of a plurality of kernels of each of the filters and the number of the kernels in each filter, wherein each of the kernels of each filter is respectively applied to a corresponding one of the input channels of the designated convolutional layer of the designated neural network.

7 . The electronic apparatus according to claim 5 ,

wherein the processing circuit loads each of the filters into the arranged CIM macros columnwisely.

8 . The electronic apparatus according to claim 5 ,

wherein the processing circuit determines whether to batch loads the weights of the plurality of filters in the arranged CIM macros based on the height and width of each kernel and a summation of a horizontal dimension of the arranged CIM macro.

9 . A non-transitory computer readable medium storing a program causing a computer to:

obtaining information of the number of a plurality of CIM macros and information of a dimension of each of the CIM macros;

obtaining information of the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network; and

determining a CIM macro arrangement for arranging the CIM macros according to the number of the CIM macros, the dimension of each of the CIM macros, the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network, for applying convolution operation to the input channels to generate the output channels,

wherein the determined CIM macro arrangement provides a summation of a vertical dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the input channels.

10 . A method for compute-in-memory (CIM) macro arrangement comprising:

obtaining information of the number of a plurality of CIM macros and information of a dimension of each of the CIM macros;

obtaining information of the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network; and

determining a CIM macro arrangement for arranging the CIM macros according to the number of the CIM macros, the dimension of each of the CIM macros, the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network, for applying convolution operation to the input channels to generate the output channels,

wherein the determined CIM macro arrangement provides a summation of a horizontal dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the filters.

11 . An electronic apparatus comprising:

a plurality of compute-in-memory (CIM) macros, wherein the CIM macros are arranged in a predetermined CIM macro arrangement based on the number of the CIM macros, the dimensions of each of the CIM macros, and the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network, wherein the determined CIM macro arrangement provides a summation of a horizontal dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the filters; and

a processing circuit, configured to:

load weights in the arranged CIM macros; and

input a plurality of input channels of one input feature map into the arranged CIM macros with the loaded weights for a convolutional operation to generate an output activation of one of a plurality of output feature maps.

12 . A non-transitory computer readable medium storing a program causing a computer to:

obtaining information of the number of a plurality of CIM macros and information of a dimension of each of the CIM macros;

obtaining information of the number of a plurality of input channels and the number of a plurality of output channels of a designated convolutional layer of a designated neural network; and

determining a CIM macro arrangement for arranging the CIM macros according to the number of the CIM macros, the dimension of each of the CIM macros, the number of the input channels and the number of the output channels of the designated convolutional layer of the designated neural network, for applying convolution operation to the input channels to generate the output channels,

wherein the determined CIM macro arrangement provides a summation of a horizontal dimension of the CIM macros adapted for performing convolution of a plurality of filters and the input channels of the designated convolution layer by a minimum number of times for batch loading the filters.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2022
From: TUNG, CHIEN TE; JUAN, CHIH FENG; LIANG, JEN-WEI
To: NOVATEK MICROELECTRONICS CORP.
Reel/Frame 059993/0041 →
Continuity (2)
Provisional Application 63187952 · May 13, 2021
Related Publication 20220366216A1 · Nov 17, 2022
References Cited (32)
US 10915298B1 · Far · 2021 [cited by applicant]
US 10943652B2 · Lu et al. · 2021 [cited by applicant]
US 11354123B2 · Chang et al. · 2022 [cited by applicant]
US 12120331B2 · Kang et al. · 2024 [cited by applicant]
US 20180285715A1 · Son · 2018 [cited by examiner]
US 20180341495A1 · Culurciello · 2018 [cited by examiner]
US 20180374541A1 · Jung et al. · 2018 [cited by applicant]
US 20190362787A1 · Lu et al. · 2019 [cited by applicant]
US 20200193293A1 · Song · 2020 [cited by examiner]
US 20200242459A1 · Manipatruni · 2020 [cited by examiner]
US 20210064234A1 · Zhang et al. · 2021 [cited by applicant]
US 20210089865A1 · Wang et al. · 2021 [cited by applicant]
US 20210117187A1 · Chang et al. · 2021 [cited by applicant]
US 20210210138A1 · Lu et al. · 2021 [cited by applicant]
US 20220223199A1 · Gu · 2022 [cited by examiner]
US 20220321900A1 · Kang et al. · 2022 [cited by applicant]
US 20220351032A1 · Chou · 2022 [cited by examiner]
US 20220366947A1 · Tung · 2022 [cited by examiner]
US 20220400222A1 · Nakagawa · 2022 [cited by examiner]
US 20230074229A1 · Jia · 2023 [cited by examiner]
US 20240061649A1 · Kwon · 2024 [cited by examiner]
US 20240111828A1 · Kim · 2024 [cited by examiner]
CN 111126579 · 2020 [cited by applicant]
CN 112580774 · 2021 [cited by applicant]
TW 202013213 · 2020 [cited by applicant]
TW 202117561 · 2021 [cited by applicant]
WO 2021027238 · 2021 [cited by applicant]
“Office Action of Taiwan Counterpart Application”, issued on Jul. 8, 2022, p. 1-p. 3. [cited by applicant]
Xiaochen Peng et al., “Optimizing Weight Mapping and Data Flow for Convolutional Neural Networks on Processing-In-Memory Architectures”, IEEE Transactions on Circuits and Systems—I: Regular Papers, Apr. 2020, pp. 1333-1… [cited by applicant]
“Office Action of China Counterpart Application”, issued on May 29, 2025, p. 1-p. 12. [cited by applicant]
“Notice of allowance of China counterpart Application”, issued on Sep. 12, 2025, p. 1-p. 4. [cited by applicant]
Li Danqing, “Design and Optimization of Reconfigurable Array for DCT and IDCT”, A thesis submitted to Southeast University, China Outstanding Master's Thesis Full-text Database, Information Technology vol. Issue No. 1 o… [cited by applicant]