IP Library › Granted Patent US 12,517,836
Granted Patent B2
US 12,517,836 · App. 18/917,369 · Granted Jan 6, 2026

Computing system and method for power-saving compute-in-memory design

Inventor: Tso Wang (Hsinchu, TW)
Assignee: MEDIATEK INC.
G06F12/0897G06F2212/27G06F2212/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,517,836
App. No.
18/917,369
Granted
Jan 6, 2026
Kind
B2
Abstract

A computing system with power-saving compute-in-memory (CIM) design that minimizes the computation energy of the matrix-matrix multiplication is shown. A processor control unit loads A blocks divided from a matrix A M×K from a second-level (L2) memory to a first-level (L1) memory, and loads B blocks divided from a matrix B K×N from the L2 memory to a CIM memory. The A blocks buffered in the L1 memory are programmed to a register file to be entered into the CIM memory. The CIM memory performs multiply-and-accumulate (MAC) calculations on the A blocks and the B blocks to generate C blocks which form a matrix C M×N (=A M×K ×B K×N ). Based on the size of A M×K and B K×N , an A block buffering capability of the L1 memory, and a B block buffering capability of the CIM memory, the reuse scheme is properly selected to reuse the buffered A blocks and B blocks.

Claims (166)

1 . A computing system with a compute-in-memory (CIM) design, comprising:

a CIM processor, including a processor control unit, a CIM memory with CIM capability, and a register file; and

a two-level memory system coupled to the CIM processor, wherein the two-level memory system includes a first-level (L1) memory and a second-level (L2) memory;

wherein:

the processor control unit loads A blocks divided from a matrix A M×K from the L2 memory to the L1 memory, and loads B blocks divided from a matrix B K×N from the L2 memory to the CIM memory;

the processor control unit further programs the A blocks buffered in the L1 memory to the register file to be entered into the CIM memory, wherein the CIM memory performs multiply-and-accumulate (MAC) calculations on the A blocks and the B blocks to generate C blocks which form a matrix C M×N that is A M×K ×B K×N ; and

based on size of the matrix A M×K , size of the matrix B K×N , an A block buffering capability, μ, of the L1 memory, and a B block buffering capability, L, of the CIM memory, the CIM processor selects a reuse scheme to reuse the A blocks buffered in the L1 memory and the B blocks buffered in the CIM memory.

2 . The computing system as claimed in claim 1 , wherein:

the number of A blocks divided from the matrix A M×K is M b ×K b , wherein M b =M/m, and K b =K/k;

the number of B blocks divided from the matrix B K×N is K b ×N b , wherein N b =N/n; and

M b , K b , and No representing the size of the matrix A M×K and the matrix B K×N are used in selecting the reuse scheme.

3 . The computing system as claimed in claim 2 , wherein:

m=n=m c , and k=αm c , where α is greater than 0.

4 . The computing system as claimed in claim 3 , wherein:

in response to a situation wherein M b >1 and N b >1, the CIM processor judges K b to select the reuse scheme based on a threshold function T(⋅), wherein the threshold function T(⋅) is a function of the A block buffering capability, μ, of the L1 memory and s s , the sparsity of the matrix B K×N , where 1≤μ≤K b , and 0≤s s <1.

5 . The computing system as claimed in claim 4 , wherein:

T

⁡

(

·

)

=

μ

(

1

-

s

s

)

⁢

or

⁢

⁢

T

⁡

(

·

)

=

T

⁡

(

M

b

,

N

b

)

=

μ

(

1

-

s

s

)

·

M

b

(

N

b

-

1

)

N

b

(

M

b

-

1

)

which further depends on M b and N b .

6 . The computing system as claimed in claim 5 , wherein:

when min(L max , K b )<T(⋅), the CIM processor sets the B block buffering capability, L, of the CIM memory to min(2, K b ), and selects a first reuse scheme first_reuse_A, where L max is an upper limit to which the CIM memory buffers the B blocks and, according to the first reuse scheme first_reuse_A, the A blocks buffered in the L1 memory are reused even if the CIM memory has updated the B blocks buffered therein.

7 . The computing system as claimed in claim 5 , wherein:

when min(L max , K b )≥T(⋅), the CIM processor sets the B block buffering capability, L, of the CIM memory to min(L max , K b ), and selects a second reuse scheme first_reuse_B, where L max is an upper limit to which the CIM memory buffers the B blocks and, according to the second reuse scheme first_reuse_B, the B blocks buffered in the CIM memory are reused even if the L1 memory has updated the A blocks buffered therein.

8 . The computing system as claimed in claim 3 , wherein:

in response to a situation wherein N b >M b =1, the CIM processor sets the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), and selects a first reuse scheme first_reuse_A, by which the A blocks buffered in the L1 memory are reused even if the CIM memory has updated the B blocks buffered therein.

9 . The computing system as claimed in claim 3 , wherein:

in response to a situation wherein M b >N b =1, the CIM processor sets the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), and selects a second reuse scheme first_reuse_B, by which the B blocks buffered in the CIM memory are reused even if the L1 memory has updated the A blocks buffered therein.

10 . The computing system as claimed in claim 3 , wherein:

in response to a situation wherein M b =N b =1, the CIM processor sets the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), without changing the previously selected reuse scheme.

11 . A method for power saving of a compute-in-memory (CIM) design, comprising:

loading A blocks divided from a matrix A M×K from a second-level (L2) memory to a first-level (L1) memory;

loading B blocks divided from a matrix B K×N from the L2 memory to a CIM memory;

programming the A blocks buffered in the L1 memory to a register file to be entered into the CIM memory, wherein the CIM memory performs multiply-and-accumulate (MAC) calculations on the A blocks and the B blocks to generate C blocks which form a matrix C M×N that is A M×K ×B K×N ; and

based on the size of the matrix A M×K , the size of the matrix B K×N , an A block buffering capability, μ, of the L1 memory, and a B block buffering capability, L, of the CIM memory, selecting a reuse scheme to reuse the A blocks buffered in the L1 memory and the B blocks buffered in the CIM memory.

12 . The method as claimed in claim 11 , wherein:

the number of A blocks divided from the matrix A M×K is M b ×K b , wherein M b =M/m, and K b =K/k;

the number of B blocks divided from the matrix B K×N is K b ×N b , wherein N b =N/n; and

M b , K b , and N b representing the size of the matrix A M×K and the matrix B K×N are used in selecting the reuse scheme.

13 . The method as claimed in claim 12 , wherein:

m=n=m c , and k=αm c , where α is greater than 0.

14 . The method as claimed in claim 13 , further comprising:

in response to a situation wherein M b >1 and N b >1, judging K b to select the reuse scheme based on a threshold function T(⋅), wherein the threshold function T(⋅) is a function of the A block buffering capability, μ, of the L1 memory, and s s , the sparsity of the matrix B K×N , where 1≤μ≤K b , and 0≤s s <1.

15 . The method as claimed in claim 14 , wherein:

T

⁡

(

·

)

=

μ

(

1

-

s

s

)

⁢

or

⁢

⁢

T

⁡

(

·

)

=

T

⁡

(

M

b

,

N

b

)

=

μ

(

1

-

s

s

)

·

M

b

(

N

b

-

1

)

N

b

(

M

b

-

1

)

which further depends on M b and N b .

16 . The method as claimed in claim 15 , further comprising:

when min(L max , K b )<T(⋅), setting the B block buffering capability, L, of the CIM memory to min(2, K b ), and selecting a first reuse scheme first_reuse_A, where L max is an upper limit to which the CIM memory buffers the B blocks and, according to the first reuse scheme first_reuse_A, the A blocks buffered in the L1 memory are reused even if the CIM memory has updated the B blocks buffered therein.

17 . The method as claimed in claim 15 , further comprising:

when min(L max , K b )≥T(⋅), setting the B block buffering capability, L, of the CIM memory to min(L max , K b ), and selecting a second reuse scheme first_reuse_B, where L max is an upper limit to which the CIM memory buffers the B blocks and, according to the second reuse scheme first_reuse_B, the B blocks buffered in the CIM memory are reused even if the L1 memory has updated the A blocks buffered therein.

18 . The method as claimed in claim 13 , further comprising:

in response to a situation wherein N b >M b =1, setting the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), and selecting a first reuse scheme first_reuse_A, by which the A blocks buffered in the L1 memory are reused even if the CIM memory has updated the B blocks buffered therein.

19 . The method as claimed in claim 13 , further comprising:

in response to a situation wherein M b >N b =1, setting the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), and selecting a second reuse scheme first_reuse_B, by which the B blocks buffered in the CIM memory are reused even if the L1 memory has updated the A blocks buffered therein.

20 . The method as claimed in claim 13 , further comprising:

in response to a situation wherein M b =N b =1, setting the B block buffering capability, L, of the CIM memory to L=min(L max , K b ), without changing the previously selected reuse scheme.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2024
From: WANG, TSO
To: MEDIATEK INC.
Reel/Frame 068915/0629 →
Continuity (2)
Provisional Application 63591468 · Oct 19, 2023
Related Publication 20250130950A1 · Apr 24, 2025
References Cited (3)
US 10831678B2 · Wang · 2020 [cited by examiner]
US 20120151232A1 · Fish, III · 2012 [cited by examiner]
US 20220197647A1 · Kayiran · 2022 [cited by examiner]