IP Library Granted Patent US 12,481,347
Granted Patent B2
US 12,481,347 · App. 18/051,263 · Granted Nov 25, 2025

Techniques for optimizing neural networks for memoization using shifted value localization

Inventors: Georgios Keramidas (Patras, GR); Iakovos Stamoulis (Patras, GR)
Assignee: Think Silicon Research and Technology Single Member S.A.
G06F1/3275G06F9/3832G06F9/3885G06F12/0808G06F12/0813G06F12/0875G06F12/0877G06F17/16G06N3/10G06V10/454
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,347
App. No.
18/051,263
Granted
Nov 25, 2025
Kind
B2
Abstract

A system and method for increasing cache hits in a value cache utilizing memoization is disclosed. The method includes receiving an input matrix for a parallel processing circuitry, wherein the parallel processing circuitry configured to process the input matrix with a second matrix; selecting a portion of the input matrix, wherein the portion includes a plurality of values; adjusting a first value of the plurality of values based on a second value of the plurality of values; generating a new input matrix based on the input matrix and the adjusted first value; and configuring the parallel processing circuitry to process the new input matrix with the second matrix.

Claims (56)

1 . A method for increasing cache hits in a value cache utilizing memoization, comprising:

receiving an input matrix for a parallel processing circuitry, wherein the parallel processing circuitry configured to process the input matrix with a second matrix;

selecting a portion of the input matrix, wherein the portion includes a plurality of values;

adjusting a first value of the plurality of values based on a second value of the plurality of values;

generating a new input matrix based on the input matrix and the adjusted first value; and

configuring the parallel processing circuitry to process the new input matrix with the second matrix.

2 . The method of claim 1 , further comprising:

determining a threshold value for comparing the first value to the second value.

3 . The method of claim 2 , further comprising:

adjusting the first value of the plurality of values in response to determining that the second value of the plurality of values is within the determined threshold value from the first value.

4 . The method of claim 2 , further comprising:

determining the threshold value based on a number of least significant bits (LSBs).

5 . The method of claim 1 , further comprising:

adjusting a third value of the plurality of values based on a target function of a neural network.

6 . The method of claim 1 , further comprising:

adjusting the first value of the plurality of values to the second value of the plurality of values.

7 . The method of claim 1 , further comprising:

performing a z-buffer depth test on the plurality of values; and

generating a threshold value for the portion of the input matrix based on a result of the z-buffer depth test.

8 . The method of claim 7 , further comprising:

adjusting the first value of the plurality of values in response to determining that the second value of the plurality of values is within the generated threshold value from the first value.

9 . The method of claim 1 , wherein the input matrix is a kernel of a convolutional neural network (CNN).

10 . The method of claim 1 , wherein the portion overlaps with a second portion, such that a value of the portion is also a value of the second portion.

11 . The method of claim 1 , wherein the second matrix is any one of: a representation of an image, a fully connected layer of a CNN, and a convolutional layer of a CNN.

12 . A non-transitory computer readable medium having stored thereon instructions for causing a processing circuitry to execute a process, the process comprising:

receiving an input matrix for a parallel processing circuitry, wherein the parallel processing circuitry configured to process the input matrix with a second matrix;

selecting a portion of the input matrix, wherein the portion includes a plurality of values;

adjusting a first value of the plurality of values based on a second value of the plurality of values;

generating a new input matrix based on the input matrix and the adjusted first value; and

configuring the parallel processing circuitry to process the new input matrix with the second matrix.

13 . A system for increasing cache hits in a value cache utilizing memoization, comprising:

a processing circuitry; and

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

receive an input matrix for a parallel processing circuitry, wherein the parallel processing circuitry configured to process the input matrix with a second matrix;

select a portion of the input matrix, wherein the portion includes a plurality of values;

adjust a first value of the plurality of values based on a second value of the plurality of values;

generate a new input matrix based on the input matrix and the adjusted first value; and

configure the parallel processing circuitry to process the new input matrix with the second matrix.

14 . The system of claim 13 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

determine a threshold value for comparing the first value to the second value.

15 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

adjust the first value of the plurality of values in response to determining that the second value of the plurality of values is within the determined threshold value from the first value.

16 . The system of claim 14 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

determine the threshold value based on a number of least significant bits (LSBs).

17 . The system of claim 13 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

adjust a third value of the plurality of values based on a target function of a neural network.

18 . The system of claim 13 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

adjust the first value of the plurality of values to the second value of the plurality of values.

19 . The system of claim 13 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

perform a z-buffer depth test on the plurality of values; and

generate a threshold value for the portion of the input matrix based on a result of the z-buffer depth test.

20 . The system of claim 19 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

adjust the first value of the plurality of values in response to determining that the second value of the plurality of values is within the generated threshold value from the first value.

21 . The system of claim 13 , wherein the input matrix is a kernel of a convolutional neural network (CNN).

22 . The system of claim 13 , wherein the portion overlaps with a second portion, such that a value of the portion is also a value of the second portion.

23 . The system of claim 13 , wherein the second matrix is any one of: a representation of an image, a fully connected layer of a CNN, and a convolutional layer of a CNN.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2026
From: THINK SILICON SINGLE MEMBER P.C. AND APPLIED MATERIALS, INC.
To: QUALCOMM INCORPORATED
Reel/Frame 075735/0803 →
CHANGE OF NAME Recorded Mar 6, 2026
From: THINK SILICON RESEARCH AND TECHNOLOGY SINGLE MEMBER S.A.
To: THINK SILICON SINGLE MEMBER P.C.
Reel/Frame 075032/0035 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2023
From: KERAMIDAS, GEORGIOS; STAMOULIS, IAKOVOS
To: THINK SILICON RESEARCH AND TECHNOLOGY SINGLE MEMBER S.A.
Reel/Frame 062398/0289 →
Continuity (2)
Provisional Application 63278747 · Nov 12, 2021
Related Publication 20230136786A1 · May 4, 2023
References Cited (19)
US 8656378B2 · Gounares et al. · 2014 [cited by applicant]
US 8752034B2 · Gounares et al. · 2014 [cited by applicant]
US 9110814B2 · Keramidas et al. · 2015 [cited by applicant]
US 20150347139A1 · Keramidas et al. · 2015 [cited by applicant]
US 20180025092A1 · Aharonov et al. · 2018 [cited by applicant]
US 20200220851A1 · Storm et al. · 2020 [cited by applicant]
US 20220237116A1 · Hood et al. · 2022 [cited by applicant]
US 20230195419A1 · Gope et al. · 2023 [cited by applicant]
CN 111738403A · 2020 [cited by examiner]
WO 2020252762A1 · 2020 [cited by applicant]
English translation of CN-111738403-A. (Year: 2020). [cited by examiner]
International Search Report for PCT Application No. PCT/IB2022/060310 dated Jan. 31, 2023. The International Bureau of WIPO. [cited by applicant]
Written Opinion of the International Searching Authority for PCT Application No. PCT/IB2022/060310 dated Jan. 31, 2023. The International Bureau of WIPO. [cited by applicant]
International Search Report for PCT Application No. PCT/IB2022/060419 dated Jan. 4, 2023. The International Bureau of WIPO. [cited by applicant]
International Search Report for PCT Application No. PCT/IB2022/060420 dated Mar. 2, 2023. The International Bureau of WIPO. [cited by applicant]
Sandhupatla Amruth: “Faster convolutional neural network training via approximate memoization”, Oct. 18, 2018, XP093010822. [cited by applicant]
Seongmin Hong et al.: “Enabling Energy Efficient Image Encryption using Approximate Memoization”, Journal of Semiconductor Technology and Science, Jan. 1, 2017, pp. 465-472, XP055680555 DOI: 10.5573/JSTS.2017.17.3.465. [cited by applicant]
Written Opinion of the International Searching Authority for PCT Application No. PCT/IB2022/060419 dated Jan. 4, 2023. The International Bureau of WIPO. [cited by applicant]
Written Opinion of the International Searching Authority for PCT Application No. PCT/IB2022/060420 dated Mar. 2, 2023. The International Bureau of WIPO. [cited by applicant]