IP Library › Granted Patent US 12,645,588
Granted Patent B1
US 12,645,588 · App. 18/980,075 · Granted Jun 2, 2026

Method of operation for a computing system

Inventors: David Garrett (Los Altos, CA); Michael Gao (Los Altos, CA); Gilbert Hendry (Los Altos, CA)
Assignee: Fabric of Truth, Inc.
G06F12/0292G06F17/16G06F2212/251
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,588
App. No.
18/980,075
Granted
Jun 2, 2026
Kind
B1
Abstract

A computing system, preferably including a scratchpad, a plurality of computation units, and a register file, and optionally including a controller. The computing system, or any suitable elements thereof, can optionally be integrated into a processor unit, wherein one or more such processor units can optionally be integrated into a larger-scale computing system. Some or all elements of the larger-scale computing system can be collocated and/or codefined on a single semiconductor chip, can be located and/or defined on separate chips, and/or can be otherwise located and/or defined. A method of operation, preferably including performing computation unit I/O operations and/or performing scratchpad I/O operations. The method of operation is preferably performed using the computing system, but can additionally or alternatively be performed using any other suitable systems.

Claims (71)

1 . A method of operation for a computing system, the method comprising, while operating the computing system in a first mode associated with a first mode word length, the first mode word length equal to W bits:

receiving a vector request indicative of a memory address A, expressed in bytes, within a scratchpad of the computing system, wherein:

the scratchpad comprises a scratchpad memory;

the scratchpad memory defines a plurality of memory rows that partition the scratchpad memory;

each memory row 0f the plurality has a memory capacity, expressed in bytes, of N=KW/8 for a positive integer K, wherein N has a prime factor greater than 2; and

each memory row 0f the plurality stores a respective set of K words, wherein each word of the respective set has a bit width of W;

based on the memory address:

determining an initial memory row 0f the scratchpad memory by performing an integer division Q=A/N, wherein the initial memory row has an index of Q;

determining an initial word position within the initial memory row; and

selecting a vector of K contiguous word positions from the scratchpad memory, the vector comprising:

the initial word position; and

at least one word position located in a subsequent memory row 0f the scratchpad memory, the subsequent memory row having an index of Q+1; and

based on the vector request and the vector, performing a memory operation between the scratchpad memory and a register file of the computing system, the register file communicatively coupled to the scratchpad, wherein performing the memory operation comprises at least one of:

reading the vector from the scratchpad memory; or

writing to the vector within the scratchpad memory.

2 . The method of claim 1 , further comprising, while operating the computing system in a second mode associated with a second mode word length, the second mode word length equal to W 2 bits, wherein N=K 2 W 2 /8 for a second positive integer K 2 :

receiving a second vector request indicative of a second memory address A 2 , expressed in bytes, within the scratchpad, wherein each memory row 0f the plurality stores a respective set of K 2 words, wherein each word of the respective set has a bit width of W 2 ;

based on the second memory address:

determining a second initial memory row 0f the scratchpad memory by performing a second integer division Q 2 =A 2 /N, wherein the second initial memory row has an index of Q 2 ;

determining a second initial word position within the second initial memory row; and

selecting a second vector of K 2 contiguous word positions from the scratchpad memory, the second vector comprising:

the second initial word position; and

at least one word position located in a second subsequent memory row 0f the scratchpad memory, the second subsequent memory row having an index of Q 2 +1; and

based on the second vector request and the second vector, performing a second memory operation between the scratchpad memory and the register file, wherein performing the second memory operation comprises at least one of:

reading the second vector from the scratchpad memory; or

writing to the second vector within the scratchpad memory.

3 . The method of claim 2 , wherein:

W has a prime factor of 3;

K does not have a prime factor of 3;

W 2 does not have a prime factor of 3;

K 2 has a prime factor of 3;

W is not divisible by W 2 ; and

W 2 is not divisible by W.

4 . The method of claim 1 , wherein the prime factor is 3.

5 . The method of claim 1 , wherein the first mode word length is 384 bits.

6 . The method of claim 1 , further comprising, while operating the computing system in the first mode:

receiving a second vector request indicative of a second memory address A 2 , expressed in bytes, within the scratchpad;

based on the second memory address:

determining a second initial memory row 0f the scratchpad memory by performing a second integer division Q 2 =A 2 /N, wherein the second initial memory row has an index of Q 2 ;

determining a second initial word position within the second initial memory row, wherein:

the initial word position is represented by a first offset from a beginning of the initial memory row;

the second initial word position is represented by a second offset from a beginning of the second initial memory row; and

the first offset is different from the second offset; and

selecting a second vector of K 2 contiguous word positions from the scratchpad memory, the second vector comprising the second initial word position; and

based on the second vector request and the second vector, performing a second memory operation between the scratchpad memory and the register file, wherein performing the second memory operation comprises at least one of:

reading the second vector from the scratchpad memory; or

writing to the second vector within the scratchpad memory.

7 . The method of claim 6 , wherein:

Q is not equal to Q 2 ;

the second offset is non-zero; and

the second vector further comprises at least one word position located in a second subsequent memory row 0f the scratchpad memory, the second subsequent memory row having an index of Q 2 +1.

8 . The method of claim 7 , wherein:

the register file is communicatively coupled to the scratchpad memory via a memory rotation datapath;

performing the memory operation comprises, at the memory rotation datapath, barrel shifting the vector based on the first offset; and

performing the second memory operation comprises, at the memory rotation datapath, barrel shifting the second vector based on the second offset.

9 . The method of claim 1 , wherein:

the scratchpad defines a plurality of memory lanes that partition the scratchpad memory;

the plurality of memory lanes consists of k memory lanes;

each memory lane of the scratchpad defines an identical bit width b=2 s+3 , wherein s is a positive integer; and

determining the initial word position comprises determining the value (A»s) mod k, where»is the right shift operator, by:

generating an intermediary result by performing an s-bit right shift on A; and

performing a modulo operation with modulus k on the intermediary result, wherein k has a prime factor greater than 2.

10 . The method of claim 1 , further comprising, while operating the computing system in the first mode:

receiving an element-wise request indicative of a plurality of memory addresses;

determining a set of element positions, comprising, for each memory address of the plurality, based on the memory address, determining a respective element position within the scratchpad memory, determining the respective element position comprising:

determining a respective memory row 0f the scratchpad memory by performing an integer division of the memory address divided by N; and

determining a respective word position within the respective memory row; and

based on the element-wise request and the set of element positions, performing a second memory operation between the scratchpad memory and the register file, wherein performing the second memory operation comprises at least one of:

concurrently reading from each element position of the set; or

concurrently writing to each element position of the set;

wherein, for each element position of the set, the respective word position differs from the word position of all other element positions of the set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 3, 2025
From: GAO, MICHAEL; GARRETT, DAVID; HENDRY, GILBERT
To: FABRIC OF TRUTH, INC.
Reel/Frame 069741/0606 →
Continuity (2)
Continuation 18777335 · Jul 18, 2024
Provisional Application 63590309 · Oct 13, 2023
References Cited (19)
US 7725641B2 · Park et al. · 2010 [cited by applicant]
US 10218494B1 · De Quehen et al. · 2019 [cited by applicant]
US 11816572B2 · Bruestle et al. · 2023 [cited by applicant]
US 20020029364A1 · Edmonston · 2002 [cited by examiner]
US 20020133688A1 · Lee · 2002 [cited by examiner]
US 20040015758A1 · Pathak · 2004 [cited by examiner]
US 20110296118A1 · Carter et al. · 2011 [cited by applicant]
US 20140053000A1 · Yap et al. · 2014 [cited by applicant]
US 20190073196A1 · Hiscock · 2019 [cited by applicant]
US 20190227981A1 · Tomishima et al. · 2019 [cited by applicant]
US 20200004506A1 · Langhammer et al. · 2020 [cited by applicant]
US 20220180187A1 · Kim · 2022 [cited by examiner]
US 20230134216A1 · Anderson · 2023 [cited by applicant]
Atlas, et al., “Multi-Precision Fast Modular Multiplication”, Ingonyama Blog, published Jan. 15, 2023. [cited by applicant]
Banakar, et al., “Scratchpad Memory : A Design Alternative for Cache On-chip memory in Embedded Systems”, Proceedings of the Tenth International Symposium on Hardware/Software Codesign. CODES 2002 (IEEE Cat. No. 02TH862… [cited by applicant]
Barrett, Paul, “Implementing the Rivest Shamir and Adleman Public Key Encryption Algorithm on a Standard Digital Signal Processor”, Advances in Cryptology—CRYPTO '86, Santa Barbara, California, USA, 1986, Proceedings, v… [cited by applicant]
Domb, Yuval, “Fast Modular Multiplication”, Ingonyama blog, published Jul. 25, 2022. [cited by applicant]
Gao, et al., “Method and System for Barrett Reduction Computation”, U.S. Appl. No. 18/825,710, filed Sep. 5, 2024. [cited by applicant]
Langhammer, et al., “Efficient FPGA Modular Multiplication Implementation”, FPGA '21, Session 3: Machine Learning and Supporting Algorithms, 2021 (Year: 2021). [cited by applicant]