IP Library Granted Patent US 11,468,002
Granted Patent B2
US 11,468,002 · App. 17/187,082 · Granted Oct 11, 2022

Computational memory with cooperation among rows of processing elements and memory thereof

Inventors: William Martin Snelgrove (Toronto, CA); Jonathan Scobbie (Toronto, CA)
Assignee: UNTETHER AI CORPORATION
G06F15/8015G06F9/3001G06F9/30047G06F9/3887G06F1/32G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,468,002
App. No.
17/187,082
Granted
Oct 11, 2022
Kind
B2
Abstract

A computing device includes an array of processing elements mutually connected to perform single instruction multiple data (SIMD) operations, memory cells connected to each processing element to store data related to the SIMD operations, and a cache connected to each processing element to cache data related to the SIMD operations. Caches of adjacent processing elements are connected. The same or another computing device includes rows of mutually connected processing elements to share data. The computing device further includes a row arithmetic logic unit (ALU) at each row of processing elements. The row ALU of a respective row is configured to perform an operation with processing elements of the respective row.

Claims (35)

1. A computing device comprising:

an array of processing elements mutually connected to perform single instruction multiple data (SIMD) operations;

memory cells connected to each processing element to store data related to the SIMD operations; and

a plurality of caches connected to each processing element to cache data related to the SIMD operations, wherein each cache is associated with a different block of the memory cells, wherein a first cache of a first processing element is connected to a second cache of a second processing element that is adjacent the first processing element in the array of processing elements.

2. The computing device of claim 1 , wherein the first cache of the first processing element is connected to a third cache of a third processing element that is adjacent the first processing element in the array of processing elements, wherein the first processing element is positioned between the second and third processing elements in the array of processing elements.

3. The computing device of claim 1 , further comprising a write multiplexer including an output connected to the first cache, wherein selectable inputs of the write multiplexer are connected to a register of the first processing element and to the second cache.

4. The computing device of claim 3 , wherein:

the array of processing elements is configured to perform a multiplying accumulation; and

the computing device further comprises a controller to control the write multiplexer to write to the first cache accumulated results and/or input values of the multiplying accumulation.

5. The computing device of claim 1 , further comprising a read multiplexer including an output connected to a register of the first processing element, wherein selectable inputs of the read multiplexer are connected to the first cache and to the second cache.

6. The computing device of claim 5 , wherein:

the array of processing elements is configured to perform a multiplying accumulation; and

the computing device further comprises a controller to control the read multiplexer to read from the first cache coefficients and/or input values of the multiplying accumulation.

7. A computing device comprising:

an array of processing elements mutually connected to perform single instruction multiple data (SIMD) operations;

memory cells connected to each processing element to store data related to the SIMD operations;

a cache connected to each processing element to cache data related to the SIMD operations, wherein a first cache of a first processing element is connected to a second cache of a second processing element that is adjacent the first processing element in the array of processing elements; and

a write multiplexer including an output connected to the first cache, wherein selectable inputs of the write multiplexer are connected to a register of the first processing element and to the second cache.

8. The computing device of claim 7 , wherein the first cache of the first processing element is connected to a third cache of a third processing element that is adjacent the first processing element in the array of processing elements, wherein the first processing element is positioned between the second and third processing elements in the array of processing elements.

9. The computing device of claim 7 , wherein:

the array of processing elements is configured to perform a multiplying accumulation; and

the computing device further comprises a controller to control the write multiplexer to write to the first cache accumulated results and/or input values of the multiplying accumulation.

10. The computing device of claim 7 , further comprising a read multiplexer including an output connected to a register of the first processing element, wherein selectable inputs of the read multiplexer are connected to the first cache and to the second cache.

11. The computing device of claim 10 , wherein:

the array of processing elements is configured to perform a multiplying accumulation; and

the computing device further comprises a controller to control the read multiplexer to read from the first cache coefficients and/or input values of the multiplying accumulation.

12. A computing device comprising:

an array of processing elements mutually connected to perform single instruction multiple data (SIMD) operations;

memory cells connected to each processing element to store data related to the SIMD operations;

a cache connected to each processing element to cache data related to the SIMD operations, wherein a first cache of a first processing element is connected to a second cache of a second processing element that is adjacent the first processing element in the array of processing elements; and

a read multiplexer including an output connected to a register of the first processing element, wherein selectable inputs of the read multiplexer are connected to the first cache and to the second cache.

13. The computing device of claim 12 , wherein the first cache of the first processing element is connected to a third cache of a third processing element that is adjacent the first processing element in the array of processing elements, wherein the first processing element is positioned between the second and third processing elements in the array of processing elements.

14. The computing device of claim 12 , wherein:

the array of processing elements is configured to perform a multiplying accumulation; and

the computing device further comprises a controller to control the read multiplexer to read from the first cache coefficients and/or input values of the multiplying accumulation.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2026
From: UNTETHER AI CORPORATION
To: AT-MEMORY COMPUTING LP
Reel/Frame 075495/0905 →
RELEASE OF SECURITY INTEREST Recorded Jun 17, 2025
From: NATIONAL BANK OF CANADA
To: UNTETHER AI CORPORATION
Reel/Frame 071655/0897 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: SNELGROVE, WILLIAM MARTIN; SCOBBIE, JONATHAN
To: UNTETHER AI CORPORATION
Reel/Frame 058406/0043 →
Continuity (2)
Provisional Application 62983076 · Feb 28, 2020
Related Publication 20210271631A1 · Sep 2, 2021