IP Library Granted Patent US 12,561,279
Granted Patent B2
US 12,561,279 · App. 18/731,952 · Granted Feb 24, 2026

Deterministic memory for tensor streaming processors

Inventor: Dennis Charles Abts (Eau Claire, WI)
Assignee: Groq, Inc.
G06F15/80G06F12/023G06F12/0207G06F12/0813G06F2212/251G06F2212/454G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,279
App. No.
18/731,952
Granted
Feb 24, 2026
Kind
B2
Abstract

Embodiments are directed to a deterministic streaming system with one or more deterministic streaming processors each having an array of processing elements and a first deterministic memory coupled to the processing elements. The deterministic streaming system further includes a second deterministic memory with multiple data banks having a global memory address space, and a controller. The controller initiates retrieval of first data from the data banks of the second deterministic memory as a first plurality of streams, each stream of the first plurality of streams streaming toward a respective group of processing elements of the array of processing elements. The controller further initiates writing of second data to the data banks of the second deterministic memory as a second plurality of streams, each stream of the second plurality of streams streaming from the respective group of processing elements toward a respective data bank of the second deterministic memory.

Claims (42)

1 . A system comprising:

a deterministic processor comprising:

a plurality of processing elements, and

a first memory communicatively coupled to the plurality of processing elements;

a second memory communicatively coupled with the plurality of processing elements, the second memory comprising a plurality of data banks having a global memory address space for the deterministic processor; and

a controller communicatively coupled with the second memory, the controller configured to:

facilitate retrieval of first data from the plurality of data banks as a first plurality of streams, wherein respective first streams of the first plurality of streams are configured to stream toward a respective group of processing elements of the plurality of processing elements, and

facilitate writing of second data to the plurality of data banks as a second plurality of streams, wherein respective second streams of the second plurality of streams are configured to stream from the respective group of processing elements toward respective data banks of the plurality of data banks;

wherein the controller is further configured to facilitate global visibility of a write request at a memory address.

2 . The system of claim 1 , wherein the second memory further comprises global memory address spaces, including the global memory address space, for one or more deterministic processors, including the deterministic processor, of the system.

3 . The system of claim 1 , further comprising:

a compiler configured to manage initiation of retrieval of the first data from the second memory at a defined first time.

4 . The system of claim 3 , wherein the defined first time is selected based on a determination that the first data from the second memory are timely placed on the first plurality of streams.

5 . The system of claim 3 , wherein the compiler is further configured to manage initiation of writing of the second data to the second memory at a defined second time.

6 . The system of claim 1 , wherein the deterministic processor comprises a plurality of functional units, wherein functional units of the plurality of functional units operate independently and comprise respective instruction control units that execute instructions, in order, from instruction buffers.

7 . The system of claim 1 , wherein, based on the global visibility and a subsequent request for retrieval of data from the memory address, the controller is further configured to retrieve a latest value written in the memory address.

8 . The system of claim 1 , further comprising:

a plurality of deterministic processors, including the deterministic processor, organized as a node of the system.

9 . The system of claim 8 , wherein the controller is further configured to:

assign respective deterministic processors of the plurality of deterministic processors to a respective pair of pseudo channels of a plurality of pseudo channels of the second memory; and

initiate streaming of data between the respective deterministic processors and a respective data bank of the plurality of data banks of the second memory associated with the respective pair of pseudo channels.

10 . The system of claim 8 , wherein a plurality of nodes, including the node, are organized as a rack of the system.

11 . The system of claim 10 , wherein the controller is further configured to:

assign respective nodes of the plurality of nodes to a respective pair of pseudo channels of a plurality of pseudo channels of the second memory; and

initiate streaming of data between the respective nodes and a respective data bank of the plurality of data banks of the second memory associated with the respective pair of pseudo channels.

12 . The system of claim 1 , wherein the system further comprises a plurality of racks, wherein respective racks of the plurality of racks comprise a plurality of nodes, and wherein respective nodes of the plurality of nodes comprise a plurality of deterministic processors, including the deterministic processor; and

the controller is further configured to:

assign each of the racks to a respective pair of pseudo channels of a plurality of pseudo channels of the second memory, and

initiate streaming of data between each of the racks and a respective data bank of the plurality of data banks of the second memory associated with the respective pair of pseudo channels.

13 . The system of claim 1 , wherein the first memory comprises a static memory and the second memory comprises a dynamic memory.

14 . The system of claim 13 , wherein the dynamic memory comprises one or more three-dimensional stacks of a high bandwidth memory device.

15 . A method, comprising:

facilitating, by a controller of a deterministic processor, retrieval of first data from a plurality of data banks of a first memory as a first plurality of streams, wherein respective first streams of the first plurality of streams are configured to stream toward a respective group of processing elements of an array of processing elements of a second memory;

facilitating, by the controller, writing of second data to the plurality of data banks as a second plurality of streams, wherein respective second streams of the second plurality of streams are configured to stream from the respective group of processing elements toward respective data banks of the plurality of data banks; and

facilitating, by the controller, global visibility of a write request at a memory address.

16 . The method of claim 15 , further comprising:

initiating, by a compiler associated with the deterministic processor, first data from the second memory at a defined first time, wherein the defined first time is selected based on a determination that the first data from the second memory are timely placed on the first plurality of streams; and

managing, by the compiler, initiation of writing of the second data to the second memory at a defined second time.

17 . The method of claim 15 , further comprising:

receiving, by the controller, a subsequent request for the memory address; and

based on the global visibility and the subsequent request, retrieving, by the controller, a latest value written in the memory address.

18 . The method of claim 15 , wherein the first memory comprises a dynamic memory and the second memory comprises a static memory, wherein the dynamic memory comprises one or more three-dimensional stacks of a high bandwidth memory device.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2024
From: ABTS, DENNIS CHARLES
To: GROQ, INC.
Reel/Frame 067600/0408 →
Continuity (3)
Continuation 17858493 · Jul 6, 2022
Provisional Application 63219145 · Jul 7, 2021
Related Publication 20240320185A1 · Sep 26, 2024
References Cited (10)
US 11243880B1 · Ross · 2022 [cited by examiner]
US 11360934B1 · Abts · 2022 [cited by examiner]
US 20200125995A1 · Wang · 2020 [cited by examiner]
US 20200234396A1 · Qi · 2020 [cited by examiner]
US 20210081806A1 · Chai · 2021 [cited by examiner]
US 20230024670A1 · Abts · 2023 [cited by examiner]
Abts et al., “A Software-defined Tensor Streaming Multiprocessor for Large-scale Machine Learning”, 49th Annual International Symposium on Computer Architecture, Jun. 18-22, 2022, 14 pages. [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/858,493, dated Sep. 20, 2023, 51 pages. [cited by applicant]
Abts et al., “Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workloads”, ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA), 2020, pp. 145-158. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/858,493, dated Jan. 31, 2024, 56 pages. [cited by applicant]