IP Library Patent Application 19307690
Patent Application
App. No. 19/307,690

SYSTEMS AND METHODS FOR AUTOREGRESSIVE INFERENCE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
19/307,690
Abstract

The present disclosure includes systems and methods for autoregressive inference using one or more compute accelerators. A method includes configuring one or more compute accelerators to implement a processing sequence of a machine learning (ML) model, wherein the ML model includes a plurality of model layers, and wherein the configuring includes mapping, based at least in part on the processing sequence, the plurality of model layers to a plurality of processing regions of the one or more compute accelerators and arranging connections between the plurality of processing regions to form a processing pipeline corresponding to the processing sequence. The method includes, based at least in part on receiving one or more queries, processing, using the ML model, the one or more queries through the processing pipeline.

Claims (53)

1 . A method comprising:

configuring one or more compute accelerators to implement a processing sequence of a machine learning (ML) model, wherein the ML model comprises a plurality of model layers, and wherein the configuring comprises:

mapping, based at least in part on the processing sequence, the plurality of model layers to a plurality of processing regions of the one or more compute accelerators; and

arranging connections between the plurality of processing regions to form a processing pipeline corresponding to the processing sequence; and

based at least in part on receiving one or more queries, processing, using the ML model, the one or more queries through the processing pipeline.

2 . The method of claim 1 , wherein each processing region of the plurality of processing regions comprises a respective plurality of processing elements, and wherein each processing element of the respective plurality of processing elements comprises one or more compute elements and memory positioned proximate to the one or more compute elements.

3 . The method of claim 2 , wherein the one or more compute accelerators comprise a fabric, wherein the fabric is to connect processing elements of the plurality of processing regions, wherein each processing element of the respective plurality of processing elements comprises a router coupled to the fabric, and wherein the arranging the connections comprises arranging fabric elements of the fabric to connect adjacent processing elements of the plurality of processing regions.

4 . The method of claim 1 , wherein a first compute accelerator, of the one or more compute accelerators, comprises one or more processing regions, of the plurality of processing regions, disposed at a substantially whole substrate.

5 . The method of claim 1 , wherein the mapping comprises mapping successive model layers of the processing sequence to adjacent processing regions of the plurality of processing regions.

6 . The method of claim 1 , wherein a first processing region of a first compute accelerator is adjacent to a second processing region of a second compute accelerator, and wherein the arranging the connections comprises:

identifying, at the first processing region, one or more processing elements neighboring the second processing region; and

configuring one or more local communication paths between the one or more identified processing elements and the second processing region.

7 . The method of claim 6 , further comprising retrieving, from local memory of the first processing region via the one or more local communication paths, model data for processing the one or more queries at the second processing region.

8 . The method of claim 1 , wherein the processing, using the ML model, the one or more queries comprises:

determining, based on the one or more queries, a respective sequence of tokens;

generating, at a respective processing region, model data associated with the respective sequence of tokens using a respective model layer; and

storing, in local memory of the respective processing region, the model data.

9 . The method of claim 1 , further comprising:

mapping a de-embedding layer associated with the ML model to a last processing region of the plurality of processing regions for the processing pipeline; and

generating model output using the de-embedding layer.

10 . The method of claim 1 , wherein the ML model is a target model, and wherein the plurality of model layers is a first plurality of model layers, the method further comprising:

determining, based at least in part on the target model, a second plurality of model layers of a draft model;

mapping the second plurality of model layers to the plurality of processing regions;

determining one or more draft tokens by processing, using the draft model concurrently with using the target model, the one or more queries through the processing pipeline; and

validating, using the target model, the one or more draft tokens.

11 . A system comprising:

one or more compute accelerators comprising a plurality of processing regions; and

processing circuitry to:

configure the one or more compute accelerators to implement a processing sequence of a machine learning (ML) model, wherein the ML model comprises a plurality of model layers, and wherein the configuring comprises:

map, based at least in part on the processing sequence, the plurality of model layers to the plurality of processing regions of the one or more compute accelerators; and

arrange connections between the plurality of processing regions to form a processing pipeline corresponding to the processing sequence; and

based at least in part on receiving one or more queries, process, using the ML model, the one or more queries through the processing pipeline.

12 . The system of claim 11 , wherein each processing region of the plurality of processing regions comprises a respective plurality of processing elements, and wherein each processing element of the respective plurality of processing elements comprises one or more compute elements and memory positioned proximate to the one or more compute elements.

13 . The system of claim 12 , wherein the one or more compute accelerators comprise one or more fabrics, wherein the one or more fabrics is to connect processing elements of the plurality of processing regions, wherein each processing element of the respective plurality of processing elements comprises a router coupled to the fabric, and wherein the processing circuitry is further to arrange fabric elements of the one or more fabrics to connect adjacent processing elements of the plurality of processing regions.

14 . The system of claim 11 , wherein a first compute accelerator, of the one or more compute accelerators, comprises one or more processing regions, of the plurality of processing regions, disposed at a substantially whole substrate.

15 . The system of claim 11 , wherein the processing circuitry is to map successive model layers of the processing sequence to adjacent processing regions of the plurality of processing regions.

16 . The system of claim 11 , wherein a first processing region of a first compute accelerator is adjacent to a second processing region of a second compute accelerator, and wherein the processing circuitry is arrange the connections by:

identifying, at the first processing region, one or more processing elements neighboring the second processing region; and

configuring one or more local communication paths between the one or more identified processing elements and the second processing region.

17 . The system of claim 16 , wherein the processing circuitry is further to retrieve, from local memory of the first processing region via the one or more local communication paths, model data for processing the one or more queries at the second processing region.

18 . The system of claim 11 , wherein the processing circuitry is to process, using the ML model, the one or more queries by:

determining, based on the one or more queries, a respective sequence of tokens;

generating, at a respective processing region, model data associated with the respective sequence of tokens using a respective model layer; and

storing, in local memory of the respective processing region, the model data.

19 . The system of claim 11 , wherein the processing circuitry is further to:

map a de-embedding layer associated with the ML model to a last processing region of the plurality of processing regions for the processing pipeline; and

generate model output using the de-embedding layer.

20 . The system of claim 11 , wherein the ML model is a target model, wherein the plurality of model layers is a first plurality of model layers, and wherein the processing circuitry is further to:

determine, based at least in part on the target model, a second plurality of model layers of a draft model;

map the second plurality of model layers to the plurality of processing regions;

determine one or more draft tokens by processing, using the draft model concurrently with using the target model, the one or more queries through the processing pipeline; and

validate, using the target model, the one or more draft tokens.

21 - 30 . (canceled)

Assignments (2)
SECURITY INTEREST Recorded Jun 18, 2026
From: CEREBRAS SYSTEMS INC.
To: MORGAN STANLEY SENIOR FUNDING, INC., AS THE COLLATERAL AGENT
Reel/Frame 075845/0844 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2025
From: JAMES, MICHAEL EDWIN; LIE, SEAN; KIBARDIN, VLADIMIR
To: CEREBRAS SYSTEMS INC.
Reel/Frame 072447/0054 →