IP Library Granted Patent US 12,664,413
Granted Patent B2
US 12,664,413 · App. 18/989,525 · Granted Jun 23, 2026

Methods for an AI accelerator integrated circuit chip with integrated cell-based fabric adapter

Inventors: Gary S. Goldman (Los Altos, CA); Ramalingam K. Anand (Los Altos Hills, CA); Kalyana S. Venkataraman (San Jose, CA); Berend Ozceri (Los Gatos, CA); Pradeep R. Joginipally (San Jose, CA); Chung Y. Lau (Milpitas, CA); Jigar K. Savla (San Jose, CA); Ashwin Radhakrishnan (Fremont, CA); Michael Davie (St Augustine, FL); Shijun Li (Southborough, MA)
Assignee: TENSORDYNE, INC.
G06N3/063G06F12/1081G06F13/28G06F13/4068G06F13/409G06F15/17331G06F2213/2806
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,413
App. No.
18/989,525
Filed
Dec 20, 2024
Granted
Jun 23, 2026
Kind
B2
Art Unit
2184
USPC
710/22
Abstract

An integrated circuit formed on (i) a single semiconductor die or (ii) a plurality semiconductor dies that are integrated into a single package. The integrated circuit may include a communication interface including a serializer/deserializer (SerDes) interface; a fabric adapter communicatively coupled to the communication interface; a plurality of inference engine clusters, each inference engine cluster including a respective memory element and/or memory interface; and a data interconnect communicatively coupling each respective memory element and/or memory interfaces of the plurality of inference engine clusters to the fabric adapter. The fabric adapter may be configured to facilitate remote direct memory access (RDMA) read and write services and/or datagram communication over a cell-based switch fabric to and from the respective memory elements and/or memory interfaces of the plurality of inference engine clusters via the data interconnect.

Claims (37)

1 . A method for transmitting data from a first artificial intelligence (AI) chip to a second AI chip over a cell-based switch fabric, the first AI chip being a leaf node of a Clos network and serving as a source endpoint, the first AI chip comprising a virtual output queue (VOQ) subsystem with a plurality of VOQs,

wherein the first AI chip comprises:

a communication interface;

a fabric adapter communicatively coupled to the communication interface, wherein the fabric adapter includes the VOQ subsystem;

an inference engine cluster including a memory element or a memory interface; and

a data interconnect communicatively coupling the memory element or the memory interface of the inference engine cluster to the fabric adapter,

the method comprising:

dividing, by the first AI chip, the data into packets; and

for each of the packets:

determining, by the first AI chip, a respective one of the VOQs to assign to the packet;

attaching, by the first AI chip, a packet header to the packet;

enqueuing, by the first AI chip, the packet to the respective one of the VOQs assigned to the packet; and

when selected for transport to the second AI chip, dequeuing, by the first AI chip, the packet from the respective one of the VOQs, cellifying, by the first AI chip, the packet, attaching, by the first AI chip, a cell header to each cell to indicate a destination OQ on the second AI chip corresponding to the assigned VOQ, and transmitting, by the first AI chip as the source endpoint, the cells to one or more cell-fabric switch chips.

2 . The method of claim 1 , wherein the packet header is used by the second AI chip to identify the data in the packet.

3 . The method of claim 1 , wherein cellifying the packet comprises dividing data contained in the packet into regular sized blocks of data.

4 . The method of claim 1 , further comprising maintaining a count of a total number of packets stored at each of the VOQs.

5 . The method of claim 1 , wherein enqueuing the packet to the respective one of the VOQs assigned to the packet comprises issuing, by a packet generator, an enqueue command to the VOQ subsystem, the enqueue command including an identifier of the packet, a length of the packet and a target VOQ number.

6 . The method of claim 1 , wherein enqueuing the packet to the respective one of the VOQs comprises communicating the packet from the memory element of the first AI chip to the respective one of the VOQs via a crossbar switch.

7 . A method for receiving data at a first artificial intelligence (AI) chip comprising an output queue (OQ) subsystem with a plurality of output queues (OQs), the first AI chip being a leaf node of a Clos network and serving as a destination endpoint, the OQs being associated with a VOQ subsystem on a second AI chip, wherein the first AI chip comprises:

a communication interface;

a fabric adapter communicatively coupled to the communication interface, wherein the fabric adapter includes the OQ subsystem;

an inference engine cluster including a memory element or a memory interface; and

a data interconnect communicatively coupling the memory element or the memory interface of the inference engine cluster to the fabric adapter,

the method comprising:

receiving, by the first AI chip, a plurality of packets from the second AI chip via one or more cell-fabric switch chips;

enqueueing, by the first AI chip, the plurality of packets to the OQs;

dequeuing, by the first AI chip, one or more of the packets from the OQs according to a scheduling policy based on a class of service assigned to each of the OQs; and

for each of the packets,

decellifying, by the first AI chip as the destination endpoint, respective cells of the packet to reconstitute packet data;

examining, by the first AI chip, a header of the packet to identify the packet data and determine its appropriate processing;

removing, by the first AI chip, the header of the packet; and

storing, by the first AI chip, the packet data to the memory element or communicating the packet data to the memory interface.

8 . The method of claim 7 , wherein the cells of a first one of the packets are decellified either before enqueueing the first packet to one of the OQs, or after the first packet has been dequeued from one of the OQs.

9 . The method of claim 7 , further comprising, prior to decellifying respective cells of the packet, reordering the respective cells of the packet.

10 . The method of claim 7 , wherein the scheduling policy is a round-robin policy, a weighted round-robin policy, a strict priority policy, or a combination of one or more of the aforementioned policies.

11 . The method of claim 7 , wherein the header of the packet includes one or more of receive and transmit remote direct memory access (RDMA) job identifiers which identify a specific RDMA transfer, a packet sequence number within a job, a last-packet-in-job flag, a packet type in order to indicate whether the packet contains RDMA data, datagram data, or a hardware message.

12 . The method of claim 7 , further comprising, prior to storing the packet data to the memory element or communicating the packet data to the memory interface, transmitting the packet data to the data interconnect configured to route the packet data to the memory element or memory interface, wherein the data interconnect comprises one or more of a crossbar switch, a ring, a torus interconnect or a mesh.

Assignments (10)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: ANAND, RAMALINGAM K.
To: RECOGNI INC.
Reel/Frame 069766/0697 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: DAVIE, MICHAEL
To: RECOGNI INC.
Reel/Frame 069766/0734 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: GOLDMAN, GARY S.
To: RECOGNI INC.
Reel/Frame 069766/0782 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: JOGINIPALLY, PRADEEP R
To: RECOGNI INC.
Reel/Frame 069766/0785 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: LAU, CHUNG Y.
To: RECOGNI INC.
Reel/Frame 069766/0832 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: LI, SHIJUN
To: RECOGNI INC.
Reel/Frame 069766/0835 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: OZCERI, BEREND
To: RECOGNI INC.
Reel/Frame 069766/0838 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: RADHAKRISHNAN, ASHWIN
To: RECOGNI INC.
Reel/Frame 069766/0852 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: SAVLA, JIGAR K.
To: RECOGNI INC.
Reel/Frame 069766/0868 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2025
From: VENKATARAMAN, KALYANA S.
To: RECOGNI INC.
Reel/Frame 069766/0908 →
Continuity (2)
Provisional Application 63694397 · Sep 13, 2024
Related Publication 20260079868A1 · Mar 19, 2026
References Cited (50)
US 8145785B1 · Finkelstein et al. · 2012 [cited by applicant]
US 10721187B1 · Goldman · 2020 [cited by examiner]
US 11025544B2 · Marolia et al. · 2021 [cited by applicant]
US 11194753B2 · Marolia · 2021 [cited by examiner]
US 11321254B2 · Kim · 2022 [cited by examiner]
US 11880289B2 · Smith · 2024 [cited by examiner]
US 11916800B2 · Arditti Ilitzky · 2024 [cited by examiner]
US 12014130B2 · Ting · 2024 [cited by examiner]
US 12309070B2 · Klenk · 2025 [cited by examiner]
US 12430547B1 · Goldman · 2025 [cited by examiner]
US 12436896B1 · Goldman · 2025 [cited by examiner]
US 20020006110A1 · Brezzo et al. · 2002 [cited by applicant]
US 20110219208A1 · Asaad · 2011 [cited by examiner]
US 20130117621A1 · Saraiya et al. · 2013 [cited by applicant]
US 20130117766A1 · Bax et al. · 2013 [cited by applicant]
US 20140281335A1 · Xu · 2014 [cited by examiner]
US 20150055649A1 · DeCusatis et al. · 2015 [cited by applicant]
US 20150339570A1 · Scheffler · 2015 [cited by examiner]
US 20180100201A1 · Garraway · 2018 [cited by examiner]
US 20190042518A1 · Marolia · 2019 [cited by examiner]
US 20190087708A1 · Goulding · 2019 [cited by examiner]
US 20190121761A1 · Yuenyongsgool · 2019 [cited by examiner]
US 20200304427A1 · Sandler · 2020 [cited by examiner]
US 20200387564A1 · Simpson · 2020 [cited by applicant]
US 20200412659A1 · Arditti Ilitzky · 2020 [cited by examiner]
US 20210089696A1 · Ting · 2021 [cited by examiner]
US 20210117246A1 · Lal · 2021 [cited by examiner]
US 20210117360A1 · Kutch · 2021 [cited by examiner]
US 20210334184A1 · Menon et al. · 2021 [cited by applicant]
US 20220200912A1 · Bataineh et al. · 2022 [cited by applicant]
US 20220276973A1 · Wei · 2022 [cited by examiner]
US 20220414443A1 · Li · 2022 [cited by examiner]
US 20230231811A1 · Dalal · 2023 [cited by examiner]
US 20230333999A1 · Omer · 2023 [cited by examiner]
US 20230409881A1 · Heitzmann · 2023 [cited by examiner]
US 20250047560A1 · Friedman · 2025 [cited by examiner]
“AI Networking”, Arista, White Paper, Jun. 5, 2024, 14 pgs. [cited by applicant]
Biglari; et al., “Designing Reconfigurable Interconnection Network of Heterogeneous Chiplets Using Kalman Filter”, Cornell University, arXiv:2406.00568v1 [cs.AR] Jun. 1, 2024, 6 pgs. [cited by applicant]
Gao; et al., “Customized High Performance and Energy Efficient Communication Networks for AI Chips”, IEEE Access, May 17, 2019 (revised Jun. 10, 2019), vol. 7, 69434-69446. [cited by applicant]
Mabavinejad; et al., “An Overview of Efficient Interconnection Networks for Deep Neural Network Accelerators”, IEEE Journal on Emerging and Selected Topics in Circuits and Systems, Sep. 2020, 10(3):268-282. [cited by applicant]
Rettkowski, Jens, M.Sc., “Design and Programming Methods for Reconfigurable Multi-Core Architectures using a Network-on-Chip-Centric Approach”, Technische Universitat Dresden, dissertation, Sep. 9, 2021, pp. 1-246. [cited by applicant]
“Season 3 Ep 11: Cell-based AI Fabric”, DriveNets, CloudNets Video accessed Jul. 1, 2024, transcription, 4 pgs. [cited by applicant]
Non-Final Office Action dated May 2, 2025, for U.S. Appl. No. 18/989,508, (filed Dec. 20, 2024), 19 pgs. [cited by applicant]
Non-Final Office Action dated Mar. 14, 2025, for U.S. Appl. No. 18/989,492, (filed Dec. 20, 2024), 10 pgs. [cited by applicant]
Amendment filed Jun. 11, 2025, for U.S. Appl. No. 18/989,492, (filed Dec. 20, 2024), 10 pgs. [cited by applicant]
Notice of Allowance mailed Aug. 11, 2025, for U.S. Appl. No. 18/989,492, (filed Dec. 20, 2024), 9 pgs. [cited by applicant]
Amendment filed Jul. 3, 2025, for U.S. Appl. No. 18/989,508, (filed Dec. 20, 2024), 9 pgs. [cited by applicant]
Notice of Allowance mailed Jul. 22, 2025, for U.S. Appl. No. 18/989,508, (filed Dec. 20, 2024), 9 pgs. [cited by applicant]
International Search Report and Written Opinion mailed Feb. 26, 2026, from the ISA/European Patent Office, for International Patent Application No. PCT/US2025/043474 (filed Aug. 26, 2025), 18 pp. [cited by applicant]
Tork; et al., “Lynx: A SmartNIC-driven Accelerator-centric Architecture for Network Servers”, ASPLOS'20, Mar. 16-20, 2020, Lausanne, Switzerland, pp. 117-131. [cited by applicant]