IP Library › Granted Patent US 12,748,715
Granted Patent B1
US 12,748,715 · App. 18/950,623 · Granted Sep 29, 2026

Digital switch between chiplet devices for handling transformer workloads

Inventors: Nithesh Kurella (Santa Clara, CA); Akhil Arunkumar (Santa Clara, CA); Jayaprakash Balachandran (Santa Clara, CA); Aayush Ankit (Santa Clara, CA)
Assignee: d-MATRIX CORPORATION
G06F13/4027G06F13/4022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,748,715
App. No.
18/950,623
Granted
Sep 29, 2026
Kind
B1
Abstract

An active bridge module is present between substrates of an apparatus that accelerates AI computation workload. In response to an input, the active bridge module is configured to dynamically configure links between packages on the different substrates (e.g., PCIe cards). The active bridge module may comprise a network of digital switches. In certain embodiments, the active bridge module employs optical conversion with links comprising optical waveguides and/or optical cables. For a certain input, links can be dynamically established between packages on either of the adjacent PCIe boards (e.g., accommodating TP-4 tensor parallelism). For a different input to the module, links can be dynamically established between individual packages on each of the PCIe boards (e.g., accommodating TP-2 tensor parallelism). Embodiments are not limited to implementing parallelism of any particular degree, and can scale generally to implement parallel processing utilizing numbers of package(s) and/or their supporting substrate(s) that are greater than two.

Claims (73)

1 . An AI Accelerator apparatus configured with in-memory compute, the apparatus comprising:

a first substrate supporting,

first N chiplets, where N is an integer greater than 1, each of the chiplets comprising a first plurality of tiles, and each of the tiles comprising:

a first plurality of slices,

a first central processing unit (CPU), and

a first hardware dispatch device;

a first plurality of die-to-die (D2D) interconnects;

a first global reduced instruction set computer (RISC) interface coupled to each of the CPUs in each of the first tiles;

a first digital in memory compute (DIMC) matrix multiplier device configured within one or more portions of each of the first plurality of slices to allow for a high throughput of one or more matrix computations provided in the first DIMC;

a second substrate supporting,

second N chiplets, where N is an integer greater than 1, each of the chiplets comprising a second plurality of tiles, and each of the tiles comprising:

a second plurality of slices,

a second central processing unit (CPU), and

a second hardware dispatch device;

a second plurality of die-to-die (D2D) interconnects;

a second global reduced instruction set computer (RISC) interface coupled to each of the CPUs in each of the second tiles;

a second digital in memory compute (DIMC) matrix multiplier device configured within one or more portions of each of the second plurality of slices to allow for a high throughput of one or more matrix computations provided in the second DIMC; and

an active bridge module comprising a network of digital switches, a first plurality of ports in communication with a plurality of chiplets on the first substrate, and a second plurality of ports in communication with a plurality of chiplets on the second substrate,

wherein in response to a first received input, the network of digital switches is dynamically configurable to a first state providing a first link between one or more of the first plurality of N chiplets on the first substrate, and one or more of the second plurality of N chiplets on the second substrate, in order to accommodate a tensor parallelism of a first type; and

wherein in response to a second received input, the network of digital switches is dynamically configurable to a second state providing first and second links between one or more of the first plurality of N chiplets on the first substrate, and one or more of the second plurality of N chiplets on the second substrate, in order to accommodate a tensor parallelism of a second type.

2 . The system of claim 1 wherein the active bridge module comprises:

an analog to digital converter in communication with the first link; and

a digital to analog converter in communication with the first link.

3 . The system of claim 1 wherein the active bridge module comprises:

a transmitter comprising a light source and an optical modulator in communication with the first link;

a receiver comprising a photodetector in communication with the first link; and

the first link comprises an optical waveguide.

4 . The system of claim 3 wherein the optical modular comprises a Mach Zehnder modulator.

5 . The system of claim 3 wherein the first link comprises an optical cable.

6 . The system of claim 3 wherein the first link comprises a silicon waveguide on the active bridge module.

7 . The system of claim 1 wherein:

the first substrate is a first Peripheral Component Interconnect express (PCIe) card and includes a first PCIe bus; and

the second substrate is a second Peripheral Component Interconnect express (PCIe) card and includes a second PCIe bus.

8 . The system of claim 1 wherein the tensor parallelism of the first type is TP-2.

9 . The system of claim 1 wherein the active bridge module is configured on a board.

10 . The system of claim 1 wherein the tensor parallelism of the second type is TP-4.

11 . The system of claim 1 further comprising a passive bridge module disposed on a board and providing a fixed link between one or more of the first plurality of N chiplets on the first substrate, and one or more of the second plurality of N chiplets on the second substrate.

12 . A method comprising:

providing an input to an active bridge module present between a first substrate and a second substrate of an AI Accelerator apparatus configured with in-memory compute,

the first substrate supporting,

first N chiplets, where N is an integer greater than 1, each of the chiplets comprising a first plurality of tiles, and each of the tiles comprising:

a first plurality of slices,

a first central processing unit (CPU), and

a first hardware dispatch device;

a first plurality of die-to-die (D2D) interconnects;

a first global reduced instruction set computer (RISC) interface coupled to each of the CPUs in each of the first tiles;

a first digital in memory compute (DIMC) matrix multiplier device configured within one or more portions of each of the first plurality of slices to allow for a high throughput of one or more matrix computations provided in the first DIMC;

the second substrate supporting,

second N chiplets, where N is an integer greater than 1, each of the chiplets comprising a second plurality of tiles, and each of the tiles comprising:

a second plurality of slices,

a second central processing unit (CPU), and

a second hardware dispatch device;

a second plurality of die-to-die (D2D) interconnects;

a second global reduced instruction set computer (RISC) interface coupled to each of the CPUs in each of the second tiles;

a second digital in memory compute (DIMC) matrix multiplier device configured within one or more portions of each of the second plurality of slices to allow for a high throughput of one or more matrix computations provided in the second DIMC,

an active bridge module comprising a network of digital switches, a first plurality of ports in communication with a plurality of chiplets on the first substrate, and a second plurality of ports in communication with a plurality of chiplets on the second substrate,

wherein in response to the received input, the network of digital switches is dynamically configurable to a first state providing a first link between at least one of the first plurality of N chiplets on the first substrate, and at least one of the second plurality of N chiplets on the second substrate, in order to accommodate a tensor parallelism of a first type; and

wherein in response to a second received input, the network of digital switches is dynamically configurable to a second state providing first and second links between one or more of the first plurality of N chiplets on the first substrate, and one or more of the second plurality of N chiplets on the second substrate, in order to accommodate a tensor parallelism of a second type.

13 . The method of claim 12 wherein the active bridge module comprises:

an analog to digital converter in communication with the first link; and

a digital to analog converter in communication with the first link.

14 . The method of claim 12 wherein the active bridge module comprises:

a transmitter comprising a light source and an optical modulator in communication with the first link;

a receiver comprising a photodetector in communication with the first link; and

the first link comprises an optical waveguide.

15 . The method of claim 14 wherein the optical modular comprises a Mach Zehnder modulator.

16 . The method of claim 12 wherein:

the first substrate is a first Peripheral Component Interconnect express (PCIe) card and includes a first PCIe bus; and

the second substrate is a second Peripheral Component Interconnect express (PCIe) card and includes a second PCIe bus.

17 . The method of claim 12 wherein the tensor parallelism of the first type is TP-2.

18 . The method of claim 12 wherein the active bridge module is configured on a board.

19 . The method of claim 12 wherein the tensor parallelism of the second type is TP-4.

20 . The method of claim 12 further comprising a passive bridge module disposed on a board and providing a fixed link between one or more of the first plurality of N chiplets on the first substrate, and one or more of the second plurality of N chiplets on the second substrate.

References Cited (3)
US 11488935B1 · Zaman · 2022 [cited by examiner]
US 12361262B1 · Uberti · 2025 [cited by examiner]
US 20230058355A1 · Hornung et al. · 2023 [cited by applicant]