IP Library › Granted Patent US 12,361,091
Granted Patent B1
US 12,361,091 · App. 18/922,976 · Granted Jul 15, 2025

Tensor parallel group

Inventor: Gavin Uberti (Kirkland, WA)
Assignee: ETCHED.AI INC.
G06F17/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,091
App. No.
18/922,976
Granted
Jul 15, 2025
Kind
B1
Abstract

A tensor parallel group including multiple processing devices separated into a first set of two or more of the processing devices and a second set of two or more of the processing devices. The tensor parallel group may also include multiple communication channels to directly communicatively couple every processing device in the first set of the processing devices with every processing device in the second set of the processing devices without communicatively coupling any of the processing devices in the same set of the processing devices. In these and other embodiments, the processing devices may be configured such that each of the processing devices may be able to communicate with any of the other of the processing devices through at most one other of the processing devices.

Claims (35)

1. A tensor parallel group comprising:

a plurality of processing devices separated into a first set of four or more of the plurality of processing devices and a second set of four or more of the plurality of processing devices; and

a plurality of communication channels to directly communicatively couple every processing device in the first set of the plurality of processing devices with every processing device in the second set of the plurality of processing devices without communicatively coupling any of the plurality of processing devices in the same set of the plurality of processing devices,

wherein the plurality of processing devices are configured such that each of the plurality of processing devices is able to communicate with any of the other of the plurality of processing devices through at most one other of the plurality of processing devices.

2. The tensor parallel group of claim 1 , wherein there is no intersection of processing devices between the first set and the second set.

3. The tensor parallel group of claim 1 , wherein each of the plurality of processing devices are communicatively coupled to a same number of the plurality of processing devices.

4. The tensor parallel group of claim 1 , wherein each of the plurality of processing devices is coupled to a same number of communication channels.

5. The tensor parallel group of claim 1 , wherein each of the plurality of communication channels is configured for a same data bandwidth.

6. The tensor parallel group of claim 1 , wherein each of the plurality of processing devices is configured to simultaneously transmit data and receive data over different ones of the plurality of communication channels.

7. The tensor parallel group of claim 1 , wherein each of the processing devices includes a systolic array of data processing units.

8. The tensor parallel group of claim 1 , wherein the plurality of communication channels are separated into a first subset of communication channels and a second subset of communication channels and each of the plurality of processing devices are coupled to at least one communication channel of the first subset of communication channels and at least one communication channel of the second subset of communication channels, and

the plurality of processing devices are configured to perform an operation that includes a first sub-operation and a second sub-operation that is different than the first sub-operation and data transfer for the first sub-operation occurs only via the first subset of communication channels and data transfer for the second sub-operation occurs only via the second subset of communication channels.

9. The tensor parallel group of claim 8 , wherein each of the plurality of processing devices are configured to perform the first sub-operation and the second sub-operation in overlapping time periods such that first data from the first sub-operation is transferred over the first subset of communication channels at the same time second data from the second sub-operation is transferred over the second subset of communication channels.

10. The tensor parallel group of claim 9 , wherein the operation is a matrix multiplication.

11. The tensor parallel group of claim 10 , wherein the first sub-operation is a reduction operation and the second sub-operation is a gather operation.

12. The tensor parallel group of claim 1 , wherein a number of the plurality of processing devices is a multiple of two.

13. The tensor parallel group of claim 12 , wherein the number of the plurality of processing devices is eight.

14. A system comprising:

a plurality of tensor parallel groups, each of the tensor parallel groups comprising:

a plurality of processing devices separated into a first set of four or more of the plurality of processing devices and a second set of four or more of the plurality of processing devices; and

a plurality of communication channels to directly communicatively couple every processing device in the first set of the plurality of processing devices with every processing device in the second set of the plurality of processing devices without communicatively coupling any of the plurality of processing devices in the same set of the plurality of processing devices,

wherein the plurality of processing devices are configured such that each of the plurality of processing devices is able to communicate with any of the other of the plurality of processing devices through at most one other of the plurality of processing devices.

15. The system of claim 14 , wherein two or more of the plurality of tensor parallel groups are arranged in a parallel pipeline configuration.

16. The system of claim 14 , wherein the plurality of processing devices groups are configured to process data in parallel.

17. The system of claim 14 , wherein the plurality of processing devices groups are configured to process data in parallel and two or more of the plurality of tensor parallel groups are arranged in a parallel pipeline configuration.

18. A tensor parallel group comprising:

a plurality of processing devices separated into a first set of two or more of the plurality of processing devices that includes a first group of processing devices and a second group of processing devices and a second set of two or more of the plurality of processing devices that includes a third group of processing devices and a fourth group of processing devices;

a plurality of first communication channels to directly communicatively couple every processing device in the first group of processing devices to every processing device in the third group of processing devices and to directly communicatively couple every processing device in the second group of processing devices to every processing device in the fourth group of processing devices; and

a plurality of second communication channels to directly communicatively couple every processing device in the first group of processing devices to every processing device in the fourth group of processing devices and to directly communicatively couple every processing device in the second group of processing devices to every processing device in the third group of processing devices,

wherein the plurality of processing devices are configured to perform an operation that includes a first sub-operation and a second sub-operation, wherein data transfer for the first sub-operation occurs only via the plurality of first communication channels and data transfer for the second sub-operation occurs only via the plurality of second communication channels.

19. The tensor parallel group of claim 18 , wherein the plurality of first and second communication channels directly communicatively couple every processing device in the first set of the plurality of processing devices with every processing device in the second set of the plurality of processing devices without communicatively coupling any of the plurality of processing devices in the same set of the plurality of processing devices.

20. The tensor parallel group of claim 18 , wherein each of the plurality of processing devices is coupled to a first number of the plurality of first communication channels and a second number of the plurality of second communication channels.

21. The tensor parallel group of claim 20 , wherein the first number and the second number are the same.

22. The tensor parallel group of claim 18 , wherein the plurality of processing devices are configured such that each of the plurality of processing devices is able to communicate with any of the other of the plurality of processing devices through at most one other of the plurality of processing devices.

23. The tensor parallel group of claim 18 , the operation is a matrix multiplication, the first sub-operation is a reduction operation, and the second sub-operation is a gather operation.

Assignments (2)
SECURITY INTEREST Recorded Jul 22, 2025
From: ETCHED.AI, INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 071792/0869 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2024
From: UBERTI, GAVIN
To: ETCHED.AI INC.
Reel/Frame 069062/0556 →
References Cited (7)
US 11500802B1 · Xu · 2022 [cited by examiner]
US 20220294848A1 · Matthews · 2022 [cited by examiner]
US 20220308890A1 · Han · 2022 [cited by examiner]
CN 116861966A · 2023 [cited by examiner]
Of H. Wang et al., PrimePar: Efficient Spatial-temporal Tensor Partitioning for Large Transformer Model Training, APLOS, ACM May 2024 (Year: 2024). [cited by examiner]
H. Kim et al., TCP: A Tensor Contraction Processor for AI Workloads, 2024 ACM/IEEE 51st Annual International Symposium on Computer Architecture (ISCA), Aug. 2024 (Year: 2024). [cited by examiner]
Shoeybi et al. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism, arXiv:1909.08053v4 [cs.CL], 15 pages, Mar. 13, 2020. [cited by applicant]