IP Library › Granted Patent US 12,361,262
Granted Patent B1
US 12,361,262 · App. 18/922,984 · Granted Jul 15, 2025

Tensor operations in AI models

Inventor: Gavin Uberti (Kirkland, WA)
Assignee: ETCHED.AI INC.
G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,262
App. No.
18/922,984
Granted
Jul 15, 2025
Kind
B1
Abstract

A method of performing computations for artificial intelligence models may include obtaining an input tensor based on an input to an artificial intelligence model and loading the input tensor into multiple processing devices. The input tensor may be split into multiple input tensor tiles that are distributed among the processing devices such that each of the processing devices does not include an entirety of the input tensor. The method may also include performing multiple tensor operations according to the artificial intelligence model to generate multiple intermediate tensors and an output tensor, one or more of the tensor operations performed using the input tensor.

Claims (32)

1. A method of performing computations for artificial intelligence models, the method comprising:

obtaining an input tensor based on an input to an artificial intelligence model;

loading the input tensor into a plurality of processing devices, the input tensor split into a plurality of input tensor tiles that are distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the input tensor; and

performing a plurality of tensor operations according to the artificial intelligence model to generate a plurality of intermediate tensors and an output tensor, one or more of the plurality of tensor operations performed using the input tensor and each of the plurality of intermediate tensors being split into tensor tiles distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the plurality of intermediate tensors during any of the plurality of tensor operations.

2. The method of claim 1 , further comprising iteratively performing the steps of obtaining, loading, and performing, wherein the input tensor for a subsequent iteration is the output tensor from a previous iteration.

3. The method of claim 1 , wherein the distribution of the tensor tiles of the one or more of the plurality of intermediate tensors among the plurality of processing devices is different than the distribution of the plurality of input tensor tiles among the plurality of processing devices.

4. The method of claim 1 , wherein the output tensor is split into tensor tiles distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the output tensor.

5. The method of claim 1 , wherein the plurality of input tensor tiles are distributed among the plurality of processing devices such that half of the plurality of processing devices each include half of the plurality of input tensor tiles.

6. The method of claim 1 , wherein the artificial intelligence model implements a transformer architecture.

7. The method of claim 6 , wherein the plurality of intermediate tensors include a self-attention tensor, and the self-attention tensor is split into a plurality of intermediate tensor tiles and distributed among the plurality of processing devices such that each of the plurality of processing devices includes a different sub-set of the plurality of intermediate tensor tiles and no duplication of the plurality of intermediate tensor tiles exists among the plurality of processing devices.

8. The method of claim 7 , wherein the plurality of intermediate tensors includes a projection tensor that is split into a plurality of projection tensor tiles and distributed among the plurality of processing devices in the same manner as the input tensor is distributed among the plurality of processing devices.

9. The method of claim 8 , wherein the distribution of the input tensor among the plurality of processing devices is different than the distribution of the self-attention tensor among the plurality of processing devices.

10. A system comprising:

one or more memory devices configured to store a plurality of tensors; and

a plurality of processing devices coupled to the one or more memory devices and configured to perform tensor operations on the plurality of tensors, the system configured to execute instructions to cause the system to perform operations, the operations comprising:

obtaining an input tensor based on an input to an artificial intelligence model;

loading the input tensor into the plurality of processing devices, the input tensor split into a plurality of input tensor tiles that are distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the input tensor; and

performing, using the plurality of processing devices, a plurality of tensor operations according to the artificial intelligence model to generate a plurality of intermediate tensors and an output tensor, one or more of the plurality of tensor operations performed using the input tensor and each of the plurality of intermediate tensors being split into tensor tiles distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the plurality of intermediate tensors during any of the plurality of tensor operations.

11. The system of claim 10 , wherein the operations further comprise iteratively performing the operations with the input tensor for a subsequent iteration being the output tensor from a previous iteration.

12. The system of claim 10 , wherein the distribution of the tensor tiles of the one or more of the plurality of intermediate tensors among the plurality of processing devices is different than the distribution of the plurality of input tensor tiles among the plurality of processing devices.

13. The system of claim 10 , wherein the output tensor is split into tensor tiles distributed among the plurality of processing devices such that each of the plurality of processing devices does not include an entirety of the output tensor.

14. The system of claim 10 , wherein the plurality of input tensor tiles are distributed among the plurality of processing devices such that half of the plurality of processing devices each include half of the plurality of input tensor tiles.

15. The system of claim 10 , wherein the artificial intelligence model implements a transformer architecture.

16. The system of claim 15 , wherein the plurality of intermediate tensors include a self-attention tensor, and the self-attention tensor is split into a plurality of intermediate tensor tiles and distributed among the plurality of processing devices such that each of the plurality of processing devices includes a different sub-set of the plurality of intermediate tensor tiles and no duplication of the plurality of intermediate tensor tiles exists among the plurality of processing devices.

17. The system of claim 16 , wherein the plurality of intermediate tensors includes a projection tensor that is split into a plurality of projection tensor tiles and distributed among the plurality of processing devices in the same manner as the input tensor is distributed among the plurality of processing devices.

18. The system of claim 17 , wherein the distribution of the input tensor among the plurality of processing devices is different than the distribution of the self-attention tensor among the plurality of processing devices.

19. A method of performing computations for artificial intelligence models, the method comprising:

obtaining an input tensor based on an input to an artificial intelligence model;

loading the input tensor into a plurality of separate processing devices, the input tensor split into a plurality of input tensor tiles that are distributed among the plurality of separate processing devices such that each of the plurality of separate processing devices includes a portion of the plurality of input tensor tiles but does not include an entirety of the input tensor; and

performing a plurality of tensor operations according to the artificial intelligence model to generate a plurality of intermediate tensors and an output tensor, one or more of the plurality of tensor operations performed using the input tensor.

20. The method of claim 19 , wherein at least one of the plurality of intermediate tensors is split into a plurality of intermediate tensor tiles and distributed among the plurality of separate processing devices and the distribution of the input tensor among the plurality of separate processing devices is different than the distribution of the at least one of the plurality of intermediate tensors among the plurality of separate processing devices.

21. The method of claim 20 , wherein the artificial intelligence model implements a transformer architecture and the at least one of the plurality of intermediate tensors includes a self-attention tensor.

Assignments (2)
SECURITY INTEREST Recorded Jul 22, 2025
From: ETCHED.AI, INC.
To: TRIPLEPOINT CAPITAL LLC
Reel/Frame 071792/0869 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2024
From: UBERTI, GAVIN
To: ETCHED.AI INC.
Reel/Frame 069062/0583 →
References Cited (38)
US 9779786B1 · Wu · 2017 [cited by examiner]
US 10956500B2 · Brevdo · 2021 [cited by examiner]
US 12182687B2 · Arthur · 2024 [cited by examiner]
US 20190130276A1 · Shiring · 2019 [cited by examiner]
US 20190356394A1 · Bunandar · 2019 [cited by examiner]
US 20190370652A1 · Shen · 2019 [cited by examiner]
US 20200034710A1 · Sidhu · 2020 [cited by examiner]
US 20200042875A1 · Shazeer · 2020 [cited by examiner]
US 20200117999A1 · Yoon · 2020 [cited by examiner]
US 20200242189A1 · Chatterjee · 2020 [cited by examiner]
US 20200250518A1 · Khoury · 2020 [cited by examiner]
US 20210182077A1 · Chen · 2021 [cited by examiner]
US 20220012575A1 · Fok · 2022 [cited by examiner]
US 20220044096A1 · Asad · 2022 [cited by examiner]
US 20220121903A1 · Zhang · 2022 [cited by examiner]
US 20220121914A1 · Huang · 2022 [cited by examiner]
US 20220366006A1 · Zhang · 2022 [cited by examiner]
US 20220405552A1 · Lichtenau · 2022 [cited by examiner]
US 20220405598A1 · Lichtenau · 2022 [cited by examiner]
US 20230023101A1 · Su · 2023 [cited by examiner]
US 20230297643A1 · Shivam · 2023 [cited by examiner]
US 20230385077A1 · Wang · 2023 [cited by examiner]
US 20240095493A1 · Lin · 2024 [cited by examiner]
US 20240135572A1 · Singh · 2024 [cited by examiner]
US 20240256475A1 · Chen · 2024 [cited by examiner]
US 20240346108A1 · Insalata · 2024 [cited by examiner]
US 20240348703A1 · Byadgi · 2024 [cited by examiner]
US 20240412334A1 · Sarokin · 2024 [cited by examiner]
US 20240428576A1 · Jiang · 2024 [cited by examiner]
Shail Dave et al. ,“Hardware Acceleration of Sparse and Irregular Tensor Computations of ML Models: A Survey and Insights,” Sep. 20, 2021, Proceedings of the IEEE,vol. 109,No. 10,Oct. 2021,pp. 1706-1734. [cited by examiner]
Size Zheng et al.,“FleXTensor: An Automatic Schedule Exploration and Optimization Framework for Tensor Computation on Heterogeneous System,” Mar. 13, 2020, ASPLOS '20: Proceedings of the Twenty-Fifth International Confe… [cited by examiner]
Dong He et al.,“Query Processing on Tensor Computation Runtimes,” Feb. 10, 2023, arXiv:2203.01877,pp. 1-10. [cited by examiner]
Minjie Wang et al.,“Unifying Data, Model and Hybrid Parallelism in Deep Learning via Tensor Tiling,” May 10, 2018, arXiv: 1805.04170v1,pp. 1-8. [cited by examiner]
Dennis Abts et al.,“Think Fast: A Tensor Streaming Processor (TSP) for Accelerating Deep Learning Workload,” Jul. 13, 2020, 2020 ACM/IEEE 47th Annual International Symposium on Computer Architecture (ISCA),pp. 145-156. [cited by examiner]
Yaoyao Ding et al.,“Hidet: Task-Mapping Programming Paradigm for Deep Learning Tensor Programs,” Jan. 30, 2023, ASPLOS 2023: Proceedings of the 28th ACM International Conference on Architectural Support for Programming … [cited by examiner]
Sergio Moreno-Alvarez et al.,“Heterogeneous model parallelism for deep neural networks,” Feb. 23, 2021, Neurocomputing 441( 2021) , pp. 1-8. [cited by examiner]
Noam Shazeer et al.,“Mesh-TensorFlow: Deep Learning for Supercomputers,” Dec. 2-8, 2018, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada,pp. 1-7. [cited by examiner]
Shoeybi et al. Megatron-LM: Training Multi-Billion Parameter Language Models Using Model Parallelism, arXiv:1909.08053v4 [cs. CL], 15 pages, Mar. 13, 2020. [cited by applicant]
Cited By (1)
US 12,748,715