IP Library Granted Patent US 11,960,982
Granted Patent B1
US 11,960,982 · App. 17/970,703 · Granted Apr 16, 2024

System and method of determining and executing deep tensor columns in neural networks

Inventors: Alexander Matveev (Cambridge, MA); Nir Shavit (Cambridge, MA); Govind Ramnarayan (Somerville, MA); Tyler Michael Smith (Somerville, MA); Sage Moore (Somerville, MA)
Assignee: NEURALMAGIC, INC.
G06N3/04G06N3/0495G06N20/10G06N3/02G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,960,982
App. No.
17/970,703
Granted
Apr 16, 2024
Kind
B1
Abstract

A system and method may partition and/or execute a NN, by, for a graph including nodes and hyper edges, each node representing a data item in the NN and each hyper edge representing an operation in the NN, identifying a deep tensor column comprising a subset of the nodes and a subset of the hyper edges, such that the operations in the deep tensor column, when executed, use only data which fits within a preselected cache.

Claims (28)

1. A method of partitioning a neural network (NN), the method comprising:

receiving a representation of a NN comprising nodes and edges; and

identifying a deep tensor column comprising a subset of the nodes and a subset of the edges, such that the operations in the deep tensor column, when executed, use only data which fits within a preselected cache and wherein identifying a deep tensor column comprises using a cost function measured by the compute to memory ratio of the deep tensor column, the compute to memory ratio determining a measure of the ratio of compute to memory accesses.

2. The method of claim 1 , where data fits within a preselected cache when, during execution in a processor according to the processor's cache policy, it is expected that a cache eviction is not expected until the end of the execution of the task.

3. The method of claim 1 , wherein partitioning comprises iteratively adding a data item in the NN to a deep tensor column until the adding of a data item causes the deep tensor column to have a cost function exceeding a threshold.

4. The method of claim 1 , wherein partitioning comprises iteratively adding a data item in the NN to a deep tensor column and minimizing the cost of a set of component partitions comprised in the deep tensor column.

5. The method of claim 1 , comprising identifying a plurality of deep tensor columns, and executing the plurality of deep tensor columns in order to execute the NN.

6. The method of claim 1 , wherein the deep tensor column comprises a set of elemental operations which form a portion of a layer of the NN.

7. The method of claim 1 , wherein identifying a deep tensor column comprises adding an operation to a deep tensor column if the compute to memory ratio of the operation is below a threshold.

8. A system for partitioning a neural network (NN), the system comprising:

a memory; and

a processor to:

receive a NN comprising nodes and edges; and

identify a deep tensor column comprising a subset of the nodes and a subset of the edges, such that the operations in the deep tensor column, when executed, use only data which fits within a preselected cache; wherein identifying a deep tensor column comprises using a cost function measured by the compute to memory ratio of the deep tensor column, the compute to memory ratio determining a measure of the ratio of compute to memory accesses.

9. The system of claim 8 , where data fits within a preselected cache when, during execution in a processor according to the processor's cache policy, it is expected that a cache eviction is not expected until the end of the execution of the task.

10. The system of claim 8 , wherein partitioning comprises iteratively adding a data item in the NN to a deep tensor column until the adding of a data item causes the deep tensor column to have a cost function exceeding a threshold.

11. The system of claim 8 , wherein partitioning comprises iteratively adding a data item in the NN to a deep tensor column and minimizing the cost of a set of component deep tensor columns comprised in the deep tensor column.

12. The system of claim 8 , wherein the processor is to identify a plurality of deep tensor columns, and the plurality of deep tensor columns are executed in order to execute the NN.

13. The system of claim 8 , wherein the deep tensor column comprises a set of elemental operations which form a portion of a layer of the NN.

14. The system of claim 8 , wherein identifying a deep tensor column comprises adding an operation to a deep tensor column if the compute to memory ratio of the operation is below a threshold.

15. A method of executing a neural network (NN), the method comprising:

receiving a set NN partitions, each NN partition comprising nodes and edges, such that the operations in each partition, when executed, use only data which fits within a preselected cache, wherein the NN partitions are identified using a cost function measured by the compute to memory ratio of a partition, the compute to memory ratio determining a measure of the ratio of compute to memory accesses; and

executing the plurality of NN partitions in order to execute the NN.

16. The method of claim 15 , where data fits within a preselected cache when, during execution in a processor according to the processor's cache policy, it is expected that a cache eviction is not expected until the end of the execution of the task.

17. The method of claim 15 , comprising partitioning the NN by iteratively adding a data item in the NN to a partition until the adding of a data item causes the partition to have a cost function exceeding a threshold.

18. The method of claim 15 , comprising partitioning the NN by iteratively adding a data item in the NN to partition and minimizing the cost of a set of component partitions comprised in the deep tensor column.

19. The method of claim 15 , wherein the partition comprises a set of elemental operations which form a portion of a layer of the NN.

20. The method of claim 15 , wherein identifying a partition comprises adding an operation to a partition if the compute to memory ratio of the operation is below a threshold.

Assignments (3)
CHANGE OF NAME Recorded Mar 3, 2026
From: RED HAT, INC.
To: RED HAT, LLC
Reel/Frame 074913/0759 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 30, 2025
From: NEURALMAGIC, INC.
To: RED HAT, INC.
Reel/Frame 072278/0309 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 2, 2025
From: MATVEEV, ALEXANDER; SHAVIT, NIR; RAMNARAYAN, GOVIND; SMITH, TYLER; MOORE, SAGE
To: NEURALMAGIC INC.
Reel/Frame 071010/0230 →
Continuity (1)
Provisional Application 63270291 · Oct 21, 2021
Cited By (1)
US 12,626,132