IP Library › Granted Patent US 12,639,582
Granted Patent B2
US 12,639,582 · App. 17/970,450 · Granted May 26, 2026

Method and apparatus for using a packet architecture to process neural networks in a neural processing unit

Inventors: Sharad Vasantrao Chole (San Jose, CA); Shang-Tse Chuang (Los Altos, CA); Siyad Chih-Hua Ma (Palo Alto, CA)
Assignee: Expedera, Inc.
G06N3/091G06F8/433G06F8/451G06F8/453G06N3/063G06F9/30036G06F9/30058G06F9/3455G06F9/3838G06F9/3842G06F9/3863G06F9/3877G06F9/545G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,639,582
App. No.
17/970,450
Granted
May 26, 2026
Kind
B2
Abstract

Artificial intelligence is an increasingly important sector of the computer industry. However, artificial intelligence is an extremely computationally intensive field such that it can be expensive, time consuming, and energy consuming field. Fortunately, many of the calculations required for artificial intelligence can be performed in parallel such that specialized processors can greatly increase computational performance for AI applications. Specifically, artificial intelligence generally requires large numbers of matrix operations such that specialized matrix processor circuits can greatly improve performance. To efficiently execute all these matrix operations, the matrix processor circuits must be quickly and efficiently supplied with a stream of data and instructions to process or else the matrix processor circuits end up idle. Thus, this document discloses packet architecture for efficiently creating and supplying neural network processors with work packets to process.

Claims (25)

1 . A method of processing a multi-layer neural network, the method comprising the steps of:

dividing, with a packet compiler, each neural network layer of a set of neural network layers into a set of individual work packets for the neural network layer, each of the neural network layers comprising a set of input data elements and a set of output data elements, each individual work packet comprising (a) a subset of the set of output data elements for the neural network layer, (b) an associated weight matrix reference identifying weights specific to the subset of the set of output data elements, and (c) metadata comprising at least scheduling information for the individual work packet, the set of individual work packets for the neural network layer collectively covering all of the set of output data elements for the neural network layer;

generating, with the packet compiler, a dependency graph identifying data dependencies between the individual work packets of different neural network layers and creating an optimized packet stream order based on the data dependencies, the optimized packet stream order minimizing sequential processing of dependent work packets; and

processing, with a neural processing unit having multiple matrix processors, the individual work packets from the optimized packet stream order as processing tasks in parallel while maintaining the data dependencies identified in the dependency graph.

2 . The method of claim 1 , wherein the metadata for the individual work packet comprises a specification for a convolution operation.

3 . The method of claim 1 , wherein the output data elements are multidimensional tensors.

4 . The method of claim 1 , wherein the processing occurs on multiple neural processing units.

5 . The method of claim 4 , wherein the processing of each of the individual work packets comprises processing each of the individual work packets as native execution on a neural processing unit.

6 . The method of claim 1 , wherein each individual work packet comprises a priority value to schedule the individual work packet.

7 . The method of claim 1 , wherein the metadata for the individual work packet comprises metadata for synchronizing individual work packet execution.

8 . The method of claim 1 , wherein the metadata for each individual work packet comprises resource metadata for allocating hardware resources associated with a neural processing unit.

9 . The method of claim 1 , wherein the metadata for the individual work packet comprises data movement metadata for efficiently moving data.

10 . The method of claim 1 , wherein each individual work packet comprises an encapsulation of one or more of the individual work packets.

11 . An apparatus for processing a multi-layer neural network having a set of neural network layers, each of the neural network layers comprising a set of input data elements and a set of output data elements, the apparatus comprising:

a packet compiler configured to divide each of the neural network layers into a set of individual work packets for the neural network layer, each individual work packet comprising (a) a subset of the set of output data elements for the neural network layer, (b) an associated weight matrix reference identifying weights specific to the subset of the set of output data elements, and (c) metadata comprising at least scheduling information for the individual work packet, the set of individual work packets for the neural network layer collectively covering all of the set of output data elements for the neural network layer, the packet compiler further configured to generate a dependency graph identifying data dependencies between the individual work packets of different neural network layers and create an optimized packet stream order based on the data dependencies, the optimized packet stream order minimizing sequential processing of dependent work packets; and

a neural processing unit having multiple matrix processors, the neural processing unit configured to process the individual work packets from the optimized packet stream order as processing tasks in parallel while maintaining the data dependencies identified in the dependency graph.

12 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises a specification for a convolution operation.

13 . The apparatus of claim 11 , wherein the output data elements are multidimensional tensors.

14 . The apparatus of claim 11 , wherein the apparatus comprises multiple neural processing units.

15 . The apparatus of claim 14 , wherein each of the multiple matrix processors is configured to execute each individual work packet as a native instruction operation.

16 . The apparatus of claim 11 , wherein each individual work packet comprises a priority value to schedule the individual work packet.

17 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises metadata for synchronizing individual work packet execution.

18 . The apparatus of claim 14 , wherein the metadata for each individual work packet comprises resource metadata for allocating hardware resources associated with a neural processing unit.

19 . The apparatus of claim 11 , wherein the metadata for the individual work packet comprises data movement metadata for efficiently moving data.

20 . The apparatus of claim 11 , wherein each individual work packet comprises an encapsulation of one or more of the individual work packets.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 15, 2023
From: CHOLE, SHARAD VASANTRAO; CHUANG, SHANG-TSE; MA, SIYAD CHIH-HUA
To: EXPEDERA, INC.
Reel/Frame 065885/0500 →
Continuity (2)
Provisional Application 63270558 · Oct 21, 2021
Related Publication 20240152761A1 · May 9, 2024
References Cited (15)
US 11467811B1 · Durakovic · 2022 [cited by examiner]
US 20180285715A1 · Son · 2018 [cited by examiner]
US 20190205736A1 · Bleiweiss · 2019 [cited by examiner]
US 20200117400A1 · Golov · 2020 [cited by examiner]
US 20200160226A1 · Ross · 2020 [cited by examiner]
US 20210035258A1 · Ray · 2021 [cited by examiner]
US 20220147797A1 · Mclelland · 2022 [cited by examiner]
US 20220342666A1 · Ju · 2022 [cited by examiner]
US 20220343137A1 · Surendran · 2022 [cited by examiner]
US 20220374723A1 · Blukis · 2022 [cited by examiner]
US 20220383082A1 · Zhang · 2022 [cited by examiner]
US 20220391678A1 · Zhang · 2022 [cited by examiner]
US 20220414455A1 · Collins · 2022 [cited by examiner]
US 20230121044A1 · Grover · 2023 [cited by examiner]
Kornilios Kourtis et al., Compiling Neural Networks for a Computational Memory Accelerator, Apr. 24, 2020, [Retrieved on Jan. 3, 2026]. Retrieved from the internet: <URL: https://arxiv.org/pdf/2003.04293> 8 Pages (1-8) … [cited by examiner]