IP Library › Granted Patent US 12,518,133
Granted Patent B2
US 12,518,133 · App. 17/238,054 · Granted Jan 6, 2026

Kernel generation for neural networks

Inventors: Rishi Surendran (Santa Clara, CA); Dz-ching Ju (Saratoga, CA); Yuan Lin (Cupertino, CA)
Assignee: NVIDIA Corporation
G06N3/044
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,133
App. No.
17/238,054
Granted
Jan 6, 2026
Kind
B2
Abstract

Apparatuses, systems, and techniques to automatically generate a reduced number of compute kernels for performing operations of one or more neural networks. In at least one embodiment, one or more operations of one or more neural network graph nodes of the one or more neural network are automatically adjusted to generate an optimized one or more operations that are compiled to generate the reduced number of compute kernels.

Claims (50)

1 . A processor, comprising:

one or more circuits to;

modify one or more operations of one or more neural network graph nodes, wherein the modified one or more operations perform a same function as prior to the modification; and

generate a first number of compute kernel modules to perform the one or more neural network graph nodes having the modified one or more operations, wherein the first number is a smaller number than would have been generated to perform one or more neural network graph nodes had the one or more operations not been modified.

2 . The processor of claim 1 , wherein the one or more neural network graph nodes comprise a plurality of neural network graph nodes for a neural network graph associated with a cell of one or more cells of a neural network.

3 . The processor of claim 2 , wherein to reduce the first number of compute kernel modules to perform the one or more neural network graph nodes, the one or more circuits are further to automatically adjust one or more operations of one or more of the plurality of neural network graph nodes to compile the first number of compute kernel modules.

4 . The processor of claim 3 , wherein to automatically adjust the one or more operations, the one or more circuits are further to:

replace a matrix-vector multiplication operation of the one or more operations with a sequence of operations comprising a reshape operation, an element-wise multiplication operation, and a sum reduction operation.

5 . The processor of claim 3 , wherein to automatically adjust the one or more operations, the one or more circuits are further to:

replace a matrix-matrix multiplication operation of the one or more operations with a sequence of operations comprising a plurality of reshape operations, an element-wise multiplication operation, and a sum reduction operation.

6 . The processor of claim 3 , wherein to automatically adjust the one or more operations, the one or more circuits are further to:

replace, within the one or more operations, a first sequence of operations comprising two sum reduction operations and a first addition operation with a second sequence of operations comprising a replicate operation, a concatenation operation, a second addition operation, and a sum reduction operation.

7 . The processor of claim 3 , wherein to automatically adjust the one or more operations, the one or more circuits are further to:

adjust a first sequence of operations comprising a sum reduction operation followed by a slice operation to generate a second sequence of operations by moving the slice operation from a first location following the sum reduction operation to a second location preceding the sum reduction operation.

8 . The processor of claim 3 , wherein the one or more circuits are further to:

remove unused operations from the one or more operations of the one or more of the plurality of neural network graph nodes.

9 . The processor of claim 3 , wherein the one or more circuits are further to:

generate, based on the adjusted one or more operations, the first number of compute kernel modules to execute the graph of the cell by compiling software code comprising the adjusted one or more operations.

10 . The processor of claim 3 , wherein the first number of compute kernel modules is a single compute kernel module.

11 . The processor of claim 2 , wherein the cell is a long short-term memory (LSTM) cell, and wherein the first number of compute kernel modules is a single compute kernel module.

12 . A method, comprising:

modifying, by a processing device, one or more operations of one or more neural network graph nodes, wherein the modified one or more operations perform a same function as prior to the modification; and

generating, by the processing device, a first number of compute kernel modules to perform the one or more neural network graph nodes having the modified one or more operations, wherein the first number is a smaller number than would have been generated to perform one or more neural network graph nodes had the one or more operations not been modified.

13 . The method of claim 12 , wherein the one or more neural network graph nodes comprise a plurality of neural network graph nodes for a neural network graph associated with a cell of one or more cells of a neural network.

14 . The method of claim 13 , wherein the modifying comprises automatically adjusting one or more operations of one or more of the plurality of neural network graph nodes to compile the first number of compute kernel modules.

15 . The method of claim 14 , wherein automatically adjusting the one or more operations further comprises replacing a matrix-vector multiplication operation of the one or more operations with a sequence of operations comprising a reshape operation, an element-wise multiplication operation, and a sum reduction operation.

16 . The method of claim 14 , wherein automatically adjusting the one or more operations further comprises replacing a matrix-matrix multiplication operation of the one or more operations with a sequence of operations comprising a plurality of reshape operations, an element-wise multiplication operation, and a sum reduction operation.

17 . The method of claim 14 , wherein automatically adjusting the one or more operations further comprises replacing, within the one or more operations, a first sequence of operations comprising two sum reduction operations and a first addition operation with a second sequence of operations comprising a replicate operation, a concatenation operation, a second addition operation, and a sum reduction operation.

18 . The method of claim 14 , wherein automatically adjusting the one or more operations further comprises adjusting a first sequence of operations comprising a sum reduction operation followed by a slice operation to generate a second sequence of operations by moving the slice operation from a first location following the sum reduction operation to a second location preceding the sum reduction operation.

19 . The method of claim 14 further comprising:

removing unused operations from the one or more operations of the of one or more of the plurality of neural network graph nodes.

20 . The method of claim 14 , wherein the cell is a long short-term memory (LSTM) cell, and wherein the first number of compute kernel modules is a single compute kernel module.

21 . A system, comprising:

a memory device; and

one or more processors to:

modify one or more operations of one or more neural network graph nodes, wherein the modified one or more operations perform a same function as prior to the modification; and

generate a first number of compute kernel modules to perform the one or more neural network graph nodes having the modified one or more operations, wherein the first number is a smaller number than would have been generated to perform one or more neural network graph nodes had the one or more operations not been modified.

22 . The system of claim 21 , wherein the one or more neural network graph nodes comprise a plurality of neural network graph nodes for a neural network graph associated with a cell of one or more cells of a neural network.

23 . The system of claim 22 , wherein to reduce the first number of compute kernel modules to perform the one or more neural network graph nodes, the one or more processors are further to automatically adjust one or more operations of one or more of the plurality of neural network graph nodes to compile the first number of compute kernel modules.

24 . The system of claim 23 , wherein to automatically adjust the one or more operations, the one or more processors are further to:

replace a matrix-vector multiplication operation of the one or more operations with a sequence of operations comprising a reshape operation, an element-wise multiplication operation, and a sum reduction operation.

25 . The system of claim 23 , wherein to automatically adjust the one or more operations, the one or more processors are further to:

replace a matrix-matrix multiplication operation of the one or more operations with a sequence of operations comprising a plurality of reshape operations, an element-wise multiplication operation, and a sum reduction operation.

26 . The system of claim 23 , wherein to automatically adjust the one or more operations, the one or more processors are further to:

replace, within the one or more operations, a first sequence of operations comprising two sum reduction operations and a first addition operation with a second sequence of operations comprising a replicate operation, a concatenation operation, a second addition operation, and a sum reduction operation.

27 . The system of claim 23 , wherein to automatically adjust the one or more operations, the one or more processors are further to:

adjust a first sequence of operations comprising a sum reduction operation followed by a slice operation to generate a second sequence of operations by moving the slice operation from a first location following the sum reduction operation to a second location preceding the sum reduction operation.

28 . The system of claim 23 , wherein the one or more processors are further to:

generate, based on the adjusted one or more operations, the first number of compute kernel modules to execute the graph of the cell by compiling software code comprising the adjusted one or more operations.

29 . The system of claim 23 , wherein the first number of compute kernel modules is a single compute kernel module.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 10, 2021
From: SURENDRAN, RISHI; JU, DZ-CHING; LIN, YUAN
To: NVIDIA CORPORATION
Reel/Frame 056191/0188 →
Continuity (1)
Related Publication 20220343137A1 · Oct 27, 2022
References Cited (19)
US 11561826B1 · Nagpal · 2023 [cited by examiner]
US 20110231510A1 · Korsunsky et al. · 2011 [cited by applicant]
US 20170228358A1 · Hirzel · 2017 [cited by examiner]
US 20180136912A1 · Venkataramani · 2018 [cited by examiner]
US 20180336453A1 · Merity · 2018 [cited by applicant]
US 20200034710A1 · Sidhu et al. · 2020 [cited by applicant]
US 20200349426A1 · Luo · 2020 [cited by examiner]
US 20210034966A1 · Qian · 2021 [cited by examiner]
US 20210312310A1 · Wu · 2021 [cited by examiner]
US 20210383188A1 · Zhou · 2021 [cited by examiner]
US 20220051104A1 · Interlandi · 2022 [cited by examiner]
US 20220343137A1 · Surendran · 2022 [cited by examiner]
CN 108875956A · 2018 [cited by applicant]
CN 111008972A · 2020 [cited by applicant]
CN 112465129A · 2021 [cited by applicant]
Bhaskaracharya et al, “Automatic Kernel Generation for Volta Tensor Cores”, 2020, Cornell University Library abstract, Sections I, IV, VI, 13 pages. [cited by applicant]
Vasilache et al., “The Next 700 Accelerated Layers: From Mathematical Expressions of Network Computation Graphs to Accelerated GPU Kernels, Automatically”, ACM Transactions on Architecture and Code Optimization, vol. 16… [cited by applicant]
Office Action for Patent Application No. GB2205788.9 dated Jan. 6, 2023, 10 pages. [cited by applicant]
Office Action mailed Apr. 25, 2025 in Chinese Patent Application No. 202210382719.0, NVIDIA Corporation, 34 pages including translation. [cited by applicant]