IP Library Granted Patent US 11,625,605
Granted Patent B2
US 11,625,605 · App. 16/723,608 · Granted Apr 11, 2023

Selecting computational kernel variants using neural networks

Inventors: Jonathan Edward Barker (Boulder, CO); Christopher Thomas Cheng (Santa Clara, CA); Paul Martin Springer (Iserlohn, DE); Wojciech Jablonski (Warsaw, PL)
Assignee: Nvidia Corporation
G06N3/08G06F7/57G06F17/16G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,625,605
App. No.
16/723,608
Granted
Apr 11, 2023
Kind
B2
Abstract

Apparatuses, systems, and techniques to optimize kernel selection for performing a computation. In at least one embodiment, a neural network is trained and utilized to generate a list of kernels so that an (e.g., optimal) kernel may be identified. The neural network receives characteristics of the input matrices and determines relevancy scores for a list of possible kernels. Based on an ordered listing of kernels by relevant score, a kernel is selected from the list and utilized to perform the computation and provide the result.

Claims (72)

1. A processor comprising:

one or more arithmetic logic units (ALUs) to be configured to perform a matrix computation by:

receiving a request for the matrix computation, the request including at least one matrix and the matrix computation to be performed;

providing, to a neural network, a candidate kernel list and characteristics of the at least one matrix and the matrix computation;

generating, based at least on the candidate kernel list, a ranked list of kernels using the neural network, the ranking based on a relevancy score determined using the neural network for at least one kernel in the ranked list;

selecting a first kernel from the ranked list of kernels;

generating a result for the matrix computation using the at least one matrix and the first kernel; and

providing the result.

2. The one or more ALUs of claim 1 , further configured to perform the matrix computation by:

generating the candidate kernel list based on the at least one matrix and the matrix computation; and

providing the candidate kernel list to the neural network with the characteristics of the at least one matrix and matrix computation, wherein the generated ranked list of kernels includes only kernels that are included in the candidate kernel list.

3. The one or more ALUs of claim 2 , further configured to perform the matrix computation by:

identifying a kernel processor configured to generate the result;

removing one or more kernels from the candidate kernel list based on one or more hardware constraints of the kernel processor.

4. The processor of claim 1 , wherein the request is received via an application programming interface (API).

5. The one or more ALUs of claim 1 , further configured to perform the matrix computation by:

identifying hardware behavior information; and

providing the hardware behavior information to the neural network with the characteristics of the at least one matrix and the matrix computation.

6. The processor of claim 1 , wherein the matrix computation is a general matrix multiply (GeMM).

7. The one or more ALUs of claim 1 , further configured to perform the matrix computation by:

determining a second kernel for the at least one matrix and the matrix computation;

comparing the ranked list of kernels by the neural network to the second kernel; and

generating additional training data for the neural network based on the at least one matrix and the second kernel.

8. The processor of claim 7 , wherein generating the additional training data includes:

determining a class for the at least one matrix;

generating input based on the class; and

providing the additional training data to the neural network.

9. A system, comprising:

one or more processors to be configured to perform, using one or more neural networks:

receiving a request for the matrix computation, the request including at least one matrix and the matrix computation to be performed;

providing, to a neural network, a candidate kernel list and characteristics of the at least one matrix and the matrix computation;

generating, based at least on the candidate kernel list, a ranked list of kernels using the neural network, the ranking based on a relevancy score determined using the neural network for at least one kernel in the ranked list;

selecting a first kernel from the ranked list of kernels;

generating a result for the matrix computation using the at least one matrix and the first kernel; and

providing the result; and

one or more memories to store parameters corresponding to the one or more neural networks.

10. The system of claim 9 , wherein the processors are further configured to perform:

generating the candidate kernel list based on the at least one matrix and the matrix computation; and

providing the candidate kernel list to the neural network with the at least one matrix and matrix computation, wherein the generated ranked list of kernels includes only kernels that are included in the candidate kernel list.

11. The system of claim 10 , wherein the processors are further configured to perform:

identifying a kernel processor configured to generate the result;

removing one or more kernels from the candidate kernel list on one or more hardware constraints of the kernel processor.

12. The system of claim 9 , wherein the processors are further configured to perform:

identifying hardware behavior information; and

providing the hardware behavior information to the neural network with the characteristics of the at least one matrix and the matrix computation.

13. The system of claim 9 , wherein the processors are further configured to perform:

determining a second kernel for the at least one matrix and the matrix computation;

comparing the ranked list of kernels by the neural network to the second kernel; and

generating additional training data for the neural network based on the at least one matrix and the second kernel.

14. The system of claim 13 , wherein the processors are further configured to perform:

determining a class for the at least one matrix;

generating input based on the class; and

providing the additional training data to the neural network.

15. A computer-readable storage medium having stored thereon a set of instructions, which if performed by one or more processors, cause the one or more processors to at least:

provide a request for the matrix computation, the request including at least one matrix and the matrix computation to be performed, the request provided to a system configured to:

provide, to a neural network, a candidate kernel list and characteristics of the at least one matrix and the matrix computation;

generate, based at least on the candidate kernel list, a ranked list of kernels using the neural network, the ranking based on a relevancy score determined using the neural network for at least one kernel in the ranked list;

select a first kernel from the ranked list of kernels;

generate a result for the matrix computation using the at least one matrix and the first kernel; and

provide the result in response to the request.

16. The computer-readable storage medium of claim 15 , wherein the request is provided via an application programming interface (API).

17. The computer-readable storage medium of claim 15 , wherein the matrix computation is a general matrix multiply (GeMM).

18. The computer-readable storage medium of claim 15 , wherein the system is further configured to:

determine a second kernel for the at least one matrix and the matrix computation;

compare the ranked list of kernels by the neural network to the second kernel; and

generate additional training data for the neural network based on the at least one matrix and the second kernel.

19. The computer-readable storage medium of claim 15 , wherein the instructions further include instructions to:

determine a class for the at least one matrix; and

provide the class with the request.

20. The computer-readable storage medium of claim 19 , wherein the system is further configured to:

identify hardware behavior information; and

provide the hardware behavior information to the neural network with the characteristics of the at least one matrix and the matrix computation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2019
From: BARKER, JON; CHENG, CHRIS; SPRINGER, PAUL; JABLONSKI, WOLCIECH
To: NVIDIA CORPORATION
Reel/Frame 051348/0638 →
Continuity (1)
Related Publication 20210192334A1 · Jun 24, 2021