IP Library › Granted Patent US 12,282,526
Granted Patent B2
US 12,282,526 · App. 18/620,228 · Granted Apr 22, 2025

Application programming interface to accelerate matrix operations

Inventors: Piotr Majcher (Sunnyvale, CA); Mostafa Hagog (Folsom, CA); Philippe Vandermersch (San Jose, CA)
Assignee: NVIDIA Corporation
G06F17/16G06F9/3001G06F9/30145G06N3/08G06N5/046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,526
App. No.
18/620,228
Granted
Apr 22, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to determine a matrix multiplication algorithm for a matrix multiplication operation. In at least one embodiment, a matrix multiplication operation is analyzed to determine an appropriate matrix multiplication algorithm to perform the matrix multiplication algorithm.

Claims (67)

1. An acceleration processor unit (APU) comprising:

one or more core complexes, wherein the one or more core complexes include one or more central processing unit (CPU) cores;

one or more graphics complexes, wherein the one or more graphics complexes include one or more compute units and include at least one L2 cache;

one or more fabric interconnects;

one or more memory controllers;

one or more input/output (I/O) bus interfaces, at least one comprising a peripheral component interconnect express (PCIe) interface;

wherein the APU is utilized to implement an application programming interface (API) call to select one or more general matrix-to-matrix multiply (GEMM) implementations from among a plurality of GEMM implementations;

wherein the API call is to use:

input matrices;

a matrix multiply operation descriptor;

a matrix layout descriptor;

a search preferences parameter;

an algorithm count parameter to specify a number of algorithms desired;

a results array parameter;

a results count parameter to store a number of algorithms returned; and

an operation status to indicate whether the API call is successful or if other errors have occurred.

2. The APU of claim 1 , wherein the matrix multiply operation descriptor is a pointer.

3. The APU of claim 1 , wherein the search preferences parameter is to indicate a search method to utilize to search for the possible algorithms.

4. The APU of claim 1 , wherein the search preferences parameter is to indicate a preferred resource utilization.

5. The APU of claim 1 , wherein the algorithm count parameter is to include an integer value of the number of algorithms desired.

6. The APU of claim 1 , wherein the input matrices comprise one or more pointers to one or more values of the input matrices.

7. The APU of claim 1 , wherein the matrix layout descriptor comprises one or more pointers.

8. The APU of claim 1 , wherein the results array parameter comprises a pointer to a location in which a result of the GEMM is to be stored.

9. A method comprising:

implementing an application programming interface (API) call to select one or more general matrix-to-matrix multiply (GEMM) implementations from among a plurality of GEMM implementations, wherein the API call is to use:

input matrices;

a matrix multiply operation descriptor;

a matrix layout descriptor;

a search preferences parameter;

an algorithm count parameter to specify a number of algorithms desired;

a results array parameter;

a results count parameter to store a number of algorithms returned; and

an operation status to indicate whether the API call is successful or if other errors have occurred; and

the API call is to be implemented by a processor comprising:

one or more core complexes, wherein the one or more core complexes include one or more central processing unit (CPU) cores;

one or more graphics complexes, wherein the one or more graphics complexes include one or more compute units and include at least one L2 cache;

one or more fabric interconnects;

one or more memory controllers;

one or more input/output (I/O) bus interfaces, at least one comprising a peripheral component interconnect express (PCIe) interface.

10. The method of claim 9 , wherein the matrix multiply operation description is a pointer.

11. The method of claim 9 , wherein the search preferences parameter is to indicate a search method.

12. The method of claim 9 , wherein the search preferences parameter is to indicate a preferred resource utilization.

13. The method of claim 9 , wherein the algorithm count parameter is to include an integer value of the number of algorithms desired.

14. The method of claim 9 , further comprising performing the selected one or more GEMM implementations.

15. A system comprising:

memory; and

an acceleration processor unit (APU) comprising:

one or more core complexes, wherein the one or more core complexes include one or more central processing unit (CPU) cores;

one or more graphics complexes, wherein the one or more graphics complexes include one or more compute units and include at least one L2 cache;

one or more fabric interconnects;

one or more memory controllers;

one or more input/output (I/O) bus interfaces, at least one comprising a peripheral component interconnect express (PCIe) interface;

wherein the APU is utilized to implement an application programming interface (API) call to select one or more general matrix-to-matrix multiply (GEMM) implementations from among a plurality of GEMM implementations;

wherein the API call is to use:

input matrices;

a matrix multiply operation descriptor;

a matrix layout descriptor;

a search preferences parameter;

an algorithm count parameter to specify a number of algorithms desired;

a results array parameter;

a results count parameter to store a number of algorithms returned; and

an operation status to indicate whether the API call is successful or if other errors have occurred.

16. The system of claim 15 , wherein the matrix multiply operation descriptor is a pointer.

17. The system of claim 15 , wherein the search preferences parameter is to indicate a search method to utilize to search for the possible algorithms.

18. The system of claim 15 , wherein the search preferences parameter is to indicate a preferred resource utilization.

19. The system of claim 15 , wherein the algorithm count parameter is to include an integer value of the number of algorithms desired.

20. The system of claim 15 , wherein the input matrices comprise one or more pointers to one or more values of the input matrices.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 28, 2024
From: MAJCHER, PIOTR; HAGOG, MOSTAFA; VANDERMERSCH, PHILIPPE
To: NVIDIA CORPORATION
Reel/Frame 066937/0970 →
Continuity (2)
Continuation 16795380 · Feb 19, 2020
Related Publication 20240256633A1 · Aug 1, 2024
References Cited (67)
US 7002591B1 · Leather et al. · 2006 [cited by applicant]
US 7307638B2 · Leather et al. · 2007 [cited by applicant]
US 9304835B1 · Ekanadham et al. · 2016 [cited by applicant]
US 9400700B2 · Ekanadham et al. · 2016 [cited by applicant]
US 9772890B2 · Ekanadham et al. · 2017 [cited by applicant]
US 9778967B2 · Ekanadham et al. · 2017 [cited by applicant]
US 10073815B2 · Zhou · 2018 [cited by applicant]
US 12020076B2 · Merrill, III · 2024 [cited by examiner]
US 20050237337A1 · Leather et al. · 2005 [cited by applicant]
US 20140289445A1 · Savich · 2014 [cited by applicant]
US 20160188385A1 · Ekanadham et al. · 2016 [cited by applicant]
US 20170344514A1 · Zhou · 2017 [cited by applicant]
US 20180000709A1 · Yumioka et al. · 2018 [cited by applicant]
US 20180157471A1 · Venkataramani et al. · 2018 [cited by applicant]
US 20180189234A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180189239A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180189638A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180189675A1 · Nurvitadhi et al. · 2018 [cited by applicant]
US 20180293777A1 · Sarel et al. · 2018 [cited by applicant]
US 20190114538A1 · Ng et al. · 2019 [cited by applicant]
US 20190278593A1 · Elango et al. · 2019 [cited by applicant]
US 20190362197A1 · Jain · 2019 [cited by applicant]
US 20200133735A1 · Zhao et al. · 2020 [cited by applicant]
US 20210048991A1 · Tanner · 2021 [cited by applicant]
US 20210110506A1 · Prakash et al. · 2021 [cited by applicant]
US 20210192334A1 · Barker · 2021 [cited by examiner]
CN 109144471A · 2019 [cited by applicant]
CN 109460533A · 2019 [cited by applicant]
CN 113240570A · 2021 [cited by applicant]
EP 3343383A1 · 2018 [cited by applicant]
EP 3343390A1 · 2018 [cited by applicant]
EP 3343391A1 · 2018 [cited by applicant]
EP 3343392A1 · 2018 [cited by applicant]
EP 3343460A1 · 2018 [cited by applicant]
WO 2018125250A1 · 2018 [cited by applicant]
WO 2019027924A1 · 2019 [cited by applicant]
WO 2019085655A1 · 2019 [cited by applicant]
WO 2020046859A1 · 2020 [cited by applicant]
WO 2020050886A1 · 2020 [cited by applicant]
WO 2021076425A1 · 2021 [cited by applicant]
Gorman, “Sorting A Vector in C++,” retrieved from http://www.gormanalysis.com/blog/sorting-a-vector-in-cpp/, Mar. 7, 2019, 4 pages. [cited by applicant]
Johansson et al, “Algorithms for Large Matrix Multiplications—Assessment of Strassen's Algorithm,” KTH Royal Institute of Technology School of Engineering Sciences, 2018, 41 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2102376.7, mailed May 8, 2024, 9 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2112943.2, mailed May 8, 2024, 8 pages. [cited by applicant]
“Guide and Reference” retrieved from cenapad.unicamp.br, archived on Jan. 14, 2016, 33 pages. [cited by applicant]
“Guide and Reference” retrieved from cenapad.unicamp.br, archived on Oct. 27, 2007, 173 pages. [cited by applicant]
Bientinesi et al., “Representing Linear Algebra Algorithms in Code: The FLAME Application Program Interfaces,” ACM Transactions on Mathematical Software, 31(1): Mar. 2005, 33 pages. [cited by applicant]
Goto et al., “Anatomy of High-Performance Matrix Multiplication,” ACM Transactions on Mathematical Software, 34(3): May 2008, 25 pages. [cited by applicant]
IEEE, “IEEE Standard 754-2008 (Revision of IEEE Standard 754-1985): IEEE Standard for Floating-Point Arithmetic,” Aug. 29, 2008, 70 pages. [cited by applicant]
NVIDIA, “Matrix Multiply (GEMM),” 2021, 1 page. [cited by applicant]
Office Action for Chinese Application No. 202110191400.5 mailed Feb. 6, 2024, 19 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2102376.7, mailed Dec. 11, 2023, 5 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2102376.7, mailed Jun. 14, 2023, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2102376.7, mailed Nov. 10, 2022, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2112943.2, mailed Dec. 11, 2023, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2112943.2, mailed Jun. 14, 2023, 4 pages. [cited by applicant]
Office Action for United Kingdom Application No. GB2112943.2, mailed Nov. 10, 2022, 4 pages. [cited by applicant]
Zhang et al., “Software Architecture for Modular Self-Reconfigurable Robots ,” 2011, 6 pages. [cited by applicant]
JUNR_0926, “cuBLAS Level-2 Function,” retrieved from <https://www.jianshu.com/p/0ee1134a528b,> Nov. 10, 2018, 5 pages. [cited by applicant]
Office Action for Chinese Application No. 202110191400.5, mailed Jul. 9, 2024, 31 pages. [cited by applicant]
Office Action for Chinese Application No. 202111061790.0, mailed Aug. 21, 2024, 33 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Application No. GB2102376.7, mailed Sep. 13, 2024, 7 pages. [cited by applicant]
United Kingdom Combined Search and Examination Report for Application No. GB2112943.2, mailed Sep. 13, 2024, 7 pages. [cited by applicant]
Notice of intention to Grant for United Kingdom Application No. GB2112943.2, mailed Dec. 23, 2024, 2 pages. [cited by applicant]
Office Action for Chinese Application No. 202110191400.5, mailed Jan. 14, 2025, 33 pages. [cited by applicant]
Wang et al., “Collaborative Transparent Programming Model of Software and Hardware for Self-Reconfigurable Systems,” Science and Technology Information, Aug. 5, 2011, 4 pages. [cited by applicant]
Notice of Intention to Grant for United Kingdom Application No. GB102376.7, mailed Oct. 15, 2024, 2 pages. [cited by applicant]