IP Library Granted Patent US 8,108,844
Granted Patent B2
US 8,108,844 · App. 11/714,654 · Granted Jan 31, 2012

Systems and methods for dynamically choosing a processing element for a compute kernel

Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,108,844
App. No.
11/714,654
Granted
Jan 31, 2012
Kind
B2
Abstract

A runtime system implemented in accordance with the present invention provides an application platform for parallel-processing computer systems. Such a runtime system enables users to leverage the computational power of parallel-processing computer systems to accelerate/optimize numeric and array-intensive computations in their application programs. This enables greatly increased performance of high-performance computing (HPC) applications.

Claims (40)

1. A computer-implemented method configured to be performed by a runtime system at a parallel-processing computer system that includes multiple types of processing elements, comprising:

at runtime:

receiving one or more operation requests issued by an application;

dynamically choosing a respective type of processing element identifying at least one of the one or more types of processing elements for at least one of the one or more operation requests, the at least one operation request corresponding to an intrinsic operation, wherein the intrinsic operation is an arithmetic operation, and the respective type of processing element is one of the multiple types of processing elements, which include sing-core/multi-core central processing units, graphics processing units and single-core/multi-core co-processors;

dynamically identifying, from a library, a precompiled, processor-specific compute kernel corresponding to the intrinsic operation of the at least one operation request, the library comprising a plurality of processor-specific compute kernels corresponding to a plurality of intrinsic operations; and

dynamically preparing one or more compute kernels for the at least one operation request, wherein the one or more compute kernels include the at least one precompiled, processor-specific compute kernel, and are configured to execute on the respective type of processing element.

2. The computer-implemented method of claim 1 , wherein the precompiled, processor-specific compute kernels in the library are compressed and/or encrypted.

3. The computer-implemented method of claim 1 , wherein the precompiled, processor-specific compute kernels are hand-coded for a specific type of processor.

4. The computer-implemented method of claim 1 , wherein the intrinsic operation corresponding to the at least one operation request is selected from a group consisting of:

matrix multiplication, matrix solvers, fast fourier transforms, convolutions, and LU decomposition.

5. The computer-implemented method of claim 1 , wherein the respective type of processing element is dynamically chosen based, at least in part, on one of a predefined computer resource requirement metric of the one or more operation requests and a predefined workload requirement metric of the parallel-processing computer system running the runtime system.

6. The computer-implemented method of claim 1 , wherein the one or more operation requests correspond to application program interface function calls.

7. The computer-implemented method of claim 1 , wherein the identified precompiled, processor-specific compute kernel corresponding to the intrinsic operation is in a GPU binary instruction set architecture format.

8. A parallel-processing computer system, comprising:

memory;

multiple types of processing elements; and

at least one program stored in the memory and executed by the multiple types of processing elements, the at least one program including a runtime system comprising instructions for:

at runtime:

receiving one or more operation requests issued by an application;

dynamically choosing a respective type of processing element identifying at least one of the one or more types of processing elements for at least one of the one or more operation requests, the at least one operation request corresponding to an intrinsic operation, wherein the intrinsic operation is an arithmetic operation, and the respective type of processing element is one of the multiple types of processing elements, which include sing-core/multi-core central processing units, graphics processing units and single-core/multi-core co-processors;

dynamically identifying, from a library, a precompiled, processor-specific compute kernel corresponding to the intrinsic operation of the at least one operation request, the library comprising a plurality of processor-specific compute kernels corresponding to a plurality of intrinsic operations; and

dynamically preparing one or more compute kernels for the at least one operation request, wherein the one or more compute kernels include the at least one precompiled, processor-specific compute kernel, and are configured to execute on the respective type of processing element.

9. The parallel-processing computer system of claim 8 , wherein the precompiled, processor-specific compute kernels in the library are compressed and/or encrypted.

10. The parallel-processing computer system of claim 8 , wherein the precompiled, processor-specific compute kernels are hand-coded for a specific type of processor.

11. The parallel-processing computer system of claim 8 , wherein the intrinsic operation corresponding to the at least one operation request is selected from a group consisting of: matrix multiplication, matrix solvers, fast fourier transforms, convolutions, and LU decomposition.

12. The parallel-processing computer system of claim 8 , wherein the respective type of processing element is dynamically chosen based, at least in part, on one of a predefined computer resource requirement metric of the one or more operation requests and a predefined workload requirement metric of the parallel-processing computer system running the runtime system.

13. The parallel-processing computer system of claim 8 , wherein the one or more operation requests correspond to application program interface function calls.

14. The parallel-processing computer system of claim 8 , wherein the identified precompiled, processor-specific compute kernel corresponding to the intrinsic operation is in a GPU binary instruction set architecture format.

15. A non-transitory computer readable storage medium storing one or more programs configured to be executed by computer program product for use in conjunction with a parallel-processing computer system that includes multiple types of processing elements, the one or more programs comprising instructions for:

at runtime:

receiving one or more operation requests issued by an application;

dynamically choosing a respective type of processing element identifying at least one of the one or more types of processing elements for at least one of the one or more operation requests, the at least one operation request corresponding to an intrinsic operation, wherein the intrinsic operation is an arithmetic operation, and the respective type of processing element is one of the multiple types of processing elements, which include sing-core/multi-core central processing units, graphics processing units and single-core/multi-core co-processors;

dynamically identifying, from a library, a precompiled, processor-specific compute kernel corresponding to the intrinsic operation of the at least one operation request, the library comprising a plurality of processor-specific compute kernels corresponding to a plurality of intrinsic operations; and

dynamically preparing one or more compute kernels for the at least one operation request, wherein the one or more compute kernels include the at least one precompiled, processor-specific compute kernel, and are configured to execute on the respective type of processing element.

16. The non-transitory computer readable storage medium of claim 15 , wherein the precompiled, processor-specific compute kernels in the library are compressed and/or encrypted.

17. The non-transitory computer readable storage medium of claim 15 , wherein the precompiled, processor-specific compute kernels are hand-coded for a specific type of processor.

18. The non-transitory computer readable storage medium of claim 15 , wherein the intrinsic operation corresponding to the at least one operation request is selected from a group consisting of: matrix multiplication, matrix solvers, fast fourier transforms, convolutions, and LU decomposition.

19. The non-transitory computer readable storage medium of claim 15 , wherein the respective type of processing element is dynamically chosen based, at least in part, on one of a predefined computer resource requirement metric of the one or more operation requests and a predefined workload requirement metric of the parallel-processing computer system running the runtime system.

20. The non-transitory computer readable storage medium of claim 15 , wherein the one or more operation requests correspond to application program interface function calls.

21. The non-transitory computer readable storage medium of claim 15 , wherein the identified precompiled, processor-specific compute kernel corresponding to the intrinsic operation is in a GPU binary instruction set architecture format.

Assignments (3)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0405 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 16, 2009
From: PEAKSTREAM, INC.
To: GOOGLE INC.
Reel/Frame 022963/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 25, 2007
From: CRUTCHFIELD, WILLIAM Y.; GRANT, BRIAN K.; PAPAKIPOS, MATTHEW N.
To: PEAKSTREAM, INC.
Reel/Frame 019346/0332 →
Continuity (3)
Provisional Application 60815532 · Jun 20, 2006
Provisional Application 60903188 · Feb 23, 2007
Related Publication 20070294512A1 · Dec 20, 2007