IP Library Granted Patent US 9,442,757
Granted Patent B2
US 9,442,757 · App. 14/163,710 · Granted Sep 13, 2016

Data parallel computing on multiple processors

Inventors: Aaftab Munshi (Los Gatos, CA); Jeremy Sandmel (San Mateo, CA)
Assignee: Apple Inc.
G06F9/4843G06F9/5044G06F2209/5018
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,442,757
App. No.
14/163,710
Granted
Sep 13, 2016
Kind
B2
Abstract

A method and an apparatus that allocate one or more physical compute devices such as central processing units or graphical processing units attached to a host processing unit running an application for executing one or more threads of the application are described. The allocation may be based on data representing a processing capability requirement from the application for executing an executable in the one or more threads. A compute device identifier may be associated with the allocated physical compute devices to schedule and execute the executable in the one or more threads concurrently in one or more of the allocated physical compute devices concurrently.

Claims (44)

1. A computer implemented method comprising:

receiving, from a host application executing on a host processor, a request to identify any compute device that matches a processing requirement for a task corresponding to source code in the host application;

sending, to the host application, a compute identifier for each compute device that matches the processing requirement;

receiving, from the host application, a request specifying a compute identifier selected by the host application; and

generating, by the host processor, a context for the compute device that corresponds to the selected compute identifier and defines an execution queue for the selected compute device.

2. The computer implemented method of claim 1 further comprising:

generating, by the host processor, a compute identifier for each compute device that matches the processing requirement.

3. The computer implemented method of claim 1 further comprising:

retrieving, by the host processor, characteristics of each compute device coupled to the host processor.

4. The computer implemented method of claim 1 , wherein generating the context further comprises allocating resources for the selected compute device.

5. The computer implemented method of claim 1 , wherein the processing requirement specifies multiple physical processors.

6. The computer implemented method of claim 5 further comprising:

defining, by the host processor, a set of physical processors as a logical compute device; and

generating, by the host processor, a compute identifier for the logical compute device.

7. The computer implemented method of claim 1 , wherein the requests from the host application are processed by an interpreter during execution of the host application.

8. A non-transitory computer readable storage medium storing instruction that when executed by a host processor cause the host processor to perform operations comprising:

receiving, from a host application executing on the host processor, a request to identify any compute device that matches a processing requirement for a task corresponding to source code in the host application;

sending, to the host application, a compute identifier for each compute device that matches the processing requirement;

receiving, from the host application, a request specifying a compute identifier selected by the host application; and

generating, by the host processor, a context for the compute device that corresponds to the selected compute identifier and defines an execution queue for the selected compute device.

9. The non-transitory computer readable storage medium of claim 8 wherein the instructions further cause the host processor to perform an operation comprising:

generating a compute identifier for each compute device that matches the processing requirement.

10. The non-transitory computer readable storage medium of claim 8 wherein the instructions further cause the host processor to perform an operation comprising:

retrieving characteristics of each compute device coupled to the host processor.

11. The non-transitory computer readable storage medium of claim 8 , wherein generating the context further comprises allocating resources for the selected compute device.

12. The non-transitory computer readable storage medium of claim 8 , wherein the processing requirement specifies multiple physical processors.

13. The non-transitory computer readable storage medium of claim 12 , wherein the instructions further cause the host processor to perform operations comprising:

defining a set of physical processors as a logical compute device; and

generating a compute identifier for the logical compute device.

14. The non-transitory computer readable storage medium of claim 8 , wherein the requests from the host application are processed by an interpreter during execution of the host application.

15. A system comprising:

a host processor coupled to a memory through a bus;

a plurality of compute devices coupled to the host processor through a bus; and

a platform layer process stored as instructions in the memory to cause the host processor to

receive, from a host application executing on the host processor, a request to identify any compute device that matches a processing requirement for a task corresponding to source code in the host application;

send, to the host application, a compute identifier for each compute device that matches the processing requirement;

receive, from the host application, a request specifying a compute identifier selected by the host application; and

generate, by the host processor, a context for the compute device that corresponds to the selected compute identifier and defines an execution queue for the selected compute device.

16. The system of claim 15 wherein the instructions further cause the host processor to generate a compute identifier for each compute device that matches the processing requirement.

17. The system of claim 15 wherein the instructions further cause the host processor to retrieve characteristics of each compute device coupled to the host processor.

18. The system of claim 15 , wherein generating the context further comprises allocating resources for the selected compute device.

19. The system of claim 15 , wherein the processing requirement specifies multiple physical processors.

20. The system of claim 19 , wherein the instructions further cause the host processor to define a set of physical processors as a logical compute device and to generate a compute identifier for the logical compute device.

21. The system of claim 15 further comprising an interpreter process stored as instructions in the memory to cause the host processor to interpret the requests from the host application during execution of the host application.

Continuity (5)
Continuation 13614975 · Sep 13, 2012
Continuation 11800185 · May 3, 2007
Provisional Application 60923030 · Apr 11, 2007
Provisional Application 60925616 · Apr 20, 2007
Related Publication 20140201755A1 · Jul 17, 2014