IP Library Granted Patent US 11,556,382
Granted Patent B1
US 11,556,382 · App. 16/508,106 · Granted Jan 17, 2023

Hardware accelerated compute kernels for heterogeneous compute environments

Inventors: Ahmad Byagowi (Fremont, CA); Michael Maroye Lambeta (Mountain View, CA); Martin Mroz (San Francisco, CA)
Assignee: Meta Platforms, Inc.
G06F9/4893G06F1/26G06F9/48G06F9/4806G06F9/4843G06F9/4881G06F9/50G06F9/5005G06F9/5027G06F9/5094
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,556,382
App. No.
16/508,106
Granted
Jan 17, 2023
Kind
B1
Abstract

A request to perform a compute task is received. A plurality of compute processor resources eligible to perform the compute task is identified, wherein the plurality of compute processor resources includes two or more of the following: a field-programmable gate array, an application-specific integrated circuit, a graphics processing unit, or a central processing unit. Based on an optimization metric, one of the compute processor resources is dynamically selected to perform the compute task.

Claims (38)

1. A system, comprising:

a processor configured to:

receive a request to perform a compute task;

identify a plurality of compute processor resources eligible to perform the compute task, wherein the plurality of compute processor resources includes a field-programmable gate array and an application-specific integrated circuit; and

based on a cost function that includes power consumption and delay in returning results of the compute task of one of the plurality of computer processor resources, dynamically select one of the compute processor resources to perform the compute task, wherein the power consumption is calculated based on an expected computation time of the compute task and an average power consumption of the one of the plurality of compute processor resources, and wherein the delay in returning results is calculated by multiplying the expected computation time of the compute task by a request queue depth for the compute task of the one of the plurality of compute processor resources; and

the compute processor resources.

2. The system of claim 1 , wherein the processor is configured to dynamically select one of the compute processor resources to perform the compute task including by being configured to select a compute processor resource associated with a lowest cost function.

3. The system of claim 1 , wherein the processor is configured to dynamically select one of the compute processor resources to perform the compute task based at least in part on minimizing a cost function associated with a plurality of compute requests.

4. The system of claim 1 , wherein the processor is configured to dynamically select one of the compute processor resources to perform the compute task including by being configured to select a first available compute processor resource capable of performing the compute task.

5. The system of claim 1 , wherein the request to perform the compute task is received via an application programming interface.

6. The system of claim 1 , wherein the request to perform the compute task originates from a software application.

7. The system of claim 1 , wherein the processor is configured to identify the plurality of compute processor resources eligible to perform the compute task based as least in part on a determination of which compute processor resources are available to perform the compute task.

8. The system of claim 1 , wherein the processor is configured to identify the plurality of compute processor resources eligible to perform the compute task based as least in part on a determination of which compute processor resources have been configured to perform the compute task.

9. The system of claim 1 , wherein the processor is configured to identify the plurality of compute processor resources eligible to perform the compute task from among compute processor resources distributed across multiple server clusters.

10. The system of claim 1 , wherein the selected compute processor resource is configured using a driver to perform the compute task.

11. The system of claim 1 , wherein the compute processor resources are located in a single server cluster.

12. The system of claim 1 , wherein the compute processor resources are located across multiple server clusters.

13. The system of claim 1 , wherein the processor is further configured to report a result associated with performing the compute task.

14. The system of claim 13 , wherein the result is formatted as an application programming interface response object.

15. The system of claim 1 , wherein field-programmable gate array compute processor resources are configured to be automatically reprogrammed to execute a different type of compute task.

16. A system, comprising:

a processor configured to:

receive a first request to perform a first compute task;

identify a first plurality of compute processor resources eligible to perform the first compute task, wherein the first plurality of compute processor resources includes a field-programmable gate array and an application-specific integrated circuit;

based on a cost function that includes power consumption and delay in returning results of the compute task of one of the plurality of computer processor resources dynamically select one of the first plurality of compute processor resources to perform the first compute task, wherein the power consumption is calculated based on an expected computation time of the first compute task and an average power consumption of the one of the plurality of compute processor resources, and wherein the delay in returning results is calculated by multiplying the expected computation time of the first compute task by a request queue depth for the first compute task of the one of the plurality of compute processor resources;

receive a second request to perform a second compute task, wherein the second compute task is a different type of compute task than the first compute task;

identify a second plurality of compute processor resources eligible to perform the second compute task, wherein the second plurality of compute processor resources includes two or more of the following: a field-programmable gate array, an application-specific integrated circuit, a graphics processing unit, or a central processing unit; and

based on a cost function that is based at least in part on one or more computational properties of the two or more compute processor resources of the second plurality of compute processor resources, dynamically select one of the second plurality of compute processor resources to perform the second compute task;

the first plurality of compute processor resources; and

the second plurality of compute processor resources.

17. The system of claim 16 , wherein the processor is configured to dynamically select one of the first plurality of compute processor resources to perform the first compute task including by being configured to select a compute processor resource associated with a lowest cost function.

18. The system of claim 16 , wherein the processor is configured to dynamically select one of the first plurality of compute processor resources to perform the first compute task including by being configured to select a first available compute processor resource capable of performing the first compute task.

19. A method, comprising:

receiving a request to perform a compute task;

identifying a plurality of compute processor resources eligible to perform the compute task, wherein the plurality of compute processor resources includes a field-programmable gate array and an application-specific integrated circuit; and

based on a cost function that includes power consumption and delay in returning results of the compute task of one of the plurality of computer processor resources, dynamically selecting one of the compute processor resources to perform the compute task, wherein the power consumption is calculated based on an expected computation time of the compute task and an average power consumption of the one of the plurality of compute processor resources, and wherein the delay in returning results is calculated by multiplying the expected computation time of the compute task by a request queue depth for the compute task of the one of the plurality of compute processor resources.

20. The method of claim 19 , further comprising:

dynamically selecting one of the compute processor resources to perform the compute task including by being configured to select a compute processor resource associated with a lowest cost function.

Assignments (2)
CHANGE OF NAME Recorded Nov 19, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058214/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 10, 2019
From: BYAGOWI, AHMAD; LAMBETA, MICHAEL MAROYE; MROZ, MARTIN
To: FACEBOOK, INC.
Reel/Frame 050331/0466 →
Cited By (3)
US 12,217,085 US 12,474,763 US 12,561,137