IP Library Granted Patent US 11,662,986
Granted Patent B1
US 11,662,986 · App. 16/825,282 · Granted May 30, 2023

Reducing data transfer to machine learning accelerator hardware

Inventors: Garret Ray Catron (Sunnyvale, CA); Jordan Samuel Fix (San Francisco, CA); Bertrand Allen Maher (Newark, CA); Nicholas Gibson (Menlo Park, CA); Nadathur Rajagopalan Satish (San Jose, CA); Roman Dzhabarov (Redwood City, CA); Hector Yuen (Palo Alto, CA)
Assignee: Meta Platforms, Inc.
G06F8/41G06T1/20
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,662,986
App. No.
16/825,282
Granted
May 30, 2023
Kind
B1
Abstract

A computer program compiled for a machine learning accelerator hardware and associated with a default input data size is received. An execution of an operation of the computer program is initiated. It is identified that a data size of an input data of the operation is smaller than the default input data size. The smaller data size of the input data of the operation rather than the default input data size is caused to be transferred to the machine learning accelerator hardware for the input data of the operation.

Claims (47)

1. A method, comprising:

receiving a computer program compiled for a machine learning accelerator hardware and associated with a default input data size;

initiating an execution of an operation of the computer program;

receiving a data size of an input data of the operation;

identifying that the data size of the input data of the operation is smaller than the default input data size; and

causing the data size of the input data of the operation that is smaller than the default input data size to be transferred to the machine learning accelerator hardware for the input data of the operation, including by:

utilizing a device manager component configured to manage the machine learning accelerator hardware and provide a direct memory transfer instruction using the data size of the input data of the operation that is smaller than the default input data size;

receiving the direct memory transfer instruction using a driver component that is configured to be an interface between the device manager component and the machine learning accelerator hardware; and

utilizing the driver component to generate a peripheral component interconnect bus compatible transfer command to transfer the data size of the input data of the operation based on the received direct memory transfer instruction.

2. The method of claim 1 , wherein the machine learning accelerator hardware includes one or more of the following components: an application-specific integrated circuit, a graphics processing unit, or a field-programmable gate array.

3. The method of claim 1 , wherein the operation is a convolution operation.

4. The method of claim 1 , wherein the operation is a personalized recommendation system operation.

5. The method of claim 1 , wherein the operation is part of a machine learning inference operation.

6. The method of claim 1 , wherein the data size of the input data is a size of a tensor of raw data.

7. The method of claim 6 , wherein the tensor of raw data includes image data or embedding table data.

8. The method of claim 6 , wherein the tensor includes data organized along one or more dimensions corresponding to one or more of the following properties: batch size, height, width, or depth.

9. The method of claim 1 , further comprising receiving a request to execute the operation.

10. The method of claim 9 , wherein the request is received via a network.

11. The method of claim 1 , further comprising receiving the data size of the input data, the default input data size, or both the data size of the input data and the default input data size from a requestor of the operation.

12. The method of claim 1 , further comprising receiving the data size of the input data, the default input data size, or both the data size of the input data and the default input data size as metadata in a container that also includes the input data.

13. The method of claim 1 , further comprising returning a result of the execution of the operation.

14. The method of claim 1 , wherein initiating the execution of the operation includes loading the computer program into a software runtime environment that is configured to communicate with the machine learning accelerator hardware.

15. A system, comprising:

a processor configured to:

receive a computer program compiled for a machine learning accelerator hardware and associated with a default input data size;

initiate an execution of an operation of the computer program;

receive a data size of an input data of the operation;

identify that the data size of the input data of the operation is smaller than the default input data size; and

cause the data size of the input data of the operation that is smaller than the default input data size to be transferred to the machine learning accelerator hardware for the input data of the operation, including by being configured to:

utilize a device manager component configured to manage the machine learning accelerator hardware and provide a direct memory transfer instruction using the data size of the input data of the operation that is smaller than the default input data size;

utilize a driver component configured to be an interface between the device manager component and the machine learning accelerator hardware to receive the direct memory transfer instruction; and

utilize the driver component to generate a peripheral component interconnect bus compatible transfer command to transfer the data size of the input data of the operation based on the received direct memory transfer instruction;

the machine learning accelerator hardware; and

a memory coupled to the machine learning accelerator hardware.

16. A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

receiving a computer program compiled for a machine learning accelerator hardware and associated with a default input data size;

initiating an execution of an operation of the computer program;

receiving a data size of an input data of the operation;

identifying that the data size of the input data of the operation is smaller than the default input data size; and

causing the data size of the input data of the operation that is smaller than the default input data size to be transferred to the machine learning accelerator hardware for the input data of the operation, including by:

utilizing a device manager component configured to manage the machine learning accelerator hardware and provide a direct memory transfer instruction using the data size of the input data of the operation that is smaller than the default input data size;

receiving the direct memory transfer instruction using a driver component that is configured to be an interface between the device manager component and the machine learning accelerator hardware; and

utilizing the driver component to generate a peripheral component interconnect bus compatible transfer command to transfer the data size of the input data of the operation based on the received direct memory transfer instruction.

17. The computer program product of claim 16 , wherein the machine learning accelerator hardware includes one or more of the following components: an application-specific integrated circuit, a graphics processing unit, or a field-programmable gate array.

18. The computer program product of claim 16 , wherein the operation is a convolution operation.

19. The computer program product of claim 16 , wherein the operation is a personalized recommendation system operation.

20. The computer program product of claim 16 , wherein the operation is part of a machine learning inference operation.

Assignments (2)
CHANGE OF NAME Recorded Nov 19, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058214/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 20, 2020
From: CATRON, GARRET RAY; FIX, JORDAN SAMUEL; MAHER, BERTRAND ALLEN; GIBSON, NICHOLAS; SATISH, NADATHUR RAJAGOPALAN; DZHABAROV, ROMAN; YUEN, HECTOR
To: FACEBOOK, INC.
Reel/Frame 052713/0936 →
Cited By (4)
US 12,236,414 US 12,327,177 US 12,561,539 US 12,602,349