IP Library Patent Application 16945679
Patent Application
App. No. 16/945,679

OPERATION-BASED PARTITIONING OF A PARALLELIZABLE MACHINE LEARNING MODEL NETWORK ON ACCELERATOR HARDWARE

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
16/945,679
Abstract

A machine learning model network is analyzed to identify types of operations and dependencies associated with different portions of the machine learning model network, including by classifying at least a portion of the types of operations as being memory bandwidth intensive or compute intensive. The machine learning model network is partitioned across a plurality of different machine learning accelerator hardware units based at least in part on the analysis. Parallelization and pipelining of an execution of the machine learning model network is allowed based on the partitioning.

Claims (31)

1 . A method, comprising:

analyzing a machine learning model network to identify types of operations and dependencies associated with different portions of the machine learning model network, including by classifying at least a portion of the types of operations as being memory bandwidth intensive or compute intensive;

partitioning the machine learning model network across a plurality of different machine learning accelerator hardware units based at least in part on the analysis; and

allowing parallelization and pipelining of an execution of the machine learning model network based on the partitioning.

2 . The method of claim 1 , wherein the machine learning model network is associated with personalized recommendation computations.

3 . The method of claim 1 , wherein at least a portion of the machine learning model network is associated with an embedding table lookup operation.

4 . The method of claim 1 , wherein at least a portion of the machine learning model network is associated with a combination operation that operates on embedding table lookup results.

5 . The method of claim 1 , wherein the types of operations classified as being compute intensive receive outputs of the types of operations classified as being memory bandwidth intensive.

6 . The method of claim 1 , wherein the machine learning model network includes a multilayer perceptron.

7 . The method of claim 1 , wherein the plurality of different machine learning accelerator hardware units includes one or more the following: an application-specific integrated circuit, a graphics processing unit, or a field-programmable gate array.

8 . The method of claim 1 , wherein each machine learning accelerator hardware unit of the plurality of different machine learning accelerator hardware units includes a compute unit and a memory unit.

9 . The method of claim 8 , wherein the compute unit include multiple computing cores.

10 . The method of claim 1 , wherein partitioning the machine learning network across the plurality of different machine learning accelerator hardware units includes partitioning the machine learning model network across a plurality of different processing units of the plurality of different machine learning accelerator hardware units.

11 . The method of claim 1 , wherein partitioning the machine learning model network across the plurality of different machine learning accelerator hardware units is based at least in part on costs associated with different portions of the machine learning model network that are tracked during one or more inference executions of the machine learning model network.

12 . The method of claim 1 , wherein allowing parallelization of the execution of the machine learning model network includes distributing a portion of the machine learning model network across multiple machine learning accelerator hardware units of the plurality of different machine learning accelerator hardware units.

13 . The method of claim 12 , wherein the distributed portion of the machine learning model network is memory bandwidth intensive.

14 . The method of claim 1 , wherein allowing pipelining of the execution of the machine learning model network includes duplicating a portion of the machine learning model network on multiple machine learning accelerator hardware units of the plurality of different machine learning accelerator hardware units.

15 . The method of claim 14 , wherein the duplicated portion of the machine learning model network is compute intensive.

16 . The method of claim 1 , wherein allowing pipelining of the execution of the machine learning model network includes concurrently executing a memory bandwidth intensive portion of the machine learning model network and a compute intensive portion of the machine learning model network on at least one machine learning accelerator hardware unit of the plurality of different machine learning accelerator hardware units.

17 . The method of claim 1 , wherein allowing pipelining of the execution of the machine learning model network includes allocating a first specified number of processing units of a machine learning accelerator hardware unit to a first portion of the machine learning model network and allocating a second specified number of processing units of the machine learning accelerator hardware unit to a second portion of the machine learning model network.

18 . The method of claim 1 , wherein the plurality of different machine learning accelerator hardware units is communicatively connected to a programmed computer system that directs partitioning of the machine learning model network across the plurality of different machine learning accelerator hardware units.

19 . A system, comprising:

a plurality of different machine learning accelerator hardware units; and

one or more processors configured to:

analyze a machine learning model network to identify types of operations and dependencies associated with different portions of the machine learning model network, including by classifying at least a portion of the types of operations as being memory bandwidth intensive or compute intensive;

partition the machine learning model network across the plurality of different machine learning accelerator hardware units based at least in part on the analysis; and

allow parallelization and pipelining of an execution of the machine learning model network based on the partitioning.

20 . A computer program product, the computer program product being embodied in a non-transitory computer readable storage medium and comprising computer instructions for:

analyzing a machine learning model network to identify types of operations and dependencies associated with different portions of the machine learning model network, including by classifying at least a portion of the types of operations as being memory bandwidth intensive or compute intensive;

partitioning the machine learning model network across a plurality of different machine learning accelerator hardware units based at least in part on the analysis; and

allowing parallelization and pipelining of an execution of the machine learning model network based on the partitioning.

Assignments (2)
CHANGE OF NAME Recorded Nov 19, 2021
From: FACEBOOK, INC.
To: META PLATFORMS, INC.
Reel/Frame 058214/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2020
From: CATRON, GARRET RAY; WANG, MAN; SATISH, NADATHUR RAJAGOPALAN; ANDERSON, MICHAEL; ZHANG, YING; MAHER, BERTRAND ALLEN
To: FACEBOOK, INC.
Reel/Frame 053722/0401 →