IP Library › Granted Patent US 11,556,381
Granted Patent B2
US 11,556,381 · App. 17/738,909 · Granted Jan 17, 2023

Asynchronous distributed data flow for machine learning workloads

Inventors: Jeffrey Adgate Dean (Palo Alto, CA); Sudip Roy (San Jose, CA); Michael Acheson Isard (San Francisco, CA); Aakanksha Chowdhery (Mountain View, CA); Brennan Saeta (Kirkland, WA); Chandramohan Amyangot Thekkath (Palo Alto, CA); Daniel William Hurt (Westminster, CO); Hyeontaek Lim (Palo Alto, CA); Laurent El Shafey (Mountain View, CA); Parker Edward Schuh (Mountain View, CA); Paul Ronald Barham (San Francisco, CA); Ruoming Pang (New York, NY); Ryan Sepassi (Palo Alto, CA); Sanjay Ghemawat (Mountain View, CA); Yonghui Wu (Fremont, CA)
Assignee: Google LLC
G06F9/4881G06N3/063G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,556,381
App. No.
17/738,909
Granted
Jan 17, 2023
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media, for distributing machine learning workloads, e.g., computations for training a neural network or computing an inference using a neural network, across multiple hardware accelerators. One of the systems comprises a plurality of accelerator islands, each hardware accelerator island comprising a respective plurality of hardware devices that include a plurality of hardware accelerators and a corresponding host for each of the plurality of hardware accelerators; and a respective scheduler for each of the accelerator islands that is configured to schedule workloads across the plurality of accelerators and corresponding hosts in the accelerator island, wherein the system is configured to: receive data representing a machine learning workload; and assign a respective portion of the machine learning workload to each of the plurality of accelerator islands for scheduling by the respective scheduler for the accelerator island.

Claims (32)

1. A system comprising:

a plurality of accelerator islands, each accelerator island comprising a respective plurality of hardware devices that include a plurality of hardware accelerators and a corresponding host for each of the plurality of hardware accelerators; and

a respective scheduler for each of the accelerator islands that is configured to schedule workloads across the plurality of accelerators and corresponding hosts in the accelerator island, wherein the system is configured to:

receive data representing a machine learning workload; and

assign a respective portion of the machine learning workload to each of the plurality of accelerator islands for scheduling by the respective scheduler for the accelerator island, comprising assigning the respective portion of the machine learning workload to each of the plurality of accelerator islands by sending a single message to the respective scheduler for the accelerator island when the respective portion of the machine learning workload is a regular computation.

2. The system of claim 1 , wherein the data representing the machine learning workload is data representing a sharded dataflow program comprising a plurality of shards.

3. The system of claim 2 , wherein assigning the respective portion of the machine learning workload to each of the plurality of accelerator islands comprises assigning one or more shards of the sharded dataflow program to each of the plurality of accelerator islands.

4. The system of claim 1 , wherein each scheduler is configured to, when the respective portion of the machine learning workload assigned to the accelerator island is a regular computation, schedule the portion of the computation using parallel asynchronous dispatch.

5. The system of claim 4 , wherein scheduling the portion of the computation using parallel asynchronous dispatch comprises:

generating a schedule that assigns, to each of a set of the hardware accelerators in the accelerator island, a respective set of one or more operations that takes as input and output of one or more respective other operations that are performed by another one of the hardware accelerators in the accelerator island;

determining, for each of the set of hardware accelerators, a respective size of the output of the one or more respective other operations; and

transmitting, in parallel and to the corresponding host for each of the set of hardware accelerators, respective future data specifying the respective size of the output of the one or more respective other operations.

6. The system of claim 5 , wherein the respective future data causes the corresponding host to (i) allocate memory on the hardware accelerator for storing the output of the one or more respective other operations and (ii) transmit data to a corresponding host of the accelerator assigned to the one or more respective other operations that identifies the allocated memory.

7. The system of claim 6 , wherein the corresponding host of the accelerator assigned to the one or more respective other operations is configured to cause the accelerator assigned to the one or more respective other operations to transmit the output of the respective other operations to the allocated memory.

8. The system of claim 7 , wherein the output is transmitted over an accelerator interconnect network.

9. One or more computer-readable storage media storing instructions that when executed by one or more computers cause the one or more computers to perform the operations comprising:

receiving data representing a machine learning workload by a plurality of accelerator islands, wherein each accelerator island comprises a respective plurality of hardware devices that include a plurality of hardware accelerators and a corresponding host for each of the plurality of hardware accelerators, and wherein each accelerator island has a respective scheduler that is configured to schedule workloads across the plurality of accelerators and corresponding hosts in the accelerator island; and

assigning a respective portion of the machine learning workload to each of the plurality of accelerator islands for scheduling by the respective scheduler for the accelerator island, comprising assigning the respective portion of the machine learning workload to each of the plurality of accelerator islands by sending a single message to the respective scheduler for the accelerator island when the respective portion of the machine learning workload is a regular computation.

10. The computer-readable storage media of claim 9 , wherein each scheduler is configured to, when the respective portion of the machine learning workload assigned to the accelerator island is a regular computation, schedule the portion of the computation using parallel asynchronous dispatch.

11. A method comprising:

receiving data representing a machine learning workload by a plurality of accelerator islands, wherein each accelerator island comprises a respective plurality of hardware devices that include a plurality of hardware accelerators and a corresponding host for each of the plurality of hardware accelerators, and wherein each accelerator island has a respective scheduler that is configured to schedule workloads across the plurality of accelerators and corresponding hosts in the accelerator island; and

assigning a respective portion of the machine learning workload to each of the plurality of accelerator islands for scheduling by the respective scheduler for the accelerator island, comprising assigning the respective portion of the machine learning workload to each of the plurality of accelerator islands by sending a single message to the respective scheduler for the accelerator island when the respective portion of the machine learning workload is a regular computation.

12. The method of claim 11 , wherein the data representing the machine learning workload is data representing a sharded dataflow program comprising a plurality of shards.

13. The method of claim 12 , wherein assigning the respective portion of the machine learning workload to each of the plurality of accelerator islands comprises assigning one or more shards of the sharded dataflow program to each of the plurality of accelerator islands.

14. The method of claim 11 , wherein each scheduler is configured to, when the respective portion of the machine learning workload assigned to the accelerator island is a regular computation, schedule the portion of the computation using parallel asynchronous dispatch.

15. The method of claim 14 , wherein scheduling the portion of the computation using parallel asynchronous dispatch comprises:

generating a schedule that assigns, to each of a set of the hardware accelerators in the accelerator island, a respective set of one or more operations that takes as input and output of one or more respective other operations that are performed by another one of the hardware accelerators in the accelerator island;

determining, for each of the set of hardware accelerators, a respective size of the output of the one or more respective other operations; and

transmitting, in parallel and to the corresponding host for each of the set of hardware accelerators, respective future data specifying the respective size of the output of the one or more respective other operations.

16. The method of claim 15 , wherein the respective future data causes the corresponding host to (i) allocate memory on the hardware accelerator for storing the output of the one or more respective other operations and (ii) transmit data to a corresponding host of the accelerator assigned to the one or more respective other operations that identifies the allocated memory.

17. The method of claim 16 , wherein the corresponding host of the accelerator assigned to the one or more respective other operations is configured to cause the accelerator assigned to the one or more respective other operations to transmit the output of the respective other operations to the allocated memory.

18. The method of claim 17 , wherein the output is transmitted over an accelerator interconnect network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2023
From: DEAN, JEFFREY ADGATE; ROY, SUDIP; ISARD, MICHAEL ACHESON; CHOWDHERY, AAKANKSHA; SAETA, BRENNAN; THEKKATH, CHANDRAMOHAN AMYANGOT; HURT, DANIEL WILLIAM; LIM, HYEONTAEK; SHAFEY, LAURENT EL; SCHUH, PARKER EDWARD; BARHAM, PAUL RONALD; PANG, RUOMING; SEPASSI, RYAN; GHEMAWAT, SANJAY; WU, YONGHUI
To: GOOGLE LLC
Reel/Frame 063633/0777 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2022
From: DEAN, JEFFREY ADGATE; ROY, SUDIP; ISARD, MICHAEL ACHESON; CHOWDHERY, AAKANKSHA; SAETA, BRENNAN; THEKKATH, CHANDRAMOHAN AMYANGOT; HURT, DANIEL WILLIAM; LIM, HYEONTAEK; SHAFEY, LAURENT EL; SCHUH, PARKER EDWARD; BARHAM, PAUL RONALD; PANG, RUOMING; SEPASSI, RYAN; GHEMAWAT, SANJAY; WU, YONGHUI
To: GOOGLE LLC
Reel/Frame 060815/0180 →
Continuity (2)
Provisional Application 63186031 · May 7, 2021
Related Publication 20220357985A1 · Nov 10, 2022