IP Library › Granted Patent US 12,159,225
Granted Patent B2
US 12,159,225 · App. 17/136,229 · Granted Dec 3, 2024

Queue allocation in machine learning accelerators

Inventors: Xiangyu Dong (San Jose, CA); Kais Belgaied (San Jose, CA); Yazhou Zu (San Francisco, CA)
Assignee: Google LLC
G06N3/08G06F9/544G06F9/547G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,159,225
App. No.
17/136,229
Granted
Dec 3, 2024
Kind
B2
Abstract

This disclosure generally provides solutions for improving the performance of a custom-built, packet-switched, TPU accelerator-side communication network. Specifically a set of solutions to improve the flow-control behavior by tuning the packet buffer queues in the on-chip router in the distributed training supercomputer network are described.

Claims (158)

1. A computer-implemented memory allocation method for a machine learning accelerator communication network, the method comprising:

accessing metadata associated with a plurality of communications ports of an application specific integrated circuit (ASIC), wherein the metadata identifies, for each communications port of the plurality of communications ports, whether a respective communications port is used in a current configuration and a communications medium associated with the respective communications port;

determining an expected latency for each communications port of the plurality of communications ports based on the accessed metadata;

allocating portions of a shared memory to each communications port of the plurality of communications ports by:

assigning a zero memory for communications ports that are not used;

determining a memory allocation for each communications port based on the expected latency; and

assigning a start address and a stop address of the shared memory to each communications port, wherein assigning the start address and the stop address of the shared memory to each communications port comprises invoking, by a different device from the ASIC, an application programming interface (API).

2. The method of claim 1 , comprising executing a process on the ASIC, the process using the machine learning accelerator communications network and the allocated shared memory.

3. The method of claim 2 , wherein the process comprises training a neural network.

4. The method of claim 1 , wherein the ASIC is a Tensor Processing unit (TPU).

5. The method of claim 1 , wherein the communications medium identified in the metadata includes at least one of:

a copper cable medium;

an optical medium; or

a printed circuit board (PCB) medium.

6. The method of claim 1 , wherein memory allocation is determined based on

Queue

⁢

⁢

Size

⁡

(

p

)

=

l

⁢

a

⁢

t

⁢

e

⁢

n

⁢

c

⁢

y

p

(

N

-

1

)

*

∑

latency

i

*

Total

⁢

⁢

Size

.

7. A computer-implemented memory allocation method for a machine learning accelerator communication network, the method comprising:

determining a network topology for a network of machine learning accelerator application specific integrated circuits (ASICs);

accessing metadata associated with a plurality of communications ports of each ASIC within the network, wherein the metadata identifies, for each communications port of the plurality of communications ports, whether a respective communications port is used in a current configuration and a communications medium associated with the respective communications port;

determining, for each used communications port in the network topology a round-trip-time (RTT) delay;

allocating portions of a shared memory to each communications port of the plurality of communications ports by:

determining a memory allocation for each communications port proportional to the RTT delay; and

assigning a start address and a stop address of the shared memory to each communications port for the determined memory allocation;

executing a process on the machine learning accelerator to send profiling traffic into the network for a predetermined duration;

determining, for each communications port, a number of received traffic packets; and

reallocating portions of the shared memory to each communications port of the plurality of communications ports by:

determining a memory allocation for each communications port proportional to the number of received traffic packets; and

reassigning the start address and the stop address of the shared memory for each communications port for the determined memory allocation.

8. The method of claim 7 , wherein assigning and reassigning the start address and the stop address of the shared memory to each communications port comprises invoking, by a different device from the ASIC, an application programming interface (API).

9. The method of claim 7 , comprising executing a process on the ASIC, the process using the machine learning accelerator communications network and the allocated shared memory.

10. The method of claim 9 , wherein the process comprises training a neural network.

11. The method of claim 7 , wherein the ASIC is a Tensor Processing unit (TPU).

12. The method of claim 7 , wherein memory allocation proportional to the determined RTT delay is determined based on

Queue

⁢

⁢

Size

⁡

(

p

)

=

l

⁢

a

⁢

t

⁢

e

⁢

n

⁢

c

⁢

y

p

(

N

-

1

)

*

∑

latency

i

*

Total

⁢

⁢

Size

.

13. The method of claim 7 , wherein memory allocation for each communications port proportional to the number of received packets is determined based on

Queue

⁢

⁢

Size

⁡

(

p

)

=

packet_num

p

(

N

-

1

)

*

∑

packet_num

i

*

Total

⁢

⁢

Size

.

14. The method of claim 7 , wherein RTT delay is calculated prior to execution by transmitting and receiving one or more timing messages to determine latency.

15. The method of claim 7 , wherein the profiling traffic is at least one of:

all-to-all traffic;

nearest-neighbor traffic; or

a synthetic traffic profile.

16. A computer-implemented memory allocation method for a machine learning accelerator communication network, the method comprising:

determining a network topology for a network of machine learning accelerator application specific integrated circuits (ASICs);

accessing metadata associated with a plurality of communications ports of each ASIC, wherein the metadata identifies, for each communications port of the plurality of communications ports, whether a respective communications port is used in a current configuration and a communications medium associated with the respective communications port;

determining, for each used communications port in the network topology a round-trip-time (RTT) delay;

allocating portions of a shared memory to each communications port of the plurality of communications ports by:

determining a memory allocation for each communications port proportional to the RTT delay; and

assigning a start address and a stop address of the shared memory to each communications port for the determined memory allocation;

executing a process on the ASIC, the process using the machine learning accelerator communications network with the allocated shared memory; and

during execution of the process:

determining a number of message packets received at each communications port of the plurality of communications ports over a first time period;

determining, based on the number of message packets received at each communications port, a desired portion size of the shared memory for each communications port of the plurality of communications ports;

pausing the process for a second time period;

determining that the shared memory is clear of pending message packets;

reassigning the start address and the stop address of the shared memory for each communications port for the desired portion size; and

resuming execution of the process.

17. The method of claim 16 , wherein assigning and reassigning the start address and the stop address of the shared memory to each communications port comprises invoking, by a different device from the ASIC, an application programming interface (API).

18. The method of claim 16 , wherein the ASIC is a tensor processing unit (TPU).

19. The method of claim 16 , wherein the process comprises training a neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2021
From: BELGAIED, KAIS; DONG, XIANGYU; ZU, YAZHOU
To: GOOGLE LLC
Reel/Frame 054995/0373 →
Continuity (2)
Provisional Application 63091708 · Oct 14, 2020
Related Publication 20220114440A1 · Apr 14, 2022