IP Library › Granted Patent US 11,436,060
Granted Patent B2
US 11,436,060 · App. 16/552,065 · Granted Sep 6, 2022

Proactive management of inter-GPU network links

Inventors: Karthik Rao (Austin, TX); Abhinav Vishnu (Austin, TX)
Assignee: Advanced Micro Devices, Inc.
G06F9/5094G06F9/3877G06F9/545G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,436,060
App. No.
16/552,065
Granted
Sep 6, 2022
Kind
B2
Abstract

Systems, apparatuses, and methods for proactively managing inter-processor network links are disclosed. A computing system includes at least a control unit and a plurality of processing units. Each processing unit of the plurality of processing units includes a compute module and a configurable link interface. The control unit dynamically adjusts a clock frequency and a link width of the configurable link interface of each processing unit based on a data transfer size and layer computation time of a plurality of layers of a neural network so as to reduce execution time of each layer. By adjusting the clock frequency and the link width of the link interface on a per-layer basis, the overlapping of communication and computation phases is closely matched, allowing layers to complete more quickly.

Claims (45)

1. A system comprising:

a plurality of processing units, wherein each processing unit of the plurality of processing units comprises:

a compute module; and

a configurable link interface configured to communicate data between a respective processing unit and one or more other processing units of the plurality of processing units; and

a control unit comprising circuitry configured to:

select, from a plurality of power combination settings, a first power combination setting for sharing power between the compute module and the configurable link interface of each processing unit, wherein the first power combination setting is selected based at least in part on an amount of data to transfer and a compute kernel execution time during a first period of time while executing a given software application; and

cause, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the first power combination setting during a period of overlap, within the first period of time, between a communication phase and a kernel computation phase while executing the given software application.

2. The system as recited in claim 1 , wherein the control unit is further configured to:

select a second power combination different from the first power combination setting based at least in part on an amount of data to transfer and compute kernel execution time during a second period of time while executing the given software application; and

cause, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the second power combination setting during overlap, within the second period of time, of the communication phase and the kernel computation phase while executing the given software application.

3. The system as recited in claim 1 , wherein:

each power combination setting specifies a given power setting for the compute module and a given power state for the configurable link interface of the plurality of processing units; and

the given power state specifies a number of active lanes and a clock frequency for the configurable link interface.

4. The system as recited in claim 1 , wherein the given software application is a neural network application, and wherein the first power combination setting is selected so as to allow a second neural network layer to begin at an earlier point in time compared to other power combination settings of the plurality of power combination settings.

5. The system as recited in claim 1 , wherein the control unit is configured to generate a table with entries for different power combination settings.

6. The system as recited in claim 5 , wherein the table includes estimates of an amount of time needed to complete data transfer for the plurality of power combination settings, and wherein the table includes estimates of an amount of time needed to complete kernel computation for the plurality of power combination settings.

7. The system as recited in claim 5 , wherein the control unit is configured to generate a plurality of different tables for a plurality of layers of a neural network, wherein each table corresponds to a different layer of the neural network.

8. A method comprising:

selecting, by a control unit comprising circuitry, a first power combination setting from a plurality of power combination settings for sharing power between a compute module and a configurable link interface, wherein the first power combination setting is selected based at least in part on an amount of data to transfer and a compute kernel execution time during a first period of time while executing a given software application; and

causing, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the first power combination setting during a period of overlap, within the first period of time, between a communication phase and a kernel computation phase while executing given software application.

9. The method as recited in claim 8 , further comprising:

selecting a second power combination different from the first power combination setting based at least in part on an amount of data to transfer and compute kernel execution time during a second period of time while executing the given software application; and

causing, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the second power combination setting during overlap, within the second period of time, of the communication phase and the kernel computation phase while executing the given software application.

10. The method as recited in claim 8 , wherein:

each power combination setting specifies a given power setting for the compute module and a given power state for the configurable link interface; and

the given power state specifies a number of active lanes and a clock frequency for the configurable link interface.

11. The method as recited in claim 8 , wherein the given software application is a neural network application, and wherein the first power combination setting is selected so as to allow a second neural network layer to begin at an earlier point in time compared to other power combination settings of the plurality of power combination settings.

12. The method as recited in claim 8 , further comprising generating a table with entries for different power combination settings.

13. The method as recited in claim 12 , wherein the table includes estimates of an amount of time needed to complete data transfer for the plurality of power combination settings, and wherein the table includes estimates of an amount of time needed to complete kernel computation for the plurality of power combination settings.

14. The method as recited in claim 12 , further comprising generating a plurality of different tables for a plurality of layers of a neural network, wherein each table corresponds to a different layer of the neural network.

15. A processing unit comprising:

a compute module; and

a configurable link interface configured to communicate data between the processing unit and one or more other processing units;

wherein the processing unit is configured to:

select, from a plurality of power combination settings, a first power combination setting for sharing power between the compute module and the configurable link interface of each processing unit, wherein the first power combination setting is selected based at least in part on an amount of data to transfer and a compute kernel execution time during a first period of time while executing a given software application; and

cause, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the first power combination setting during a period of overlap, within the first period of time, between a communication phase and a kernel computation phase while executing the given software application.

16. The processing unit as recited in claim 15 , wherein the processing unit is further configured to:

select a second power combination different from the first power combination setting based at least in part on an amount of data to transfer and compute kernel execution time during a second period of time while executing the given software application; and

cause, in a proactive manner, the compute module and the configurable link interface of each processing unit to operate with the second power combination setting during overlap, within the second period of time, of the communication phase and the kernel computation phase while executing the given software application.

17. The processing unit as recited in claim 15 , wherein:

each power combination setting specifies a given power setting for the compute module and a given power state for the configurable link interface; and

the given power state specifies a number of active lanes and a clock frequency for the configurable link interface.

18. The processing unit as recited in claim 15 , wherein the given software application is a neural network application, and wherein the first power combination setting is selected so as to allow a second neural network layer to begin at an earlier point in time compared to other power combination settings of the plurality of power combination settings.

19. The processing unit as recited in claim 15 , wherein the processing unit is configured to generate a table with entries for different power combination settings.

20. The processing unit as recited in claim 19 , wherein the table includes estimates of an amount of time needed to complete data transfer for the plurality of power combination settings, and wherein the table includes estimates of an amount of time needed to complete kernel computation for the plurality of power combination settings.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 25, 2019
From: RAO, KARTHIK; VISHNU, ABHINAV
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 050825/0378 →
Continuity (1)
Related Publication 20210064444A1 · Mar 4, 2021