IP Library Granted Patent US 12,416,962
Granted Patent B2
US 12,416,962 · App. 17/031,739 · Granted Sep 16, 2025

Mechanism for performing distributed power management of a multi-GPU system by powering down links based on previously detected idle conditions

Inventor: Benjamin Tsien (Santa Clara, CA)
Assignee: Advanced Micro Devices, Inc.
G06F1/3228G06F9/4893
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,416,962
App. No.
17/031,739
Granted
Sep 16, 2025
Kind
B2
Abstract

Systems, apparatuses, and methods for efficient power management of a multi-node computing system are disclosed. A computing system includes multiple nodes that receive tasks to process. The nodes include a processor, local memory, a power controller, and multiple link interfaces for transferring messages with other nodes across links. Using a distributed approach for power management, negotiation for powering down components of the computing system occurs without performing a centralized system-wide power down. Each node is able to power down its links, its processor and other components regardless of whether other components of the computing system are still active or powered up. A link interface initiates power down of a link with delay or without delay based on a prediction of whether a link idle condition leads to the link interface remaining idle for at least a target idle threshold period of time.

Claims (71)

1. A computing system comprising:

a first partition comprising:

a plurality of processing nodes, wherein each processing node comprises:

a memory; and

a plurality of clients, each comprising circuitry configured to:

execute tasks assigned to the processing node by a host processing node; and

send memory requests to the memory; and

a plurality of links between the plurality of processing nodes; and

wherein a first processing node of the plurality of processing nodes comprises circuitry configured to:

transfer data, via a first link interface of the first processing node, on a first link of the plurality of links between the first processing node and a second processing node of the plurality of processing nodes, wherein the data comprises at least tasks assigned by the second processing node for execution by the plurality of clients of the first processing node that cause the first processing node to be in an active operational state; and

initiate power down of the first link interface, in response to:

an idle condition of the first link interface; and

a prediction, based on one or more previously detected idle conditions, that the first link interface will remain idle for at least a target idle threshold period of time; and

wherein in response to each link interface of the first processing node being powered down and at least one processor of the first processing node being active, the circuitry of the first processing node is further configured to change a power management state of the at least one processor to a higher performance power management state.

2. The computing system as recited in claim 1 , wherein the circuitry of the first processing node is further configured to initiate power down of the first link interface after a wait threshold period of time greater than the target idle threshold period of time, in response to:

the idle condition of the first link interface after the wait threshold period of time; and

a prediction before the target idle threshold period of time, based on the one or more previously detected idle conditions, that the first link interface will become active prior to the target idle threshold period of time.

3. The computing system as recited in claim 1 , wherein the circuitry of the first processing node is further configured to:

update a power down prediction value to indicate a higher confidence that the first link interface will remain idle for at least the target idle threshold period of time, in response to:

no interruption has occurred prior to the target idle threshold period of time elapsing that prevents the first link interface from remaining idle for the target idle threshold period.

4. The computing system as recited in claim 1 , wherein the circuitry of the first processing node is further configured to:

update a power down prediction value to indicate a lower confidence that the first link interface will remain idle for at least the target idle threshold period of time, in response to:

an interruption has occurred prior to the target idle threshold period of time elapsing that prevents the first link interface from remaining idle for the target idle threshold period.

5. The computing system as recited in claim 1 , wherein the first link interface comprises circuitry configured to:

update a prediction to indicate that a next detected idle condition of the first link interface will not lead to the first link interface remaining idle for the target idle threshold period of time, in response to determining a power down prediction value is less than a threshold.

6. The computing system as recited in claim 1 , wherein the first link interface comprises circuitry configured to:

update a prediction to indicate that a next detected idle condition of the first link interface will lead to the first link interface remaining idle for the target idle threshold period of time, in response to determining a power down prediction value is greater than or equal to a threshold.

7. The computing system as recited in claim 1 , wherein:

the computing system further comprises a second partition; and

the first partition further comprises circuitry configured to power down remaining components of the first partition while the second partition has one or more active processing nodes, in response to determining each processing node of the first partition is powered down.

8. A method, comprising:

processing, by a first partition, a plurality of tasks, wherein the first partition comprises:

a plurality of processing nodes configured to process tasks, wherein each processing node comprises:

a memory; and

a plurality of clients, each comprising circuitry configured to execute program instructions of tasks assigned to the processing node by a host processing node that includes sending memory requests to the memory; and

a plurality of links between the plurality of processing nodes; and

transferring data, via a first link interface of a first processing node, on a first link of the plurality of links between the first processing node and a second processing node, wherein the data comprises at least tasks assigned by the second processing node for execution by the plurality of clients of the first processing node that cause the first processing node to be in an active operational state; and

initiating power down of the first link interface, in response to:

an idle condition of the first link interface; and

a prediction, based on one or more previously detected idle conditions, that the idle condition leads to the first link interface will remain idle for at least a target idle threshold period of time; and

wherein in response to each link interface of the first processing node being powered down and at least one processor of the first processing node being active, the method comprises changing, by circuitry of the first processing node, a power management state of the at least one processor to a higher performance power management state.

9. The method as recited in claim 8 , further comprising initiating power down of the first link interface, by the circuitry of the first processing node, after a wait threshold period of time greater than the target idle threshold period of time, in response to:

the idle condition of the first link interface after the wait threshold period of time; and

a prediction before the target idle threshold period of time, based on the one or more previously detected idle conditions, that first link interface will become active prior to the target idle threshold period of time.

10. The method as recited in claim 8 , further comprising:

updating, by the circuitry of the first processing node, a power down prediction value to indicate a higher confidence that the first link interface will remain idle for at least the target idle threshold period of time, in response to:

no interruption has occurred prior to the target idle threshold period of time elapsing that prevents the first link interface from remaining idle for the target idle threshold period.

11. The method as recited in claim 8 , further comprising:

updating, by the circuitry of the first processing node, a power down prediction value to indicate a lower confidence that the first link interface remaining idle for at least the target idle threshold period of time, in response to determining:

an interruption has occurred prior to the target idle threshold period of time elapsing that prevents the first link interface from remaining idle for the target idle threshold period.

12. The method as recited in claim 8 , further comprising:

updating, by the circuitry of the first processing node, a prediction to indicate that a next detected idle condition of the first link interface will not lead to the first link interface remaining idle for the target idle threshold period of time, in response to determining a power down prediction value is less than a threshold.

13. The method as recited in claim 8 , further comprising:

updating, by the circuitry of the first processing node, a prediction to indicate that a next detected idle condition of the first link interface will lead to the first link interface remaining idle for the target idle threshold period of time, in response to determining a power down prediction value is greater than or equal to a threshold.

14. An apparatus comprising:

a physical unit comprising circuitry configured to manage transfer of data on a link between a first processing node and a second processing node of a plurality of processing nodes, wherein the data comprises at least tasks assigned by the second processing node for execution by a plurality of clients of the first processing node that cause the first processing node to be in an active operational state, wherein each processing node comprises:

a memory; and

a plurality of clients, each comprising circuitry configured to execute program instructions of tasks assigned to the processing node by a host processing node that includes sending memory requests to the memory; and

a power down unit; and

wherein the power down unit comprises circuitry configured to initiate power down of the apparatus, based at least in part on:

an idle condition of the apparatus; and

a prediction, based on one or more previously detected idle conditions, that the apparatus will remain idle for at least a target idle threshold period of time; and

wherein in response to each link interface of the first processing node being powered down and at least one processor of the first processing node being active, the circuitry of the first processing node is configured to change a power management state of the at least one processor to a higher performance power management state.

15. The apparatus as recited in claim 14 , wherein the circuitry of the power down unit is further configured to initiate power down of the apparatus after a wait threshold period of time greater than the target idle threshold period of time, in response to:

the idle condition of the apparatus after the wait threshold period of time; and

a prediction before the target idle threshold period of time, based on the one or more previously detected idle conditions, that the apparatus will become active prior to the target idle threshold period of time.

16. The apparatus as recited in claim 14 , wherein the circuitry of the power down unit is further configured to:

update a power down prediction value to indicate a higher confidence that the apparatus will remain idle for at least the target idle threshold period of time, in response to:

no interruption has occurred prior to the target idle threshold period of time elapsing that prevents the apparatus from remaining idle for the target idle threshold period.

17. The apparatus as recited in claim 14 wherein the circuitry of the power down unit is further configured to:

update a prediction to indicate that a next detected idle condition of the apparatus will not lead to the apparatus remaining idle for the target idle threshold period of time, in response to determining a power down prediction value is less than a threshold.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 29, 2021
From: TSIEN, BENJAMIN
To: ADVANCED MICRO DEVICES, INC.
Reel/Frame 056709/0897 →
Continuity (1)
Related Publication 20220091657A1 · Mar 24, 2022
References Cited (73)
US 4980836A · Carter et al. · 1990 [cited by applicant]
US 5396635A · Fung · 1995 [cited by applicant]
US 5617572A · Pearce et al. · 1997 [cited by applicant]
US 5692202A · Kardach et al. · 1997 [cited by applicant]
US 6334167B1 · Gerchman et al. · 2001 [cited by applicant]
US 6657534B1 · Beer · 2003 [cited by applicant]
US 6657634B1 · Sinclair et al. · 2003 [cited by applicant]
US 7028200B2 · Ma · 2006 [cited by applicant]
US 7085941B2 · Li · 2006 [cited by applicant]
US 7394288B1 · Agarwal · 2008 [cited by applicant]
US 7428644B2 · Jeddeloh et al. · 2008 [cited by applicant]
US 7437579B2 · Jeddeloh et al. · 2008 [cited by applicant]
US 7496777B2 · Kapil · 2009 [cited by applicant]
US 7613941B2 · Samson et al. · 2009 [cited by applicant]
US 7743267B2 · Snyder et al. · 2010 [cited by applicant]
US 7800621B2 · Fry · 2010 [cited by applicant]
US 7802060B2 · Hildebrand · 2010 [cited by applicant]
US 7840827B2 · Dahan et al. · 2010 [cited by applicant]
US 7868479B2 · Subramaniam · 2011 [cited by applicant]
US 7873850B2 · Cepulis et al. · 2011 [cited by applicant]
US 7899990B2 · Moll et al. · 2011 [cited by applicant]
US 8181046B2 · Marcu et al. · 2012 [cited by applicant]
US 8402232B2 · Avudaiyappan et al. · 2013 [cited by applicant]
US 8438416B2 · Kocev et al. · 2013 [cited by applicant]
US 8656198B2 · Branover et al. · 2014 [cited by applicant]
US 8924758B2 · Steinman et al. · 2014 [cited by applicant]
US 8949644B2 · Ma · 2015 [cited by applicant]
US 9563257B2 · Fang · 2017 [cited by applicant]
US 9983652B2 · Piga et al. · 2018 [cited by applicant]
US 10671148B2 · Tsien et al. · 2020 [cited by applicant]
US 20040015628A1 · Glasco et al. · 2004 [cited by applicant]
US 20050283523A1 · Almeida · 2005 [cited by examiner]
US 20060271649A1 · Tseng · 2006 [cited by applicant]
US 20080288798A1 · Cooper et al. · 2008 [cited by applicant]
US 20090235105A1 · Branover et al. · 2009 [cited by applicant]
US 20100077243A1 · Wang · 2010 [cited by examiner]
US 20100106876A1 · Nakahashi et al. · 2010 [cited by applicant]
US 20110083023A1 · Dickens · 2011 [cited by applicant]
US 20110153924A1 · Vash et al. · 2011 [cited by applicant]
US 20110264934A1 · Branover et al. · 2011 [cited by applicant]
US 20120254526A1 · Kalyanasundharam · 2012 [cited by applicant]
US 20130003559A1 · Matthews · 2013 [cited by examiner]
US 20130007483A1 · Diefenbaugh · 2013 [cited by examiner]
US 20130007491A1 · Iyer · 2013 [cited by examiner]
US 20130132755A1 · Cooper · 2013 [cited by examiner]
US 20130179621A1 · Smith · 2013 [cited by applicant]
US 20130311804A1 · Garg et al. · 2013 [cited by applicant]
US 20130332764A1 · Juang · 2013 [cited by examiner]
US 20140095801A1 · Bodas et al. · 2014 [cited by applicant]
US 20140122833A1 · Davis · 2014 [cited by applicant]
US 20140281275A1 · Kruckemyer et al. · 2014 [cited by applicant]
US 20150373566A1 · Pius · 2015 [cited by examiner]
US 20160147285A1 · Alshinnawi et al. · 2016 [cited by applicant]
US 20160314024A1 · Chang et al. · 2016 [cited by applicant]
US 20170160781A1 · Piga et al. · 2017 [cited by applicant]
US 20170310492A1 · Wang · 2017 [cited by examiner]
US 20170353926A1 · Zhu · 2017 [cited by applicant]
US 20180004273A1 · Leucht-Roth · 2018 [cited by examiner]
US 20180081420A1 · Jones et al. · 2018 [cited by applicant]
US 20180157311A1 · Maisuria · 2018 [cited by applicant]
US 20180188797A1 · Wang · 2018 [cited by examiner]
US 20190196574A1 · Tsien · 2019 [cited by examiner]
US 20190204899A1 · Tsien et al. · 2019 [cited by applicant]
US 20190384733A1 · Jen · 2019 [cited by examiner]
US 20190387074A1 · Seibert · 2019 [cited by examiner]
Invitation to Pay Additional Fees, Communication Relating to the Results of the Partial International Search and Provisional Opinion Accompanying the Partial Search Result in International Application No. PCT/US2021/051… [cited by applicant]
Li et al., “Compiler-Directed Proactive Power Management for Networks”, CASES '05: Proceedings of the 2005 International Conference on Compilers, Architectures and Synthesis for Embedded Systems, Sep. 24, 2005, pp. 137-… [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2018/051789, mailed Jan. 2, 2019, 14 pages. (5800-70901). [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2018/051916, mailed Jan. 31, 2019, 10 pages. [cited by applicant]
Yuan et al., “Buffering Approach for Energy Saving in Video Sensors”, 2003 International Conference on Multimedia and Expo, Jul. 2003, 4 pages. [cited by applicant]
“Intel Power Management Technologies for Processor Graphics, Display, and Memory: White Paper for 2010-2011 Desktop and Notebook Platforms”, Intel Corporation, Aug. 2010, 10 pages. [cited by applicant]
International Search Report and Written Opinion in International Application No. PCT/US2021/051814, mailed Apr. 4, 2022, 20 pages. [cited by applicant]
Office Action in Japan Patent Application No. 2023-518226, dated Jul. 8, 2025, 6 pages. [cited by applicant]