IP Library › Granted Patent US 12,399,546
Granted Patent B2
US 12,399,546 · App. 18/601,001 · Granted Aug 26, 2025

System, apparatus and method for increasing performance in a processor during a voltage ramp

Inventors: Altug Koker (El Dorado Hills, CA); Abhishek R. Appu (El Dorado Hills, CA); Bhushan M. Borole (Rancho Cordova, CA); Wenyin Fu (Folsom, CA); Kamal Sinha (Rancho Cordova, CA); Joydeep Ray (Folsom, CA)
Assignee: Intel Corporation
G06F1/3234G06F1/3237G06F1/324G06F1/3296G09G5/363G09G5/366G09G2310/066G09G2310/08G09G2340/02G09G2360/06G09G2360/08G09G2370/022G09G2370/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,399,546
App. No.
18/601,001
Granted
Aug 26, 2025
Kind
B2
Abstract

In one embodiment, a processor includes: a graphics processor to execute a workload; and a power controller coupled to the graphics processor. The power controller may include a voltage ramp circuit to receive a request for the graphics processor to operate at a first performance state having a first operating voltage and a first operating frequency and cause an output voltage of a voltage regulator to increase to the first operating voltage. The voltage ramp circuit may be configured to enable the graphics processor to execute the workload at an interim performance state having an interim operating voltage and an interim operating frequency when the output voltage reaches a minimum operating voltage. Other embodiments are described and claimed.

Claims (43)

1. An apparatus comprising:

a graphics processing unit (GPU) comprising:

a plurality of texture units;

a shared memory coupled to the plurality of texture units;

a plurality of register files coupled to the shared memory;

a plurality of load/store units coupled to the shared memory;

a security engine;

a compression circuit to compress and decompress data;

a plurality of graphics processing cores coupled to the shared memory;

a plurality of clock generators, each of the plurality of clock generators to provide a clock signal to at least one of the plurality of graphics processing cores; and

a power controller to control power consumption of the plurality of graphics processing cores, wherein the power controller, in response to a request for the GPU to exit a low power state in which the GPU is powered down and operate at a first performance state having a first operating voltage and a first operating frequency, is to cause an output voltage of a voltage regulator to increase to the first operating voltage,

wherein the GPU is to exit the low power state and execute a workload at a plurality of interim performance states before the output voltage reaches the first operating voltage, wherein each of the plurality of interim performance states has an operating voltage less than the first operating voltage and an operating frequency less than the first operating frequency.

2. The apparatus of claim 1 , further comprising the voltage regulator comprising at least one integrated voltage regulator.

3. The apparatus of claim 1 , wherein the plurality of clock generators comprises at least one phase locked loop to generate the clock signal.

4. The apparatus of claim 3 , wherein the at least one phase locked loop is to dynamically drift output of the clock signal from a first interim operating frequency less than the first operating frequency to another interim operating frequency less than the first operating frequency during the execution of the workload.

5. The apparatus of claim 1 , wherein the GPU is to exit the low power state and execute the workload at a first interim performance state having a second operating voltage, the second operating voltage sufficient to enable the GPU to operate at a second operating frequency comprising a minimum operating frequency.

6. The apparatus of claim 5 , further comprising storage to store a table having a plurality of entries, each of the plurality of entries to associate a voltage ramp value with a time duration.

7. The apparatus of claim 6 , wherein the GPU is to execute the workload at the first interim performance state after the time duration of an entry of the table associated with the second operating voltage.

8. The apparatus of claim 6 , wherein the storage comprises a non-volatile memory.

9. The apparatus of claim 1 , wherein the power controller is to cause a first graphics processing core to operate at a higher performance state based on availability of a power budget.

10. The apparatus of claim 1 , further comprising at least one special function unit.

11. The apparatus of claim 1 , wherein the GPU is to couple to a central processing unit (CPU) via a high speed interconnect.

12. The apparatus of claim 11 , further comprising:

a CPU domain comprising the CPU, at least one first integrated voltage regulator, and a first power controller; and

a GPU domain comprising the GPU, at least one second integrated voltage regulator, and a second power controller.

13. The apparatus of claim 1 , wherein the GPU is to execute the workload at the plurality of interim performance states according to a step function.

14. A graphics processing unit comprising:

a security engine;

a compression circuit to compress and decompress data;

a plurality of texture units;

a shared memory coupled to the plurality of texture units;

a plurality of register files coupled to the shared memory;

a plurality of load/store units coupled to the shared memory;

a plurality of graphics processing cores coupled to the plurality of register files; and

a power controller to cause an output voltage of a voltage regulator to increase to a first operating voltage of a first performance state at which a workload is to be executed, the first performance state comprising the first operating voltage and a first operating frequency, wherein the graphics processing unit is to exit a low power state and execute the workload at a plurality of interim performance states before the output voltage reaches the first operating voltage, wherein each of the plurality of interim performance states has an operating voltage less than the first operating voltage and an operating frequency less than the first operating frequency.

15. The graphics processing unit of claim 14 , wherein the power controller is to cause the graphics processing unit to exit the low power state in response to a request for the execution of the workload.

16. The graphics processing unit of claim 14 , further comprising a plurality of clock generators comprising at least one phase locked loop to generate a clock signal and dynamically drift output of the clock signal from a first interim operating frequency less than the first operating frequency to another interim operating frequency less than the first operating frequency during the execution of the workload.

17. A non-transitory storage medium comprising instructions that when executed cause a power controller to:

receive a request for request for a graphics processing unit to exit a low power state in which the graphics processing unit is powered down and operate at a first performance state comprising a first operating voltage and a first operating frequency;

in response to the request, issue a command to a voltage regulator to increase an output voltage to the first operating voltage; and

cause the graphics processing unit to exit the low power state and execute a workload at a plurality of interim performance states, each of the plurality of interim performance states comprising an interim operating voltage and an interim operating frequency, before the output voltage reaches the first operating voltage, the plurality of interim performance states less than the first performance state.

18. The non-transitory storage medium of claim 17 , wherein the instructions further cause the power controller to allocate a power budget between the graphics processing unit and a central processing unit coupled to the graphics processing unit.

19. The non-transitory storage medium of claim 18 , wherein the instructions further cause the power controller to cause a first core of the central processing unit to operate at a higher performance state when there is available power budget.

Continuity (4)
Continuation 17517090 · Nov 2, 2021
Continuation 16595543 · Oct 8, 2019
Continuation 15488662 · Apr 17, 2017
Related Publication 20240264657A1 · Aug 8, 2024
References Cited (70)
US 6574738B2 · Jain et al. · 2003 [cited by applicant]
US 7376854B2 · Lehwalder et al. · 2008 [cited by applicant]
US 7827424B2 · Bounitch · 2010 [cited by applicant]
US 8134543B1 · Han · 2012 [cited by applicant]
US 8700925B2 · Wyatt · 2014 [cited by applicant]
US 8756451B2 · Kurd et al. · 2014 [cited by applicant]
US 8839012B2 · Khodorkovsky · 2014 [cited by applicant]
US 9269120B2 · Kaburlasos et al. · 2016 [cited by applicant]
US 9305324B2 · McGuire et al. · 2016 [cited by applicant]
US 9367114B2 · Wells et al. · 2016 [cited by applicant]
US 9459689B2 · Ganpule et al. · 2016 [cited by applicant]
US 9626576B2 · Grujic et al. · 2017 [cited by applicant]
US 9760160B2 · Weissman et al. · 2017 [cited by applicant]
US 9965019B2 · Ganpule et al. · 2018 [cited by applicant]
US 9996135B2 · Wells et al. · 2018 [cited by applicant]
US 10223333B2 · Chetlur et al. · 2019 [cited by applicant]
US 10275010B2 · Mair et al. · 2019 [cited by applicant]
US 10324519B2 · Weissman et al. · 2019 [cited by applicant]
US 10372198B2 · Weissman et al. · 2019 [cited by applicant]
US 10394300B2 · Wells et al. · 2019 [cited by applicant]
US 10444817B2 · Koker et al. · 2019 [cited by applicant]
US 11175719B2 · Koker et al. · 2021 [cited by applicant]
US 20020188884A1 · Jain et al. · 2002 [cited by applicant]
US 20050195181A1 · Khodorkovsky · 2005 [cited by applicant]
US 20050223259A1 · Lehwalder et al. · 2005 [cited by applicant]
US 20060026450A1 · Bounitch · 2006 [cited by applicant]
US 20110057937A1 · Wu et al. · 2011 [cited by applicant]
US 20110060924A1 · Khodorkovsky · 2011 [cited by applicant]
US 20110126056A1 · Kelleher et al. · 2011 [cited by applicant]
US 20120331321A1 · Kaburlasos · 2012 [cited by examiner]
US 20130086410A1 · Kurd et al. · 2013 [cited by applicant]
US 20130155045A1 · Khodorkovsky et al. · 2013 [cited by applicant]
US 20140089699A1 · O'Connor et al. · 2014 [cited by applicant]
US 20140125679A1 · Kaburlasos et al. · 2014 [cited by applicant]
US 20140176575A1 · McGuire et al. · 2014 [cited by applicant]
US 20140189399A1 · Govindaraju · 2014 [cited by examiner]
US 20140258760A1 · Wells et al. · 2014 [cited by applicant]
US 20140286579A1 · Grujic et al. · 2014 [cited by applicant]
US 20150149800A1 · Gendler · 2015 [cited by applicant]
US 20150177823A1 · Maiyuran et al. · 2015 [cited by applicant]
US 20150177824A1 · Ganpule et al. · 2015 [cited by applicant]
US 20150317764A1 · Govindaraju et al. · 2015 [cited by applicant]
US 20160062947A1 · Chetlur et al. · 2016 [cited by applicant]
US 20160154677A1 · Barik et al. · 2016 [cited by applicant]
US 20160259389A1 · Wells et al. · 2016 [cited by applicant]
US 20160349828A1 · Weissmann et al. · 2016 [cited by applicant]
US 20160370839A1 · Ganpule et al. · 2016 [cited by applicant]
US 20170068296A1 · Mair et al. · 2017 [cited by applicant]
US 20170371399A1 · Weissmann et al. · 2017 [cited by applicant]
US 20170371400A1 · Weissmann et al. · 2017 [cited by applicant]
US 20180301119A1 · Koker et al. · 2018 [cited by applicant]
US 20180314307A1 · Wells et al. · 2018 [cited by applicant]
US 20200004319A1 · Mizuno · 2020 [cited by examiner]
US 20200110458A1 · Koker et al. · 2020 [cited by applicant]
US 20210073249A1 · Chang et al. · 2021 [cited by applicant]
US 20220197362A1 · Koker et al. · 2022 [cited by applicant]
US 20230162660A1 · Zou et al. · 2023 [cited by applicant]
Goodfellow, I., et al., “Adaptive Computation and Machine Learning Series,” Chapter 5, Nov. 18, 2016, p. 98-165. [cited by applicant]
Junkins, S., “The Compute Architecture of Intel Processor Graphics Gen9,” Ver. 1, Aug. 14, 2015 (22 pages). [cited by applicant]
Ross, J., et al., “Intel Processor Graphics: Architecture & Programming” PowerPoint presentation, Aug. 2015 (78 pages). [cited by applicant]
Wilt, N., The CUDA Handbook: A Comprehensive Guide to GPU Programming, Pearson Education, 2013, pp. 41-57. [cited by applicant]
Cook, S., “Cuda Programming,” CUDA Hardware Overview, Chapter 3, http://dx.doi.org/10.1016/B978-0-12-415933-4.00003-X, Elsevier, Inc. 2013, pp. 37-52. [cited by applicant]
Final Office Action, U.S. Appl. No. 15/488,662, Nov. 26, 2018, 29 pages. [cited by applicant]
Final Office Action, U.S. Appl. No. 16/595,543, Apr. 16, 2021, 29 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 17/517,090, Sep. 21, 2023, 39 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 15/488,662, Jul. 25, 2018, 19 pages. [cited by applicant]
Non-Final Office Action, U.S. Appl. No. 16/595,543, Oct. 8, 2020, 34 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/488,662, Jun. 5, 2019, 12 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/595,543, Jul. 15, 2021, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 17/517,090, Jan. 30, 2024, 17 pages. [cited by applicant]