IP Library › Granted Patent US 12,265,844
Granted Patent B2
US 12,265,844 · App. 17/468,312 · Granted Apr 1, 2025

Quality of service techniques in distributed graphics processor

Inventors: Benjamin Bowman (London, GB); Fergus W. MacGarry (Cambridge, GB); Kutty Banerjee (Santa Clara, CA); Pratik Chandresh Shah (Santa Clara, CA)
Assignee: Apple Inc.
G06F9/4881
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,844
App. No.
17/468,312
Filed
Sep 7, 2021
Granted
Apr 1, 2025
Kind
B2
Art Unit
2194
USPC
718/103
Abstract

Disclosed techniques relate to circuitry configured to aggregate and report usage information in a distributed processor (e.g., a GPU). In some embodiments, graphics processor circuitry that includes at least first and second portions that are respectively configured to execute sets of graphics work. First utilization circuitry may track execution time for sets of graphics work on the first portion of the graphics processor circuitry and second utilization circuitry may track execution time for sets of graphics work on the second portion of the graphics processor circuitry. Command queue circuitry may store multiple different command queues. Control circuitry may access the first and second utilization circuitry and aggregate utilization data on a per-command-queue basis, where for a given command queue, the aggregated utilization data indicates respective utilization of the first and second portions of the graphics processor circuitry. The control circuitry may provide the aggregated per-command-queue utilization data in software-accessible registers.

Claims (54)

1. An apparatus, comprising:

a graphics processor on an integrated circuit die, wherein the graphics processor includes:

at least first and second portions that respectively include shader pipeline circuitry configured to execute sets of graphics work;

first utilization circuitry configured to track execution time for sets of graphics work on the first portion of the graphics processor;

second utilization circuitry configured to track execution time for sets of graphics work on the second portion of the graphics processor;

command queue circuitry configured to store multiple different command queues, wherein the command queues include entries that store sets of graphics work;

control circuitry in the graphics processor configured to:

access the first and second utilization circuitry and aggregate utilization data on a per-command-queue basis, wherein, for a given command queue, the aggregated utilization data separately indicates utilization of the first portion of the graphics processor and utilization of the second portion of the graphics processor by the given command queue; and

provide the aggregated per-command-queue utilization data in software-accessible registers; and

schedule work, from the different command queues for execution by the first and second portions of the graphics processor, based on software-specified adjustments to one or more scheduler parameters for the command queues generated in response to the aggregated per-command-queue utilization data.

2. The apparatus of claim 1 , wherein to schedule the work, the control circuitry is further configured to:

independently adjust utilization target weights for the first and second portions of the graphics processor, for different types of work, based on differences between historical aggregated per-commend queue utilization data and utilization target information.

3. The apparatus of claim 1 , wherein to schedule the work, the control circuitry is configured to adjust a utilization target for a first command queue based on historical tracking of aggregate utilization data for the first command queue.

4. The apparatus of claim 1 , wherein to schedule the work, the control circuitry is further configured to:

independently adjust, based on the aggregated per-command-queue utilization data, priority values for different types of graphics work processed by the first and second portions of the graphics processor.

5. The apparatus of claim 1 , wherein to schedule the work, the control circuitry is further configured to:

independently adjust, based on the aggregated per-command-queue utilization data, stall thresholds for different types of graphics work processed by the first and second portions of the graphics processor, wherein the stall thresholds indicate a number of cycles that lower priority work is allowed to stall higher priority work before pausing the lower priority work.

6. The apparatus of claim 1 , wherein the apparatus is configured to distribute sets of graphics work according to multiple distribution modes, including to distribute a set of graphics work to only one of the first and second portions in a first distribution mode and to distribute a set of graphics work to both the first and second portions in a second distribution mode.

7. The apparatus of claim 1 , wherein the first and second portions each include arbitration circuitry configured to arbitrate among assigned graphics work based on priority and utilization weight information.

8. The apparatus of claim 1 , wherein the first and second portions respectively include:

distributed control circuitry configured to receive work assignments from the control circuitry and assign received work to the shader pipeline circuitry.

9. The apparatus of claim 1 , wherein the control circuitry is configured to determine a priority value and a stall threshold value for multiple pipelined sets of graphics work, wherein the determination is based on a software specified rule to determine based on an oldest set of graphics work or a highest-priority set of graphics work of the multiple pipelined sets of graphics work.

10. The apparatus of claim 1 , wherein the utilization data indicates a number of processor cycles used of a given portion of the graphics processor.

11. The apparatus of claim 1 , wherein the apparatus further includes:

a central processing unit;

a memory interface; and

a communication fabric configured to transfer information between the graphics processor, the central processing unit, and the memory interface.

12. The apparatus of claim 1 , wherein the apparatus is a mobile device that includes:

a display; and

network interface circuitry.

13. A non-transitory computer-readable medium having instructions stored thereon that are executable by a computing device to perform operations comprising:

accessing aggregated per-command-queue utilization data that indicates, for a graphics processor that includes at least first and second portions that respectively include shader pipeline circuitry configured to execute sets of graphics work, respective utilization of the first and second portions of the graphics processor by respective command queues processed by the graphics processor; and

controlling execution of the instructions based on the aggregated per-command-queue utilization data, including adjusting a scheduler parameter for a first command queue based on the aggregated per-command-queue utilization data, to increase hardware utilization of the first portion or the second portion of the graphics processor by the first command queue.

14. The non-transitory computer-readable medium of claim 13 , wherein the controlling includes:

adjusting priority values for one or more command queues based on the aggregated per-command-queue utilization data.

15. The non-transitory computer-readable medium of claim 13 , wherein the controlling includes:

context switching out a set of graphics work based on the aggregated per-command-queue utilization data.

16. The non-transitory computer-readable medium of claim 13 , wherein the controlling includes:

asserting a retain slots signal that indicates to maintain performance information for a tracking slot implemented by the graphics processor.

17. A method, comprising:

executing, by at least first and second portions of a processor, sets of work, wherein the first and second portions of the processor are included on the same integrated circuit die and respectively implement execution pipelines;

tracking, by the processor, execution time for sets of work on the first portion of the processor;

tracking, by the processor, execution time for sets of work on the second portion of the processor;

storing, by the processor, multiple different command queues, wherein the command queues include entries that store sets of work;

aggregating, by the processor, utilization data on a per-command-queue basis, wherein, for a given command queue, the aggregated utilization data indicates utilization of the first and second portions of the processor;

providing, by the processor, the aggregated per-command-queue utilization data in software-accessible registers;

determining, by software executed by the processor, adjustments to one or more scheduler parameters in response to the aggregated per-command-queue utilization data; and

scheduling work, by control circuitry the processor from the different command queues for execution by the first and second portions of the processor, based on the determined adjustments to one or more scheduler parameters.

18. The method of claim 17 , further comprising:

independently adjusting utilization target weights for the first and second portions of the processor, for different types of work, based on differences between historical aggregated per-commend queue utilization data and utilization target information.

19. The method of claim 17 , further comprising:

adjusting a utilization target for a first command queue based on historical tracking of aggregate utilization data for the first command queue.

20. The method of claim 17 , further comprising:

independently adjusting, based on the aggregated per-command-queue utilization data, stall thresholds for different types of work processed by the first and second portions of the processor, wherein the stall thresholds indicate a number of cycles that lower priority work is allowed to stall higher priority work before pausing the lower priority work.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2021
From: BOWMAN, BENJAMIN; MACGARRY, FERGUS W.; BANERJEE, KUTTY; SHAH, PRATIK CHANDRESH
To: APPLE INC.
Reel/Frame 057402/0924 →
Continuity (1)
Related Publication 20230075531A1 · Mar 9, 2023
References Cited (46)
US 5664200A · Barlow et al. · 1997 [cited by applicant]
US 6560628B1 · Murata · 2003 [cited by applicant]
US 7984447B1 · Markov · 2011 [cited by applicant]
US 10109030B1 · Sun · 2018 [cited by examiner]
US 10262390B1 · Sun · 2019 [cited by examiner]
US 10275851B1 · Zhao · 2019 [cited by examiner]
US 11281498B1 · Kinney, Jr. et al. · 2022 [cited by applicant]
US 20080033696A1 · Aguaviva et al. · 2008 [cited by applicant]
US 20080049031A1 · Liao et al. · 2008 [cited by applicant]
US 20110072435A1 · Yasutake · 2011 [cited by applicant]
US 20120222035A1 · Plondke · 2012 [cited by applicant]
US 20130162661A1 · Bolz et al. · 2013 [cited by applicant]
US 20130275988A1 · Divirgilio et al. · 2013 [cited by applicant]
US 20130305258A1 · Durant · 2013 [cited by applicant]
US 20150347327A1 · Blaine et al. · 2015 [cited by applicant]
US 20160267621A1 · Liao et al. · 2016 [cited by applicant]
US 20160357600A1 · Chimene et al. · 2016 [cited by applicant]
US 20170161099A1 · Rashid et al. · 2017 [cited by applicant]
US 20170220384A1 · Anderson et al. · 2017 [cited by applicant]
US 20180165786A1 · Bourd · 2018 [cited by examiner]
US 20180173560A1 · Avkarogullari et al. · 2018 [cited by applicant]
US 20180188983A1 · Fleming, Jr. · 2018 [cited by examiner]
US 20180188997A1 · Fleming, Jr. · 2018 [cited by examiner]
US 20180307487A1 · Maiyurau et al. · 2018 [cited by applicant]
US 20180308202A1 · Appu et al. · 2018 [cited by applicant]
US 20180314547A1 · Bak et al. · 2018 [cited by applicant]
US 20180341525A1 · Gupta et al. · 2018 [cited by applicant]
US 20180349146A1 · Iwamoto · 2018 [cited by examiner]
US 20180374187A1 · Dong · 2018 [cited by examiner]
US 20190102859A1 · Hux et al. · 2019 [cited by applicant]
US 20190132257A1 · Zhao · 2019 [cited by examiner]
US 20190171489A1 · Guo · 2019 [cited by examiner]
US 20200020384A1 · Zhao et al. · 2020 [cited by applicant]
US 20200104180A1 · Banerjee · 2020 [cited by examiner]
US 20200334084A1 · Jacobson · 2020 [cited by applicant]
US 20200387393A1 · Xu et al. · 2020 [cited by applicant]
US 20210133123A1 · Feehrer et al. · 2021 [cited by applicant]
US 20210247825A1 · Kelemen et al. · 2021 [cited by applicant]
US 20210256754A1 · Guo · 2021 [cited by applicant]
US 20210271606A1 · Hensley et al. · 2021 [cited by applicant]
US 20220050790A1 · Goodman et al. · 2022 [cited by applicant]
US 20230281051A1 · Martin · 2023 [cited by examiner]
Z. Wang, J. Yang, R. Melhem, B. Childers, Y. Zhang and M. Guo, “Quality of service support for fine-grained sharing on GPUs,” 2017 ACM/IEEE 44th Annual International Symposium on Computer Architecture (ISCA), Toronto, O… [cited by examiner]
Office Action in U.S. Appl. No. 17/468,328 mailed Jan. 3, 2024, 23 pages. [cited by applicant]
International Search Report and Written Opinion in PCT Appl. No. PCT/US2022/040099 mailed Nov. 24, 2022, 8 pages. [cited by applicant]
Office Action in CN Appl. No. 202280060239.0 mailed Sep. 12, 2024, 19 pages. [cited by applicant]