IP Library › Granted Patent US 9,405,575
Granted Patent B2
US 9,405,575 · App. 14/021,895 · Granted Aug 2, 2016

Use of multi-thread hardware for efficient sampling

Inventors: Zachary Burka (San Francisco, CA); Serge Metral (San Jose, CA)
Assignee: Apple Inc.
G06F9/48G06F9/485G06F9/4806G06F9/4843G06F9/50G06F9/5005G06F9/5011G06F9/5022G06F9/5027G06F11/3003G06F11/3024G06F11/3096G06F11/3466G06F11/3409G06F2201/865G06F2201/88
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,405,575
App. No.
14/021,895
Granted
Aug 2, 2016
Kind
B2
Abstract

This disclosure pertains to systems, methods, and computer readable media for utilizing an unused hardware thread of a multi-core microcontroller of a graphical processing unit (GPU) to gather sampling data of commands being executed by the GPU. The multi-core microcontroller may include two or more hardware threads and may be responsible for managing the scheduling of commands on the GPU. In one embodiment, the firmware code of the multi-core microcontroller which is responsible for running the GPU may run entirely on one hardware thread of the microcontroller, while the second hardware thread is kept in a dormant state. This second hardware thread may be used for gathering sampling data of the commands run on the GPU. The sampling data can be used to assist developers identify bottlenecks and to help them optimize their software programs.

Claims (34)

1. A non-transitory program storage device, readable by a processing unit and comprising instructions stored thereon to cause one or more processing units to:

process commands received for a graphics processing unit (GPU) by a first hardware thread of a multicore microcontroller of the GPU;

enable a second hardware thread of the multicore microcontroller to obtain performance analysis data of the GPU while the first hardware thread processes the commands; and

disable the second hardware thread from obtaining the performance analysis data while the first hardware thread continues to processes the commands by having the first hardware thread instruct the second hardware thread to stop obtaining the performance analysis data.

2. The non-transitory program storage device of claim 1 , wherein the first and the second hardware threads are duplicates of each other, with each hardware thread having a separate execution unit.

3. The non-transitory program storage device of claim 1 , wherein the instructions are stored in a firmware of the processing unit.

4. The non-transitory program storage device of claim 1 , wherein the instructions to cause the one or more processing units to disable the second hardware thread from obtaining the performance analysis data comprise instructions by the first hardware thread to clean up one or more resources used by the second hardware thread.

5. The non-transitory program storage device of claim 4 , wherein the instructions to cause the one or more processing units to disable the second hardware thread from obtaining the performance analysis data comprise instructions by the first hardware thread to turn the second hardware thread to a dormant state.

6. The non-transitory program storage device of claim 1 , wherein obtaining the performance analysis data by the second hardware thread does not impact the processing of commands by the first hardware thread.

7. A method, comprising:

processing commands received for a graphics processing unit (GPU), the commands being processed by a first hardware thread of a multicore microcontroller of the GPU;

receiving a command to start obtaining performance analysis data of the performance of the GPU;

enabling a second hardware thread of the multicore microcontroller to start obtaining the performance analysis data;

receiving a command to stop obtaining the performance analysis data; and

stopping the second hardware thread of the multicore microcontroller from obtaining the performance analysis data by receiving a command from the first hardware thread for the second hardware thread to stop obtaining the performance analysis data,

wherein the first hardware thread of the multicore microcontroller continues processing commands while the second hardware thread is obtaining performance analysis data.

8. The method of claim 7 , wherein obtaining the performance analysis data by the second hardware thread does not impact the processing of commands by the first hardware thread.

9. The method of claim 7 , wherein processing commands by the first hardware thread comprises scheduling the commands for execution by the GPU.

10. The method of claim 7 , further comprising receiving a command to enable the second hardware thread of the multicore microcontroller to obtain performance analysis data, wherein the command comprises of a sampling bit set to a specific state.

11. A device, comprising:

a processing device having integrated firmware; and

a processor embedded in the processing device which is configured to execute program code stored in the firmware to:

process commands received for a graphics processing unit (GPU), the commands being processed by a first hardware thread of a multicore microcontroller of the GPU;

receive a command to start obtaining performance analysis data of the performance of the GPU;

enable a second hardware thread of the multicore microcontroller to start obtaining the performance analysis data;

receive a command to stop obtaining the performance analysis data; and

stop the second hardware thread of the multicore microcontroller from obtaining the performance analysis data by receiving a command from the first hardware thread for the second hardware thread to stop obtaining the performance analysis data,

wherein the first hardware thread of the multicore microcontroller continues processing commands while the second hardware thread is obtaining performance analysis data.

12. The device of claim 11 , wherein obtaining the performance analysis data by the second hardware thread does not impact the processing of commands by the first hardware thread.

13. The non-transitory program storage device of claim 1 , wherein the commands are associated with a device application, and wherein the instructions to cause the one or more processing units to process commands received for a GPU by a first hardware thread of a multicore microcontroller of the GPU further comprise instructions by the first hardware thread to schedule the commands associated with a device application for execution by the GPU.

14. The method of claim 7 , wherein stopping the second hardware thread of the multicore microcontroller from obtaining the performance analysis data further comprises receiving a command from the first hardware thread to clean up one or more resources used by the second hardware thread.

15. The method of claim 7 , wherein stopping the second hardware thread of the multicore microcontroller from obtaining the performance analysis data further comprises returning the second hardware thread to a dormant state.

16. The device of claim 11 , wherein the program code, when executed, further causes the processor to stop the second hardware thread of the multicore microcontroller from obtaining the performance analysis data by at least receiving a command from the first hardware thread to clean up one or more resources used by the second hardware thread.

17. The device of claim 11 , wherein the program code, when executed, further causes the processor to stop the second hardware thread of the multicore microcontroller from obtaining the performance analysis data by at least returning the second hardware thread to a dormant state.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2013
From: BURKA, ZACHARY; METRAL, SERGE
To: APPLE INC.
Reel/Frame 031243/0312 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 9, 2013
From: BURKA, ZACHARY; METRAL, SERGE
To: APPLE INC.
Reel/Frame 031168/0263 →
Continuity (1)
Related Publication 20150074668A1 · Mar 12, 2015