IP Library Granted Patent US 12,373,329
Granted Patent B2
US 12,373,329 · App. 18/064,225 · Granted Jul 29, 2025

Deterministic replay of a multi-threaded trace on a multi-threaded processor

Inventor: Konstantin Levit-Gurevich (Kiryat Byalik, IL)
Assignee: INTEL CORPORATION
G06F11/3636G06F9/3851G06F9/3887G06F9/3888G06T15/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,373,329
App. No.
18/064,225
Granted
Jul 29, 2025
Kind
B2
Abstract

At least one computer-readable storage medium comprising instructions for execution by at least one graphics processing unit (GPU) that, when executed, cause the at least one GPU to: obtain program code for tracing, the program code including a plurality of instructions; identify from the plurality of instructions of the program code events to be synchronized; instrument the program code corresponding to one or more of the events identified, by inserting instructions that support monitoring code; execute the instrumented program code on at least a plurality of hardware threads of the GPU and generate trace data; replay the identified events according to an order of occurrence of the events identified; and report a GPU state indicating a utilization of the GPU based; and wherein to report the GPU state includes to indicate when the GPU executes non-graphics related tasks.

Claims (44)

1. At least one non-transitory computer-readable storage medium comprising instructions for execution by at least one graphics processing unit (GPU) that, when executed, cause the at least one GPU to:

obtain program code for tracing, the program code including a plurality of instructions;

identify, from the plurality of instructions of the program code, events to be synchronized;

instrument the program code corresponding to one or more of the events identified, by inserting instructions that support monitoring code;

execute the instrumented program code on at least a plurality of hardware threads of the GPU and generate trace data;

replay the identified events according to an order of occurrence of the events identified; and

report a GPU state indicating a utilization of the GPU; and

wherein to report the GPU state includes to indicate when the GPU executes non-graphics related tasks.

2. The at least one non-transitory computer-readable storage medium according to claim 1 , further comprising instructions for execution by the at least one GPU that, when executed, cause the at least one GPU to:

report runtime shader errors.

3. The at least one non-transitory computer-readable storage medium of claim 1 , wherein the events include one or more of a code dispatch, a code end-of-thread event, a read or write access to global memory, a read or write access to local memory, a waiting state, or a memory fence instruction.

4. The at least one non-transitory computer-readable storage medium of claim 1 , wherein the program code is a kernel or a shader.

5. The at least one non-transitory computer-readable storage medium of claim 1 , wherein instrumenting the program code includes:

dividing the program code into a sequence of basic blocks; and

inserting a trace instruction into each basic block of the sequence of basic blocks that contains an event.

6. The at least non-transitory one computer-readable storage medium of claim 5 , wherein instrumenting the program code further includes:

inserting a dynamic instruction count relating to an original instruction in each basic block where a tracing instruction is added.

7. A method comprising:

obtaining program code for tracing, the program code including a plurality of instructions;

identifying, from the plurality of instructions of the program code, events to be synchronized;

instrumenting the program code corresponding to one or more of the events identified, by inserting instructions that support monitoring code;

executing the instrumented program code on at least a plurality of hardware threads of a graphics processing unit (GPU) and generating trace data;

replaying the identified events according to an order of occurrence of the events identified; and

reporting a GPU state indicating a utilization of the GPU; and

wherein to report the GPU state includes to indicate when the GPU executes non-graphics related tasks.

8. The method of claim 7 , further comprising:

reporting runtime shader errors.

9. The method of claim 7 , wherein the events include one or more of a code dispatch, a code end-of-thread event, a read or write access to global memory, a read or write access to local memory, a waiting state, or a memory fence instruction.

10. The method of claim 7 , wherein the program code is a kernel or a shader.

11. A system comprising:

one or more processors including at least one graphics processing unit (GPU); and

a memory to store data including instructions; and

wherein the instructions includes instructions to cause the at least one GPU to perform operations including the following:

obtaining program code for tracing, the program code including a plurality of instructions;

identifying, from the plurality of instructions of the program code, events to be synchronized;

instrumenting the program code corresponding to one or more of the events identified, by inserting instructions that support monitoring code;

executing the instrumented program code on at least a plurality of hardware threads of the GPU and generate trace data;

replaying the identified events according to an order of occurrence of the events identified; and

reporting a GPU state indicating a utilization of the GPU; and

wherein to report the GPU state includes to indicate when the GPU executes non-graphics related tasks.

12. The system of claim 11 , wherein the instructions includes instructions to cause the at least one GPU to perform operations including:

reporting runtime shader errors.

13. The system of claim 11 , wherein the events include one or more of a code dispatch, a code end-of-thread event, a read or write access to global memory, a read or write access to local memory, a waiting state, or a memory fence instruction.

14. The system of claim 11 , wherein the program code is a kernel or a shader.

Continuity (3)
Continuation In Part 17547765 · Dec 10, 2021
Continuation In Part 17111136 · Dec 3, 2020
Related Publication 20230109752A1 · Apr 13, 2023
References Cited (20)
US 9898385B1 · O'Dowd et al. · 2018 [cited by applicant]
US 10963367B2 · Mola · 2021 [cited by applicant]
US 11281562B2 · Fahim et al. · 2022 [cited by applicant]
US 20160179714A1 · Acharya · 2016 [cited by examiner]
US 20170337145A1 · Rozas et al. · 2017 [cited by applicant]
US 20180349119A1 · Zaidi · 2018 [cited by examiner]
US 20190102860A1 · Koker · 2019 [cited by examiner]
US 20200034276A1 · O'Dowd et al. · 2020 [cited by applicant]
US 20200210315A1 · Fahim et al. · 2020 [cited by applicant]
US 20210089429A1 · Mola · 2021 [cited by applicant]
US 20210117202A1 · Levit-Gurevich · 2021 [cited by applicant]
US 20220100512A1 · Levit-Gurevich et al. · 2022 [cited by applicant]
CN 117546139A · 2024 [cited by applicant]
EP 4445253A1 · 2024 [cited by applicant]
IN 202347085667 · 2024 [cited by applicant]
WO 2023107789A1 · 2023 [cited by applicant]
Cheng-Kung, et al., “Fast profiling framework and race detection for heterogeneous system”, Journal of Systems Architecture, vol. 81, pp. 83-91, Nov. 2017. [cited by applicant]
PCT International Search Report International Application No. PCT/US2022/079165, mailed Mar. 13, 2023, 4 pages. [cited by applicant]
Written Opinion of the International Searching Authority, International Application No. PCT/US2022/079165, mailed Mar. 13, 2023, 4 pages. [cited by applicant]
Non-Final Office Action U.S. Appl. No. 17/547,765 mailed Jan. 17, 2025, 5 pages. [cited by applicant]