IP Library › Granted Patent US 11,915,000
Granted Patent B2
US 11,915,000 · App. 18/160,600 · Granted Feb 27, 2024

Apparatuses, methods, and systems to precisely monitor memory store accesses

Inventors: Ahmad Yasin (Haifa, IL); Raanan Sade (Kibutz Sarid, IL); Liron Zur (Haifa, IL); Igor Yanover (Yokneam Illit, IL); Joseph Nuzman (Haifa, IL)
Assignee: Intel Corporation
G06F9/30145G06F9/30098G06F9/544G06F9/546G06F11/3037G06F11/348
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,915,000
App. No.
18/160,600
Granted
Feb 27, 2024
Kind
B2
Abstract

Systems, methods, and apparatuses relating to circuitry to precisely monitor memory store accesses are described. In one embodiment, a system includes a memory, a hardware processor core comprising a decoder to decode an instruction into a decoded instruction, an execution circuit to execute the decoded instruction to produce a resultant, a store buffer, and a retirement circuit to retire the instruction when a store request for the resultant from the execution circuit is queued into the store buffer for storage into the memory, and a performance monitoring circuit to mark the retired instruction for monitoring of post-retirement performance information between being queued in the store buffer and being stored in the memory, enable a store fence after the retired instruction to be inserted that causes previous store requests to complete within the memory, and on detection of completion of the store request for the instruction in the memory, store the post-retirement performance information in storage of the performance monitoring circuit.

Claims (49)

1. An apparatus comprising:

a decoder to decode an instruction into a decoded instruction;

an execution circuit to execute the decoded instruction to produce a resultant;

a first level data cache;

a store buffer;

a retirement circuit to retire the instruction when a store request for the resultant is queued into the store buffer for storage into a memory but is not yet completed to the memory; and

a performance monitoring circuit to:

monitor post-retirement performance information of the retired instruction between the store request being accepted into the first level data cache and being completed to the memory, and

store the post-retirement performance information in storage of the performance monitoring circuit.

2. The apparatus of claim 1 , wherein the performance monitoring circuit is to monitor and store in response to the performance monitoring circuit being in precise event-based sampling mode.

3. The apparatus of claim 1 , wherein, when the store request is accepted into the first level data cache, the performance monitoring circuit is to enable a counter to measure a latency between the store request being accepted into the first level data cache and being completed in the memory, and the post-retirement performance information comprises the latency from the counter.

4. The apparatus of claim 1 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

5. The apparatus of claim 1 , wherein the memory is another level of cache.

6. The apparatus of claim 1 , wherein the memory is a system memory separate from any cache of the apparatus.

7. The apparatus of claim 1 , wherein the translation lookaside buffer is a second level translation lookaside buffer.

8. The apparatus of claim 7 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

9. The apparatus of claim 1 , wherein the post-retirement performance information comprises a linear address of a destination in the memory of the store request.

10. A method comprising:

decoding an instruction into a decoded instruction with a decoder of a hardware processor comprising a first level data cache and a store buffer;

executing the decoded instruction with an execution circuit of the hardware processor to produce a resultant;

retiring the instruction with a retirement circuit of the hardware processor when a store request for the resultant is queued into the store buffer for storage into a memory but is not yet completed to the memory;

monitoring, by a performance monitoring circuit of the hardware processor, post-retirement performance information of the retired instruction between the store request being accepted into the first level data cache and being completed to the memory, wherein the post-retirement performance information comprises a first value to indicate the store request missed in the first level data cache and a second value to indicate the store request missed in a translation lookaside buffer; and

storing, by the performance monitoring circuit, the post-retirement performance information in storage of the performance monitoring circuit.

11. The method of claim 10 , wherein the monitoring and the storing occur in response to the performance monitoring circuit being in precise event-based sampling mode.

12. The method of claim 10 , further comprising, when the store request is accepted into the first level data cache, enabling, by the performance monitoring circuit, a counter to measure a latency between the store request being accepted into the first level data cache and being completed in the memory, wherein the post-retirement performance information comprises the latency from the counter.

13. The method of claim 10 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

14. The method of claim 10 , wherein the memory is another level of cache.

15. The method of claim 10 , wherein the memory is a system memory separate from any cache of the hardware processor.

16. The method of claim 10 , wherein the translation lookaside buffer is a second level translation lookaside buffer.

17. The method of claim 16 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

18. The method of claim 10 , wherein the post-retirement performance information comprises a linear address of a destination in the memory of the store request.

19. A system comprising:

a memory;

a decoder to decode an instruction into a decoded instruction;

an execution circuit to execute the decoded instruction to produce a resultant;

a first level data cache;

a store buffer;

a retirement circuit to retire the instruction when a store request for the resultant is queued into the store buffer for storage into the memory but is not yet completed to the memory; and

a performance monitoring circuit to:

monitor post-retirement performance information of the retired instruction between the store request being accepted into the first level data cache and being completed to the memory, wherein the post-retirement performance information comprises a first value to indicate the store request missed in the first level data cache and a second value to indicate the store request missed in a translation lookaside buffer, and

store the post-retirement performance information in storage of the performance monitoring circuit.

20. The system of claim 19 , wherein the performance monitoring circuit is to monitor and store in response to the performance monitoring circuit being in precise event-based sampling mode.

21. The system of claim 19 , wherein, when the store request is accepted into the first level data cache, the performance monitoring circuit is to enable a counter to measure a latency between the store request being accepted into the first level data cache and being completed in the memory, and the post-retirement performance information comprises the latency from the counter.

22. The system of claim 19 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

23. The system of claim 19 , wherein the memory is another level of cache.

24. The system of claim 19 , wherein the memory is a system memory separate from any cache.

25. The system of claim 19 , wherein the translation lookaside buffer is a second level translation lookaside buffer.

26. The system of claim 25 , wherein the post-retirement performance information comprises a value to indicate the store request is a locked access.

27. The system of claim 19 , wherein the post-retirement performance information comprises a linear address of a destination in the memory of the store request.

Continuity (3)
Continuation 17862708 · Jul 12, 2022
Continuation 16729374 · Dec 28, 2019
Related Publication 20230176870A1 · Jun 8, 2023
Cited By (1)
US 12,271,735