IP Library › Granted Patent US 12,216,932
Granted Patent B2
US 12,216,932 · App. 18/327,474 · Granted Feb 4, 2025

Precise longitudinal monitoring of memory operations

Inventors: Ahmad Yasin (Haifa, IL); Michael Chynoweth (Placitas, NM); Rajshree Chabukswar (Sunnyvale, CA); Muhammad Taher (Umm El Fahm, IL)
Assignee: Intel Corporation
G06F3/0656G06F3/0604G06F3/0653G06F3/0673G06F11/3466
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,216,932
App. No.
18/327,474
Granted
Feb 4, 2025
Kind
B2
Abstract

A processor includes a memory subunit that includes a status register and an execution engine unit to: randomly select a load operation to monitor; determine a re-order buffer identifier of the load operation; and transmit the re-order buffer identifier to the memory subsystem. Responsive to receipt of the re-order buffer identifier, the first memory subunit is to store a piece of information, related to a status of the load operation, in the status register. The processor also includes logic to, responsive to detection of retirement of the load operation, store memory information in memory-related fields of a record of a memory buffer. The memory information includes auxiliary information (AUX) and access latency information, wherein one of the auxiliary information or the access latency information includes the piece of information, from the status register, stored in a particular field of the memory-related fields.

Claims (20)

1. An apparatus comprising:

a memory subsystem including a data cache, the memory subsystem to process a load operation; and

load latency hardware including a latency counter, the load latency hardware to record performance monitoring information including an access address associated with the load operation, a data block indicator to indicate whether the load operation was blocked since its data could not be forwarded from a preceding store, and an address block indicator to indicate whether the load operation was blocked due to a potential address conflict with a preceding store.

2. The apparatus of claim 1 , wherein the performance monitoring information also includes an instruction access latency associated with the load operation.

3. The apparatus of claim 1 , wherein the performance monitoring information also includes a cache access latency associated with the load operation.

4. The apparatus of claim 1 , wherein the load operation is to be randomly selected.

5. A method comprising:

processing, by a memory subsystem of a hardware processor, the memory subsystem including a data cache, a load operation;

recording, by load latency hardware of the processor, the load latency hardware including a latency counter, performance monitoring information including an access address associated with the load operation, a data block indicator to indicate whether the load operation was blocked since its data could not be forwarded from a preceding store, and an address block indicator to indicate whether the load operation was blocked due to a potential address conflict with a preceding store.

6. The method of claim 5 , wherein the performance monitoring information also includes an instruction access latency associated with the load operation.

7. The method of claim 5 , wherein the performance monitoring information also includes a cache access latency associated with the load operation.

8. The method of claim 5 , further comprising randomly selecting the load operation.

9. A system comprising:

a memory; and

a processor including:

a memory subsystem including a data cache, the memory subsystem to process a load operation; and

load latency hardware including a latency counter, the load latency hardware to record performance monitoring information including an access address associated with the load operation, a data block indicator to indicate whether the load operation was blocked since its data could not be forwarded from a preceding store, and an address block indicator to indicate whether the load operation was blocked due to a potential address conflict with a preceding store.

10. The system of claim 9 , wherein the performance monitoring information also includes an instruction access latency associated with the load operation.

11. The system of claim 9 , wherein the performance monitoring information also includes a cache access latency associated with the load operation.

12. The system of claim 9 , wherein the load operation is to be randomly selected.

Continuity (3)
Continuation 15929272 · Apr 21, 2020
Continuation 16177642 · Nov 1, 2018
Related Publication 20230305742A1 · Sep 28, 2023
References Cited (35)
US 5918005A · Moreno et al. · 1999 [cited by applicant]
US 6216215B1 · Palanca et al. · 2001 [cited by applicant]
US 7181723B2 · Luk et al. · 2007 [cited by applicant]
US 7653727B2 · Durham et al. · 2010 [cited by applicant]
US 8037465B2 · Tian et al. · 2011 [cited by applicant]
US 8782629B2 · Sun · 2014 [cited by applicant]
US 8819699B2 · Cota-Robles et al. · 2014 [cited by applicant]
US 9069690B2 · Hildesheim et al. · 2015 [cited by applicant]
US 9268611B2 · Yer et al. · 2016 [cited by applicant]
US 9304811B2 · Yao · 2016 [cited by applicant]
US 9558006B2 · Sasanka · 2017 [cited by applicant]
US 10338834B1 · Dighe · 2019 [cited by examiner]
US 20070226447A1 · Shimozono · 2007 [cited by examiner]
US 20080071939A1 · Tanaka · 2008 [cited by examiner]
US 20080141002A1 · Bhargava · 2008 [cited by examiner]
US 20090019317A1 · Quach et al. · 2009 [cited by applicant]
US 20090077350A1 · Saraswati · 2009 [cited by examiner]
US 20140075123A1 · Hildesheim et al. · 2014 [cited by applicant]
US 20140189302A1 · Subbareddy et al. · 2014 [cited by applicant]
US 20150046506A1 · Park et al. · 2015 [cited by applicant]
US 20150212822A1 · Hooker et al. · 2015 [cited by applicant]
US 20160179541A1 · Gramunt et al. · 2016 [cited by applicant]
US 20160232103A1 · Schmisseur et al. · 2016 [cited by applicant]
US 20170093669A1 · Nortman · 2017 [cited by examiner]
US 20180285003A1 · Richardson · 2018 [cited by examiner]
US 20190171515A1 · Sperber et al. · 2019 [cited by applicant]
Eranian, Stéphane. “What can performance counters do for memory subsystem analysis?.” Proceedings of the 2008 ACM SIGPLAN workshop on Memory systems performance and correctness. (ASPLOS'08). 2008. (Year: 2008). [cited by examiner]
Ayers, Grant, et al., “Memory Hierarchy for Web Search,” In High Performace Computer Architecture (HPCA), IEEE International Symposium on High Performance Computer Architecture (HPCA), Vienna, pp. 643-656, 2018. [cited by applicant]
Dean, Jeffrey, et al., “ProfileMe: Hardware Support for Instruction-Level Profiling on Out-of-Order Processors,” MICRO, Dec. 1997, 12 pages. [cited by applicant]
Gwennap, Linley, “Server Processor Competition Heats Up”, Microprocessor Report, Dec. 18, 2017, pp. 1-4. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/929,272, Apr. 29, 2022, 10 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 15/929,272, Mar. 14, 2023, 9 pages. [cited by applicant]
Notice of Allowance, U.S. Appl. No. 16/177,642, Jan. 10, 2020, 9 pages. [cited by applicant]
Shimpi, Anand Lal, “AMD's Phenom II X4 965 Black Edition” AnandTech, http://www.anandtech.com/show/2819/6, 9 pages, Aug. 13, 2009. [cited by applicant]
Yasin, Ahmad, “A Top-Down Method for Performance Analysis and Counters Architecture,” International Symposium for Performance Analysis of System and Software (ISPASS), v1.02, Mar. 2014, 10 pages. [cited by applicant]