IP Library Granted Patent US 12,436,897
Granted Patent B2
US 12,436,897 · App. 18/180,008 · Granted Oct 7, 2025

Cache management using eviction priority based on memory reuse

Inventors: Noam Dor Korem (Tal-El, IL); Brian Scott Pharris (Cary, NC); Jacob Subag (Haifa, IL)
Assignee: NVIDIA Corporation
G06F12/126G06N3/0464
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,436,897
App. No.
18/180,008
Granted
Oct 7, 2025
Kind
B2
Abstract

Apparatuses, systems, and techniques to manage a cache located on a processor of a computing system using eviction priority based on based on memory reuse. Memory addresses associated with a workload of an application executing using the processor are identified. An amount of reuse of the memory addresses corresponding to the workload is determined. A cache management policy for the workload is determined based on the amount of reuse. The cache management policy is applied to the cache.

Claims (37)

1. A method of managing a cache located on a processor, the method comprising:

identifying a plurality of memory addresses associated with a workload of an application executing using the processor;

determining a characteristic of the workload that corresponds to an amount of traffic between the cache and off-chip memory generated at one or more memory addresses of the plurality of memory addresses;

determining an amount of reuse of the plurality of memory addresses using the characteristic;

determining a cache management policy for the workload based on the amount of reuse; and

applying the cache management policy to the cache.

2. The method of claim 1 , wherein the application is a neural network (NN) inference application, and the plurality of memory addresses are a plurality of virtual addresses.

3. The method of claim 2 , wherein one or more virtual addresses of the plurality of virtual addresses are associated with activation data of the workload of the NN inference application.

4. The method of claim 3 , wherein the determining the amount of reuse includes computing the amount of reuse according to a quantitative metric, the quantitative metric comprising a number of layers associated with the NN inference application that access the one or more virtual addresses associated with the activation data of the workload.

5. The method of claim 4 , wherein applying the cache management policy to the cache comprises applying one or more eviction priority controls to one or more memory addresses of the plurality of memory addresses based on the quantitative metric.

6. The method of claim 5 , wherein applying the one or more eviction priority controls to the one or more memory addresses of the plurality of memory addresses comprises designating at least one of an evict-last eviction priority control, an evict-normal eviction priority control, or an evict-first eviction priority control for at least one memory address of the one or more memory addresses.

7. The method of claim 2 , wherein one or more virtual addresses of the plurality of virtual addresses are associated with weight data of the workload of the NN inference application.

8. A system comprising:

a processor, having a cache located thereon, to perform operations comprising:

identifying a plurality of memory addresses associated with a workload of an application executing using the processor;

determining a characteristic of the workload that corresponds to an amount of traffic between the cache and off-chip memory generated at one or more memory addresses of the plurality of memory addresses;

determining an amount of reuse of the memory addresses using the characteristic;

determining a cache management policy for the workload based on the amount of reuse; and

applying the cache management policy to the cache.

9. The system of claim 8 , wherein the application is a neural network (NN) inference application, and the plurality of memory addresses are a plurality of virtual addresses.

10. The system of claim 9 , wherein one or more virtual addresses of the plurality of virtual addresses are associated with activation data of the workload of the NN inference application.

11. The system of claim 10 , wherein the determining the amount of reuse includes computing the amount of reuse according to a quantitative metric, the quantitative metric comprising a number of layers associated with the NN inference application that access the one or more virtual addresses associated with the activation data of the workload.

12. The system of claim 11 , wherein applying the cache management policy for the workload based on the quantitative metric comprises applying one or more eviction priority controls to one or more memory addresses of the plurality of memory addresses based on the quantitative metric.

13. The system of claim 12 , wherein applying the eviction priority controls to the one or more memory addresses of the plurality of memory addresses comprises designating at least one of an evict-last eviction priority control, an evict-normal eviction priority control, or an evict first eviction priority control for at least one memory address of the one or more memory addresses.

14. The system of claim 9 , wherein one or more virtual addresses of the plurality of virtual addresses are associated with weight data of the workload of the NN inference application.

15. A processor comprising:

one or more processing units to:

identify a plurality of memory addresses associated with a workload of an application;

determine a characteristic of the workload that corresponds to an amount of traffic between a cache corresponding to the processor and off-chip memory generated at one or more memory addresses of the plurality of memory addresses;

determine an amount of reuse of the plurality of memory addresses using the characteristic;

determine a cache management policy to apply to the cache for the workload based at least on the amount of reuse; and

apply the cache management policy to the cache.

16. The processor of claim 15 , wherein the application is a neural network (NN) inference application, and the plurality of memory addresses are a plurality of virtual addresses.

17. The processor of claim 16 , wherein one or more virtual addresses of the plurality of virtual addresses are associated with activation data of the workload of the NN inference application.

18. The processor of claim 17 , wherein the one or more processing units are to determine the amount of reuse according to a quantitative metric, the quantitative metric comprising a number of layers associated with the NN inference application that access the one or more virtual addresses associated with the activation data of the workload.

19. The processor of claim 15 , wherein the one or more processing units are to apply the cache management policy to the cache by applying one or more eviction priority controls to at least one memory address of the one or more memory addresses of the plurality of memory addresses based on the amount of reuse of the plurality of memory addresses corresponding to the workload.

20. The processor of claim 19 , wherein the applying the one or more eviction priority controls to the at least one memory address of the one or more memory addresses comprises designating at least one of an evict-last eviction priority control, an evict-normal eviction priority control, or an evict-first eviction priority control for at least one memory address of the one or more memory addresses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2023
From: KOREM, NOAM DOR; PHARRIS, BRIAN SCOTT; SUBAG, JACOB
To: NVIDIA CORPORATION
Reel/Frame 062912/0121 →
Continuity (1)
Related Publication 20240303203A1 · Sep 12, 2024
References Cited (19)
US 10019368B2 · Hagersten · 2018 [cited by examiner]
US 11182207B2 · Hirota et al. · 2021 [cited by applicant]
US 20160062916A1 · Das · 2016 [cited by examiner]
US 20170357600A1 · Moon · 2017 [cited by examiner]
US 20180349292A1 · Tal · 2018 [cited by examiner]
US 20190102302A1 · Taht · 2019 [cited by examiner]
US 20220121985A1 · Lloyd · 2022 [cited by examiner]
US 20220398038A1 · Anchi · 2022 [cited by examiner]
US 20230102767A1 · Emberling · 2023 [cited by examiner]
US 20230137205A1 · Fu · 2023 [cited by examiner]
US 20230144662A1 · Tasinga · 2023 [cited by examiner]
US 20230418765A1 · Tune · 2023 [cited by examiner]
US 20240355044A1 · Emberling · 2024 [cited by examiner]
WO WO2023055532A1 · 2023 [cited by examiner]
L. Liu, Y. Li, Z. Cui, Y. Bao, M. Chen and C. Wu, “Going vertical in memory management: Handling multiplicity by multi-policy,” 2014 ACM/IEEE 41st International Symposium on Computer Architecture (ISCA), Minneapolis, MN… [cited by examiner]
A. Arunkumar, S. -y. Lee and C.-j. Wu, “ID-cache: instruction and memory divergence based cache management for GPUs,” 2016 IEEE International Symposium on Workload Characterization (IISWC), Providence, RI, USA, 2016, pp… [cited by examiner]
A. Pan and V. S. Pai, “Runtime-driven shared last-level cache management for task-parallel programs,” SC '15: Proceedings of the International Conference for High Performance Computing, Networking, Storage and Analysis,… [cited by examiner]
Y. Yu, C. Zhang, W. Wang, J. Zhang and K. B. Letaief, “Towards Dependency-Aware Cache Management for Data Analytics Applications,” in IEEE Transactions on Cloud Computing, vol. 10, No. 1, pp. 706-723, Jan. 1-Mar. 2022. [cited by examiner]
I. jibaja and K. A. Shaw, “Understanding the applicability of CMP performance optimizations on data mining applications,” 2009 IEEE International Symposium on Workload Characterization (IISWC), Austin, TX, USA, 2009, pp… [cited by examiner]