IP Library › Granted Patent US 12,360,910
Granted Patent B1
US 12,360,910 · App. 18/435,557 · Granted Jul 15, 2025

Method, device, and computer program product for caching

Inventors: Qiang Chen (Shanghai, CN); Pedro Fernandez Orellana (Shanghai, CN)
Assignee: Dell Products L.P.
G06F12/0886
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,360,910
App. No.
18/435,557
Granted
Jul 15, 2025
Kind
B1
Abstract

Embodiments of the present disclosure relate to a method, a device, and a computer program product for caching. The method includes determining whether a length of a to-be-processed token sequence is greater than a cache length threshold associated with a key-value (KV) data cache. The method further includes determining, in response to the length of the to-be-processed token sequence being greater than the cache length threshold, target KV data based on a token sequence exceeding the cache length threshold in the to-be-processed token sequence. In addition, the method further includes determining a target token based on the to-be-processed token sequence, the KV data cache, and the target KV data.

Claims (51)

1. A method for caching, comprising:

determining whether a length of a to-be-processed token sequence is greater than a cache length threshold associated with a key-value (KV) data cache, wherein the KV data cache is related to a processed token sequence;

determining, in response to the length of the to-be-processed token sequence being greater than the cache length threshold, target KV data based on a token sequence exceeding the cache length threshold in the to-be-processed token sequence; and

determining a target token based on the to-be-processed token sequence, the KV data cache, and the target KV data, wherein the target token is a next token of the to-be-processed token sequence.

2. The method according to claim 1 , further comprising:

determining, in response to the length of the to-be-processed token sequence being smaller than or equal to the cache length threshold, the target KV data based on a last token in the to-be-processed token sequence; and

determining the target token based on the to-be-processed token sequence, the KV data cache, and the target KV data.

3. The method according to claim 2 , further comprising:

adding, in response to the length of the KV data cache being smaller than the cache length threshold, the target KV data into the KV data cache.

4. The method according to claim 3 , further comprising:

generating an updated to-be-processed token sequence by adding the target token into the to-be-processed token sequence.

5. The method according to claim 1 , wherein the KV data cache is cached on a graphics processing unit, and the cache length threshold is determined based on a memory of the graphics processing unit.

6. The method according to claim 5 , wherein the cache length threshold is determined by:

determining an available memory space of the graphics processing unit; and

determining the cache length threshold based on the available memory space and occupied memory of the target KV data.

7. The method according to claim 6 , further comprising:

allocating a cache space for the KV data cache based on the cache length threshold.

8. The method according to claim 1 , wherein the KV data cache comprises KV data associated with a token of prompt information.

9. The method according to claim 8 , further comprising:

determining the to-be-processed token sequence and the target token based on the prompt information, wherein the prompt information is text information; and

determining output text associated with the prompt information based on the to-be-processed token sequence and the target token.

10. The method according to claim 9 , wherein determining the output text associated with the prompt information comprises:

determining an updated to-be-processed token sequence and a next target token based on the to-be-processed token sequence and the target token; and

determining the output text associated with the prompt information based on the updated to-be-processed token sequence and the next target token.

11. An electronic device, comprising:

at least one processing unit; and

a memory, coupled to the at least one processing unit and storing instructions, wherein the instructions, when executed by at least one the processing unit, cause the electronic device to perform actions comprising:

determining whether a length of a to-be-processed token sequence is greater than a cache length threshold associated with a key-value (KV) data cache, wherein the KV data cache is related to a processed token sequence;

determining, in response to the length of the to-be-processed token sequence being greater than the cache length threshold, target KV data based on a token sequence exceeding the cache length threshold in the to-be-processed token sequence; and

determining a target token based on the to-be-processed token sequence, the KV data cache, and the target KV data, wherein the target token is a next token of the to-be-processed token sequence.

12. The electronic device according to claim 11 , wherein the actions further comprise:

determining, in response to the length of the to-be-processed token sequence being smaller than or equal to the cache length threshold, the target KV data based on a last token in the to-be-processed token sequence; and

determining the target token based on the to-be-processed token sequence, the KV data cache, and the target KV data.

13. The electronic device according to claim 12 , wherein the actions further comprise:

adding, in response to the length of the KV data cache being smaller than the cache length threshold, the target KV data into the KV data cache.

14. The electronic device according to claim 13 , wherein the actions further comprise:

generating an updated to-be-processed token sequence by adding the target token into the to-be-processed token sequence.

15. The electronic device according to claim 11 , wherein the KV data cache is cached on a graphics processing unit, and the cache length threshold is determined based on a memory of the graphics processing unit.

16. The electronic device according to claim 15 , wherein the cache length threshold is determined by:

determining an available memory space of the graphics processing unit; and

determining the cache length threshold based on the available memory space and occupied memory of the target KV data.

17. The electronic device according to claim 16 , wherein the actions further comprise:

allocating a cache space for the KV data cache based on the cache length threshold.

18. The electronic device according to claim 11 , wherein the KV data cache comprises KV data associated with a token of prompt information.

19. The electronic device according to claim 18 , wherein the actions further comprise:

determining the to-be-processed token sequence and the target token based on the prompt information, wherein the prompt information is text information; and

determining output text associated with the prompt information based on the to-be-processed token sequence and the target token.

20. A computer program product, wherein the computer program product is tangibly stored on a non-transitory computer-readable medium and comprises machine-executable instructions, and the machine-executable instructions, when executed by a machine, cause the machine to perform actions comprising:

determining whether a length of a to-be-processed token sequence is greater than a cache length threshold associated with a key-value (KV) data cache, wherein the KV data cache is related to a processed token sequence;

determining, in response to the length of the to-be-processed token sequence being greater than the cache length threshold, target KV data based on a token sequence exceeding the cache length threshold in the to-be-processed token sequence; and

determining a target token based on the to-be-processed token sequence, the KV data cache, and the target KV data, wherein the target token is a next token of the to-be-processed token sequence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 8, 2024
From: CHEN, QIANG; FERNANDEZ ORELLANA, PEDRO
To: DELL PRODUCTS L.P.
Reel/Frame 066413/0180 →
Priority Claims (1)
CN 202410054866.4 · Jan 12, 2024 · national
References Cited (6)
US 11747998B1 · Indupuru · 2023 [cited by examiner]
US 20240195877A1 · Xue · 2024 [cited by examiner]
US 20240419493A1 · Ramanujan · 2024 [cited by examiner]
Yunho Jin et ali: “SA3: Increasing GPU Utilization during Generative Inference for Higher Throughput”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853, Jun. 9, 2023 (Jun. 9, 2… [cited by examiner]
D. Timonin et al., “Accelerated Inference for Large Transformer Models Using NVIDIA Triton Inference Server,” https://developer.nvidia.com/blog/accelerated-inference-for-large-transformer-models-using-nvidia-fastertrans… [cited by applicant]
H. Barad et al., “Leveraging Speculative Sampling and KV-Cache Optimizations Together for Generative AI using OpenVINO,” arXiv:2311.04951v1, Nov. 8, 2023, 5 pages. [cited by applicant]