IP Library › Granted Patent US 12,561,322
Granted Patent B2
US 12,561,322 · App. 18/742,043 · Granted Feb 24, 2026

Systems and methods for semantic caching

Inventors: Ashok Ganesan (Leander, TX); Peng Wang (Bellevue, WA)
Assignee: ServiceNow, Inc.
G06F16/24539G06F16/248G06F16/3347
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,561,322
App. No.
18/742,043
Granted
Feb 24, 2026
Kind
B2
Abstract

Systems and methods are provided to improve data retrieval from a cache memory by using semantic matching to retrieve data from the cache memory. The system includes a two-tiered cache system, with a first tier implementing “key-value” pairs, and a second tier that includes a table that is configured as an artificial intelligence (AI) search indexed source. When a new input does not have a matching “key” at the first tier, the system performs a semantic search at the second tier of the cache to determine if relevant data is stored in the cache. The current systems and methods increase the likelihood of obtaining data for queries from the cache memory, reduce the response time to the queries, improve search consistency, reduce computing resource utilization, improve system performance, and reduce costs.

Claims (49)

1 . A method comprising:

receiving, via processing circuitry, a query from a client device;

determining that the query does not match any record of a first plurality of records stored in a first cache, wherein a first set of policies is used for managing the first cache;

in response to determining that the query does not match any record of the first plurality of records,

determining a semantic value of the query; and

identifying, within a second plurality of records stored in a second cache, a particular record comprising a particular query term corresponding to a particular semantic value that matches the semantic value of the query within an error threshold, wherein a second set of policies is used for managing the second cache, and wherein the second set of policies comprises a different update policy, a different eviction policy, or both, compared to the first set of policies.

2 . The method of claim 1 , wherein each record of the first plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

3 . The method of claim 2 , wherein determining that the query does not match any of the first plurality of records comprises:

comparing the query with the respective query terms of the first plurality of records; and

identifying that none of the first plurality of records comprises a query term that matches the query.

4 . The method of claim 1 , wherein each record of the second plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

5 . The method of claim 4 , wherein identifying the particular record comprises,

determining a respective match score for each record of the second plurality of records based on a comparison of a respective semantic value of the respective query term and the semantic value of the query;

identifying a matching record from the second plurality of records having a match score that satisfies a threshold match score; and

providing, in response to the query, a cached value of the particular record.

6 . The method of claim 1 , wherein the semantic value of the query is determined based on an intent of the query, and the particular semantic value is determined based on a particular intent of the particular query term.

7 . The method of claim 6 , wherein determining the intent of the query comprises generating a semantic word vector for the query.

8 . The method of claim 1 , comprising:

receiving an additional query; and

in response to an additional record corresponding to the additional query not being found in either the first plurality of records or the second plurality of records, determining a response for the additional query based on data stored in a database.

9 . The method of claim 8 , wherein the determining the response for the additional query comprises:

providing the additional query to a large language model (LLM);

receiving an output from the LLM based on the data stored in the database; and

providing the output in response to the additional query.

10 . The method of claim 1 , wherein the second cache comprises a data index table comprising a plurality of data entries for storing the second plurality of records.

11 . A system comprising:

processing circuitry; and

memory accessible by the processing circuitry, the memory storing

instructions that, when executed by the processing circuitry, cause the processing circuitry to perform operations comprising:

receiving a query from a client device;

determining that the query does not match any record of a first plurality of records stored in a first cache, wherein a first set of policies is used for managing the first cache;

in response to determining that the query does not match any record of the first plurality of records,

determining a semantic value of the query; and

identifying, within a second plurality of records stored in a second cache, a particular record comprising a particular query term corresponding to a particular semantic value that matches the semantic value of the query within an error threshold, wherein a second set of policies is used for managing the second cache, and wherein the second set of policies comprises a different update policy, a different eviction policy, or both, compared to the first set of policies.

12 . The system of claim 11 , wherein each record of the first plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

13 . The system of claim 11 , wherein each record of the second plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

14 . The system of claim 11 , wherein the first cache is stored in a local network.

15 . The system of claim 11 , wherein the second cache is stored in a data center.

16 . A tangible, non-transitory computer readable storage media storing instructions that, when executed by processing circuitry, cause the processing circuitry to perform operations comprising:

receiving a query from a client device;

determining that the query does not match any record of a first plurality of records stored in a first cache, wherein a first set of policies is used for managing the first cache;

in response to determining that the query does not match any record of the first plurality of records,

determining a semantic value of the query; and

identifying, within a second plurality of records stored in a second cache, a particular record comprising a particular query term corresponding to a particular semantic value that matches the semantic value of the query within an error threshold, wherein a second set of policies is used for managing the second cache, and wherein the second set of policies comprises a different update policy, a different eviction policy, or both, compared to the first set of policies.

17 . The non-transitory computer readable storage media of claim 16 , wherein each record of the first plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

18 . The non-transitory computer readable storage media of claim 16 , wherein each record of the second plurality of records comprises a respective cache key corresponding to a respective cache value, wherein the respective cache key comprises a respective query term and the respective cache value corresponds to a respective unique identifier used to identify a respective result for the respective query term.

19 . The method of claim 8 , further comprising:

in response to the additional record corresponding to the additional query not being found in either the first plurality of records or the second plurality of records, updating the first cache to include a record comprising the response.

20 . The system of claim 11 , wherein the first plurality of records is stored in the first cache for a first time period based on the first set of policies, and wherein the second plurality of records is stored in the second cache for a second time period based on the second set of policies, and wherein the second time period is longer than the first time period.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2024
From: GANESAN, ASHOK; WANG, PENG
To: SERVICENOW, INC.
Reel/Frame 067716/0211 →
Continuity (1)
Related Publication 20250384036A1 · Dec 18, 2025
References Cited (12)
US 11816121B2 · Ng · 2023 [cited by examiner]
US 20200320153A1 · Luz Xavier Da Costa · 2020 [cited by examiner]
US 20200409945A1 · Chen · 2020 [cited by examiner]
US 20230051025A1 · Pan · 2023 [cited by examiner]
US 20230195735A1 · Al-Qurishi · 2023 [cited by examiner]
US 20230259705A1 · Tunstall-Pedoe et al. · 2023 [cited by applicant]
US 20240411737A1 · Sivaraj · 2024 [cited by examiner]
Zilliz, “Caching LLM Queries forperformance & cost improvements,” Medium, Nov. 17, 2023, accessed on Feb. 15, 2024 via https://medium.com/@zilliz_learn/caching-llm-queries-for-performance-cost-improvements-52346fade9cd#… [cited by applicant]
Ferrarotti Flavio et al.; “A Last-Resort Semantic Cache for Web Queries”; Aug. 25, 2009, SAT 2015 18th International Conference; Sep. 24-27, 2015; pp. 310-321 (XP047380356). [cited by applicant]
Heloir Yan Han; “Maximizing AI Efficiency in Production with Cacheing: a Cost-Efficient Performance Booster”; Mar. 19, 2024 (XP093306771) [retrieved from the internet—https://medium.com/data-science/maximizing-ai-effici… [cited by applicant]
Fu Bang; “GPTCache: an Open-Source Semantic Cache for LLM Applications Enabling Faster Answers and Cost Savings”, Oct. 24, 2023; pp. 1-7; (XP093256823) [retried from the Internet—https://openreview.net/pdf?id=ivwM8NwM4z… [cited by applicant]
Shankar Arun; Implementing Semantic Caching: a Step-by-Step Guide to Faster, Cost-eEffective GenAI Workflows; Jun. 13, 2024 (XP093306769) [retrieved from the internet—https://medium.com/google-cloud/implementing-semanti… [cited by applicant]