IP Library › Granted Patent US 12,554,763
Granted Patent B2
US 12,554,763 · App. 18/797,802 · Granted Feb 17, 2026

Using a knowledge graph to determine re-prompts in a retrieval-augmentation generation (RAG) framework

Inventors: Derek William Engi (Pleasant Ridge, MI); Bradley Michael Wise (Jersey City, NJ); M. David Hanes (Lewisville, NC); Vivek Kumar Singh (Cary, NC); Ambrose D. Taylor (Durham, NC)
Assignee: CISCO TECHNOLOGY, INC.
G06F16/383G06F16/3325
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,554,763
App. No.
18/797,802
Granted
Feb 17, 2026
Kind
B2
Abstract

According to one aspect, a method includes obtaining, at an interface to a system that includes a large language model (LLM) arrangement, a first prompt, and identifying a plurality of candidate chunks of documents that substantially match the first prompt. The method also includes analyzing the plurality of candidate chunks to identify a chunk node associated with the plurality of candidate chunks, and generating a query arranged to solicit information associated with the plurality of candidate chunks. A second prompt is obtained in response to the query, and the plurality of candidate chunks is analyzed. Analyzing the plurality of candidate chunks using the second prompt includes identifying at least a first candidate chunk of the plurality of candidate chunks that is associated with the second prompt, wherein the first candidate chunk has a context. Finally, the method includes generating a response to the first prompt using the context.

Claims (64)

1 . A method comprising:

obtaining, at an interface to a system that includes a large language model (LLM) arrangement, a first prompt, the first prompt being obtained from a user;

identifying a plurality of candidate chunks of documents that match the first prompt;

analyzing the plurality of candidate chunks to identify a chunk node, the chunk node being associated with the plurality of candidate chunks;

generating a query, the query being arranged to solicit information associated with the plurality of candidate chunks;

providing the query to the interface;

obtaining, at the interface, a second prompt, the second prompt being obtained in response to the query, wherein the second prompt includes information associated with the plurality of candidate chunks;

analyzing the plurality of candidate chunks using the second prompt, wherein analyzing the plurality of candidate chunks using the second prompt includes identifying at least a first candidate chunk of the plurality of candidate chunks that is associated with the second prompt and dropping a second candidate chunk of the plurality of candidate chunks, and wherein the first candidate chunk has a context;

generating a response to the first prompt using the context; and

providing the response to the interface, wherein providing the response to the interface includes enabling the user to obtain the response through the interface.

2 . The method of claim 1 wherein analyzing the plurality of candidate chunks to identify a chunk node includes determining a centrality of the chunk node with respect to the plurality of candidate chunks, and wherein dropping the second candidate chunk of the plurality of candidate chunks includes dropping the second candidate chunk as being unrelated to the second prompt.

3 . The method of claim 1 further including:

identifying a plurality of pathways in a knowledge graph to the plurality of candidate chunks using the chunk node, wherein the plurality of candidate chunks contains similar information.

4 . The method of claim 1 wherein identifying the plurality of candidate chunks includes applying a cosine similarity analysis to the first prompt.

5 . The method of claim 1 further including:

determining at least one topic associated with the plurality of candidate chunks; and

providing the at least one topic to the LLM arrangement, wherein generating the query includes processing the at least one topic and the plurality of candidate chunks using the LLM arrangement.

6 . The method of claim 1 further including:

providing the first candidate chunk to the LLM arrangement; and

generating the response using the LLM arrangement.

7 . The method of claim 6 wherein the interface is a chatbot interface, and wherein the system includes a document chunking arrangement that is configured to ingest the documents and to create the plurality of candidate chunks.

8 . An apparatus comprising:

one or more network processor units to communicate with devices in a network; and

a processor coupled to the one or more network processor units and configured to perform:

obtaining, at an interface to a system that includes a large language model (LLM) arrangement, a first prompt, the first prompt being obtained from a user;

identifying a plurality of candidate chunks of documents that match the first prompt;

analyzing the plurality of candidate chunks to identify a chunk node, the chunk node being associated with the plurality of candidate chunks;

generating a query, the query being arranged to solicit information associated with the plurality of candidate chunks;

providing the query to the interface;

obtaining, at the interface, a second prompt, the second prompt being obtained in response to the query, wherein the second prompt includes information associated with the plurality of candidate chunks;

analyzing the plurality of candidate chunks using the second prompt, wherein analyzing the plurality of candidate chunks using the second prompt includes identifying at least a first candidate chunk of the plurality of candidate chunks that is associated with the second prompt and dropping a second candidate chunk of the plurality of candidate chunks, and wherein the first candidate chunk has a context;

generating a response to the first prompt using the context; and

providing the response to the interface, wherein providing the response to the interface includes enabling the user to obtain the response through the interface.

9 . The apparatus of claim 8 wherein analyzing the plurality of candidate chunks to identify a chunk node includes determining a centrality of the chunk node with respect to the plurality of candidate chunks, and wherein dropping the second candidate chunk of the plurality of candidate chunks includes dropping the second candidate chunk as being unrelated to the second prompt.

10 . The apparatus of claim 8 wherein the processor is further configured to perform:

identifying a plurality of pathways in a knowledge graph to the plurality of candidate chunks using the chunk node, wherein the plurality of candidate chunks contains similar information.

11 . The apparatus of claim 8 wherein identifying the plurality of candidate chunks includes applying a cosine similarity analysis to the first prompt.

12 . The apparatus of claim 8 wherein the processor is further configured to perform:

determining at least one topic associated with the plurality of candidate chunks; and

providing the at least one topic to the LLM arrangement, wherein generating the query includes processing the at least one topic and the plurality of candidate chunks using the LLM arrangement.

13 . The apparatus of claim 8 wherein the processor is further configured to perform:

providing the first candidate chunk to the LLM arrangement; and

generating the response using the LLM arrangement.

14 . A non-transitory computer readable medium encoded with instructions that, when executed by a processor configured to communicate with devices over a network, causes the processor to perform:

obtaining, at an interface to a system that includes a large language model (LLM) arrangement, a first prompt, the first prompt being obtained from a user;

identifying a plurality of candidate chunks of documents that match the first prompt;

analyzing the plurality of candidate chunks to identify a chunk node, the chunk node being associated with the plurality of candidate chunks;

generating a query, the query being arranged to solicit information associated with the plurality of candidate chunks;

providing the query to the interface;

obtaining, at the interface, a second prompt, the second prompt being obtained in response to the query, wherein the second prompt includes information associated with the plurality of candidate chunks;

analyzing the plurality of candidate chunks using the second prompt, wherein analyzing the plurality of candidate chunks using the second prompt includes identifying at least a first candidate chunk of the plurality of candidate chunks that is associated with the second prompt and dropping a second candidate chunk of the plurality of candidate chunks, and wherein the first candidate chunk has a context;

generating a response to the first prompt using the context; and

providing the response to the interface, wherein providing the response to the interface includes enabling the user to obtain the response through the interface.

15 . The non-transitory computer readable medium of claim 14 wherein analyzing the plurality of candidate chunks to identify a chunk node includes determining a centrality of the chunk node with respect to the plurality of candidate chunks, and wherein dropping the second candidate chunk of the plurality of candidate chunks includes dropping the second candidate chunk as being unrelated to the second prompt.

16 . The non-transitory computer readable medium of claim 14 wherein the instructions are further configured to cause the processor to perform:

identifying a plurality of pathways in a knowledge graph to the plurality of candidate chunks using the chunk node, wherein the plurality of candidate chunks contains similar information.

17 . The non-transitory computer readable medium of claim 16 identifying the plurality of candidate chunks includes applying a cosine similarity analysis to the first prompt.

18 . The non-transitory computer readable medium of claim 16 wherein the instructions are further configured to cause the processor to perform:

determining at least one topic associated with the plurality of candidate chunks; and

providing the at least one topic to the LLM arrangement, wherein generating the query includes processing the at least one topic and the plurality of candidate chunks using the LLM arrangement.

19 . The non-transitory computer readable medium of claim 16 wherein the instructions are further configured to cause the processor to perform:

providing the first candidate chunk to the LLM arrangement; and

generating the response using the LLM arrangement.

20 . The non-transitory computer readable medium of claim 19 wherein the interface is a chatbot interface, and wherein the system includes a document chunking arrangement that is configured to ingest the documents and to create the plurality of candidate chunks.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2024
From: ENGI, DEREK WILLIAM; WISE, BRADLEY MICHAEL; HANES, M. DAVID; SINGH, VIVEK KUMAR; TAYLOR, AMBROSE D.
To: CISCO TECHNOLOGY, INC.
Reel/Frame 068229/0305 →
Continuity (2)
Provisional Application 63653377 · May 30, 2024
Related Publication 20250371066A1 · Dec 4, 2025
References Cited (11)
US 10459989B1 · Haahr et al. · 2019 [cited by applicant]
US 12020140B1 · Mondlock · 2024 [cited by examiner]
US 12282504B1 · Ganesh · 2025 [cited by examiner]
US 20230376537A1 · Bansal et al. · 2023 [cited by applicant]
US 20250111192A1 · Bayless · 2025 [cited by examiner]
CN 117708308A · 2024 [cited by applicant]
Chan C-M., et al., “RQ-RAG: Learning to Refine Queries for Retrieval Augmented Generation”, arXiv:2404.00610v1, Mar. 31, 2024, 18 Pages. [cited by applicant]
Hu Z., et al., “Prompt Perturbation in Retrieval-Augmented Generation Based Large Language Models”, arXiv:2402.07179v1, Feb. 11, 2024, pp. 1-12. [cited by applicant]
Kuan J., “From RAGs to Riches: Helping AI Connect the Dots with a Knowledge Vector Graph”, Medium, Jan. 1, 2024, Retrieved from https://medium.com/@johnson.h.kuan/from-rags-to-riches-helping-ai-connect-the-dots-with-a-k… [cited by applicant]
Wiebeler A., et al., “Improve LLM Responses in RAG Use Cases by Interacting with the User”, AWS Machine Learning Blog, Nov. 13, 2023, Retrieved from https://aws.amazon.com/blogs/machine-learning/improve-llm-responses-in… [cited by applicant]
Wiebeler A., “Leveraging the User to Improve Agents in RAG Use Cases”, Github, aws-samples/rag-with-human-support, Retrieved from https://github.com/aws-samples/rag-with-human-support/tree/main on May 16, 2024, pp. 1-5. [cited by applicant]