IP Library › Granted Patent US 12,481,666
Granted Patent B1
US 12,481,666 · App. 18/791,219 · Granted Nov 25, 2025

Enhanced document retrieval with semantic depth and syntactic structure

Inventors: Siddharth Jain (Mountain View, CA); Sivashanker Thiruchittampalam (Toronto, CA); Jonathan Lin (Gloucester, CA); Venkat Narayan Vedam (Mountain View, CA)
Assignee: Intuit Inc.
G06F16/24578G06F16/9024G06F16/9027G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,666
App. No.
18/791,219
Granted
Nov 25, 2025
Kind
B1
Abstract

Certain aspects of the present disclosure describe a method of information retrieval. In certain aspects, the method includes identifying a set of relevant nodes of a document graph embedding semantic units associated with the document based on a document search query. The method further includes reconstructing a structural context for each relevant node in the set of relevant nodes. The method further includes processing the set of relevant nodes and the structural context of each relevant node with a large language model to generate a contextual response to the document search query.

Claims (61)

1 . A method of integrated search in structurally complex documents, comprising:

receiving a document search query;

identifying a set of relevant nodes of a document graph comprising a plurality of embedded semantic units of a document, comprising:

performing a semantic search of the document graph based on the document search query to identify a first relevant node; and

performing a lexical search of the document graph based on the document search query to identify a second relevant node;

reconstructing a structural context for each relevant node in the set of relevant nodes, comprising:

backtracking the document graph based on the each relevant node to identify a hierarchical structure of the each relevant node;

forward tracking the document graph based on the each relevant node to identify one or more details associated with the each relevant node; and

constructing a path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node, wherein the path of the document graph represents the structural context of the document; and

processing the set of relevant nodes and the structural context of the each relevant node with a large language model to generate a contextual response to the document search query.

2 . The method of claim 1 , wherein performing the semantic search of the document graph based on the document search query comprises determining a semantic similarity based on a distance between an embedded semantic unit of the plurality of embedded semantic units and an embedding of the document search query embedded in the document graph.

3 . The method of claim 1 , wherein performing the lexical search of the document graph based on the document search query comprises determining a lexical match between one or more terms of an embedded semantic unit of the plurality of embedded semantic units and the document search query.

4 . The method of claim 1 , wherein identifying the set of relevant nodes of the document graph comprising the plurality of embedded semantic units of the document, further comprises:

assigning a relevance score of the first relevant node of the document graph based on the semantic search and the lexical search;

assigning a relevance score of the second relevant node of the document graph based on the semantic search and the lexical search; and

selecting the set of relevant nodes based on the relevance score of the first relevant node and the relevance score of the second relevant node.

5 . The method of claim 4 , wherein assigning the relevance score of the first relevant node of the document graph based on the semantic search and the lexical search comprises adjusting a weighting of the semantic search of the relevance score of the first relevant node to increase semantic depth of the relevance score.

6 . The method of claim 4 , wherein assigning the relevance score of the first relevant node of the document graph based on the semantic search and the lexical search comprises adjusting a weighting of the lexical search of the relevance score of the first relevant node to increase lexical precision of the relevance score.

7 . The method of claim 1 , wherein backtracking the document graph based on the each relevant node to identify the hierarchical structure of the each relevant node, comprises:

traversing the document graph to identify one or more parent nodes of the each relevant node; and

terminating the traversing of the document graph when a root node of the document graph is identified as one of the one or more parent nodes of the each relevant node, wherein each of the one or more parent nodes indicates a hierarchical and context framework of the each relevant node within the document.

8 . The method of claim 1 , wherein forward tracking the document graph based on the each relevant node to identify the one or more details associated with the each relevant node, comprises:

traversing the document graph to identify one or more child nodes of the each relevant node; and

terminating the traversing of the document graph when a leaf node of the document graph is identified as one of the one or more child nodes of the each relevant node, wherein each of the one or more child nodes indicates an additional detail associated with a semantic unit embedded as the each relevant node.

9 . The method of claim 1 , wherein constructing the path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node, comprises determining an optimal path to traverse one or more additional nodes identified based on backtracking and forward tracking of the document graph, wherein the optimal path indicates the each relevant node's structural placement and context within the document.

10 . A processing system, comprising: a memory comprising computer-executable instructions; and a processor configured to execute the computer-executable instructions and cause the processing system to:

receive a document search query;

identify a set of relevant nodes of a document graph comprising a plurality of embedded semantic units of a document, comprising:

perform a semantic search of the document graph based on the document search query to identify a first relevant node; and

perform a lexical search of the document graph based on the document search query to identify a second relevant node;

reconstruct a structural context for each relevant node in the set of relevant nodes, comprising:

backtrack the document graph based on the each relevant node to identify a hierarchical structure of the each relevant node;

forward track the document graph based on the each relevant node to identify one or more details associated with the each relevant node; and

construct a path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node; and

process the set of relevant nodes and the structural context of the each relevant node with a large language model to generate a contextual response to the document search query.

11 . The processing system of claim 10 , wherein to perform the semantic search of the document graph based on the document search query, the processor is further configured to cause the processing system to determine a semantic similarity based on a distance between an embedded semantic unit of the plurality of embedded semantic units and an embedding of the document search query embedded in the document graph.

12 . The processing system of claim 10 , wherein to perform the lexical search of the document graph based on the document search query, the processor is further configured to cause the processing system to determine a lexical match between one or more terms of an embedded semantic unit of the plurality of embedded semantic units and the document search query.

13 . The processing system of claim 10 , wherein to identify the set of relevant nodes of the document graph comprising the plurality of embedded semantic units of the document, the processor is further configured to cause the processing system to:

assign a relevance score of the first relevant node of the document graph based on the semantic search and the lexical search;

assign a relevance score of the second relevant node of the document graph based on the semantic search and the lexical search; and

select the set of relevant nodes based on the relevance score of the first relevant node and the relevance score of the second relevant node.

14 . The processing system of claim 13 , wherein to assign the relevance score of the first relevant node of the document graph based on the semantic search and the lexical search comprises adjusting a weighting of the semantic search of the relevance score of the first relevant node to increase semantic depth of the relevance score.

15 . The processing system of claim 13 , wherein assigning the relevance score of the first relevant node of the document graph based on the semantic search and the lexical search the processor is further configured to cause the processing system to adjusting a weighting of the lexical search of the relevance score of the first relevant node to increase lexical precision of the relevance score.

16 . The processing system of claim 10 , wherein to backtrack the document graph based on the each relevant node to identify the hierarchical structure of the each relevant node, the processor is further configured to cause the processing system to:

traverse the document graph to identify one or more parent nodes of the each relevant node; and

terminate the traversing of the document graph when a root node of the document graph is identified as one of the one or more parent nodes of the each relevant node, wherein each of the one or more parent nodes indicates a hierarchical and context framework of the each relevant node within the document.

17 . The processing system of claim 10 , wherein to forward track the document graph based on the each relevant node to identify the one or more details associated with the each relevant node, the processor is further configured to cause the processing system to:

traverse the document graph to identify one or more child nodes of the each relevant node; and

terminate the traversing of the document graph when a leaf node of the document graph is identified as one of the one or more child nodes of the each relevant node, wherein each of the one or more child nodes indicates an additional detail associated with a semantic unit embedded as the each relevant node.

18 . The processing system of claim 10 , wherein to construct the path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node, the processor is further configured to cause the processing system to determine an optimal path to traverse one or more additional nodes identified based on backtracking and forward tracking of the document graph, wherein the optimal path indicates the each relevant node's structural placement and context within the document.

19 . A method of integrated search in structurally complex documents, comprising:

receiving a document search query to augment a prompt to a large language model (LLM);

identifying a set of relevant nodes of a document graph comprising a plurality of embedded semantic units of a document, comprising:

performing a semantic search of the document graph based on the document search query to identify a first relevant node; and

performing a lexical search of the document graph based on the document search query to identify a second relevant node;

reconstructing a structural context for each relevant node in the set of relevant nodes, comprising:

backtracking the document graph based on the each relevant node to identify a hierarchical structure of the each relevant node;

forward tracking the document graph based on the each relevant node to identify one or more details associated with the each relevant node; and

constructing a path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node; and

providing the set of relevant nodes and the structural context of the each relevant node with the prompt to the LLM to generate a contextual response.

20 . The method of claim 19 , wherein constructing the path of the document graph based on the hierarchical structure of the each relevant node and the one or more details associated with the each relevant node, comprises determining an optimal path to traverse one or more additional nodes identified based on backtracking and forward tracking of the document graph, wherein the optimal path indicates the each relevant node's structural placement and context within the document.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2024
From: JAIN, SIDDHARTH; THIRUCHITTAMPALAM, SIVASHANKER; LIN, JONATHAN; VEDAM, VENKAT NARAYAN
To: INTUIT INC.
Reel/Frame 069234/0309 →
References Cited (3)
US 12204524B1 · Birru · 2025 [cited by examiner]
US 12411896B1 · Selman · 2025 [cited by examiner]
US 20240037128A1 · Koneru · 2024 [cited by examiner]
Cited By (2)
US 12,645,676 US 12,664,168