IP Library › Granted Patent US 12,724,807
Granted Patent B1
US 12,724,807 · App. 19/420,230 · Granted Sep 1, 2026

Systems and methods for hybrid lexical-vector retrieval in retrieval-augmented generation models

Inventors: Yu Zhang (Hempstead, NY); Jing Shen (Great Neck, NY); Huifeng Jason Li (Highland Park, NJ); Taotao Jiang (Weehawken, NJ); Monika Nica (Ridgefield, CT)
Assignee: Morgan Stanley Services Group Inc.
G06F16/3347G06F16/335
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,807
App. No.
19/420,230
Granted
Sep 1, 2026
Kind
B1
Abstract

Systems and methods perform hybrid lexical-vector retrieval in a retrieval-augmented generation (RAG) framework. A user query is pre-processed using named-entity recognition, date standardization, and optional out-of-domain detection. The system performs both lexical search and semantic vector similarity search over a corpus of document chunks, generating ranked candidate sets that are merged by elevating overlapping results and interleaving remaining items according to a predetermined rule. A top subset of chunks is selected based on the combined ranking and provided to a large language model (LLM), together with intent-specific instructions determined through query-classification logic. The LLM generates an answer grounded in the retrieved material and formatted according to a standardized template associated with the detected intent.

Claims (68)

1 . A computer-implemented system for generating an answer to a user query, the system comprising:

a vector database storing a plurality of document chunks and corresponding vector embeddings;

a lexical indexing system storing lexical index entries for the plurality of document chunks;

a retrieval-augmented generation (RAG) system, comprising one or more programmed processors, wherein the RAG system is in communication with the vector database and with the lexical indexing system, and wherein the RAG system is configured, via programming, to:

receive the user query from a user;

generate an embedded representation of the user query;

generate a set of M ranked candidate document chunks based at least in part on a lexical search of the lexical indexing system using the user query;

generate a set of N ranked candidate document chunks based at least in part on a vector similarity search of the vector database using the embedded representation of the user query;

generate a combined rank list of document chunks by:

identifying one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks;

ranking the one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks higher in the combined rank list than document chunks not common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks;

adding to the combined rank list by interleaving, according to a predetermined merging rule and based on ranking, document chunks not common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks; and

selecting a set of K document chunks from the combined rank list, wherein the set of K document chunks comprises a highest-ranked K document chunks from the combined rank list;

provide the set of K document chunks to a large language model (LLM) of the RAG system to generate an answer to the user query based at least in part on the set of K document chunks; and

provide the answer to user.

2 . The system of claim 1 , wherein adding document chunks to the combined ranked list comprises alternating, according to the predetermined merging rule, between a highest-ranked remaining lexical document chunk and a highest-ranked remaining vector document chunk.

3 . The system of claim 1 , wherein generating the combined rank list further comprises ranking the one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks based on based on underlying lexical and semantic scoring associated with the ranked candidate document chunks.

4 . The system of claim 1 , wherein:

the lexical search of the lexical indexing system produces M candidate document chunks; and

generation of the set of M ranked candidate document chunks is further based on a vector similarity search restricted to the M candidate document chunks.

5 . The system of claim 1 , wherein a value for K is selected such that a total textual content of the K document chunks fits within a maximum context window of the LLM.

6 . The system of claim 1 , wherein the lexical search is performed using Apache Solr.

7 . The system of claim 1 , wherein the RAG system is further configured to rerank the N ranked candidate document chunks based at least in part on semantic similarity and document recency.

8 . The system of claim 1 , wherein the RAG system is configured:

for use with a specified use case domain;

to perform named entity recognition (NER) on the user query to identify one or more identified entities referenced in the user query;

to use the one or more identified entities as part of the lexical search and the vector similarity search performed by the RAG system; and

to perform NER using domain-specific entity definitions associated with the specified use case domain.

9 . The system of claim 8 , wherein the RAG system is further configured to determine whether the user query is out-of-domain for the specified use case domain and, responsive to a determination that the query is out-of-domain, to restrict answer-generation operations.

10 . The system of claim 9 , wherein:

the specified use case domain comprises sell-side research; and

the RAG system is further configured to perform date standardization on the user query to identify a standardized date range, and to use the standardized date range in both the lexical search and the vector similarity search performed by the RAG system.

11 . The system of claim 1 , wherein the RAG system is further configured to perform date standardization on the user query to identify a standardized date range, and to use the standardized date range in both the lexical search and the vector similarity search performed by the RAG system.

12 . The system of claim 1 , wherein the RAG system is further configured, via programming, make a classification of the user query into one of a predetermined plurality of intent categories and to determine one or more intent-specific retrieval parameters based on the classification.

13 . The system of claim 12 , wherein the RAG system is further configured to cause the LLM to generate the answer using a standardized answer format associated with the classification.

14 . The system of claim 13 , wherein the standardized answer format comprises a template specifying a required ordering, section structure, or presentation format, and wherein the RAG system provides the template to the LLM together with the set of K document chunks.

15 . The system of claim 1 , wherein the LLM is deployed within a private computing environment with the RAG system.

16 . The system of claim 15 , wherein the private computing environment comprises a containerized deployment executed within an isolated virtual private cloud, and wherein the LLM is executed on compute nodes of the private computing environment without external network access during inference.

17 . A computer-implemented method for generating an answer to a user query, the method comprising:

storing, in a vector database, a plurality of document chunks and corresponding vector embeddings;

storing, in a lexical indexing system, lexical index entries for the plurality of document chunks; and

by a retrieval-augmented generation (RAG) system, comprising one or more programmed processors, and in communication with the vector database and with the lexical indexing system:

receiving the user query from a user;

generating an embedded representation of the user query;

generating a set of M ranked candidate document chunks based at least in part on a lexical search of the lexical indexing system using the user query;

generating a set of N ranked candidate document chunks based at least in part on a vector similarity search of the vector database using the embedded representation of the user query;

generating a combined rank list of document chunks by:

identifying one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks;

ranking the one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks higher in the combined rank list than document chunks not common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks;

adding to the combined rank list by interleaving, according to a predetermined merging rule and based on ranking, document chunks not common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks; and

selecting a set of K document chunks from the combined rank list, wherein the set of K document chunks comprises a highest-ranked K document chunks from the combined rank list;

providing the set of K document chunks to a large language model (LLM) of the RAG system to generate an answer to the user query based at least in part on the set of K document chunks; and

providing the answer to user.

18 . The method of claim 17 , wherein generating the combined rank list comprises alternating, according to the predetermined merging rule, between a highest-ranked remaining lexical document chunk and a highest-ranked remaining vector document chunk.

19 . The method of claim 17 , wherein generating the combined rank list further comprises ranking the one or more document chunks common to both the set of M ranked candidate document chunks and the set of N ranked candidate document chunks based on an overlap weight.

20 . The method of claim 17 , wherein:

the lexical search of the lexical indexing system produces M candidate document chunks; and

generating the set of M ranked candidate document chunks is further based on a vector similarity search restricted to the M candidate document chunks.

21 . The method of claim 17 , wherein selecting the set of K document chunks comprises selecting a value of K such that a total textual content of the K document chunks fits within a maximum context window of the LLM.

22 . The method of claim 17 , wherein:

the RAG system is for use with a specified use case domain, the specified use case domain comprising sell-side research;

the method further comprises:

performing named entity recognition (NER) on the user query to identify one or more identified entities referenced in the user query, wherein the NER uses domain-specific entity definitions associated with the specified use case domain;

performing date standardization on the user query to identify a standardized publication-date range; and

using the one or more identified entities and the standardized publication-date range in both the lexical search and the vector similarity search performed by the RAG system.

23 . The method of claim 22 , wherein:

the method further comprises making a classification of the user query into one of a predetermined plurality of intent categories and determining one or more intent-specific retrieval parameters based on the classification; and

providing the answer comprises causing, by the RAG system, the LLM to generate the answer using a standardized answer format associated with the classification.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2026
From: ZHANG, YU; SHEN, JING; LI, HUIFENG JASON; JIANG, TAOTAO; NICA, MONIKA
To: MORGAN STANLEY SERVICES GROUP INC.
Reel/Frame 073920/0200 →
References Cited (15)
US 11580301B2 · Narayan et al. · 2023 [cited by applicant]
US 12182125B1 · Buniatyan · 2024 [cited by applicant]
US 12346337B1 · Palatnik De Sousa · 2025 [cited by examiner]
US 12450274B1 · Jain · 2025 [cited by examiner]
US 20100153315A1 · Gao · 2010 [cited by examiner]
US 20250131289A1 · Larson et al. · 2025 [cited by applicant]
US 20250165480A1 · Raudaschl · 2025 [cited by examiner]
US 20250190460A1 · Madisetti et al. · 2025 [cited by applicant]
US 20250209108A1 · Yang · 2025 [cited by examiner]
US 20250265249A1 · Buniatyan · 2025 [cited by examiner]
US 20250292110A1 · Jain et al. · 2025 [cited by applicant]
US 20250322087A1 · Salarian et al. · 2025 [cited by applicant]
US 20260030274A1 · Croskey · 2026 [cited by examiner]
US 20260072918A1 · Meghwani · 2026 [cited by examiner]
US 20260072960A1 · Zhang · 2026 [cited by examiner]