IP Library › Granted Patent US 12,380,145
Granted Patent B2
US 12,380,145 · App. 17/710,009 · Granted Aug 5, 2025

Methods, mediums, and systems for reusable intelligent search workflows

Inventors: Paul Cho (Boston, MA); Gary B. Williams (Williamsburg, VA); Alexander Bussan (Arlington, VA); Eric Campbell (McLean, VA); Piper Alexandra Coble (Richmond, VA); Ralph Lozano (Brooklyn, NY); Mukund Manikantan (McLean, VA); Talyne Derderian Walsh (Rockville, MD); Talha Koc (Jersey City, NJ); Prarthana Bhattarai (Quincy, MA)
Assignee: Capital One Services, LLC
G06F16/3338G06F16/3326G06F16/338
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,380,145
App. No.
17/710,009
Filed
Mar 31, 2022
Granted
Aug 5, 2025
Kind
B2
Art Unit
2156
USPC
707/722
Abstract

Exemplary embodiments provide methods, mediums, and systems for performing a reusable, intelligent semantic search across a potentially large number of records. Embodiments may be particularly useful for responding to requests for information from regulatory agencies. In one embodiment, an embedding model is trained to embed queries in an embedding space. When a new query is received, the new query is embedded with the embedding model. A set of documents (e.g., previous responses to regulatory inquiries) may be searched using the embedded query and an indexing model that allows for efficient searches of embedding spaces. A number of results may be returned from the document store, and the results may be ranked by a ranking model. User feedback about the quality of the results may be received, and the ranking model may be retrained based on the feedback.

Claims (72)

1. A computer-implemented method comprising, via at least one computing device comprising at least one processor and a memory storing instructions that, when executed by the at least one processor, cause the computing device to perform the computer-implemented method via:

receiving a natural language search query pertaining to a request for information;

accessing a set of documents, each document of the set of documents comprising a document query and a document response, each document query embedded in an embedding space according to an embedding model, wherein the set of documents are clustered into a plurality of cells based on the document query;

creating a search query embedding of the search query using the embedding model;

assigning the search query embedding to a cell of the plurality of cells;

using the search query embedding to perform a semantic search on the embedding space of documents clustered in the cell; and

returning at least one returned document of the set of documents as a search result based on a proximity of the document query from the at least one returned document to the search query in the embedding space, wherein the returned document along with metadata associated with the returned document includes an identifier for an individual that generated or approved the returned document is displayed in a graphical user interface (GUI).

2. The computer-implemented method of claim 1 , wherein the request for information is from a regulatory entity, and the set of documents comprises previous responses to at least one previous query from the regulatory entity; and

wherein the natural language search query is also displayed on the GUI.

3. The computer-implemented method of claim 1 , wherein the at least one returned document comprises a plurality of returned documents, each of the plurality of returned documents being ranked based on the proximity; and

wherein the GUI is configured to display the plurality of returned documents in a ranked order based on the proximity, the method further comprising:

submitting the document query from each of the plurality of returned documents to a reranking model configured to rerank the plurality of returned documents based on user feedback provided in the GUI; and

reranking the plurality of returned documents based on the reranking model.

4. The computer-implemented method of claim 3 , further comprising:

receiving the user feedback for a selected search result; and

retraining the reranking model based on the user feedback.

5. The computer-implemented method of claim 1 , further comprising:

training the embedding model with a set of labeled data; and

training an indexing model configured to perform the semantic search on the embedding space.

6. The computer-implemented method of claim 5 , wherein the set of documents is a first set of documents and the semantic search is a first semantic

swapping the first set of documents for a second set of documents different from the first set of documents; and

applying the indexing model to perform a second semantic search based on a new natural language search query.

7. The computer-implemented method of claim 1 , wherein the proximity of the document query from the returned document to the search query in the embedding space is represented as a cosine similarity.

8. A non-transitory computer-readable storage medium, the computer-readable storage medium storing instructions that,-when executed by a computer, cause the computer to:

receive a natural language search query pertaining to a request for information;

access a set of documents, each document of the set of documents comprising a document query and a document response, each document query embedded in an embedding space according to an embedding model, wherein the set of documents are clustered into a plurality of cells based on the document query;

create a search query embedding of the search query using the embedding model;

assign the search query embedding to a cell of the plurality of cells;

use the search query embedding to perform a semantic search on the embedding space of documents clustered in the cell; and

return at least one returned document of the set of documents as a search result based on a proximity of the document query from the at least one returned document to the search query in the embedding space, wherein the returned document along with metadata associated with the returned document includes an identifier for an individual that generated or approved the returned document is displayed in a graphical user interface (GUI).

9. The computer-readable storage medium of claim 8 , wherein the request for information is from a regulatory entity, and the set of documents comprises previous responses to at least one previous query from the regulatory entity; and

wherein the natural language search query is also displayed on the GUI.

10. The computer-readable storage medium of claim 8 , wherein the at least one returned document comprises a plurality of returned documents, each of the plurality of returned documents being ranked based on the proximity;

wherein the GUI is configured to display the plurality of returned documents in a ranked order based on the proximity; and

wherein the instructions further configure the computer to:

submit the document query from each of the plurality of returned documents to a reranking model configured to rerank the plurality of returned documents based on user feedback provided on the GUI; and

rerank the plurality of returned documents based on the reranking model.

11. The computer-readable storage medium of claim 10 , wherein the instructions further configure the computer to:

receive the user feedback for a selected search result; and

retrain the reranking model based on the user feedback.

12. The computer-readable storage medium of claim 8 , wherein the instructions further configure the computer to:

train the embedding model with a set of labeled data; and

train an indexing model configured to perform the semantic search on the embedding space.

13. The computer-readable storage medium of claim 12 , wherein the set of documents is a first set of documents and the semantic search is a first semantic search, and wherein the instructions further configure the computer to:

swap the first set of documents for a second set of documents different from the first set of documents; and

apply the indexing model to perform a second semantic search based on a new natural language search query.

14. The computer-readable storage medium of claim 8 , wherein the proximity of the document query from the returned document to the search query in the embedding space is represented as a cosine similarity.

15. A computing apparatus comprising:

a processor; and

a memory storing instructions that, when executed by the processor, cause the apparatus to:

receive a natural language search query pertaining to a request for information;

access a set of documents, each document of the set of documents comprising a document query and a document response, each document query embedded in an embedding space according to an embedding model, wherein the set of documents are clustered into a plurality of cells based on the document query;

create a search query embedding of the search query using the embedding model;

assign the search query embedding to a cell of the plurality of cells;

use the search query embedding to perform a semantic search on the embedding space of documents clustered in the cell; and

return at least one returned document of the set of documents as a search result based on a proximity of the document query from the at least one returned document to the search query in the embedding space, wherein the returned document along with metadata associated the returned document includes an identifier for an individual that generated or approved the returned document is displayed in a graphical user interface (GUI).

16. The computing apparatus of claim 15 , wherein the request for information is from a regulatory entity, and the set of documents comprises previous responses to at least one previous query from the regulatory entity; and

wherein the natural language search query is also displayed on the GUI.

17. The computing apparatus of claim 15 , wherein the at least one returned document comprises a plurality of returned documents, each of the plurality of returned documents being ranked based on the proximity;

wherein the GUI is configured to display the plurality of returned documents in a ranked order based on the proximity; and

wherein the instructions further configure the apparatus to:

submit the document query from each of the plurality of returned documents to a reranking model configured to rerank the plurality of returned documents based on user feedback provided on the GUI; and

rerank the plurality of returned documents based on the reranking model.

18. The computing apparatus of claim 17 , wherein the instructions further configure the apparatus to:

receive the user feedback for a selected search result; and

retrain the reranking model based on the user feedback.

19. The computing apparatus of claim 15 , wherein the instructions further configure the apparatus to:

train the embedding model with a set of labeled data; and

train an indexing model configured to perform the semantic search on the embedding space.

20. The computing apparatus of claim 19 , wherein the set of documents is a first set of documents and the semantic search is a first semantic search, and wherein the instructions further configure the apparatus to:

swap the first set of documents for a second set of documents different from the first set of documents; and

apply the indexing model to perform a second semantic search based on a new natural language search query.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2022
From: CHO, PAUL; WILLIAMS, GARY B.; BUSSAN, ALEXANDER; CAMPBELL, ERIC; COBLE, PIPER ALEXANDRA; LOZANO, RALPH; MANIKANTAN, MUKUND; WALSH, TALYNE DERDERIAN; KOC, TALHA; BHATTARAI, PRARTHANA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 060291/0183 →
Continuity (1)
Related Publication 20230315766A1 · Oct 5, 2023
References Cited (9)
US 20180293483A1 · Abramson · 2018 [cited by examiner]
US 20190065576A1 · Peng · 2019 [cited by examiner]
US 20200250537A1 · Li · 2020 [cited by examiner]
US 20220019671A1 · Boone · 2022 [cited by examiner]
US 20220138433A1 · Divakaran · 2022 [cited by examiner]
Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding” arXiv:1810.04805v2 [cs.CL] May 24, 2019. [cited by applicant]
Conneau et al. “Supervised Learning of Universal Sentence Representations from Natural Language Inference Data” arXiv:1705.02364v5 [cs.CL] Jul. 8, 2018. [cited by applicant]
Jegou et al., “Faiss: A library for efficient similarity search” Engineering at Meta—Mar. 29, 2017. [cited by applicant]
Author Unknown, “Welcome to Elastic Docs” Elastic—retrieved Jul. 27, 2022—URL: https://www.elastic.co/guide/index.html. [cited by applicant]