IP Library › Granted Patent US 11,900,064
Granted Patent B2
US 11,900,064 · App. 17/532,896 · Granted Feb 13, 2024

Neural network-based semantic information retrieval

Inventors: Aaron Sisto (Oakland, CA); Nick Martin (Saint Augustine, FL); Brian Shin (Reno, NV); Hung Nguyen (St. Rafael, CA)
Assignee: Searchable AI Corp
G06F40/30G06F16/90332G06N3/04G06V30/153
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,900,064
App. No.
17/532,896
Granted
Feb 13, 2024
Kind
B2
Abstract

A question and answer (Q&A) system is enhanced to support natural language queries into any document format regardless of where the underlying documents are stored. The Q&A system may be implemented “as-a-service,” e.g., a network-accessible information retrieval platform. Preferably, the techniques herein enable a user to quickly and reliably locate a document, page, chart, or data point that he or she is looking for across many different datasets. This provides for a unified view of all of the user's (or, more generally, an enterprise's) information assets (such as Adobe® PDFs, Microsoft® Word documents, Microsoft Excel spreadsheets, Microsoft PowerPoint presentations, Google Docs, scanned materials, etc.), and to be able to deeply search all of these sources for the right document, page, sheet, chart, or even answer to a question.

Claims (19)

1. A computer program product in a non-transitory computer-readable medium for use in a data processing system for information and retrieval, the computer program product holding computer program instructions that, when executed by the data processing system, are configured to:

receive a corpus of documents associated with a user, wherein the documents are structured in two or more distinct formats;

for each document in the corpus, process the document to identify a set of information strings and, for each information string, encode at least a portion of the information string into an n-dimensional semantic vector;

store the n-dimensional semantic vectors for each document;

upon receipt of a query, process the query into an n-dimensional semantic query vector;

compare the n-dimensional semantic query vector against the stored n-dimensional vectors for each document and, in response, identifying a set of candidate n-dimensional vectors that represent a possible answer to the query, wherein identifying the set of candidate n-dimensional vectors applies a neural filter that has been trained against a dataset of question-answer data structured as groupings of candidate sentences for an example query, wherein for a given training example the neural filter is trained to identify a particular candidate sentence that includes an answer to the example query while remaining candidate sentences that do not include the answer are characterized by the neural filter as contrasting;

rank the candidate n-dimensional vectors; and

return as an answer to the query a data string represented by a given highest ranked candidate n-dimensional vector.

2. The computer program product as described in claim 1 wherein the information string is a sentence, and wherein the n-dimensional semantic vector also includes context information associated with the sentence.

3. The computer program product as described in claim 1 wherein the computer program instructions are further configured to also return as a response to the query the document in which the data string occurs.

4. The computer program product as described in claim 1 wherein the document is processed using a neural network.

5. The computer program product as described in claim 1 wherein the candidate n-dimensional vectors are ranked using a neural network.

6. The computer program product as described in claim 1 wherein the highest ranked candidate n-dimensional vector is identified using a neural network.

7. The computer program product as described in claim 1 wherein there is a large number of n-dimensional semantic vectors per document.

8. The computer program product as described in claim 1 wherein the query is received as a natural language query.

9. The computer program product as described in claim 1 wherein the computer program instructions are further configured to cluster groups of documents.

10. The computer program product as described in claim 1 wherein the documents are text, spreadsheets, slide presentations, and information output from an Optical Character Reader (OCR).

11. The computer program product as described in claim 1 wherein the neural filter comprises a hybrid transformer-LSTM (long short-term memory) architecture.

12. The computer program product as described in claim 1 wherein the neural filter executes against the groupings of candidate sentences in parallel.

Continuity (2)
Continuation 17323465 · May 18, 2021
Related Publication 20220083603A1 · Mar 17, 2022