IP Library Granted Patent US 11,526,692
Granted Patent B2
US 11,526,692 · App. 16/853,194 · Granted Dec 13, 2022

Systems and methods for domain agnostic document extraction with zero-shot task transfer

Inventor: Prithiviraj Damodaran (Chennai, IN)
Assignee: UST Global (Singapore) Pte. Ltd.
G06K9/6257G06N5/02G06N5/04G06N20/00G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,526,692
App. No.
16/853,194
Granted
Dec 13, 2022
Kind
B2
Abstract

A system for performing document extraction is configured to: (a) receive a first document; (b) extract the first document into document elements, the document elements including pages, lines, paragraphs, or any combination thereof; (c) determine a first set of fields of interest for the first document, wherein the first set of fields of interest are determined via a type of the first document or via a first set of queries for probing the first document; (d) determine, from a plurality of closed domain question answering (CDQA) models, a first set of CDQA models that provides answers to each field of interest included in the first set of fields of interest; and (e) provide answers to the first set of fields of interest to the client device.

Claims (35)

1. A system for performing document extraction, the system including a non-transitory computer-readable medium storing computer-executable instructions thereon such that when the instructions are executed, the system is configured to:

receive a first document;

extract the first document into document elements, the document elements including pages, lines, paragraphs, or any combination thereof;

determine a first set of fields of interest for the first document, wherein the first set of fields of interest are determined via a type of the first document or via a first set of queries for probing the first document;

determine, from a plurality of closed domain question answering (CDQA) models, a first set of CDQA models that provides answers to each field of interest included in the first set of fields of interest; and

provide answers to the first set of fields of interest to the client device.

2. The system of claim 1 , wherein the answers to the first set of fields of interest are stored in a knowledge graph.

3. The system of claim 1 , wherein the type of the first document includes an invoice statement, an annual report, a statement of work, a master service agreement, or any combination thereof.

4. The system of claim 3 , wherein a respective field of interest in the first set of fields of interest is configurable via a graphical user interface with at least one configurable option indicating that the respective field of interest is alphabetic, numeric, or alphanumeric.

5. The system of claim 1 , wherein the first set of CDQA models includes at least two CDQA models selected from the group consisting of: Bidirectional Encoder Representations from Transformers (BERT) trained on Stanford Question Answering Dataset (SQuAD), Simple Bi-Directional Attention Flow (BiDAF), ELMo-BIDAF.

6. The system of claim 1 , wherein the answers to the first set of fields of interest include dates, and the system is further configured to normalize the dates to a locale of the client device.

7. The system of claim 1 , further configured to provide a first set of indexes to the client device, wherein a respective answer in the answers to the first set of fields of interest is contained in one of the document elements referenced by a respective index in the first set of indexes.

8. The system of claim 1 , further configured to:

receive a second document;

extract the second document into document elements;

determine a second set of fields of interest for the second document, wherein the second set of fields of interest are determined via a type of the second document or via a second set of queries for probing the second document;

determine, from the plurality of CDQA models, a second set of CDQA models that provides answers to each field of interest included in the second set of fields of interest; and

based at least in part on the first set of fields of interest for the first document and the second set of fields of interest for the second document sharing common answers for common fields, linking answers to the second set of fields of interest and the answers to the first set of fields of interest in a knowledge graph.

9. The system of claim 8 , wherein the first document is a master service agreement and the second document is a statement of work.

10. The system of claim 9 , wherein the common fields include a First Party field, a Second Party field, and a Master Service Agreement effective date field.

11. The system of claim 1 , further configured to:

store the answers to the first set of fields of interest in a knowledge graph;

receive corrected answers from the client device; and

replace, in the knowledge graph, at least one of the answers to the first set of fields of interest with the corrected answers.

12. The system of claim 11 , further configured to:

generate training data based at least in part on the corrected answers; and

train a machine learning model with the training data, the trained machine learning model being a document specific model for extracting documents of the type of the first document.

13. The system of claim 12 , wherein the machine learning model is a named entity recognition model or a semantic slot filling model.

14. The system of claim 1 , wherein the first set of CDQA models includes a first CDQA model and a second CDQA model different from the first CDQA model, wherein the provided answers to the first set of fields of interest include a first answer from the first CDQA model and a second answer from the second CDQA model, the first answer and the second answer directed at different fields of interest in the first set of fields of interest.

15. The system of claim 14 , wherein the first CDQA model is Bidirectional Encoder Representations from Transformers (BERT) trained on Stanford Question Answering Dataset (SQuAD) and the second CDQA model is Simple Bi-Directional Attention Flow (BiDAF).

16. The system of claim 3 , wherein a respective field of interest in the first set of fields of interest is configurable via a graphical user interface with at least one configurable option indicating a page affinity.

17. The system of claim 3 , wherein a respective field of interest in the first set of fields of interest is configurable via a graphical user interface with at least one configurable option indicating a type of question for probing to obtain a respective answer for the respective field of interest.

18. The system of claim 1 , wherein the answers to the first set of fields of interest are provided in an email, a chatbox, or a voice recording.

19. The system of claim 1 , wherein the answers to the first set of fields of interest are provided in sentence format based on grammar rules for a given locale.

20. The system of claim 1 , wherein the first set of fields of interest is preconfigured, and the determine the first set of fields of interest for the first document includes retrieving the first set of fields of interest from a document extraction configuration file.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 13, 2025
From: UST GLOBAL (SINGAPORE) PTE. LIMITED
To: UST GLOBAL PRIVATE LIMITED
Reel/Frame 072012/0778 →
SECURITY INTEREST Recorded Aug 13, 2025
From: UST GLOBAL PRIVATE LIMITED
To: CITIBANK, N.A., AS AGENT
Reel/Frame 072012/0804 →
SECURITY INTEREST Recorded Dec 6, 2021
From: UST GLOBAL (SINGAPORE) PTE. LIMITED
To: CITIBANK, N.A., AS AGENT
Reel/Frame 058309/0929 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2020
From: DAMODARAN, PRITHIVIRAJ
To: UST GLOBAL (SINGAPORE) PTE. LTD.
Reel/Frame 052444/0623 →
Priority Claims (1)
IN 202011007948 · Feb 25, 2020 · national
Continuity (1)
Related Publication 20210264208A1 · Aug 26, 2021