IP Library Granted Patent US 11,907,650
Granted Patent B2
US 11,907,650 · App. 17/861,805 · Granted Feb 20, 2024

Methods and systems for artificial intelligence- assisted document annotation

Inventors: Jacob T. Wilson (Castle Pines, CO); Joseph D. Harrington (New York, NY); Vinston Sundara Pandiyan Sigamani (Tampa, FL); Abhishek Sanghavi (Dallas, TX); Jayakumar Pillai (Odessa, FL); Benjamin Cunningham (New York, NY); Lindsey P. Lewis (Dallas, TX)
Assignee: PwC Product Sales LLC
G06F40/169G06F3/0482G06F16/93
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,907,650
App. No.
17/861,805
Granted
Feb 20, 2024
Kind
B2
Abstract

Methods and systems for artificial intelligence (AI)-assisted document annotation and training of machine learning-based models for document data extraction are described. The methods and systems described herein take advantage of a continuous machine learning approach to create document processing pipelines that provide accurate and efficient data extraction from documents that include structured text, semi-structured text, unstructured text, or any combination thereof.

Claims (55)

1. A computer-implemented method for annotating an electronic document comprising:

displaying an electronic document, or a page therefrom, on an electronic display device;

displaying a suggestion of labels that may be applicable to categories of text within the electronic document;

receiving a first input from a user indicating a first selection of text in the electronic document;

receiving a second input from the user to assign a first label from the suggested labels to the selected text;

storing the assigned first label, the first selection of text, and the location of the first selection of text for one or more instances of the first selection of text within the electronic document as an annotated electronic document;

receiving a third input from the user indicating a second selection of text in the electronic document;

receiving a fourth input from the user to assign a second label from the suggested labels to the second selection of text;

storing the assigned second label, the second selection of text, and the location of the second selection of text for one or more instances of the second selection of text within the annotated electronic document; and

using the annotated electronic document to train or re-train a first machine learning model to extract text corresponding to the first label and to train or retrain a second machine learning model to extract text corresponding to the second label, wherein the first and second machine learning models are deployed in a continuous learning-based document data extraction pipeline.

2. The computer-implemented method of claim 1 , wherein the suggestion of labels is based on an artificial intelligence-based prediction.

3. The computer-implemented method of claim 1 , wherein the first and second machine learning models are stored in a repository of user-selectable machine learning models for performing data extraction from electronic documents.

4. The computer-implemented method of claim 1 , further comprising receiving a fifth input from the user to assign a custom label to a selection of text.

5. The computer-implemented method of claim 1 , wherein the selection of text comprises a word, a phrase, a sentence, a paragraph, a section, or a table.

6. The computer-implemented method of claim 1 , wherein the list of labels comprises a list of text categories that includes name, date, execution date, effective date, expiration date, delivery date, due date, date of sale, order date, invoice date, issuance data, address, address line 1 , street address, quantity, amount, cost, cost of goods sold, signature, or any combination thereof.

7. The computer-implemented method of claim 1 , further comprising repeating the method for one or more additional electronic documents and storing the one or more additional annotated electronic documents.

8. The computer-implemented method of claim 7 , further comprising using the stored one or more additional annotated electronic documents as training data to train or retrain a machine learning model to automatically predict and extract selections of text corresponding to a label from non-annotated electronic documents.

9. The computer-implemented method of claim 8 , further comprising:

using the trained machine learning model to predict selections of text corresponding to the label from one or more non-annotated validation electronic documents;

sequentially displaying each of the one or more validation electronic documents, or pages therefrom, on an electronic display device, wherein the predictions of text corresponding to the label are graphically highlighted;

sequentially receiving feedback from the user on accuracy of the predicted selections of text corresponding to the label in each of the one or more validation electronic documents; and

approving or correcting each of the one or more validation electronic documents according to the feedback from the user.

10. The computer-implemented method of claim 9 , further comprising retraining the machine learning model using the one or more approved or corrected validation electronic documents.

11. The computer-implemented method of claim 1 , wherein the electronic document, or the page therefrom, is displayed within a first region of a graphical user interface on the electronic display device.

12. The computer-implemented method of claim 11 , wherein the suggestion of labels is displayed within a second region of the graphical user interface.

13. The computer-implemented method of claim 11 , further comprising displaying, within the first region of the graphical user interface, a first graphic element comprising the assigned first label and the first selection of text, wherein the first graphic element is adjacent to, or overlaid on, a location of the first selection of text.

14. The computer-implemented method of claim 11 , further comprising displaying, within the first region of the graphical user interface, a second graphic element comprising the assigned second label and the second selection of text, wherein the second graphic element is adjacent to, or overlaid on, a location of the first selection of text.

15. A computer-implemented method for electronic document annotation comprising:

displaying an electronic document, or a page therefrom, on an electronic display device;

receiving a first input from a user indicating a selection of a label that may be applicable to categories of text within the electronic document;

displaying a selection of text in the electronic document that may match a category of text corresponding to the user-selected label;

receiving a second input from the user to confirm a match between the selection of text and the category of text corresponding to the user-selected label, thereby assigning the label to the selection of text; and

storing the assigned label, the selection of text, and a location of the selected text for one or more instances of selected text within the electronic document as an annotated electronic document;

wherein the annotated electronic document is used to train or retrain one or more term-based machine learning models, and wherein the one or more term-based models are used for document data extraction in a continuous learning-based document data extraction pipeline.

16. The computer-implemented method of claim 15 , wherein the selection of text is based on an artificial intelligence-based prediction.

17. The computer-implemented method of claim 15 , further comprising repeating the steps of receiving a first input from the user indicating a selection of a label, displaying a selection of text that may match a category of text corresponding to the user-selected label, and receiving a second input from the user to confirm a match and assign the label to the selection of text for one or more additional labels.

18. A system comprising:

one or more processors;

a memory;

an electronic display device; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

displaying an electronic document, or a page therefrom, on the electronic display device;

receiving a first input from a user indicating a selection of a label that may be applicable to categories of text within the electronic document;

displaying a selection of text that may match a category of text corresponding to the user-selected label, wherein the selection of text is based on an artificial intelligence-based prediction by the system;

receiving a second input from the user to confirm a match between the selection of text and the category of text corresponding to the user-selected label, thereby assigning the label to the selection of text; and

storing the assigned label, the selection of text, and a location of the selected text for one or more instances of selected text within the electronic document as an annotated electronic document;

wherein the annotated electronic document is used to train or retrain one or more term-based machine learning models, and wherein the one or more term-based models are used for document data extraction in a continuous learning-based document data extraction pipeline.

19. The system of claim 18 , further comprising instructions for repeating the steps of receiving a first input from the user indicating a selection of a label, displaying a selection of text that may match a category of text corresponding to the user-selected label, and receiving a second input from the user to confirm a match and assign the label to the selection of text for one or more additional labels.

20. A non-transitory computer-readable storage medium storing one or more programs, the one or more programs comprising instructions, the instructions when executed by one or more processors of a computing platform, cause the computing platform to:

display an electronic document, or a page therefrom, on an electronic display device;

receive a first input from a user indicating a selection of a label that may be applicable to categories of text within the electronic document;

display a selection of text that may match a category of text corresponding to the user-selected label, wherein the selection of text is based on an artificial intelligence-based prediction;

receive a second input from the user to confirm a match between the selection of text and the category of text corresponding to the user-selected label, thereby assigning the label to the selection of text; and

store the assigned label, the selection of text, and a location of the selected text for one or more instances of selected text within the electronic document as an annotated electronic document;

wherein the annotated electronic document is used to train or retrain one or more term-based machine learning models, and wherein the one or more term-based models are used for document data extraction in a continuous learning-based document data extraction pipeline.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: PRICEWATERHOUSECOOPERS LLP
To: PWC PRODUCT SALES LLC
Reel/Frame 065532/0034 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 13, 2023
From: WILSON, JACOB T.; HARRINGTON, JOSEPH D.; SIGAMANI, VINSTON SUNDARA PANDIYAN; SANGHAVI, ABHISHEK; PILLAI, JAYAKUMAR; CUNNINGHAM, BENJAMIN; LEWIS, LINDSEY P.
To: PRICEWATERHOUSECOOPERS LLP
Reel/Frame 062741/0627 →
Continuity (2)
Continuation 17402338 · Aug 13, 2021
Related Publication 20230052372A1 · Feb 16, 2023
Cited By (2)
US 12,632,905 US 12,711,449