IP Library › Granted Patent US 12,217,524
Granted Patent B2
US 12,217,524 · App. 17/850,618 · Granted Feb 4, 2025

Systems and methods for automated end-to-end text extraction of electronic documents

Inventors: Keerthan Ramnath (Chennai, IN); Punitha Chandrasekar (Bangalore, IN); Hui Su (West Roxbury, MA); Shyam Subramanian (Norwood, MA); Rachna Saxena (Bangalore, IN); Mohamed Mahdi Alouane (Toronto, CA); Vinay Iyengar (Westwood, MA)
Assignee: FMR LLC
G06V30/414G06F40/232G06F40/263G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,524
App. No.
17/850,618
Filed
Jun 27, 2022
Granted
Feb 4, 2025
Kind
B2
Art Unit
2677
USPC
382/181
Abstract

Systems and methods for extracting data from electronic documents using optical character recognition (OCR) and non-OCR based text extraction. A server computing device initiates non-OCR based text extraction for each page of an electronic document. The server calculates a document text coverage percentage corresponding to the non-OCR based text extraction for the whole document and, in response to determining that the document text coverage percentage is below a first threshold, initiates OCR for the document. The server calculates a page text coverage percentage corresponding to the non-OCR based text extraction for one or more pages of the electronic document and, in response to determining that the page text coverage percentage is below a second threshold, initiates OCR for the pages. The server combines first text extracted from the electronic document using non-OCR based text extraction and second text extracted from the electronic document using OCR.

Claims (50)

1. A computerized method for extracting data from electronic documents using optical character recognition (OCR) and non-OCR based text extraction, the method comprising:

initiating, by a server computing device, non-OCR based text extraction for each of a plurality of pages of an electronic document;

calculating, by the server computing device, a document text coverage percentage corresponding to the non-OCR based text extraction for the electronic document as a whole;

in response to determining that the document text coverage percentage for the electronic document as a whole is below a first threshold, initiating, by the server computing device, OCR for the electronic document as a whole;

calculating, by the server computing device, a page text coverage percentage corresponding to the non-OCR based text extraction for one or more pages of the electronic document;

in response to determining that the page text coverage percentage for one or more pages of the electronic document is below a second threshold, initiating, by the server computing device, OCR for the one or more pages of the electronic document; and

combining, by the server computing device, first text extracted from the electronic document using non-OCR based text extraction and second text extracted from the electronic document using OCR.

2. The computerized method of claim 1 , wherein the server computing device is further configured to:

generate a data structure comprising the combined first text and second text extracted from the electronic document; and

store the data structure in a database.

3. The computerized method of claim 1 , wherein the server computing device determines, for each of the plurality of pages of the electronic document, whether the page of the plurality of pages of the electronic document comprises a non-searchable image.

4. The computerized method of claim 3 , wherein the server computing device, in response to determining that the page of the plurality of pages of the electronic document comprises the non-searchable image, initiates optical character recognition for the page of the plurality of pages of the electronic document.

5. The computerized method of claim 1 , wherein the server computing device calculates the accuracy of the non-OCR based text extraction by calculating a percentage of extracted words that are English words.

6. The computerized method of claim 1 , wherein the server computing device removes one or more watermarks from the electronic document prior to initiating optical character recognition.

7. The computerized method of claim 1 , wherein the server computing device corrects the combined first text and second text extracted from the electronic document by replacing misspelled words.

8. The computerized method of claim 1 , wherein the first threshold comprises 90% of the plurality of pages of the electronic document.

9. The computerized method of claim 1 , wherein the second threshold comprises an image coverage of 50%.

10. The computerized method of claim 1 , wherein the third threshold comprises an accuracy of 85%.

11. A system for extracting data from electronic documents using non-optical character recognition (OCR) based text extraction and OCR, the system comprising a server computing device communicatively coupled to a database over a network, the server computing device having a memory for storing computer-executable instructions and a processor that executes the computer-executable instructions to:

initiate non-OCR based text extraction for each of a plurality of pages of an electronic document;

determine an amount of pages of the plurality of pages comprising at least one image;

receive a runtime exception during non-OCR based text extraction or determine that the amount of pages exceeds a first threshold;

initiate optical character recognition for each of the plurality of pages of the electronic document;

determine an image coverage percentage for each of the plurality of pages of the electronic document;

determine that the image coverage percentage exceeds a second threshold for a page of the plurality of pages of the electronic document;

initiate optical character recognition for the page of the plurality of pages of the electronic document;

calculate an accuracy of the non-OCR based text extraction performed for each of the plurality of pages of the electronic document;

determine that the accuracy of the non-OCR based text extraction is less than a third threshold for a page of the plurality of pages of the electronic document;

initiate optical character recognition for the page of the plurality of pages of the electronic document; and

combine first text extracted from the electronic document using non-OCR based text extraction and second text extracted from the electronic document using optical character recognition.

12. The system of claim 11 , wherein the server computing device determines, for each of the plurality of pages of the electronic document, whether the page of the plurality of pages of the electronic document comprises a non-searchable image.

13. The system of claim 12 , wherein the server computing device, in response to determining that the page of the plurality of pages of the electronic document comprises the non-searchable image, initiates optical character recognition for the page of the plurality of pages of the electronic document.

14. The system of claim 11 , wherein the server computing device calculates the accuracy of the non-OCR based text extraction by calculating a percentage of extracted words that are English words.

15. The system of claim 11 , wherein the server computing device removes one or more watermarks from the electronic document prior to initiating optical character recognition.

16. The system of claim 11 , wherein the server computing device corrects the combined first text and second text extracted from the electronic document by replacing misspelled words.

17. The system of claim 11 , wherein the first threshold comprises 90% of the plurality of pages of the electronic document.

18. The system of claim 11 , wherein the second threshold comprises an image coverage of 50%.

19. The system of claim 11 , wherein the third threshold comprises an accuracy of 85%.

20. A computerized method for extracting data from electronic documents using non-optical character recognition (OCR) based text extraction and OCR, the method comprising:

initiating, by a server computing device, non-OCR based text extraction for each of a plurality of pages of an electronic document;

determining, by the server computing device, an amount of pages of the plurality of pages comprising at least one image;

receiving, by the server computing device, a runtime exception during non-OCR based text extraction or determining, by the server computing device, that the amount of pages exceeds a first threshold;

initiating, by the server computing device, optical character recognition for each of the plurality of pages of the electronic document;

determining, by the server computing device, an image coverage percentage for each of the plurality of pages of the electronic document;

determining, by the server computing device, that the image coverage percentage exceeds a second threshold for a page of the plurality of pages of the electronic document;

initiating, by the server computing device, optical character recognition for the page of the plurality of pages of the electronic document;

calculating, by the server computing device, an accuracy of the non-OCR based text extraction performed for each of the plurality of pages of the electronic document;

determining, by the server computing device, that the accuracy of the non-OCR based text extraction is less than a third threshold for a page of the plurality of pages of the electronic document;

initiating, by the server computing device, optical character recognition for the page of the plurality of pages of the electronic document; and

combining, by the server computing device, first text extracted from the electronic document using non-OCR based text extraction and second text extracted from the electronic document using optical character recognition.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2022
From: RAMNATH, KEERTHAN; CHANDRASEKAR, PUNITHA; SU, HUI; SUBRAMANIAN, SHYAM; SAXENA, RACHNA; ALOUANE, MOHAMED MAHDI; IYENGAR, VINAY
To: FMR LLC
Reel/Frame 061669/0322 →
Continuity (1)
Related Publication 20230419711A1 · Dec 28, 2023
References Cited (10)
US 8503781B2 · Chen et al. · 2013 [cited by applicant]
US 10489682B1 · Kumar et al. · 2019 [cited by applicant]
US 20020012465A1 · Fujimoto · 2002 [cited by examiner]
US 20080170810A1 · Wu · 2008 [cited by examiner]
US 20170124413A1 · Deng · 2017 [cited by examiner]
US 20220067275A1 · Zeng · 2022 [cited by examiner]
Junker et al., “Evaluating OCR and non-OCR text representations for learning document classifiers,” Proceedings of the Fourth International Conference on Document Analysis and Recognition, Ulm, Germany, 1997, pp. 1060-1… [cited by examiner]
Tan et al., “Imaged document text retrieval without OCR,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 24, No. 6, pp. 838-844, Jun. 2002, doi: 10.1109/TPAMI.2002.1008389. [cited by examiner]
H. Bast and C. Korzen, “A Benchmark and Evaluation for Text Extraction from PDF,” JCDL '17, Toronto, Ontario, Canada, 2017, 10 pages. [cited by applicant]
C. Yu et al., “Extracting Body Text from Academic PDF Documents for Text Mining,” arXiv:2010.12647v1 [cs.IR] Oct. 23, 2020, available at https://arxiv.org/pdf/2010.12647v1, 8 pages. [cited by applicant]