IP Library › Granted Patent US 12,511,346
Granted Patent B2
US 12,511,346 · App. 17/501,681 · Granted Dec 30, 2025

Systems and methods for enabling relevant data to be extracted from a plurality of documents

Inventors: Ahmed Farouk Shaaban (South Barrington, IL); Venkat Thandra (South Barrington, IL); Dino Eliopulos (South Barrington, IL); Andrew Kenneth Blazaitis (Rochester, MN); Kennedy Muthukrishnan (South Barrington, IL)
Assignee: Fulcrum Global Technologies Inc.
G06F18/2148G06F16/93G06V30/19173G06V30/414G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,346
App. No.
17/501,681
Granted
Dec 30, 2025
Kind
B2
Abstract

Systems and methods for enabling target data to be extracted from documents are disclosed herein. In an embodiment, a method of enabling target data to be extracted from documents includes accessing a database including a plurality of documents including target data, for each of multiple of the documents, creating a region tensor based on extracted text including the target data, for each of the multiple of the documents, creating a label tensor based on an area including the target data, and using the region tensor and the label tensor, training an extraction algorithm to extract the target data from additional documents.

Claims (50)

1 . A method for enabling target data to be extracted from documents, the method comprising:

accessing a database including a plurality of documents including target data;

converting each of multiple of the documents into one or more images not readable by a computer;

using the one or more images for each of the multiple of the documents, assigning a label to a target area of a document corresponding to a target category and determining label coordinate data based on coordinates of the labeled target area in a region label extraction step;

using the one or more images for each of the multiple of the documents, creating a region tensor;

using the one or more images for each of the multiple of the documents, performing a text extraction to create extracted text;

for each of the multiple of the documents, adjusting the region tensor based on an identified fixed region surrounding the extracted text of the document including the target data in a region extraction step;

for each of the multiple of the documents, creating a label tensor based on an overlapping region of the label coordinate data from the labeled target area of the target category and text coordinate data related to the extracted text in a label merging step that merges the results of the region label extraction step and the region extraction step; and

using the region tensor and the label tensor, training an extraction algorithm to extract the target data from additional documents.

2 . The method of claim 1 , comprising

enabling extraction of the target data from the additional documents using the extraction algorithm.

3 . The method of claim 1 , wherein

at least one of the region tensor and the label tensor includes a data matrix.

4 . The method of claim 1 , wherein

creating the region tensor includes identifying the fixed region surrounding the extracted text and creating the region tensor based on the fixed region.

5 . The method of claim 1 , wherein

creating the label tensor includes converting the target area to the label coordinate data.

6 . The method of claim 1 , comprising

training the extraction algorithm to extract the target data from the additional documents by outputting new label tensors corresponding to the additional documents based on new inputted region tensors corresponding to the additional documents.

7 . A memory storing instructions configured to cause a processor to perform the method of claim 1 .

8 . A method for enabling target data to be extracted from documents, the method comprising:

accessing a database including a plurality of documents including target data;

converting each of multiple of the documents into one or more images not readable by a computer;

using the one or more images for each of the multiple of the documents, assigning a label to a target area of a document in a region label assignment step;

for each of the multiple of the documents, determining first coordinate data based on coordinates of the labeled target area in a region label extraction step;

using the one or more images for each of the multiple of the documents, extracting target text of the document including the target data in a text extraction step;

for each of the multiple of the documents, preparing a region tensor based on an identified fixed region surrounding the target text in a region extraction step;

for each of the multiple of the documents, creating a label tensor based on an overlapping region of the first coordinate data related to the label and second coordinate data related to the extracted target text in a label merging step that merges results of the region label extraction step and the region extraction step; and

using the region tensor and the label tensor, training an extraction algorithm to extract the target data from additional documents.

9 . The method of claim 8 , comprising

enabling extraction of the target data from the additional documents using the extraction algorithm.

10 . The method of claim 8 , wherein

the region tensor includes a data matrix.

11 . The method of claim 8 , comprising

creating the region tensor using third coordinate data corresponding to the fixed region.

12 . A memory storing instructions configured to cause a processor to perform the method of claim 8 .

13 . A method for enabling target data to be extracted from documents, the method comprising:

accessing a database including a plurality of documents including target data;

converting each of multiple of the documents into one or more images not readable by a computer;

using the one or more images for each of the multiple of the documents, assigning a label to a target area of a document and converting the target area to first coordinate data in one or more first coordinate steps;

using the one or more images for each of the multiple of the documents, determining second coordinate data for a piece of extracted text of the document in one or more a second coordinate steps;

for each of the multiple of the documents, identifying an overlapping region of the first coordinate data and the second coordinate data to create a label tensor based on the first coordinate data and the second coordinate data in a label merging step that merges results of the one or more first coordinate steps and the one or more second coordinate steps; and

using the label tensor, training an extraction algorithm to extract the target data from additional documents.

14 . The method of claim 13 , comprising

enabling extraction of the target data from the additional documents using the extraction algorithm.

15 . The method of claim 13 , wherein

the label tensor includes a data matrix.

16 . The method of claim 13 , comprising

training the extraction algorithm to extract the target data from the additional documents by outputting new label tensors corresponding to the additional documents.

17 . A memory storing instructions configured to cause a processor to perform the method of claim 13 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2025
From: SHAABAN, AHMED FAROUK; THANDRA, VENKAT; ELIOPULOS, DINO; BLAZAITIS, ANDREW KENNETH; MUTHUKRISHNAN, KENNEDY
To: FULCRUM GLOBAL TECHNOLOGIES INC.
Reel/Frame 073071/0212 →
Continuity (2)
Provisional Application 63093425 · Oct 19, 2020
Related Publication 20220121881A1 · Apr 21, 2022
References Cited (53)
US 6597808B1 · Guo · 2003 [cited by examiner]
US 10796383B2 · Shaaban et al. · 2020 [cited by applicant]
US 10896470B2 · Shaaban et al. · 2021 [cited by applicant]
US 11138566B2 · Shaaban et al. · 2021 [cited by applicant]
US 11144583B2 · Shaaban et al. · 2021 [cited by applicant]
US 20030177000A1 · Mao et al. · 2003 [cited by applicant]
US 20110255782A1 · Welling · 2011 [cited by examiner]
US 20120330977A1 · Inagaki · 2012 [cited by applicant]
US 20160125511A1 · Shaaban et al. · 2016 [cited by applicant]
US 20160140528A1 · Shaaban et al. · 2016 [cited by applicant]
US 20160203530A1 · Shaaban et al. · 2016 [cited by applicant]
US 20160210572A1 · Shaaban et al. · 2016 [cited by applicant]
US 20170004550A1 · Shaaban et al. · 2017 [cited by applicant]
US 20170109834A1 · Shaaban et al. · 2017 [cited by applicant]
US 20190066057A1 · Shaaban et al. · 2019 [cited by applicant]
US 20190156385A1 · Shaaban et al. · 2019 [cited by applicant]
US 20190279161A1 · Shaaban et al. · 2019 [cited by applicant]
US 20190279179A1 · Shaaban et al. · 2019 [cited by applicant]
US 20210012290A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210042378A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210043100A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210049535A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210049714A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210075615A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210100375A1 · Shaaban et al. · 2021 [cited by applicant]
US 20210201018A1 · Patel · 2021 [cited by examiner]
US 20210319039A1 · Gerber, Jr. · 2021 [cited by examiner]
US 20220121881A1 · Shaaban et al. · 2022 [cited by applicant]
CN 110427488A · 2019 [cited by examiner]
WO 2016004123A1 · 2016 [cited by applicant]
WO 2016004126A1 · 2016 [cited by applicant]
WO 2016004127A1 · 2016 [cited by applicant]
WO 2016004128A1 · 2016 [cited by applicant]
WO 2016004129A1 · 2016 [cited by applicant]
WO 2016004132A1 · 2016 [cited by applicant]
WO 2016004133A1 · 2016 [cited by applicant]
WO 2016004135A1 · 2016 [cited by applicant]
WO 2016004138A2 · 2016 [cited by applicant]
WO 2016004445A1 · 2016 [cited by applicant]
WO 2016007334A1 · 2016 [cited by applicant]
WO 2016018632A1 · 2016 [cited by applicant]
WO 2016029194A1 · 2016 [cited by applicant]
WO 2018045126A1 · 2018 [cited by applicant]
WO 2019036310A1 · 2019 [cited by applicant]
WO 2022086813A1 · 2022 [cited by applicant]
Zheng, L., Wang, S., Guo, P., Liang, H., & Tian, Q. (2015). Tensor index for large scale image retrieval. Multimedia Systems, 21, 569-579. (Year: 2015). [cited by examiner]
Majumder, B. P., Potti, N., Tata, S., Wendt, J. B., Zhao, Q., & Najork, M. (Jul. 2020). Representation learning for information extraction from form-like documents. In proceedings of the 58th annual meeting of the Assoc… [cited by examiner]
Paliwal, S. S., Vishwanath, D., Rahul, R., Sharma, M., & Vig, L. (Sep. 2019). Tablenet: Deep learning model for end-to-end table detection and tabular data extraction from scanned document images. In 2019 international … [cited by examiner]
Rashid, S. F., Akmal, A., Adnan, M., Aslam, A. A., & Dengel, A. (Nov. 2017). Table recognition in heterogeneous documents using machine learning. In 2017 14th IAPR International conference on document analysis and recog… [cited by examiner]
A Search Report in the corresponding International Patent Application No. PCT/US21/55198, dated Jan. 25, 2022. [cited by applicant]
An extended European Search Report in the corresponding European Patent Application No. 21883603.9, dated Oct. 22, 2024. [cited by applicant]
Jin Xiongnan Wnkim@ICL Yonsei AC KR et al: “Learning Region Similarity over Spatial Knowledge Graphs with Hierarchical Types and Semantic Relations”, Proceedings of the 28th ACM Joint Meeting on European Software Engine… [cited by applicant]
Cui Yuba et al: 11 3D Object Tracking with Transformer, Oct. 28, 2021 (Oct. 28, 2021), XP093210277, Retrieved from the Internet: URL:https://arxiv.org/pdf/2110.14921. [cited by applicant]