IP Library › Granted Patent US 12,405,970
Granted Patent B2
US 12,405,970 · App. 18/482,671 · Granted Sep 2, 2025

Multi-layer approach to improving generation of field extraction models

Inventors: Shalin Avlani (San Jose, CA); Rajesh M. Desai (San Jose, CA); Mayank Vipin Shah (San Jose, CA); Xiaoying Gao (San Jose, CA)
Assignee: International Business Machines Corporation
G06F16/285G06F16/24573
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,405,970
App. No.
18/482,671
Granted
Sep 2, 2025
Kind
B2
Abstract

A computer-implemented process for generating cluster templates used for creating extraction models includes the following operations. A plurality of training files associated with a selected class are received. An automated visual analysis is performed on each of the plurality of training files. An automated contextual analysis is performed on each of the plurality of training files. A first clustering of the plurality of training files into a first plurality of clusters using results from the automated visual analysis is performed. A second clustering of one of the plurality of clusters into a second plurality of clusters is performed using results from the automated contextual analysis. Cluster templates for the first and second plurality of clusters are generated.

Claims (89)

1. A computer-implemented method for generating cluster templates used for creating extraction models, comprising:

receiving a plurality of training files associated with a selected class;

performing an automated visual analysis on each of the plurality of training files;

performing an automated contextual analysis on each of the plurality of training files;

performing a first clustering of the plurality of training files into a first plurality of clusters using results from the automated visual analysis;

performing a second clustering of one of the first plurality of clusters into a second plurality of clusters using results from the automated contextual analysis; and

generating cluster templates for the first and second plurality of clusters, wherein

the first and the second plurality of clusters are clusters of the plurality of training files.

2. The method of claim 1 , further comprising:

receiving, for a first training file within a particular one of the first and second plurality of clusters, annotations;

automatically propagating the annotations to each of a plurality of other training files within the particular one of the first and second plurality of clusters; and

generating a cluster annotation template for the particular one of the first and second plurality of clusters using the annotations.

3. The method of claim 2 , further comprising:

receiving, from a user, an instruction to modify the annotations; and

modifying the cluster annotation template based upon the instruction.

4. The method of claim 2 , further comprising:

generating field extraction models for each of the first and second plurality of clusters using the cluster templates and cluster annotation templates.

5. The method of claim 1 , wherein

the automated visual analysis includes:

generating an image hash for each of the plurality of training files, and

extracting structural information for each of the plurality of training files.

6. The method of claim 5 , wherein

the image hash includes generating a shingle.

7. The method of claim 1 , wherein

the automated contextual analysis includes generating a Term Frequency-Inverse Document Frequency vector for each of the plurality of training files.

8. The method of claim 1 , further comprising:

receiving a document for data extraction;

identifying a classification of the document;

selecting a field extraction model based upon the classification; and

extracting data from the document using the field extraction model.

9. A computer hardware system for generating cluster templates used for creating extraction models, comprising:

a hardware processor configured to perform the following executable operations:

receiving a plurality of training files associated with a selected class;

performing an automated visual analysis on each of the plurality of training files;

performing an automated contextual analysis on each of the plurality of training files;

performing a first clustering of the plurality of training files into a first plurality of clusters using results from the automated visual analysis;

performing a second clustering of one of the first plurality of clusters into a second plurality of clusters using results from the automated contextual analysis; and

generating the cluster templates for the first and second plurality of clusters, wherein

the first and the second plurality of clusters are clusters of the plurality of training files.

10. The system of claim 9 , wherein the hardware processor is further configured to perform the following:

receiving, for a first training file within a particular one of the first and second plurality of clusters, annotations;

automatically propagating the annotations to each of a plurality of other training files within the particular one of the first and second plurality of clusters; and

generating a cluster annotation template for the particular one of the first and second plurality of clusters using the annotations.

11. The system of claim 10 , wherein the hardware processor is further configured to perform the following:

receiving, from a user, an instruction to modify the annotations; and

modifying the cluster annotation template based upon the instruction.

12. The system of claim 10 , wherein the hardware processor is further configured to perform the following:

generating field extraction models for each of the first and second plurality of clusters using the cluster templates and cluster annotation templates.

13. The system of claim 9 , wherein

the automated visual analysis includes:

generating an image hash for each of the plurality of training files, and

extracting structural information for each of the plurality of training files.

14. The system of claim 13 , wherein

the image hash includes generating a shingle.

15. The system of claim 9 , wherein

the automated contextual analysis includes generating a Term Frequency-Inverse Document Frequency vector for each of the plurality of training files.

16. The system of claim 9 , wherein the hardware processor is further configured to perform the following:

receiving a document for data extraction;

identifying a classification of the document;

selecting a field extraction model based upon the classification; and

extracting data from the document using the field extraction model.

17. A computer program product, comprising:

a computer readable storage medium having stored therein program code for generating cluster templates used for creating extraction models,

the program code, which when executed by a computer hardware system, cause the computer hardware system to perform:

receiving a plurality of training files associated with a selected class;

performing an automated visual analysis on each of the plurality of training files;

performing an automated contextual analysis on each of the plurality of training files;

performing a first clustering of the plurality of training files into a first plurality of clusters using results from the automated visual analysis;

performing a second clustering of one of the first plurality of clusters into a second plurality of clusters using results from the automated contextual analysis; and

generating the cluster templates for the first and second plurality of clusters, wherein

the first and the second plurality of clusters are clusters of the plurality of training files.

18. The computer program product of claim 17 , wherein the computer hardware system is further configured to perform the following:

receiving, for a first training file within a particular one of the first and second plurality of clusters, annotations;

automatically propagating the annotations to each of a plurality of other training files within the particular one of the first and second plurality of clusters;

generating a cluster annotation template for the particular one of the first and second plurality of clusters using the annotations;

receiving, from a user, an instruction to modify the annotations;

modifying the cluster annotation template based upon the instruction; and

generating field extraction models for each of the first and second plurality of clusters using the cluster templates and cluster annotation templates.

19. The computer program product of claim 17 , wherein

the automated visual analysis includes:

generating an image hash for each of the plurality of training files, and

extracting structural information for each of the plurality of training files,

the image hash includes generating a shingle, and

the automated contextual analysis includes generating a Term Frequency-Inverse Document Frequency vector for each of the plurality of training files.

20. The computer program product of claim 17 , wherein the program code further causes the computer hardware system to perform:

receiving a document for data extraction;

identifying a classification of the document;

selecting a field extraction model based upon the classification; and

extracting data from the document using the field extraction model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 6, 2023
From: AVLANI, SHALIN; DESAI, RAJESH M.; SHAH, MAYANK VIPIN; GAO, XIAOYING
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 065152/0353 →
Continuity (1)
Related Publication 20250117405A1 · Apr 10, 2025
References Cited (28)
US 8620842B1 · Cormack · 2013 [cited by examiner]
US 10657158B2 · Sheng et al. · 2020 [cited by applicant]
US 10747994B2 · Sridharan · 2020 [cited by applicant]
US 11087077B2 · Wheaton et al. · 2021 [cited by applicant]
US 11514238B2 · Begun et al. · 2022 [cited by applicant]
US 11514489B2 · Jiang et al. · 2022 [cited by applicant]
US 20080219495A1 · Hulten · 2008 [cited by examiner]
US 20160359740A1 · Parandehgheibi · 2016 [cited by examiner]
US 20180067910A1 · Alonso · 2018 [cited by examiner]
US 20190188312A1 · Pandit · 2019 [cited by examiner]
US 20190205195A1 · Tee · 2019 [cited by examiner]
US 20200334486A1 · Joseph · 2020 [cited by examiner]
US 20200401935A1 · Malhotra · 2020 [cited by examiner]
US 20210081613A1 · Begun et al. · 2021 [cited by applicant]
US 20210200937A1 · Wheaton · 2021 [cited by examiner]
US 20220215446A1 · Jiang et al. · 2022 [cited by applicant]
US 20240232539A1 · Venkateshwaran · 2024 [cited by examiner]
CN 102567711A · 2012 [cited by applicant]
CN 108108387B · 2021 [cited by applicant]
“A Method to Intelligently Recognize, Interact and Fulfill Business Requests in Emails” [online] IP.com Prior Art Database, Technical Disclosure No. IPCOM000253683D, Apr. 23, 2018, 12 pg. [cited by applicant]
“System and Method for Generatively Extracting Pertinent Structured Information from Documents using Contextual Information,” [online] IP.com Prior Art Database, Technical Disclosure No. IPCOM000269025D, Mar. 16, 2022, … [cited by applicant]
“Contract Information Extraction Based on Knowledge Graph,” [online] IP.com Prior Art Database, Technical Disclosure No. IPCOM000267254D, Oct. 11, 2021, 4 pg. [cited by applicant]
“System and Method for Extracting the Business Objects Utilizing Optical Region Recognition (ORR),” [online] IP.com Prior Art Database, Technical Disclosure No. IPCOM000266133D, Jun. 16, 2021, 8 pg. [cited by applicant]
Nakai, T. et al., “Accuracy improvement and objective evaluation of annotation extraction from printed documents,” In The Eighth IAPR International Workshop on Document Analysis Systems, Sep. 16, 2008 (pp. 329-336). IEE… [cited by applicant]
Narasimhan K, Yala A, Barzilay R. Improving information extraction by acquiring external evidence with reinforcement learning. arXiv preprint arXiv: 1603.07954. Mar. 25, 2016. [cited by applicant]
Skalický M, Šimsa Š, Uřičàř M, šulc M. Business Document Information Extraction: Towards Practical Benchmarks. InInternational Conference of the Cross-Language Evaluation Forum for European Languages Aug. 25, 2022 (pp. … [cited by applicant]
Grace Period Disclosure: “IBM Automation Document Processing,” [online] IBM.com, Automation Document Processing for ICP4BA, Jun. 30, 2023, [retrieved Aug. 28, 2023], retrieved from the Internet: <https://www.ibm.com/pro… [cited by applicant]
Mell, P. et al., The NIST Definition of Cloud Computing, National Institute of Standards and Technology, U.S. Dept. of Commerce, Special Publication 800-145, Sep. 2011, 7 pg. [cited by applicant]