IP Library › Granted Patent US 12,259,920
Granted Patent B1
US 12,259,920 · App. 18/462,765 · Granted Mar 25, 2025

Dynamic optimization of key value pair extractors for document data extraction

Inventors: Ang Yi (Beijing, CN); Jing Zhang (Beijing, CN); Hai Cheng Wang (Beijing, CN); Jun Hong Zhao (Beijing, CN); Yang Zhong Li (Beijing, CN); Rajesh M. Desai (San Jose, CA); Xue Lan Zhang (Beijing, CN)
Assignee: International Business Machines Corporation
G06F16/383G06F16/316
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,259,920
App. No.
18/462,765
Granted
Mar 25, 2025
Kind
B1
Abstract

Disclosed embodiments provide techniques for monitoring and evaluating the effectiveness of key value pairs (KVPs) used in a document processing system. In embodiments, KVPs are obtained from multiple extractors of a document processing system. A score is computed for the KVPs by computing an effectiveness metric for each KVP from the multiple KVPs. In response to the computed score being below a predetermined threshold, a model retraining process is performed to generate a new set of KVP extractors, and provide the new set of KVPs to the document processing system.

Claims (46)

1. A computer-implemented method for optimized document processing, comprising:

obtaining a plurality of key value pairs (KVPs) from a plurality of KVP extractors of a document processing system;

computing a score for the plurality of KVPs by computing an effectiveness metric for each KVP from the plurality of KVPs;

wherein the effectiveness metric factors an importance metric associated with the document processing system;

in response to the computed score being below a predetermined threshold, performing a model retraining process to generate a new set of KVP extractors;

providing the new set of KVP extractors to the document processing system; and

dynamically adjusting the effectiveness metric based on the retraining via the document processing system.

2. The method of claim 1 , wherein the effectiveness metrics are based on a level of user correction applied to the plurality of KVPs.

3. The method of claim 1 , further comprising performing an extractor analysis, based on the effectiveness metrics.

4. The method of claim 3 , wherein the extractor analysis utilizes at least one classifier selected from the group consisting of: Forest Classifier, Support Vector Machine (SVM) classifier, logistic classifier, and gradient boost classifier.

5. The method of claim 4 , wherein the gradient boost classifier includes at least one classifier selected from the group consisting of XGB classifier, and LightGBM classifier.

6. The method of claim 1 , wherein the model retraining process includes:

generating a plurality of extractor combinations; and

ranking each extractor combination of the plurality of extractor combinations based on the importance metric.

7. The method of claim 6 , wherein the importance metric is computed using at least one function selected from the group consisting of: confusion matrix, F1 score, Receiver Operator Characteristic curve, and precision-recall curve.

8. The method of claim 1 , further comprising providing a KVP ranking model to the document processing system.

9. An electronic computation device comprising:

a processor;

a memory coupled to the processor, the memory containing instructions, that when executed by the processor, cause the electronic computation device to:

obtain a plurality of key value pairs (KVPs) from a plurality of KVP extractors of a document processing system;

compute a score for the plurality of KVPs by computing an effectiveness metric for each KVP from the plurality of KVPs;

wherein the effectiveness metric factors an importance metric associated with the document processing system;

in response to the computed score being below a predetermined threshold, perform a model retraining process to generate a new set of KVP extractors;

provide the new set of KVP extractors to the document processing system; and

dynamically adjusting the effectiveness metric based on the retraining via the document processing system.

10. The electronic computation device of claim 9 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to compute the effectiveness metrics based on a level of user correction applied to the plurality of KVPs.

11. The electronic computation device of claim 9 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to perform an extractor analysis, based on the effectiveness metrics.

12. The electronic computation device of claim 11 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to compute the extractor analysis utilizing at least one classifier from the group consisting of: Forest Classifier, Support Vector Machine (SVM) classifier, logistic classifier, and gradient boost classifier.

13. The electronic computation device of claim 12 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to compute the extractor analysis utilizing a gradient boost classifier that includes at least one of XGB classifier, and LightGBM classifier.

14. The electronic computation device of claim 9 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to:

generate a plurality of extractor combinations; and

rank each extractor combination of the plurality of extractor combinations based on the importance metric.

15. The electronic computation device of claim 14 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to compute the importance metric using at least one of, a confusion matrix, F1 score, Receiver Operator Characteristic curve, and precision-recall curve.

16. The electronic computation device of claim 9 , wherein the memory further comprises instructions, that when executed by the processor, cause the electronic computation device to provide a KVP ranking model to the document processing system.

17. A computer program product for an electronic computation device comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the electronic computation device to:

obtain a plurality of key value pairs (KVPs) from a plurality of KVP extractors of a document processing system;

compute a score for the plurality of KVPs by computing an effectiveness metric for each KVP from the plurality of KVPs;

wherein the effectiveness metric factors an importance metric associated with the document processing system;

in response to the computed score being below a predetermined threshold, perform a model retraining process to generate a new set of KVP extractors; and

provide the new set of KVP extractors to the document processing system; and

dynamically adjust the effectiveness metric based on the retraining via the document processing system.

18. The computer program product of claim 17 , wherein the computer readable storage medium further comprises program instructions, that when executed by the processor, cause the electronic computation device to compute the effectiveness metrics are based on a level of user correction applied to the plurality of KVPs.

19. The computer program product of claim 17 , wherein the computer readable storage medium further comprises program instructions, that when executed by the processor, cause the electronic computation device to:

generate a plurality of extractor combinations; and

rank each extractor combination of the plurality of extractor combinations based on the importance metric.

20. The computer program product of claim 19 , wherein the computer readable storage medium further comprises program instructions, that when executed by the processor, cause the electronic computation device to compute the importance metric using at least one function selected from the group consisting of: confusion matrix, F1 score, Receiver Operator Characteristic curve, and precision-recall curve.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2023
From: YI, ANG; ZHANG, JING; WANG, HAI CHENG; ZHAO, JUN HONG; LI, YANG ZHONG; DESAI, RAJESH M; ZHANG, XUE LAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064829/0628 →
References Cited (18)
US 9292579B2 · Madhani et al. · 2016 [cited by applicant]
US 10853638B2 · Mukhopadhyay et al. · 2020 [cited by applicant]
US 10896357B1 · Corcoran et al. · 2021 [cited by applicant]
US 20150127659A1 · Madhani et al. · 2015 [cited by applicant]
US 20180113920A1 · Klein · 2018 [cited by examiner]
US 20180114060A1 · Lozano · 2018 [cited by examiner]
US 20180232204A1 · Ghatage · 2018 [cited by examiner]
US 20190205636A1 · Saraswat · 2019 [cited by examiner]
US 20200074169A1 · Mukhopadhyay et al. · 2020 [cited by applicant]
US 20200279017A1 · Norton · 2020 [cited by examiner]
US 20210166074A1 · Tecuci et al. · 2021 [cited by applicant]
US 20210350252A1 · Alexander · 2021 [cited by examiner]
US 20220207268A1 · Gligan · 2022 [cited by examiner]
US 20220351088A1 · Kumar · 2022 [cited by examiner]
CN 110889310B · 2023 [cited by applicant]
Daniel Akinbade et al., “An Adaptive Thresholding Algorithm-Based Optical Character Recognition System for Information Extraction in Complex Images”, Journal of Computer Science 2020, pp. 784-801. [cited by applicant]
Suzan Verberne et al., “Evaluation and analysis of term scoring methods for term extraction”, Information Retrieval Journal, 2016, pp. 510-545. [cited by applicant]
Henning Wachsmuth et al., “Learning Efficient Information Extraction on Heterogeneous Texts”, International Joint Conference on Natural Language Processing, Oct. 2013, pp. 534-542. [cited by applicant]