IP Library › Granted Patent US 11,194,841
Granted Patent B2
US 11,194,841 · App. 16/698,986 · Granted Dec 7, 2021

Value classification by contextual classification of similar values in additional documents

Inventors: Sigal Asaf (Zichron Yaakov, IL); Ariel Farkash (Shinshit, IL); Micha Gideon Moffie (Zichron Yaakov, IL)
Assignee: International Business Machines Corporation
G06F16/285G06F16/93G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,194,841
App. No.
16/698,986
Granted
Dec 7, 2021
Kind
B2
Abstract

Automated classification, by: Obtaining an examined document having an examined value appearing therein. Identifying: a location in the examined document at which the examined value appears, and a structure of the examined value. Identifying additional documents of a same type as the examined document, in which values having a same structure as the examined value appear at a same location as in the examined document. Applying a classifier to the examined value and the values in the additional documents, to output a single class to which the examined value and the values in the additional documents belong.

Claims (56)

1. A method comprising:

automatically obtaining an examined document having an examined value appearing therein;

automatically identifying: a location in the examined document at which the examined value appears, and a structure of the examined value;

automatically identifying additional documents of a same type as the examined document, in which values having a same structure as the examined value appear at a same location as in the examined document; and

automatically applying a classifier to the examined value and the values in the additional documents, to output a single class to which the examined value and the values in the additional documents belong, wherein;

the classifier is a machine learning classifier whose application produces classification scores for multiple classes,

any value belonging to each of the multiple classes is subject to one or more content restrictions,

the method further comprises, prior to outputting the single class: automatically adjusting the classification scores based on how they deviate from theoretical classification scores statistically expected for random values that are subject to the one or more content restrictions, and

the single class outputted is one of the multiple classes having a highest one of the adjusted classification scores.

2. The method according to claim 1 , wherein said automatic identification of the additional documents of the same type comprises identifying documents which are stored, in a document repository, under a same category or folder as the examined document.

3. The method according to claim 1 , wherein the examined value, the values in the additional documents, and the value belonging to each of the multiple classes, are textual, numeric, or alphanumeric.

4. The method according to claim 3 , wherein the one or more content restrictions are selected from the group consisting of:

adherence of a number included in the respective value to a mathematical rule; and

adherence of a text included in the respective value to certain structure.

5. The method according to claim 1 , performed by at least one hardware processor.

6. A system comprising:

(a) at least one hardware processor; and

(b) a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by said at least one hardware processor to:

automatically obtain an examined document having an examined value appearing therein,

automatically identify: a location in the examined document at which the examined value appears, and a structure of the examined value,

automatically identify additional documents of a same type as the examined document, in which values having a same structure as the examined value appear at a same location as in the examined document, and

automatically apply a classifier to the examined value and the values in the additional documents, to output a single class to which the examined value and the values in the additional documents belong, wherein:

the classifier is a machine learning classifier whose application produces classification scores for multiple classes,

any value belonging to each of the multiple classes is subject to one or more content restrictions,

the program code is further executable, prior to outputting the single class, to: automatically adjust the classification scores based on how they deviate from theoretical classification scores statistically expected for random values that are subject to the one or more content restrictions, and

the single class outputted is one of the multiple classes having a highest one of the adjusted classification scores.

7. The system according to claim 6 , wherein said automatic identification of the additional documents of the same type comprises identifying documents which are stored, in a document repository, under a same category or folder as the examined document.

8. The system according to claim 6 , wherein the examined value, the values in the additional documents, and the value belonging to each of the multiple classes, are textual, numeric, or alphanumeric.

9. The system according to claim 8 , wherein the one or more content restrictions are selected from the group consisting of:

adherence of a number included in the respective value to a mathematical rule; and

adherence of a text included in the respective value to certain structure.

10. A method comprising:

automatically obtaining an examined document having an examined value appearing therein;

automatically identifying: a location in the examined document at which the examined value appears, and a structure of the examined value;

automatically identifying additional documents of a same type as the examined document, in which values having a same structure as the examined value appear at a same location as in the examined document; and

automatically applying a classifier to the examined value and the values in the additional documents, to output a single class to which the examined value and the values in the additional documents belong, wherein the classifier is a statistical analysis classifier configured to:

automatically determine multiple possible classes of the examined value and the values in the additional documents, wherein each of the multiple possible classes has a known distribution of values,

automatically calculate a distribution of the examined value together with the values in the additional documents, and

automatically calculate divergence of the calculated distribution from each of the known distributions of the multiple possible classes,

wherein the single class outputted is one of the multiple possible classes whose known distribution has the least divergence from the calculated distribution.

11. The method according to claim 10 , wherein the divergence is calculated as a Kullback-Leibler divergence.

12. The method according to claim 10 , wherein the examined value, the values in the additional documents, and the value belonging to each of the multiple classes, are textual, numeric, or alphanumeric.

13. The method according to claim 10 , performed by at least one hardware processor.

14. A system comprising:

(a) at least one hardware processor; and

(b) a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by said at least one hardware processor to:

automatically obtain an examined document having an examined value appearing therein,

automatically identify: a location in the examined document at which the examined value appears, and a structure of the examined value,

automatically identify additional documents of a same type as the examined document, in which values having a same structure as the examined value appear at a same location as in the examined document, and

automatically apply a classifier to the examined value and the values in the additional documents, to output a single class to which the examined value and the values in the additional documents belong, wherein the classifier is a statistical analysis classifier configured to:

automatically determine multiple possible classes of the examined value and the values in the additional documents, wherein each of the multiple possible classes has a known distribution of values,

automatically calculate a distribution of the examined value together with the values in the additional documents, and

automatically calculate divergence of the calculated distribution from each of the known distributions of the multiple possible classes,

wherein the single class outputted is one of the multiple possible classes whose known distribution has the least divergence from the calculated distribution.

15. The system according to claim 14 , wherein the divergence is calculated as a Kullback-Leibler divergence.

16. The system according to claim 14 , wherein the examined value, the values in the additional documents, and the value belonging to each of the multiple classes, are textual, numeric, or alphanumeric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 28, 2019
From: ASAF, SIGAL; FARKASH, ARIEL; MOFFIE, MICHA GIDEON
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 051134/0759 →
Continuity (1)
Related Publication 20210165807A1 · Jun 3, 2021