IP Library Granted Patent US 9,633,257
Granted Patent B2
US 9,633,257 · App. 14/314,892 · Granted Apr 25, 2017

Method and system of pre-analysis and automated classification of documents

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,633,257
App. No.
14/314,892
Granted
Apr 25, 2017
Kind
B2
Abstract

Automatic classification of different types of documents is disclosed. An image of a form or document is captured. The document is assigned to one or more type definitions by identifying one or more objects within the image of the document. A matching model is selected via identification of the document image. In the case of multiple identifications, a profound analysis of the document type is performed—either automatically or manually. An automatic classifier may be trained with document samples of each of a plurality of document classes or document types where the types are known in advance or a system of classes may be formed automatically without a priori information about types of samples. An automatic classifier determines possible features and calculates a range of feature values and possible other feature parameters for each type or class of document. A decision tree, based on rules specified by a user, may be used for classifying documents. Processing, such as optical character recognition (OCR), may be used in the classification process.

Claims (75)

1. A non-transitory machine-readable storage medium having instructions that, when executed by a processing device, cause the processing device to perform operations comprising:

identifying a plurality of document features in a document image;

correlating each of the plurality of document features with one or more of different document classes;

selecting a first group of document features and a second group of document features from the plurality of document features based on document feature types of the plurality of document features;

for the first group of document features, generating a first decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the first group;

for the second group of document features, generating a second decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the second group; and

associating the document image with one of the different document classes based at least in part on the first decision tree and the second decision tree.

2. The non-transitory machine-readable storage medium of claim 1 , wherein:

the feature is a feature that was previously determined to be a feature capable of distinguishing the document image or a document that comprises the document image, and

the feature was previously identified by analysis of a plurality of training documents each having at least one unique feature different from at least one of the other training documents.

3. The non-transitory machine-readable storage medium of claim 2 , wherein:

the feature is associated with a feature type,

identifying the feature in the plurality of training documents further comprises identifying a corresponding feature type for the feature, and

a feature type decision tree is generated that comprises the feature type.

4. The non-transitory machine-readable storage medium of claim 3 , the operations further comprising generating a second node corresponding to the class associated with the decision tree.

5. The non-transitory machine-readable storage medium of claim 3 , wherein the feature type is: a raster, a title, an image object, a text string, a word, a unique mark, a unique character, a numeric code, or a non-human readable marking.

6. The non-transitory machine-readable storage medium of claim 2 , wherein:

the feature is predefined, and

the identifying the feature in the document image further comprises performing optical character recognition (OCR) on the feature in the document image.

7. The non-transitory machine-readable storage medium of claim 2 , wherein:

the associating the document image with the class further comprises:

determining a value from the document image, and

determining a reliability index of the decision tree, wherein the reliability index is based on the identified feature of the document image.

8. The non-transitory machine-readable storage medium of claim 2 , wherein the operations further comprise identifying an object in the document image prior to the identifying the feature in the document image, wherein the identifying the feature in the document image comprises identifying the feature in the object.

9. A method comprising:

identifying a plurality of document features in a document image;

correlating each of the plurality of document features with one or more of different document classes;

selecting a first group of document features and a second group of document features from the plurality of document features based on document feature types of the plurality of document features;

for the first group of document features, generating a first decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the first group;

for the second group of document features, generating a second decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the second group; and

associating the document image with one of the different document classes based at least in part on the first decision tree and the second decision tree.

10. The method of claim 9 , wherein:

the feature is a feature that was previously determined to be a feature capable of distinguishing the document image or a document that comprises the document image, and

the feature was previously identified by analysis of a plurality of training documents each having at least one unique feature different from at least one of the other training documents.

11. The method of claim 10 , wherein:

the feature is associated with a feature type,

identifying the feature in the plurality of training documents further comprises identifying a corresponding feature type for the feature, and

a feature type decision tree is generated that comprises the feature type.

12. The method of claim 11 , further comprising generating a second node corresponding to the class associated with the first decision tree.

13. The method of claim 11 , wherein the feature type is: a raster, a title, an image object, a text string, a word, a unique mark, a unique character, a numeric code, or a non-human readable marking.

14. The method of claim 10 , wherein:

the feature is predefined, and

the identifying the feature in the document image further comprises performing optical character recognition (OCR) on the feature in the document image.

15. The method of claim 10 , wherein:

the associating the document image with the class further comprises:

determining a value from the document image, and

determining a reliability index of the first decision tree, wherein the reliability index is based on the feature of the document image.

16. The method of claim 9 , wherein the method further comprises:

identifying an object in the document image prior to the identifying the feature in the document image, wherein the identifying the feature in the document image comprises identifying the feature in the document object.

17. A system comprising:

a storage device;

a processor coupled to the storage device, the processor to:

identifying a plurality of document features in a document image;

correlate each of the plurality of document features with one or more of different document classes;

selecting a first group of document features and a second group of document features from the plurality of document features based on document feature types of the plurality of document features;

for the first group of document features, generate a first decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the first group;

for the second group of document features, generate a second decision tree that includes one or more nodes corresponding to document features of a document feature type corresponding to the second group; and

associate the document image with one of the different document classes based at least in part on the first decision tree and the second decision tree.

18. The system of claim 17 , wherein:

the feature is a feature that was previously determined to be a feature capable of distinguishing the document image or a document that comprises the document image, and

the feature was previously identified by analysis of a plurality of training documents each having at least one unique feature different from at least one of the other training documents.

19. The system of claim 18 , wherein:

the feature is associated with a feature type,

identify the feature in the plurality of training documents further identifying a corresponding feature type for the feature, and

a feature type decision tree is generated that comprises the feature type.

20. The system of claim 19 , further comprising generate a second node corresponding to the class associated with the decision tree.

21. The system of claim 19 , wherein the feature type is: a raster, a title, an image object, a text string, a word, a unique mark, a unique character, a numeric code, or a non-human readable marking.

22. The system of claim 18 , wherein:

the feature is predefined, and

to identify the feature in the document image, the processor is further to perform optical character recognition (OCR) on the feature in the document image.

23. The system of claim 18 , wherein:

to associate the document image with the class, the processor is further to:

determine a value from the document image and

determine a reliability index of the decision tree, wherein the reliability index is based on the feature of the document image.

24. The system of claim 17 , wherein the processor is further to: identify an object in the document image prior to the identifying the feature in the document image, wherein to identify the feature in the document image, the processor is further to identify the feature in the object.

Assignments (3)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
MERGER Recorded Jan 24, 2019
From: ABBYY DEVELOPMENT LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 048129/0558 →