IP Library Granted Patent US 11,798,258
Granted Patent B2
US 11,798,258 · App. 17/306,495 · Granted Oct 24, 2023

Automated categorization and assembly of low-quality images into electronic documents

Inventors: Van Nguyen (Plano, TX); Sean Michael Byrne (Tampa, FL); Syed Talha (McKinney, TX); Aftab Khan (Richardson, TX); Beena Khushalani (Moorpark, CA); Sharad K. Kalyani (Coppell, TX)
Assignee: Bank of America Corporation
G06V10/464G06F16/93G06F40/20G06N20/00G06V10/30G06V30/413G06V30/416
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,798,258
App. No.
17/306,495
Granted
Oct 24, 2023
Kind
B2
Abstract

An apparatus includes a memory and processor. The memory stores document categories, text generated from an image a physical document page, and a machine learning algorithm. The text includes errors associated with noise in the image. The machine learning algorithm is configured to extract features associated with natural language processing and features associated with the errors from the text. The machine learning algorithm is also configured to generate a feature vector that includes the first and second pluralities of features, and to generate, based on the feature vector, a set of probabilities, each of which is associated with a document category and indicates a probability that the physical document from which the text was generated belongs to that document category. The processor applies the machine learning algorithm to the text, to generate the set of probabilities, identifies a largest probability, and assigns the image to the associated document category.

Claims (148)

1. An apparatus comprising:

a memory configured to store:

a set of document categories;

a first set of text generated from a first image of a page of a physical document, the page comprising a second set of text, wherein the first set of text is different from the second set of text by at least a set of errors, the set of errors associated with noise in the first image; and

a machine learning algorithm configured, when applied to the first set of text and executed by a hardware processor, to:

extract a set of features from the first set of text, the set of features comprising:

a first plurality of features obtained by performing natural language processing feature extraction operations on the first set of text; and

a second plurality of features, each feature of the second plurality of features assigned to an error type of a set of error types and associated with one or more errors of the set of errors, the one or more errors belonging to the assigned error type of the set of error types;

generate a first feature vector comprising the first plurality of features and the second plurality of features; and

generate, based on the first feature vector, a first set of probabilities, each probability of the first set of probabilities associated with a document category of the set of document categories and indicating a probability that the physical document from which the first set of text was generated belongs to the associated document category; and

the hardware processor communicatively coupled to the memory, the hardware processor configured to:

apply the machine learning algorithm to the first set of text, to generate the first set of probabilities;

identify a largest probability of the first set of probabilities; and

assign the first image to the document category associated with the largest probability of the first set of probabilities.

2. The apparatus of claim 1 , wherein at least one error type of the set of error types is associated with at least one of:

non-ascii characters;

non-English words;

stray letters other than “a” and “I”;

numbers comprising one or more letters; or

misplaced punctuation marks.

3. The apparatus of claim 1 , wherein:

the image comprises at least one of:

a scanned image generated by scanning the page of the physical document; or

a faxed image generated by faxing the page of the physical document; and

the noise in the image is associated with at least one of:

ruled lines on the page of the physical document;

hole punches on the page of the physical document;

uneven contrast in the image;

one or more folds in the page of the physical document;

handwritten text in a margin of the page of the physical document;

a background color of the image; or

interfering strokes in the image, wherein the physical document comprises double-sided pages.

4. The apparatus of claim 1 , wherein the machine learning algorithm comprises at least one of a neural network algorithm, a decision tree algorithm, a naive Bayes algorithm, and a logistic regression algorithm.

5. The apparatus of claim 1 , wherein:

the memory is further configured to store a third set of text generated from a second image of the page of the physical document, wherein the third set of text is different from the second set of text by at least a second set of errors, the second set of errors associated with noise in the second image, the noise in the second image different from the noise in the first image;

the machine learning algorithm is further configured, when applied to the third set of text and executed by the hardware processor, to:

extract a second set of features from the third set of text, the second set of features comprising:

a first plurality of features obtained by performing the natural language processing feature extraction operations on the third set of text; and

a second plurality of features, each feature of the second plurality of features assigned to an error type of the set of error types and associated with one or more errors of the set of errors in the third set of text, the one or more errors belonging to the assigned error type of the set of error types;

generate a second feature vector comprising the first plurality of features of the second set of features and the second plurality of features of the second set of features, wherein the second feature vector is different from the first feature vector; and

generate, based on the second feature vector, a second set of probabilities, each probability of the second set of probabilities associated with a document category of the set of document categories and indicating a probability that the physical document from which the third set of text was generated belongs to the associated document category; and

the hardware processor is further configured to:

apply the machine learning algorithm to the third set of text, to generate the second set of probabilities;

identify a largest probability of the second set of probabilities; and

assign the second image to the document category associated with the largest probability of the second set of probabilities, wherein the document category assigned to the second image is the same as the document category assigned to the first image.

6. The apparatus of claim 1 , wherein:

the memory is further configured to store training data comprising:

a plurality of sets of text, each set of text of the plurality of sets of text generated from an image of a page of a known physical document of a set of known physical documents; and

a plurality of labels, each label of the plurality of labels assigned to a set of text of the plurality of sets of text, and identifying a document category of the set of document categories into which the known physical document of the set of known physical documents has previously been categorized; and

the hardware processor is further configured to use the training data to train the machine learning algorithm.

7. The apparatus of claim 6 , wherein:

the machine learning algorithm comprises a set of adjustable weights; and

training the machine learning algorithm comprises adjusting the set of adjustable weights.

8. A method comprising:

applying a machine learning algorithm to a first set of text, to generate a first set of probabilities, wherein:

the first set of text was generated from a first image of a page of a physical document, wherein:

the page comprises a second set of text, the first set of text different from the second set of text by at least a set of errors, the set of errors associated with noise in the first image; and

the physical document is assigned to a first document category of a set of document categories;

the machine learning algorithm is configured, when applied to the first set of text, to:

extract a set of features from the first set of text, the set of features comprising:

a first plurality of features obtained by performing natural language processing feature extraction operations on the first set of text; and

a second plurality of features, each feature of the second plurality of features assigned to an error type of a set of error types and associated with one or more errors of the set of errors, the one or more errors belonging to the assigned error type of the set of error types;

generate a first feature vector comprising the first plurality of features and the second plurality of features; and

generate, based on the first feature vector, the first set of probabilities; and

each probability of the first set of probabilities is associated with a document category of the set of document categories and indicates a probability that the physical document from which the first set of text was generated belongs to the associated document category;

identifying a largest probability of the first set of probabilities; and

assigning the first image to the document category associated with the largest probability of the first set of probabilities.

9. The method of claim 8 , wherein at least one error type of the set of error types is associated with at least one of:

non-ascii characters;

non-English words;

stray letters other than “a” and “I”;

numbers comprising one or more letters; or

misplaced punctuation marks.

10. The method of claim 8 , wherein:

the image comprises at least one of:

a scanned image generated by scanning the page of the physical document; or

a faxed image generated by faxing the page of the physical document; and

the noise in the image is associated with at least one of:

ruled lines on the page of the physical document;

hole punches on the page of the physical document;

uneven contrast in the image;

one or more folds in the page of the physical document;

handwritten text in a margin of the page of the physical document;

a background color of the image; or

interfering strokes in the image, wherein the physical document comprises double-sided pages.

11. The method of claim 8 , wherein the machine learning algorithm comprises at least one of a neural network algorithm, a decision tree algorithm, a naive Bayes algorithm, or a logistic regression algorithm.

12. The method of claim 8 , further comprising:

applying the machine learning algorithm to a third set of text, to generate a second set of probabilities, wherein:

the third set of text was generated from a second image of the page of the physical document, wherein the third set of text is different from the second set of text by at least a second set of errors, the second set of errors associated with noise in the second image, the noise in the second image different from the noise in the first image;

the machine learning algorithm is configured, when applied to the third set of text, to:

extract a second set of features from the third set of text, the second set of features comprising:

a first plurality of features obtained by performing the natural language processing feature extraction operations on the third set of text; and

a second plurality of features, each feature of the second plurality of features assigned to an error type of the set of error types and associated with one or more errors of the second set of errors, the one or more errors belonging to the assigned error type of the set of error types;

generate a second feature vector comprising the first plurality of features of the second set of features and the second plurality of features of the second set of features, wherein the second feature vector is different from the first feature vector; and

generate, based on the second feature vector, the second set of probabilities; and

each probability of the second set of probabilities is associated with a document category of the set of document categories and indicates a probability that the physical document from which the third set of text was generated belongs to the associated document category;

identifying a largest probability of the second set of probabilities; and

assigning the second image to the document category associated with the largest probability of the second set of probabilities, wherein the document category assigned to the second image is the same as the document category assigned to the first image.

13. The method of claim 8 , further comprising using training data to train the machine learning algorithm, wherein the training data comprises:

a plurality of sets of text, each set of text of the plurality of sets of text generated from an image of a page of a known physical document of a set of known physical documents; and

a plurality of labels, each label of the plurality of labels assigned to a set of text of the plurality of sets of text, and identifying a document category of the set of document categories into which the known physical document of the set of known physical documents has previously been categorized.

14. The method of claim 13 , wherein:

the machine learning algorithm comprises a set of adjustable weights; and

training the machine learning algorithm comprises adjusting the set of adjustable weights.

15. A system comprising:

a memory configured to store:

a first image of a page of a physical document;

a first machine learning algorithm configured, when executed by a hardware processor, to convert a first image of a page of a physical document into a first set of text, wherein:

the page comprises a second set of text; and

the first set of text is different from the second set of text by at least a set of errors, the set of errors associated with noise in the first image; and

a second machine learning algorithm configured, when applied to the first set of text and executed by the hardware processor, to:

extract a set of features from the first set of text, the set of features comprising:

a first plurality of features obtained by performing natural language processing feature extraction operations on the first set of text; and

a second plurality of features, each feature of the second plurality of features assigned to an error type of a set of error types and associated with one or more errors of the set of errors, the one or more errors of the set of errors belonging to the assigned error type of the set of error types;

generate a first feature vector comprising the first plurality of features and the second plurality of features; and

generate, based on the first feature vector, a first set of probabilities, each probability of the first set of probabilities associated with a document category of the set of document categories and indicating a probability that the physical document from which the first set of text was generated belongs to the associated document category; and

the hardware processor communicatively coupled to the memory, the hardware processor configured to:

apply the first machine learning algorithm to the first image, to generate the first set of text;

apply the second machine learning algorithm to the first set of text, to generate the first set of probabilities;

identify a largest probability of the first set of probabilities; and

assign the first image to the document category associated with the largest probability of the first set of probabilities.

16. The system of claim 15 , wherein at least one error type of the set of error types is associated with at least one of:

non-ascii characters;

non-English words;

stray letters other than “a” and “I”;

numbers comprising one or more letters; or

misplaced punctuation marks.

17. The system of claim 15 , wherein:

the image comprises at least one of:

a scanned image generated by scanning the page of the physical document; or

a faxed image generated by faxing the page of the physical document; and

the noise in the image is associated with at least one of:

ruled lines on the page of the physical document;

hole punches on the page of the physical document;

uneven contrast in the image;

one or more folds in the page of the physical document;

handwritten text in a margin of the page of the physical document;

a background color of the image; or

interfering strokes in the image, wherein the physical document comprises double-sided pages.

18. The system of claim 15 , wherein the second machine learning algorithm comprises at least one of a neural network algorithm, a decision tree algorithm, a naive Bayes algorithm, and a logistic regression algorithm.

19. The system of claim 15 , wherein:

the memory is further configured to store training data comprising:

a plurality of sets of text, each set of text of the plurality of sets of text generated from an image of a page of a known physical document of a set of known physical documents; and

a plurality of labels, each label of the plurality of labels assigned to a set of text of the plurality of sets of text, and identifying a document category of the set of document categories into which the known physical document of the set of known physical documents has previously been categorized; and

the hardware processor is further configured to use the training data to train the second machine learning algorithm.

20. The system of claim 19 , wherein:

the second machine learning algorithm comprises a set of adjustable weights; and

training the second machine learning algorithm comprises adjusting the set of adjustable weights.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 3, 2021
From: NGUYEN, VAN; BYRNE, SEAN MICHAEL; TALHA, SYED; KHAN, AFTAB; KHUSHALANI, BEENA; KALYANI, SHARAD K.
To: BANK OF AMERICA CORPORATION
Reel/Frame 056118/0150 →
Continuity (1)
Related Publication 20220350999A1 · Nov 3, 2022