IP Library › Granted Patent US 10,963,691
Granted Patent B2
US 10,963,691 · App. 16/707,697 · Granted Mar 30, 2021

Platform for document classification

Inventors: Steven Dang (Plano, TX); Jason Gould (San Jose, CA); Jennifer Jiang (Plano, TX); Christopher Akatsuka (Frisco, TX); Douglas Slattery (McKinney, TX); Vijaya Pasam (Frisco, TX)
Assignee: Capital One Services, LLC
G06K9/00456G06F16/93G06F17/18G06K9/628G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,963,691
App. No.
16/707,697
Granted
Mar 30, 2021
Kind
B2
Abstract

A device obtains image data associated with a document. Using a first machine learning model, the device determines, for the document, a first classification of one of a plurality of document types and a first confidence score associated with the first classification, and a second classification of one of the plurality of document types and a second confidence score associated with the second classification based on the image data. The device determines a difference between the first confidence score and the second confidence score, compares the difference and a threshold value, and accept the first classification of the document when the difference satisfies the threshold value.

Claims (73)

1. A device, comprising:

one or more memories; and

one or more processors, communicatively coupled to the one or more memories, configured to:

obtain data associated with a document;

determine, for the document and using a machine learning model, a first classification of one of a plurality of document types and a first confidence score associated with the first classification, and a second classification of the one of the plurality of document types and a second confidence score associated with the second classification based on the data,

the first confidence score being greater than the second confidence score;

determine a difference between the first confidence score and the second confidence score; and

accept the first classification of the document when the difference satisfies a threshold spread value.

2. The device of claim 1 , where the one or more processors, when obtaining data associated with the document, are to:

obtain image data or text data associated with the document.

3. The device of claim 1 , where the first confidence score is a highest confidence score associated with the document and the second confidence score is a second-highest confidence score associated with the document.

4. The device of claim 1 , where the one or more processors are further to:

determine the threshold spread value based on using one or more artificial intelligence techniques.

5. The device of claim 1 , where the first classification is different than the second classification.

6. The device of claim 1 , where the first confidence score does not satisfy a confidence threshold value.

7. The device of claim 1 , where the one or more processors are further to:

determine the threshold spread value based on a type of data.

8. A method, comprising:

determining, by a device, for a document, and using a machine learning model, a first classification of a document type, of a plurality of document types, and a first confidence score associated with the first classification,

the first confidence score indicating a confidence level that the document type corresponds to the first classification;

determining, by the device, for the document, and using the machine learning model, a second classification of the one of the plurality of document types and a second confidence score associated with the second classification,

the first confidence score being greater than the second confidence score;

comparing, by the device, the first confidence score and the second confidence score to determine a difference between the first confidence score and the second confidence score; and

accepting, by the device, the first classification of the document of a particular document type when the difference satisfies a threshold spread value.

9. The method of claim 8 , where determining for the document the first classification of the document type comprises:

determining the first classification using an image classifying engine.

10. The method of claim 8 , where determining for the document the first classification of the document type comprises:

determining the first classification using a text classifying engine.

11. The method of claim 8 , where the first confidence score is a highest confidence score associated with the document and the second confidence score is a second-highest confidence score associated with the document.

12. The method of claim 8 , where the document type is one of:

a contract,

a driver's license,

an odometer reading,

a paystub,

an insurance statement,

a service contract,

title information,

a personal financial document, or

an income reporting form.

13. The method of claim 8 , further comprising:

determining the threshold spread value using deep learning.

14. The method of claim 8 , further comprising:

determining the threshold spread value based on whether the first classification and the second classification are determined using a text classification model or an image classification model,

where the threshold spread value is lower when the first classification and second classification are determined using the text classification model.

15. A non-transitory computer-readable medium storing instructions, the instructions comprising:

one or more instructions that, when executed by one or more processors, cause the one or more processors to:

obtain data associated with a document;

determine, for the document and using a machine learning model, a first classification of one of a plurality of document types and a first confidence score associated with the first classification,

determine, for the document and using the machine learning mode, a second classification of the one of the plurality of document types and a second confidence score associated with the second classification based on the data,

where the first confidence score is greater than the second confidence score;

determine a difference between the first confidence score and the second confidence score; and

accept the first classification of the document when the difference satisfies a threshold spread value.

16. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions when executed by the one or more processors, further cause the one or more processors to:

train the machine learning model to use a computer vision technique to:

identify a plurality of features of the document, and

perform a dimensionality reduction to reduce the plurality of features to a particular feature set; and

where the first classification is determined based on the trained machine learning model.

17. The non-transitory computer-readable medium of claim 15 , where the first confidence score represents a minimum confidence score that produces a reliable document classification.

18. The non-transitory computer-readable medium of claim 15 , where the one or more instructions, that cause the one or more processors to obtain the data associated with the document, cause the one or more processors to:

determine the threshold spread value based on whether the first classification and the second classification are determined using a text classification model or an image classification model,

where the threshold spread value is lower when the first classification and second classification are determined using the text classification model.

19. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions when executed by the one or more processors, further cause the one or more processors to:

preprocess the document by one or more of:

an image contrast adjustment,

an image noise filter,

an image histogram modification,

cropping an image of the document, or

a skew correction operation; and

where the first classification is determined based on the preprocessing.

20. The non-transitory computer-readable medium of claim 15 , wherein the one or more instructions when executed by the one or more processors, further cause the one or more processors to:

receive a multiple-page document;

use a splitting technique to form multiple single-page documents from the multiple-page document; and

determine the document from the multiple single-page documents.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2019
From: DANG, STEVEN; GOULD, JASON; JIANG, JENNIFER; AKATSUKA, CHRISTOPHER; SLATTERY, DOUGLAS; PASAM, VIJAYA
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 051223/0696 →
Continuity (3)
Continuation 16540287 · Aug 14, 2019
Continuation 16358046 · Mar 19, 2019
Related Publication 20200302165A1 · Sep 24, 2020