IP Library Granted Patent US 9,406,030
Granted Patent B2
US 9,406,030 · App. 14/235,658 · Granted Aug 2, 2016

System and methods for computerized machine-learning based authentication of electronic documents including use of linear programming for classification

Inventors: Guy Dolev (Herzliya, IL); Sergey Markin (Herzeliya, IL); Avi Bar-Nissim (Hod Hasharon, IL); Asher Uziel (Modiin, IL)
Assignee: AU10TIX LIMITED
G06N99/005G06K9/00442G06K9/6276
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,406,030
App. No.
14/235,658
Granted
Aug 2, 2016
Kind
B2
Abstract

Electronic document classification comprising providing training documents sorted into classes; linear programming including selecting inputs which maximize an output, given constraints on inputs, the output maximized being a difference between: a. first estimated probability that a document instance will be correctly classified, by a classifier corresponding to given inputs, as belonging to its own class, and b. second estimated probability that document instance will be classified, by the classifier, as not belonging to its own class; and classifying electronic document instances into classes, using a preferred classifier corresponding, to the inputs selected by the linear programming. A computerized electronic document forgery detection method provides training documents and uses a processor to select value-ranges of non-trivial parameters, such that selected values-range(s) of parameters are typical to an authentic document of given class, and atypical to a forged document of same class.

Claims (56)

1. A computerized method for electronic document classification, the method comprising:

providing training documents sorted into a plurality of classes;

using a processor to perform linear programming including selecting input values which maximize an output value, given specific constraints on the input values,

wherein the output value maximized is a difference between:

a. a first estimated probability that a document instance will be correctly classified, by a given classifier corresponding to given input values, as belonging to its own class, and

b. a second estimated probability that the document instance will be classified, by the given classifier, as belonging to a class other than its own class; and

classifying electronic document instances into the plurality of classes, using at least one preferred classifier corresponding to the input values selected by said linear programming including storing an indication of said classifying in computer memory,

wherein some electronic document instances are classified as belonging to none of the plurality of classes.

2. A method according to claim h wherein said training documents sorted into a plurality of classes are sorted by a human supervisor and said own class comprises a class to which a training document belongs, as determined by the human supervisor.

3. A method according to claim 1 , wherein each electronic document instance includes at least one digital scan, using at least one illumination, of a physical document.

4. A method according to claim 1 , wherein said classifying uses said preferred classifier in conjunction with available partial information regarding correspondence between electronic document instances and the plurality of classes.

5. A method according to claim 4 , wherein said partial information includes information read from an electronic document instance's machine readable zone.

6. A computerized method for electronic document classification, the method comprising:

providing training documents sorted into a plurality of classes;

using a processor to perform linear programming including selecting input values which maximize an output value, given specific constraints on the input values,

wherein the output value maximized is a difference between:

a. a first estimated probability that a document instance will be correctly classified, by a given classifier corresponding to given input values, as belonging to its own class, and

b. a second estimated probability that the document instance will be classified, by the given classifier, as belonging to a class other than its own class; and

classifying electronic document instances into the plurality of classes, using at least one preferred classifier corresponding to the input values selected by said linear programming including storing an indication of said classifying in computer memory,

wherein said input values comprise weights used to compute linear combinations of functions of features derived from individual electronic document instances.

7. A method according to claim 6 , wherein at least one feature derived from at least one individual electronic document instance characterizes a local patch within the individual electronic document instance.

8. A method according to claim 7 , wherein at least one feature derived from at least one individual electronic document instance comprises a texture feature.

9. A method according to claim 7 , wherein at least one feature derived from at least one individual electronic document instance comprises a color moment feature.

10. A method according to claim 7 , wherein at least one feature derived from at least one individual electronic document instance comprises a ratio between a central tendency of at a color characterizing at least a portion of the electronic document instance, and a measure of spread of the color.

11. A method according to claim 10 , wherein said color is expressed in terms of least one channel in a color space.

12. A method according to claim 1 , wherein each feature is associated with at least one k-nearest-neighbors weak classifier.

13. A computerized method for electronic document classification, the method comprising:

providing training documents sorted into a plurality of classes;

using a processor to perform linear programming including selecting input values which maximize an output value, given specific constraints on the input values,

wherein the output value maximized is a difference between:

a. a first estimated probability that a document instance will be correctly classified, by a given classifier corresponding to given input values, as belonging to its own class, and

b. a second estimated probability that the document instance will be classified, by the given classifier, as belonging to a class other than its own class; and

classifying electronic document instances into the plurality of classes, using at least one preferred classifier corresponding to the input values selected by said linear programming including storing an indication of said classifying in computer memory; and

electronically determining whether each of a stream of electronic document instances are forgeries, by performing electronic forgery tests specific to individual classes from among said plurality of classes, on individual electronic document instances in said stream which have been classified by said preferred classifier, as belonging to said individual classes respectively.

14. A method according to claim 6 , wherein said functions include probabilities that an individual document instance belongs to a given class given that the individual document instance is characterized by a particular feature derived from individual electronic document instances.

15. A method according to claim 14 , wherein said constraints include at least one constraint whereby a pair of said linear combinations, corresponding to different classes, differ by at least a predetermined margin.

16. A method according to claim 14 wherein said constraints include at least one constraint whereby a pair of said linear combinations, corresponding to different classes, differ by at least a predetermined margin but for a slack variable characterizing an individual electronic document and selected to be large if the individual electronic document is an outlier in its class.

17. A method according to claim 3 , wherein each electronic document instance includes a plurality of scans, using a plurality of illuminations, of a physical document.

18. A method according to claim 1 , wherein at least one classifier for at least one document is obtained by:

tiling a visible (VIS) image of said document to patches, and

from each of said patches, extracting values of parameters including at least color moments; and

performing forgery testing of said instances using, for at least one individual document classified into an individual class from among said plurality of classes, at least one forgery test specific to said individual class.

19. A method according to claim 1 , wherein at least one classifier for at least one document is obtained by:

tiling a visible (VIS) image of said document to patches, and

from each of said patches, extracting values of parameters including at least one texture parameter generated by transforming each patch to grey level, and computing at least one linear combination of the resulting gray image's highest fourier-transform coefficients; and

performing forgery testing of said instances using, for at least one individual document classified into an individual class from among said plurality of classes, at least one forgery test specific to said individual class.

20. A method according to claim 1 , wherein at least one classifier for at least one document is obtained by:

tiling a visible (VIS) image of said document to patches, and

from each of said patches, extracting values of parameters including at least an std2mean parameter generated by computing a ratio between average and standard deviation in a grey-level transformed image of said document; and

performing forgery testing of said instances using, for at least one individual document classified into an individual class from among said plurality of classes, at least one forgery test specific to said individual class.

21. A method according to claim 1 , wherein at least one estimated probability that an electronic document belongs to a particular class of documents, or to no known class thereof is computed by:

finding K documents that are the nearest neighbors to the current document instance;

computing an average distance to said K documents; and

using said average distance to compute an estimated probability to be in any of C classes and an estimated probability to belong to none of the C classes.

22. A method according to claim 7 , wherein at least one feature derived from at least one individual electronic document instance comprises at least one color moment feature including at least averages for at least H and S channels in hue-saturation-value (HSV) color space.

23. A method according to claim 13 , wherein said classes include versions of individual document types and forgery testing is differentially performed for different versions.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2020
From: AU10TIX LIMITED
To: AU10TIX LTD.
Reel/Frame 054380/0980 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2014
From: DOLEV, GUY; MARKIN, SERGEY; BAR-NISSIM, AVI; UZIEL, ASHER
To: AU10TIX LIMITED
Reel/Frame 032065/0254 →
Continuity (2)
Provisional Application 61512487 · Jul 28, 2011
Related Publication 20140180981A1 · Jun 26, 2014