IP Library Granted Patent US 11,961,094
Granted Patent B2
US 11,961,094 · App. 17/098,443 · Granted Apr 16, 2024

Fraud detection via automated handwriting clustering

Inventors: Sujit Eapen (Plainsboro, NJ); Sonil Trivedi (Jersey City, NJ)
Assignee: Morgan Stanley Services Group Inc.
G06Q30/0185G06F18/22G06F18/23G06F21/32G06F21/64G06N3/04G06Q10/103G06V30/18086G06V30/274G06V30/347G06V30/36
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,961,094
App. No.
17/098,443
Granted
Apr 16, 2024
Kind
B2
Abstract

A computer-implemented method for automatically analyzing handwritten text to determine a mismatch between a purported writer and an actual writer is disclosed. The method comprises receiving two samples of digitized handwriting each allegedly created by one individual and received and entered into a digital system by another. The method further comprises performing a series of feature extractions to convert the samples into two vectors of extracted features; automatically clustering a set of vectors such that the first vector and the second vector are assigned to the same cluster among multiple clusters, based on vector similarity; and automatically determining that a same individual being associated with both the first and second samples indicates a heightened probability that the individual fraudulently created both samples. Finally, the method comprises automatically transmitting a message to flag additional samples of digitized handwriting entered into a digital system as possibly fraudulent.

Claims (33)

1. A system for automatically analyzing a corpus of handwritten text to determine a mismatch between a purported writer and an actual writer, comprising:

a text sample database, storing a plurality of digitized images comprising handwritten text;

one or more processors, and

non-transient memory storing instructions that, when executed by the one or more processors, cause the one or more processors to:

receive a first sample of digitized handwriting and metadata associating the first sample with a first individual who allegedly created the sample and with a second individual who allegedly received the sample from the first individual and entered it into a digital system;

receive a second sample of digitized handwriting and metadata associating the second sample with a third individual who allegedly created the sample and with the second individual, who also allegedly received the sample from the third individual and entered it into the digital system;

automatically perform a series of feature extractions to convert the first sample and the second sample into a first vector and a second vector, respectively, of extracted features, wherein the extracted features comprise at least one of: a histogram of oriented gradients, an energy-entropy comparison, and a Pearson coefficient between sets of extracted waveforms from tiles;

automatically cluster a set of vectors comprising the first vector and the second vector such that the first vector and the second vector are assigned to the same cluster among multiple clusters, based on vector similarity;

automatically determine that the metadata associating the second individual with both the first and second samples indicates a heightened probability that the first individual and third individual did not create the first and second samples, and rather that the second individual created both samples, based at least in part on pairing regions with same semantic contents or functions for pairwise similarity analysis;

automatically transmit a message to flag additional samples of digitized handwriting entered into the digital system by the second individual as possibly fraudulent.

2. The system of claim 1 , wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: compute a similarity score between each of the additional samples and the first or second samples; and responsive to a score above a predetermined threshold, transmit a message reporting that an additional sample may be fraudulent.

3. The system of claim 1 , further comprising a workflow management server, and wherein the instructions, when executed by the one or more processors, further cause the one or more processors to: transmit a message to the workflow management server to require authorization before proceeding with a workflow that a signature in the first or second sample was intended to authorize.

4. The system of claim 1 , wherein the metadata also stores, for each sample, a plurality of indicators of semantic content or function of handwritten text in regions of the sample.

5. The system of claim 4 , wherein the function of the handwritten text in one region of the sample is a signature.

6. The system of claim 1 , wherein the automatically determining that the metadata indicates a heightened probability that the second individual created both samples comprises is based at least in part on a convolutional neural network classifier into which the samples are input.

7. The system of claim 1 , wherein the extracted features comprise a histogram of oriented gradients.

8. The system of claim 1 , wherein the extracted features comprise an energy-entropy comparison.

9. The system of claim 1 , wherein the extracted features comprise a Pearson coefficient between sets of extracted waveforms from tiles.

10. A computer-implemented method for automatically analyzing a corpus of handwritten text to determine a mismatch between a purported writer and an actual writer, comprising:

receiving a first sample of digitized handwriting and metadata associating the first sample with a first individual who allegedly created the sample and with a second individual who allegedly received the sample from the first individual and entered it into a digital system;

receiving a second sample of digitized handwriting and metadata associating the second sample with a third individual who allegedly created the sample and with the second individual, who also allegedly received the sample from the third individual and entered it into the digital system;

automatically performing a series of feature extractions to convert the first sample and the second sample into a first vector and a second vector, respectively, of extracted features, wherein the extracted features comprise at least one of: a histogram of oriented gradients, an energy-entropy comparison, and a Pearson coefficient between sets of extracted waveforms from tiles;

automatically clustering a set of vectors comprising the first vector and the second vector such that the first vector and the second vector are assigned to the same cluster among multiple clusters, based on vector similarity;

automatically determining that the metadata associating the second individual with both the first and second samples indicates a heightened probability that the first individual and third individual did not create the first and second samples, and rather that the second individual created both samples, based at least in part on pairing regions with same semantic contents or functions for pairwise similarity analysis;

automatically transmitting a message to flag additional samples of digitized handwriting entered into the digital system by the second individual as possibly fraudulent.

11. The method of claim 10 , further comprising: computing a similarity score between each of the additional samples and the first or second samples; and responsive to a score above a predetermined threshold, transmitting a message reporting that an additional sample may be fraudulent.

12. The method of claim 10 , further comprising: transmitting a message to a workflow management server to require authorization before proceeding with a workflow that a signature in the first or second sample was intended to authorize.

13. The method of claim 10 , wherein the metadata also stores, for each sample, a plurality of indicators of semantic content or function of handwritten text in regions of the sample.

14. The method of claim 13 , wherein the function of the handwritten text in one region of the sample is a signature.

15. The method of claim 10 , wherein the automatically determining that the metadata indicates a heightened probability that the second individual created both samples comprises is based at least in part on a convolutional neural network classifier into which the samples are input.

16. The method of claim 10 , wherein the extracted features comprise a histogram of oriented gradients.

17. The method of claim 10 , wherein the extracted features comprise an energy-entropy comparison.

18. The method of claim 10 , wherein the extracted features comprise a Pearson coefficient between sets of extracted waveforms from tiles.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 16, 2020
From: EAPEN, SUJIT; TRIVEDI, SONIL
To: MORGAN STANLEY SERVICES GROUP INC.
Reel/Frame 054369/0366 →
Continuity (1)
Related Publication 20220156756A1 · May 19, 2022
Cited By (1)
US 12,198,459