IP Library Granted Patent US 9,792,520
Granted Patent B2
US 9,792,520 · App. 15/451,781 · Granted Oct 17, 2017

System and method for transcribing handwritten records using word grouping with assigned centroids

Inventors: Jack Reese (Lindon, UT); Michael Murdock (Lehi, UT); Shawn Reid (Orem, UT); Laryn Brown (Highland, UT)
Assignee: Ancestry.com Operations Inc.
G06K9/344G06K9/00456G06K9/00852G06K9/03G06K9/52G06K9/6215G06K9/6218G06T7/70G06K2209/01G06T2207/30176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,792,520
App. No.
15/451,781
Granted
Oct 17, 2017
Kind
B2
Abstract

A handwriting recognition system converts word images on documents, such as document images of historical records, into computer searchable text. Word images (snippets) on the document are located, and have multiple word features identified. For each word image, a word feature vector is created representing multiple word features. Based on the similarity of word features (e.g., the distance between feature vectors), similar words are grouped together in clusters, and a centroid that has features most representative of words in the cluster is selected. A digitized text word is selected for each cluster based on review of a centroid in the cluster, and is assigned to all words in that cluster and is used as computer searchable text for those word images where they appear in documents. An analyst may review clusters to permit refinement of the parameters used for grouping words in clusters, including the adjustment of weights and other factors used for determining the distance between feature vectors.

Claims (78)

1. A method for assigning text to a record from an image of the record, comprising:

obtaining a scanned image of a record;

determining at an optical character recognition system that at least some words in the scanned image are unidentified;

evaluating the record image in order to locate each of multiple word images corresponding to the unidentified words;

for each located word image, identifying multiple word features of that word image;

assigning each of the multiple word images that have similar word features to one of a plurality of word clusters;

selecting a representative word image in each of the word clusters as a centroid;

reviewing, by an analyst, the centroid in each of the word clusters, and entering data representing text for the centroid; and

assigning the representing text for the centroid to all other word images in the same word cluster as the centroid.

2. The method of claim 1 , further comprising:

reviewing, by the analyst, at least one sampling of word images in at least one word cluster;

determining, based on judgment of the analyst, whether the sampled word images are the same word as the centroid for the word cluster and whether the sampled words have been correctly included in the word cluster;

determining that a threshold number of the sampled word images have not been correctly included in the word cluster; and

in response to determining that a threshold number of words have not been correctly included, marking the cluster as suspicious.

3. The method of claim 2 , further comprising:

determining that a threshold number of the sample word images have been correctly included in the cluster; and

in response to determining that a threshold number of words have been correctly included in the cluster, maintaining the cluster.

4. The method of claim 2 , wherein each of the word images have corresponding word features, and wherein the method further comprises:

assigning a value to each of the word features;

assigning a weight to each of the word features;

assigning each of the multiple word images that have similar word features to one of a plurality of word clusters, based at least partially on the weight; and

in response to determining that a threshold number of words have not been correctly included, adjusting the assigned weight by the analyst.

5. The method of claim 1 , wherein the record is a historical record having handwritten words, and wherein the multiple word images are each an image of one of the handwritten words.

6. The method of claim 1 , wherein assigning each of the multiple word images to one of a plurality of clusters comprises:

assigning one or more values to each of the multiple word features for each word image in order to create a feature vector for that word image; and

assigning each word image to a word cluster based on its feature vector.

7. The method of claim 1 , wherein the step of assigning each word image to a word cluster based on its feature vector, comprises:

calculating a distance between each one of the multiple word images and every other one of the multiple word images, based on feature vectors associated with those word images;

selecting, from among the multiple word images, two of the word images that are closest in distance to each other; and

assigning the two of the word images to the word cluster.

8. The method of claim 7 , further comprising:

selecting, from among the multiple word images other than the assigned word images, an additional one of the multiple word images that is closest to the representative word image;

assigning the additional one of the word images to the word cluster; and

repeating the foregoing steps until a predetermined number of the multiple word images have been assigned to the word cluster.

9. The method of claim 8 , wherein the step of selecting a representative word image as a centroid comprises:

determining a mean of the values in feature vectors for the word images that are assigned to the word cluster; and

selecting, as the representative word image, one of the word images having values in its associated feature vector closest to the mean.

10. The method of claim 1 , wherein the multiple word features are selected from a group comprising: top line profile, bottom-line profile, left line profile, right line profile, vertical projection profile, horizontal projection profile, peaks, valleys, watershed cup areas, watershed cap areas, loops and holes, intersections and crossings, stroke orientation, word aspect ratio, and convex whole.

11. A system for assigning text to a record from an image of the record, comprising:

one or more processors; and

a memory, the memory storing instructions that are executable by the one or more processors and configure the system to:

obtain a scanned image of a record;

determine at an optical character recognition system that at least some words in the scanned image are unidentified;

evaluate the record image in order to locate each of multiple word images corresponding to the unidentified words;

for each located word image, identify multiple word features of that word image;

assign each of the multiple word images that have similar word features to one of a plurality of word clusters;

select a representative word image in each of the word clusters as a centroid;

receive, from an analyst, the centroid in each of the word clusters, and corresponding data representing text for the centroid; and

assign the representing text for the centroid to all other word images in the same word cluster as the centroid.

12. The system of claim 11 , wherein the stored instructions further configure the system to:

receive, from the analyst, at least one sampling of word images in at least one word cluster;

determine, based on judgment of the analyst, whether the sampled word images are the same word as the centroid for the word cluster and whether the sampled words have been correctly included in the word cluster;

determine that a threshold number of the sampled word images have not been correctly included in the word cluster; and

in response to determining that a threshold number of words have not been correctly included, mark the cluster as suspicious.

13. The system of claim 12 , wherein the stored instructions further configure the system to:

determine that a threshold number of the sample word images have been correctly included in the cluster; and

in response to determining that a threshold number of words have been correctly included in the cluster, maintain the cluster.

14. The system of claim 12 , wherein each of the word images have corresponding word features, and wherein the stored instructions further configure the system to:

assign a value to each of the word features;

assign a weight to each of the word features;

assign each of the multiple word images that have similar word features to one of a plurality of word clusters, based at least partially on the weight; and

in response to determining that a threshold number of words have not been correctly included, adjust the assigned weight by the analyst.

15. The system of claim 11 , wherein the record is a historical record having handwritten words, and wherein the multiple word images are each an image of one of the handwritten words.

16. The system of claim 11 , wherein each of the multiple word images is assigned to one of a plurality of clusters by:

assigning one or more values to each of the multiple word features for each word image in order to create a feature vector for that word image; and

assigning each word image to a word cluster based on its feature vector.

17. The system of claim 11 , wherein each word image is assigned to a word cluster based on its feature vector, by:

calculating a distance between each one of the multiple word images and every other one of the multiple word images, based on feature vectors associated with those word images;

selecting, from among the multiple word images, two of the word images that are closest in distance to each other; and

assigning the two of the word images to the word cluster.

18. The system of claim 17 , wherein the stored instructions further configure the system to:

select, from among the multiple word images other than the assigned word images, an additional one of the multiple word images that is closest to the representative word image;

assign the additional one of the word images to the word cluster; and

repeat the foregoing elements until a predetermined number of the multiple word images have been assigned to the word cluster.

19. The system of claim 18 , wherein a representative word image is selected as a centroid by:

determining a mean of the values in feature vectors for the word images that are assigned to the word cluster; and

selecting, as the representative word image, one of the word images having values in its associated feature vector closest to the mean.

20. The system of claim 11 , wherein the multiple word features are selected from a group comprising: top line profile, bottom-line profile, left line profile, right line profile, vertical projection profile, horizontal projection profile, peaks, valleys, watershed cup areas, watershed cap areas, loops and holes, intersections and crossings, stroke orientation, word aspect ratio, and convex whole.

Assignments (5)
RELEASE OF FIRST LIEN SECURITY INTEREST Recorded Dec 7, 2020
From: JPMORGAN CHASE BANK, N.A.
To: ANCESTRY.COM OPERATIONS INC.; ANCESTRY.COM DNA, LLC
Reel/Frame 054618/0243 →
SECURITY INTEREST Recorded Dec 7, 2020
From: ANCESTRY.COM DNA, LLC; ANCESTRY.COM OPERATIONS INC.; IARCHIVES, INC.; ANCESTRYHEALTH.COM, LLC
To: CREDIT SUISSE AG, CAYMAN ISLANDS BRANCH
Reel/Frame 054627/0212 →
SECURITY INTEREST Recorded Dec 7, 2020
From: ANCESTRY.COM DNA, LLC; ANCESTRY.COM OPERATIONS INC.; IARCHIVES, INC.; ANCESTRYHEALTH.COM, LLC
To: WILMINGTON TRUST, NATIONAL ASSOCIATION
Reel/Frame 054627/0237 →
FIRST LIEN SECURITY AGREEMENT Recorded Nov 30, 2017
From: ANCESTRY.COM DNA, LLC; ANCESTRY.COM OPERATIONS INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 044552/0538 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 14, 2017
From: REESE, JACK; MURDOCK, MICHAEL; REID, SHAWN; BROWN, LARYN
To: ANCESTRY.COM OPERATIONS INC.
Reel/Frame 041567/0969 →
Continuity (3)
Continuation 14841542 · Aug 31, 2015
Provisional Application 62044076 · Aug 29, 2014
Related Publication 20170193323A1 · Jul 6, 2017