IP Library › Granted Patent US 10,558,893
Granted Patent B2
US 10,558,893 · App. 16/421,952 · Granted Feb 11, 2020

Systems and methods for recognizing characters in digitized documents

Inventor: Theodore Damien Christian Bluche (Paris, FR)
Assignee: A2IA S.A.S.
G06K9/6256G06K9/00409G06K9/00429G06K9/4628G06N3/0445G06N3/0454G06K2209/01G06N3/082G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,558,893
App. No.
16/421,952
Granted
Feb 11, 2020
Kind
B2
Abstract

Methods and systems are provided for end-to-end text recognition in digitized documents of handwritten characters over multiple lines without explicit line segmentation. An image is received. Based on the image, one or more feature maps are determined. Each of the one or more feature maps include one or more feature vectors. Based at least in part on the one or more feature maps, one or more scalar scores are determined. Based on the one or more scalar scores, one or more attention weights are determined. By applying the one or more attention weights to each of the one or more feature vectors, one or more image summary vectors are determined. Based at least in part on the one or more image summary vectors, one or more handwritten characters are determined.

Claims (33)

1. A system for recognizing a plurality of handwritten characters over multiple lines in an image, the system comprising:

a neural network configured to receive the image, the neural network including:

a cascade of a plurality of pairs of a first long short-term memory (LSTM) layer and a convolution layer, wherein each first LSTM layer is configured to generate a first output according to a scanning direction, each convolution layer is configured to generate a feature map based on the first output from a corresponding first LSTM layer in the pair, and feature maps generated by a plurality of pairs are inputted to a next plurality of pairs in the cascade;

a second LSTM layer configured to generate a second output from a plurality of features maps generated by a last plurality of pairs in the cascade; and

a linear layer configured to generate final feature maps based on the second output, wherein the final feature maps include a feature vector at each grid thereof;

a weight calculator configured to calculate a weight vector for each grid of the final feature maps to generate an image summary; and

a decoder configured to determine a probability of each character in the image based on the image summary and the final feature maps.

2. The system according to claim 1 , wherein a dimension of the feature maps is less than or equal to a dimension of the image.

3. The system according to claim 1 , wherein a number of scanning directions is four.

4. The system according to claim 3 , wherein the scanning directions include upward and downward in a vertical direction and backward and forward in a horizontal direction.

5. The system according to claim 1 , wherein a number of convolution layers in the plurality of pairs is four.

6. The system according to claim 1 , wherein a dimension of the weight vector is a number of the final feature maps.

7. The system according to claim 1 , wherein the neural network is further configured to generate an attention vector for each grid based on a corresponding weight vector and a corresponding feature vector of the final feature maps.

8. The system according to claim 7 , wherein attention vectors for all grids of the final feature maps are the image summary.

9. The system according to claim 1 , wherein the neural network further includes a collapse layer configured to concatenate sequences of attention vectors to generate a concatenated sequence of image vectors.

10. The system according to claim 9 , wherein the decoder is further configured to decode the concatenated sequence to identify line beginnings and endings of whole paragraphs in the image.

11. A method for recognizing a plurality of handwritten characters over multiple lines in an image, the method comprising:

generating, by each of a plurality of first long short-term memory (LSTM) layers, a first output according to a scanning direction, wherein each first LSTM layer is paired with a convolution layer;

generating, by each of a plurality of convolution layers, a feature map based on the first output from a corresponding first LSTM layer in the pair;

iterating generating the first output and generating the feature map in a cascade manner;

generating, by a second LSTM layer, a second output from a plurality of feature maps generated by a plurality of pairs of the first LSTM layers and the convolution layers;

generating, by a linear layer, final feature maps based on the second output, wherein the final feature maps include a feature vector at each grid thereof;

calculating a weight vector for each grid of the final feature maps to generate an image summary; and

determining a probability of each character in the image based on the image summary and the final feature maps.

12. The method according to claim 11 , wherein a dimension of the feature maps is less than or equal to a dimension of the image.

13. The method according to claim 11 , wherein a number of scanning directions is four.

14. The method according to claim 13 , wherein the scanning directions include upward and downward in a vertical direction and backward and forward in a horizontal direction.

15. The method according to claim 11 , wherein a number of convolution layers in the plurality of pairs is four.

16. The method according to claim 11 , wherein a dimension of the weight vector is a number of the final feature maps.

17. The method according to claim 11 , further comprising generating an attention vector for each grid based on a corresponding weight vector and a corresponding feature vector of the final feature maps.

18. The method according to claim 17 , wherein attention vectors for all grids of the final feature maps are the image summary.

19. The method according to claim 11 , further comprising concatenating sequences of attention vectors to generate a concatenated sequence of image vectors.

20. The method according to claim 19 , wherein the concatenated sequence is decoded to identify line beginnings and endings of whole paragraphs in the image.

Continuity (3)
Continuation 15481754 · Apr 7, 2017
Provisional Application 62320912 · Apr 11, 2016
Related Publication 20190279035A1 · Sep 12, 2019
Cited By (26)
US 12,197,712 US 12,197,817 US 12,200,297 US 12,204,932 US 12,211,502 US 12,216,894 US 12,219,314 US 12,223,282 US 12,236,952 US 12,254,887 US 12,260,234 US 12,277,954 US 12,293,763 US 12,301,635 US 12,333,404 US 12,361,943 US 12,367,879 US 12,386,434 US 12,386,491 US 12,431,128 US 12,477,470 US 12,556,890 US 12,608,171 US 12,613,730 US 12,619,452 US 12,748,568