IP Library › Granted Patent US 10,671,878
Granted Patent B1
US 10,671,878 · App. 16/457,346 · Granted Jun 2, 2020

Systems and methods for text localization and recognition in an image of a document

Inventors: Mohammad Reza Sarshogh (Arlington, VA); Keegan Hines (Washington, DC)
Assignee: Capital One Services, LLC
G06K9/46G06K9/3233G06K9/6256G06K9/6262G06K2209/01
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,671,878
App. No.
16/457,346
Granted
Jun 2, 2020
Kind
B1
Abstract

Disclosed are methods, systems, and non-transitory computer-readable medium for localization and recognition of text from images. For instance, a first method may include: receiving an image; processing the image through a convolutional backbone to obtain feature maps(s); processing the feature maps through a region of interest (RoI) network to obtain RoIs; filtering the RoIs through a filtering block to obtain final RoIs; and processing the final RoIs through a text recognition stack to obtain predicted character sequences for the final RoIs. A second method may include: constructing a text localization and recognition neural network (TLaRNN); obtaining training data; training the TLaRNN on the training data; and storing trained weights of the TLaRNN. The constructing the TLaRNN may include: connecting a convolutional backbone to a region of interest (RoI) network; connecting the RoI network to a filtering block; and connecting the filtering block to a text recognition network.

Claims (85)

1. A method for training a text localization and recognition neural network (TLaRNN), comprising:

constructing the TLaRNN, the constructing the TLaRNN including:

connecting a convolutional backbone to a region of interest (RoI) network;

connecting the RoI network to a filtering block; and

connecting the filtering block to a text recognition network;

obtaining training data;

training the TLaRNN on the training data to obtain trained weights of the TLaRNN; and

storing the trained weights of the TLaRNN,

wherein

the convolutional backbone includes, in series, a first convolutional stack and a feature pyramid network,

the RoI network includes, in series, a region proposal network and a bounding box and classifier network, the bounding box and classifier network includes, in parallel, a bounding box regression head and a classifier head,

the filtering block includes a RoI selection network, and

the text recognition network includes, in series, a feature extraction mechanism and a text recognition stack, and

wherein proposals from the region proposal network and at least one pyramid feature map from the feature pyramid network are input to the bounding box regression head and the classifier head of the bounding box and classifier network.

2. The method of claim 1 , wherein

images of the training data are input to the first convolutional stack,

the first convolutional stack outputs convolutional feature maps,

the convolutional feature maps are input to the feature pyramid network, and

the feature pyramid network outputs the pyramid feature maps.

3. The method of claim 2 , wherein the first convolutional stack is a residual neural network (ResNet), a densely connected convolutional network (DenseNet), or a customized DenseNet.

4. The method of claim 2 , wherein

the pyramid feature maps are input to the region proposal network,

the region proposal network outputs the proposals,

the bounding box regression head outputs deltas on predicted regions of interest (RoIs) for the proposals, and

the classifier head outputs probability of containing text for the proposals.

5. The method of claim 4 , wherein

the bounding box regression head is a bounding box regression network, and

the classifier head is a classification network.

6. The method of claim 4 , wherein the at least one of the pyramid feature maps is from a block closer to an input side than an output side of the feature pyramid network or from a middle block of the pyramid network.

7. The method of claim 4 , wherein

the RoIs, deltas, and the classifications are input to the RoI filtering block, and

the RoI filtering block outputs final RoIs.

8. The method of claim 7 , wherein

the final RoIs and at least one of the convolutional feature maps are input to the feature extraction mechanism,

the feature extraction mechanism outputs extracted feature maps,

the extracted feature maps are input to the text recognition stack, and

the text recognition stack outputs predicted character sequences for the final RoIs.

9. The method of claim 8 , wherein the at least one of the convolutional feature maps is from a block closer to an input side than an output side of the first convolutional neural network or from a middle block of the first convolutional neural network.

10. The method of claim 8 , wherein the text recognition network is a recurrent neural network, a second convolutional neural network, or an attention assisted convolutional stack.

11. A system for extraction of text from images, the system comprising:

a memory storing instructions; and

a processor executing the instructions to perform a process including:

receiving an image;

processing the image through a convolutional backbone to obtain feature maps;

processing the feature maps through a region of interest (RoI) network to obtain RoIs;

filtering the RoIs through a filtering block to obtain final RoIs; and

processing the final RoIs through a text recognition network to obtain predicted character sequences for the final RoIs,

wherein

the convolutional backbone includes, in series, a first convolutional stack and a feature pyramid network,

the RoI network includes, in series, a region proposal network and a bounding box and classifier network, the bounding box and classifier network includes, in parallel, a bounding box regression head and a classifier head,

the filtering block includes a RoI selection network, and

the text recognition network includes, in series, a feature extraction mechanism and a text recognition stack, and

wherein proposals from the region proposal network and at least one of the pyramid feature maps from the feature pyramid network are input to the bounding box regression head and the classifier head of the bounding box and classifier network.

12. The system of claim 11 , wherein to process the image through the convolutional backbone to obtain the feature maps, the process further includes:

inputting the image to the convolutional stack,

processing the image through the convolutional stack to output convolutional feature maps,

inputting the convolutional feature maps to the feature pyramid network, and

processing the convolutional feature maps through the feature pyramid network to output the pyramid feature maps.

13. The system of claim 12 , wherein to process the feature maps through the RoI network to obtain the RoIs, the process further includes:

inputting the pyramid feature maps to the region proposal network,

processing the pyramid feature maps through the region proposal network to output the proposals,

inputting the proposals and the at least one of the pyramid feature maps to the bounding box regression head and the classifier head of the bounding box and classifier network,

processing the proposals and the at least one of the pyramid feature maps through the bounding box regression head to output deltas on the RoIs, and

processing the proposals and the at least one of the pyramid feature maps through the classifier head to output probability of containing text for the proposals.

14. The system of claim 13 , wherein the at least one of the pyramid feature maps is from a block closer to an input side than an output side of the pyramid network or from a middle block of the pyramid network.

15. The system of claim 13 , wherein to filter the RoIs through the filtering block to obtain the final RoIs, the process further includes:

inputting the RoIs, deltas, and the probabilities to the RoI selection network, and

processing the deltas, the RoIs, and the probabilities through the RoI filtering block to output the final RoIs.

16. The system of claim 15 , wherein to process the final RoIs through the text recognition network to obtain the predicted character sequences for the final RoIs, the process further includes:

inputting the final RoIs and at least one of the convolutional feature maps to the feature extraction mechanism,

processing the final RoIs and the at least one of the convolutional feature maps through the feature extraction mechanism to output corresponding feature crops,

inputting the feature crops to the text recognition stack, and

processing the feature crops through the text recognition stack to output the predicted character sequences for the final RoIs.

17. The system of claim 16 , wherein the at least one of the convolutional feature maps is from a block closer to an input side than an output side of the convolutional stack or from a middle block of the convolutional stack.

18. A method for localization and recognition of text from images, comprising:

receiving an image;

processing the image through a convolutional backbone to obtain feature maps(s);

processing the feature maps through a region of interest (RoI) network to obtain RoIs;

filtering the RoIs through a filtering block to obtain final RoIs; and

processing the final RoIs through a text recognition network to obtain predicted character sequences for the final RoIs,

wherein the convolutional stack includes, in series, a convolutional stack and a feature pyramid network,

wherein the RoI network includes, in series, a region proposal network and a bounding box and classifier network, the bounding box and classifier network includes, in parallel, a bounding box regression head and a classifier head,

wherein the filtering block includes an RoI filtering logic,

wherein the text recognition network includes, in series, a feature extraction mechanism and a text recognition stack, and

wherein proposals from the region proposal network and at least one pyramid feature map from the feature pyramid network are input to the bounding box regression head and the classifier head of the bounding box and classifier network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2019
From: SARSHOGH, MOHAMMAD REZA; HINES, KEEGAN
To: CAPITAL ONE SERVICES, LLC
Reel/Frame 049637/0451 →
Continuity (1)
Provisional Application 62791535 · Jan 11, 2019
Cited By (5)
US 12,417,611 US 12,423,858 US 12,437,571 US 12,700,207 US 12,718,082