IP Library › Granted Patent US 12,131,563
Granted Patent B2
US 12,131,563 · App. 17/689,124 · Granted Oct 29, 2024

System and method for zero-shot learning with deep image neural network and natural language processing (NLP) for optical character recognition (OCR)

Inventor: Tianhao Wu (Princeton Junction, NJ)
G06V30/19147G06V10/82G06V30/1908G06V30/19173
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,131,563
App. No.
17/689,124
Granted
Oct 29, 2024
Kind
B2
Abstract

A system and method for constructing a training dataset and training a neural network include obtaining a searchable portable document format (PDF) document, identifying a bounding box defining a region in a background image that is associated with an overlaying text object defined in the PDF document, determining an image crop of the PDF document according to the bounding box, and generating a training data sample for the training dataset, the training data sample comprising a data pair of the image crop and the associated text object.

Claims (40)

1. A system implemented by one or more computers for constructing a training dataset, the one or more computers comprising:

a storage device; and

A hardware processor, communicatively connected to the storage device, to:

obtain a searchable portable document format (PDF) document;

identify a bounding box defining a region in a background image that is associated with an overlaying text object defined in the PDF document;

determine an image crop of the PDF document according to the bounding box; and

generate a training data sample for the training dataset, the training data sample comprising a data pair of the image crop and the associated text object,

wherein the hardware processor is further to train a neural network model using the training dataset, wherein the neural network model comprises at least one subnetwork for recognizing text content in the image crop and at least one subnetwork for correcting errors in the recognized text content according to a natural language model, and wherein the neural network model is a unified OCR neural network model comprising a pipeline of subnetworks, and the pipeline of subnetworks comprise a spatial transformer network for receiving the training data sample from the training dataset, a first neural network for recognizing text content in the image crop of the training data sample, and a second neural network for correcting errors in the recognized text content according to a natural language model.

2. The system of claim 1 , wherein the hardware processor is further to augment the image crop of the training data sample to generate an augmented training data sample.

3. The system of claim 2 , wherein to augment the image crop of the training data sample, the hardware processor is further to at least one of:

add a random value to at least one pixel in the image crop, or

perform at least one of customization of color schema, random rotation, random blurry, random torsion, random transparency generation to the background image.

4. The system of claim 1 , wherein the image crop comprises an array of pixel values corresponding to the bounding box having a height and a length, and wherein the array of pixel values contains a rendered representation of the overlaying text object.

5. The system of claim 1 , wherein the text object comprises at least one of a character, a word, a sentence, a paragraph, or an article of a natural language.

6. The system of claim 1 , wherein to train the neural network model using the training dataset, the hardware processor is to:

determine a training error by comparing the text object of the training data sample with a result generated by the second neural network; and

adjust at least one parameter of the spatial transformer network, the first neural network, or the second neural network.

7. The system of claim 1 , wherein responsive to training the neural network, the hardware processor is to provide the trained neural network to perform OCR tasks on document images.

8. A method for constructing a training dataset, the method comprising:

obtaining, by a processing device, a searchable portable document format (PDF) document;

identifying a bounding box defining a region in a background image that is associated with an overlaying text object defined in the PDF document;

determining an image crop of the PDF document according to the bounding box;

generating a training data sample for the training dataset, the training data sample comprising a data pair of the image crop and the associated text object;

training a neural network model using the training dataset, wherein the neural network model comprises at least one subnetwork for recognizing text content in the image crop and at least one subnetwork for correcting errors in the recognized text content according to a natural language model, wherein the neural network model is a unified OCR neural network model comprising a pipeline of subnetworks, and wherein the pipeline of subnetworks comprise a spatial transformer network for receiving the training data sample from the training dataset, a first neural network for recognizing text content in the image crop of the training data sample, and a second neural network for correcting errors in the recognized text content according to a natural language model.

9. The method of claim 8 , further comprising augmenting the image crop of the training data sample to generate an augmented training data sample.

10. The method of claim 9 , wherein augmenting the image crop of the training data sample to generate an augmented training data sample further comprises:

adding a random value to at least one pixel in the image crop, or

performing at least one of customization of color schema, random rotation, random blurry, random torsion, random transparency generation to the background image.

11. The method of claim 8 , wherein the image crop comprises an array of pixel values corresponding to the bounding box having a height and a length, wherein the array of pixel values contains a rendered representation of the overlaying text object, and wherein the text object comprises at least one of a character, a word, a sentence, a paragraph, or an article of a natural language.

12. The method of claim 8 , wherein training a neural network model using the training dataset comprises:

determining a training error by comparing the text object of the training data sample with a result generated by the second neural network; and

adjusting at least one parameter of the spatial transformer network, the first neural network, or the second neural network.

13. The method of claim 12 , further comprising:

responsive to training the neural network, providing the trained neural network to perform OCR tasks on document images.

14. A machine-readable non-transitory storage media encoded with instructions that, when executed by one or more computers, cause the one or more computer to construct a training dataset, to:

obtain a searchable portable document format (PDF) document;

identify a bounding box defining a region in a background image that is associated with an overlaying text object defined in the PDF document;

determine an image crop of the PDF document according to the bounding box; and

generate a training data sample for the training dataset, the training data sample comprising a data pair of the image crop and the associated text object,

wherein the hardware processor is further to train a neural network model using the training dataset, wherein the neural network model comprises at least one subnetwork for recognizing text content in the image crop and at least one subnetwork for correcting errors in the recognized text content according to a natural language model, and wherein the neural network model is a unified OCR neural network model comprising a pipeline of subnetworks, and the pipeline of subnetworks comprise a spatial transformer network for receiving the training data sample from the training dataset, a first neural network for recognizing text content in the image crop of the training data sample, and a second neural network for correcting errors in the recognized text content according to a natural language model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 5, 2025
From: SINGULARITY SYSTEMS INC.
To: OPAIDA, INC.
Reel/Frame 072789/0075 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: WU, TIANHAO
To: SINGULARITY SYSTEMS INC.
Reel/Frame 059194/0754 →
Continuity (2)
Provisional Application 63157988 · Mar 8, 2021
Related Publication 20220284721A1 · Sep 8, 2022