IP Library Granted Patent US 10,872,271
Granted Patent B2
US 10,872,271 · App. 16/137,981 · Granted Dec 22, 2020

Training image-processing neural networks by synthetic photorealistic indicia-bearing images

Inventors: Ivan Germanovich Zagaynov (Dolgoprudniy, RU); Pavel Valeryevich Borin (Tomsk, RU)
Assignee: ABBYY PRODUCTION LLC
G06K9/6257G06K9/54G06K9/628G06N3/04G06N3/08G06T3/40G06T5/002G06T5/007G06T5/30G06T5/50G06T11/60G06K2209/01G06T5/004G06T2207/20024G06T2207/20081G06T2207/20084G06T2207/20212
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,872,271
App. No.
16/137,981
Granted
Dec 22, 2020
Kind
B2
Abstract

Systems and methods for training image processing neural networks by synthetic photorealistic indicia-bearing images. An example method comprises: generating an initial set of images, wherein each image of the initial set of images comprises a rendering of a text string; producing an augmented set of images by processing the initial set of images to introduce, into each image of the initial set of image, at least one simulated image defect; generating a training dataset comprising a plurality of pairs of images, wherein each pair of images comprises a first image selected from the initial set of images and a second image selected from the augmented set of images; and training, using the training dataset, a convolutional neural network for image processing.

Claims (43)

1. A method, comprising:

generating, by a computer system, an initial set of images, wherein each image of the initial set of images comprises a rendering of a text string;

producing an augmented set of images by processing the initial set of images to introduce, into each image of the initial set of images, at least one simulated image defect, by applying, to at least a subset of pixels of the image, a simulated digital noise;

generating a training dataset comprising a plurality of pairs of images, wherein each pair of images comprises a first image selected from the initial set of images and a second image selected from the augmented set of images; and

training, using the training dataset, a convolutional neural network for image processing.

2. The method of claim 1 , wherein processing the initial set of images further comprises:

superimposing, on a generated image, a transparent image of a pre-defined or randomly generated text.

3. The method of claim 1 , wherein processing the initial set of images further comprises:

de-contrasting a generated image to reduce a maximum difference in luminance of pixels of the generated image by a pre-defined value.

4. The method of claim 1 , wherein processing the initial set of images further comprises:

simulating an additional light source in a scene of a generated image by additively applying, to at least a subset of pixels of the generated image, low frequency Gaussian noise of a low amplitude.

5. The method of claim 1 , wherein processing the initial set of images further comprises:

de-focusing a generated image by applying, to at least a subset of pixels of a generated image, Gaussian blur.

6. The method of claim 1 , wherein processing the initial set of images further comprises:

simulating movement of imaged objects in a generated image by superimposing a motion blur on the generated image.

7. The method of claim 1 , wherein processing the initial set of images further comprises:

simulating camera pre-processing of a generated image by applying a filter to at least a subset of pixels of the generated image.

8. The method of claim 1 , wherein processing the initial set of images further comprises:

simulating de-mosaicing of a generated image by applying Gaussian blur to at least a subset of pixels of the generated image.

9. The method of claim 1 , wherein the convolutional neural network comprises multiple convolution layers which implement dilated convolution operators with various dilation parameter values.

10. The method of claim 9 , wherein training the convolutional neural network further comprises:

determining a per-parameter training rate by calculating an exponential moving average of a gradient and a squared gradient of an input signal.

11. The method of claim 1 , further comprising:

utilizing the convolution neural network for image pre-processing for an optical character recognition (OCR) application.

12. The method of claim 1 , further comprising:

generating a classification convolutional neural network for classifying a set of input images into a first class comprising synthetic images and a second class comprising real photo images.

13. The method of claim 12 , wherein generating a classification convolutional neural network comprises modifying a convolution neural network utilized for image pre-processing.

14. The method of claim 12 , further comprising:

utilizing the classification convolutional neural network for filtering the training data set.

15. A system, comprising:

a memory;

a processor, coupled to the memory, the processor configured to:

generate an initial set of images, wherein each image of the initial set of images comprises a rendering of a text string;

produce an augmented set of images by processing the initial set of images to introduce, into each image of the initial set of images, at least one simulated image defect;

generate a training dataset comprising a plurality of pairs of images, wherein each pair of images comprises a first image selected from the initial set of images and a second image selected from the augmented set of images; and

train, using the training dataset, a convolutional neural network for image processing, wherein the convolutional neural network comprises a preprocessing branch including a set of convolution filters performing local transformations of an input image and a context branch including multiple convolution layers which reduce the input image by a scaling factor and multiple trans-convolution layers for enlarging the image by the scaling factor.

16. A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a processing device, cause the processing device to:

generate an initial set of images, wherein each image of the initial set of images comprises a rendering of a text string;

produce an augmented set of images by processing the initial set of images to introduce, into each image of the initial set of image, at least one simulated image defect;

generate a training dataset comprising a plurality of pairs of images, wherein each pair of images comprises a first image selected from the initial set of images and a second image selected from the augmented set of images; and

train, using the training dataset, a convolutional neural network for image processing, wherein training the convolutional neural network is performed using a hinge loss function.

17. The computer-readable non-transitory storage medium of claim 16 , further comprising executable instructions causing the processing device to:

generate a classification convolutional neural network for classifying a set of input images into a first class comprising synthetic images and a second class comprising real photo images.

Assignments (3)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 15, 2018
From: ZAGAYNOV, IVAN GERMANOVICH; BORIN, PAVEL VALERYEVICH
To: ABBYY PRODUCTION LLC
Reel/Frame 047165/0334 →
Cited By (1)
US 12,639,553