IP Library Granted Patent US 10,592,764
Granted Patent B2
US 10,592,764 · App. 15/717,155 · Granted Mar 17, 2020

Reconstructing document from series of document images

Inventors: Vasily Loginov (Moscow, RU); Ivan Zagaynov (Dolgoprudniy, RU); Irina Karatsapova (Balashiha, RU)
Assignee: ABBYY Production LLC
G06K9/3283G06K9/325G06K9/4628G06K9/6274G06K2009/2045G06K2209/01G06T7/97
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,592,764
App. No.
15/717,155
Granted
Mar 17, 2020
Kind
B2
Abstract

Systems and methods for reconstructing a document from a series of document images. An example method comprises: receiving a plurality of image frames, wherein each image frame of the plurality of image frames contains at least a part of an image of an original document; identifying a plurality of visual features in the plurality of image frames; performing spatial alignment of the plurality of image frames based on matching the identified visual features; splitting each of the plurality of image frames into a plurality of image fragments; identifying one or more text-depicting image fragments among the plurality of image fragments; associating each identified text-depicting image fragment with an image frame in which that image fragment has an optimal value of a pre-defined quality metric among values of the quality metric for that image fragment in the plurality of image frames; and producing a reconstructed image frame by blending image fragments from the associated image frames.

Claims (55)

1. A method, comprising:

receiving, by a computer system, a plurality of image frames, wherein each image frame of the plurality of image frames contains at least a part of an image of an original document;

identifying a plurality of visual features in the plurality of image frames;

performing spatial alignment of the plurality of image frames based on matching the identified visual features;

splitting each of the plurality of image frames into a plurality of image fragments;

identifying one or more text-depicting image fragments among the plurality of image fragments;

associating each identified text-depicting image fragment with an image frame in which that image fragment has an optimal value of a pre-defined quality metric among values of the quality metric for that image fragment in the plurality of image frames; and

producing a reconstructed image frame by blending image fragments from the associated image frames.

2. The method of claim 1 , further comprising:

performing optical character recognition (OCR) of the reconstructed image frame.

3. The method of claim 1 , further comprising:

post-processing the reconstructed image frame in an image area residing in a proximity of an inter-fragment border.

4. The method of claim 3 , wherein post-processing the reconstructed image frame further comprises:

identifying a connected components residing in a proximity of an inter-fragment border;

associating each pixel of the connected component with a source image in which that pixel has a maximum sharpness;

selecting a source image that has been associated with a maximum number of connected component pixels; and

copying, from the identified source image, an image fragment covered by the connected component into the reconstructed image.

5. The method of claim 1 , further comprising:

post-processing the reconstructed image frame to reduce a level of noise.

6. The method of claim 1 , further comprising:

post-processing the reconstructed image frame to reduce differences in at least one of: image brightness or image contrast.

7. The method of claim 1 , further comprising:

acquiring the plurality of image frames by a camera controlled by the computer system.

8. The method of claim 1 , wherein the plurality of image frames comprises at least one of: a video frame or a still image frame.

9. The method of claim 1 , wherein splitting each of the plurality of image frames into a plurality of image fragments further comprises:

determining one or more dimensions of an image fragment of the plurality of image fragments such that the image fragment comprises at least one of: at least a pre-defined number of text lines or at least a pre-defined number of symbols per line.

10. The method of claim 1 , wherein the plurality of image fragments comprises two or more non-overlapping image fragments.

11. The method of claim 1 , wherein the plurality of image fragments comprises two or more overlapping image fragments.

12. The method of claim 1 , wherein the plurality of image frames comprises two or more image frames that differ by a position of an image acquiring device with respect to the original document.

13. The method of claim 1 , wherein the plurality of image frames comprises two or more image frames that have been acquired under different image acquiring conditions.

14. The method of claim 1 , further comprising:

performing binarization of the plurality of image frames.

15. The method of claim 1 , wherein performing spatial alignment of the plurality of image frames further comprises:

applying a projective transformation to one or more image frames of the plurality of image frames.

16. The method of claim 1 , wherein the image quality metric reflects at least one of: image sharpness, a level of noise within the image, image contrast, an optical distortion level, or presence of a certain visual artifact.

17. The method of claim 1 , wherein the image quality metric is evaluated by a trainable classifier function that yields a degree of a certain aberration type detected within an image frame.

18. The method of claim 1 , wherein the image quality metric is evaluated by a convolutional neural network that estimates OCR accuracy of an image.

19. A system, comprising:

a memory;

a processor, coupled to the memory, the processor configured to:

receive a plurality of image frames, wherein each image frame of the plurality of image frames contains at least a part of an image of an original document;

identify a plurality of visual features in the plurality of image frames;

perform spatial alignment of the plurality of image frames based on matching the identified visual features;

split each of the plurality of image frames into a plurality of image fragments;

identify one or more text-depicting image fragments among the plurality of image fragments;

associate each identified text-depicting image fragment with an image frame in which that image fragment has an optimal value of a pre-defined quality metric among values of the quality metric for that image fragment in the plurality of image frames; and

produce a reconstructed image frame by blending image fragments from the associated image frames.

20. A computer-readable non-transitory storage medium comprising executable instructions that, when executed by a processing device, cause the processing device to:

receive a plurality of image frames, wherein each image frame of the plurality of image frames contains at least a part of an image of an original document;

identify a plurality of visual features in the plurality of image frames;

perform spatial alignment of the plurality of image frames based on matching the identified visual features;

split each of the plurality of image frames into a plurality of image fragments;

identify one or more text-depicting image fragments among the plurality of image fragments;

associate each identified text-depicting image fragment with an image frame in which that image fragment has an optimal value of a pre-defined quality metric among values of the quality metric for that image fragment in the plurality of image frames; and

produce a reconstructed image frame by blending image fragments from the associated image frames.

Assignments (4)
SECURITY INTEREST Recorded Aug 14, 2023
From: ABBYY INC.; ABBYY USA SOFTWARE HOUSE INC.; ABBYY DEVELOPMENT INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 064730/0964 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 25, 2022
From: ABBYY PRODUCTION LLC
To: ABBYY DEVELOPMENT INC.
Reel/Frame 059249/0873 →
MERGER Recorded Jan 24, 2019
From: ABBYY DEVELOPMENT LLC
To: ABBYY PRODUCTION LLC
Reel/Frame 048129/0558 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2017
From: LOGINOV, VASILY; ZAGAYNOV, IVAN; KARATSAPOVA, IRINA
To: ABBYY DEVELOPMENT LLC
Reel/Frame 043716/0851 →