IP Library › Granted Patent US 11,776,287
Granted Patent B2
US 11,776,287 · App. 17/241,784 · Granted Oct 3, 2023

Document segmentation for optical character recognition

Inventors: Udi Barzelay (Haifa, IL); Ophir Azulai (Tivon, IL); Inbar Shapira (Givat Ada, IL)
Assignee: International Business Machines Corporation
G06V30/153G06N3/08G06T3/40G06V30/18057G06V30/413G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,776,287
App. No.
17/241,784
Filed
Apr 27, 2021
Granted
Oct 3, 2023
Kind
B2
Art Unit
2665
USPC
382/181
Abstract

An approach to identifying text within an image may be presented. The approach can receive an image. The approach can classify an image on a pixel-by-pixel basis whether the pixel is text. The approach can generate bounding boxes around groups of pixels that are classified as text. The approach can mask sections of an image that where pixels are not classified as text. The approach may be used as a pre-processing technique for optical character recognition in documents, scanned images, or still images.

Claims (38)

1. A computer-implemented method for detecting text within an image, the computer-implemented method comprising:

training, by a processor, a model to detect text within the image, with a plurality of synthesized noisy images containing text, wherein the model comprises a Unet convolutional neural network architecture;

receiving, by the processor, an input image;

generating, by the processor, a pixel probability estimate for each pixel in the input image, based at least in part on the Unet convolutional neural network architecture, wherein the probability estimate is the probability a pixel is part of a segmented text;

generating, by the processor, a segmentation map based at least in part on the per pixel probability estimate for each pixel in the input image;

generating, by the processor, one or more bounding boxes around the detected text, based at least in part on the segmentation map; and

masking, by the processor, one or more sections of the input image outside of the bounding boxes.

2. The computer-implemented method of claim 1 , wherein the model is based on a neural network architecture.

3. The computer-implemented method of claim 2 , wherein the neural network architecture is a convolutional neural network architecture.

4. The computer-implemented method of claim 1 , further comprising:

sending, by the processor, the masked image to an optical character recognition (OCR) engine.

5. The computer-implemented method of claim 4 , wherein the OCR engine is based on a tesseract OCR engine.

6. A computer program product for detecting text within an image, the computer program product comprising:

one or more non-transitory computer readable storage media and program instructions stored on the one or more non-transitory computer readable storage media, the program instructions comprising program instructions to:

train a model to detect text within the image, with a plurality of synthesized noisy images containing text, wherein the model comprises a Unet convolutional neural network architecture;

receive an input image;

generate a pixel probability estimate for each pixel in the input image, based at least in part on the Unet convolutional neural network architecture, wherein the probability estimate is the probability a pixel is part of a segmented text;

generate a segmentation map based at least in part on the per pixel probability estimate for each pixel in the input image;

generate one or more bounding boxes around the detected text, based at least in part on the segmentation map; and

mask one or more sections of the input image outside of the bounding boxes.

7. The computer program product of claim 6 , wherein the model is based on a neural network architecture.

8. The computer program product of claim 7 , wherein the neural network architecture is a convolutional neural network architecture.

9. The computer program product of claim 6 , further comprising instructions to:

send the masked image to an optical character recognition (OCR) engine.

10. The computer program product of claim 9 , wherein the OCR engine is based on a tesseract OCR engine.

11. A computer system for detecting text within an image, the computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more processors, the program instructions comprising:

train a model to detect text within the image, with a plurality of synthesized noisy images containing text, wherein the model comprises a Unet convolutional neural network architecture;

receive an input image;

generate a pixel probability estimate for each pixel in the input image, based at least in part on the Unet convolutional neural network architecture, wherein the probability estimate is the probability a pixel is part of a segmented text generate a segmentation map based at least in part on the per pixel probability estimate for each pixel in the input image;

generate one or more bounding boxes around the detected text, based at least in part on the segmentation map; and

mask one or more sections of the input image outside of the bounding boxes.

12. The computer system of claim 11 , wherein the model is based on a neural network architecture.

13. The computer system of claim 11 , further comprising instructions to:

send the masked image to an optical character recognition (OCR) engine.

14. The computer system of claim 13 , wherein the OCR engine is based on a tesseract OCR engine.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 27, 2021
From: BARZELAY, UDI; AZULAI, OPHIR; SHAPIRA, INBAR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 056056/0815 →
Continuity (1)
Related Publication 20220343103A1 · Oct 27, 2022
Cited By (1)
US 12,430,468