IP Library Granted Patent US 11,568,623
Granted Patent B2
US 11,568,623 · App. 16/986,325 · Granted Jan 31, 2023

Image processing apparatus, image processing method, and storage medium

Inventors: Motoki Ikeda (Tokyo, JP); Yusuke Muramatsu (Kawasaki, JP)
Assignee: CANON KABUSHIKI KAISHA
G06V10/23G06K9/6253G06K9/6256G06N3/0454G06N3/08G06V10/44G06V30/412G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,568,623
App. No.
16/986,325
Granted
Jan 31, 2023
Kind
B2
Abstract

An image processing apparatus obtains a read image of a document including a handwritten character, generates a first image formed by pixels of the handwritten character by extracting the pixels of the handwritten character from pixels of the read image using a first learning model for extracting the pixels of the handwritten character, estimates a handwriting area including the handwritten character using a second learning model for estimating the handwriting area, and performs handwriting OCR processing based on the generated first image and the estimated handwriting area.

Claims (44)

1. An image processing apparatus comprising:

at least one processor and at least one memory configured to perform:

obtaining a read image of a document including a handwritten character;

generating a first image formed by pixels of the handwritten character by extracting the pixels of the handwritten character from pixels of the read image using a first learning model for extracting the pixels of the handwritten character;

estimating a handwriting area including the handwritten character using a second learning model for estimating the handwriting area; and

performing handwriting OCR processing based on the generated first image and the estimated handwriting area.

2. The apparatus according to claim 1 , wherein the handwriting area is estimated using the second learning model for the read image.

3. The apparatus according to claim 1 , wherein the second learning model is generated by performing learning based on ground truth data in which an area within an input field of a sample image is set as a handwriting area.

4. The apparatus according to claim 1 , wherein the second learning model is generated by performing learning based on ground truth data in which an area including a handwritten character entered in an input field of a sample image is set as a handwriting area.

5. The apparatus according to claim 1 , wherein the second learning model is generated by performing learning based on ground truth data in which an area including a plurality of handwritten characters separated by a digit line is set as a handwriting area.

6. The apparatus according to claim 1 , wherein the at least one processor and the at least one memory are configured to further perform:

generating composite image data by compositing an image of the document including the handwritten character and an image of a document including not a handwritten character but only a background,

wherein the first learning model is learned based on the generated composite image data.

7. The apparatus according to claim 1 , wherein the handwriting area is estimated using the second learning model for the generated first image.

8. The apparatus according to claim 1 , wherein the second learning model is a learning model for estimating, as a handwriting area, an area including at least one handwritten character line in one column.

9. The apparatus according to claim 1 , wherein the at least one processor and the at least one memory are configured to further perform:

generating a second image by erasing the pixels of the handwritten character from the read image and extracting a printed character area including pixels of a printed character from the second image; and

performing printed character OCR processing based on the second image and the printed character area.

10. The apparatus according to claim 9 , wherein the at least one processor and the at least one memory are configured to further perform:

converting the read image into text based on a result of performing the handwriting OCR processing and a result of performing the printed character OCR processing.

11. The apparatus according to claim 10 , wherein in the converting, the printed character area having a predetermined positional relationship with the handwriting area is specified, and a character recognition result of the handwriting area and a character recognition result of the specified printed character area are saved in association with each other.

12. The apparatus according to claim 11 , wherein the printed character area having the predetermined positional relationship with the handwriting area is a printed character area at a closest position in a left direction or an upper direction with respect to the handwriting area.

13. The apparatus according to claim 11 , wherein the printed character area having the predetermined positional relationship with the handwriting area is a printed character area at a closest position in a left direction with respect to the handwriting area if the handwriting area is close to another handwriting area in a vertical direction, and is a printed character area at a closest position in an upper direction with respect to the handwriting area if the handwriting area is close to another handwriting area in a horizontal direction.

14. The apparatus according to claim 1 , wherein the first learning model is a learning model of a first neural network for extracting pixels of a handwritten character.

15. The apparatus according to claim 1 , wherein the second learning model is a learning model of a second neural network for estimating a handwriting area.

16. The apparatus according to claim 14 , wherein the at least one processor and the at least one memory are configured to further perform:

obtaining, based on a handwriting area instructed by a user in a sample image obtained by reading a document, ground truth data of pixels of a handwritten character to be used in learning of the first neural network; and

obtaining the first learning model by performing learning processing of the first neural network based on the sample image and the obtained ground truth data.

17. The apparatus according to claim 1 , wherein the at least one processor and the at least one memory are configured to further perform:

generating composite image data by compositing a handwriting image generated by reading a document obtained by entering only a handwritten character in a white sheet and a background image generated by reading a document including only a background,

wherein the first learning model is obtained by performing learning based on the generated composite image data.

18. The apparatus according to claim 1 , wherein the at least one processor and the at least one memory are configured to further perform:

generating learning data from a handwriting image generated by reading a document obtained by entering only a handwritten character in a white sheet,

wherein the second learning model is obtained by performing learning based on the generated image data.

19. An image processing method comprising:

obtaining a read image of a document including a handwritten character;

generating a first image formed by pixels of the handwritten character by extracting the pixels of the handwritten character from pixels of the read image using a first learning model for extracting the pixels of the handwritten character;

estimating a handwriting area including the handwritten character using a second learning model for estimating the handwriting area; and

performing handwriting OCR processing based on the generated first image and the estimated handwriting area.

20. A non-transitory computer-readable storage medium storing a program that causes a computer to perform:

obtaining a read image of a document including a handwritten character;

generating a first image formed by pixels of the handwritten character by extracting the pixels of the handwritten character from pixels of the read image using a first learning model for extracting the pixels of the handwritten character;

estimating a handwriting area including the handwritten character using a second learning model for estimating the handwriting area; and

performing handwriting OCR processing based on the generated first image and the estimated handwriting area.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2020
From: IKEDA, MOTOKI; MURAMATSU, YUSUKE
To: CANON KABUSHIKI KAISHA
Reel/Frame 054583/0419 →
Priority Claims (3)
JP JP2019-152169 · Aug 22, 2019 · national
JP JP2019-183923 · Oct 4, 2019 · national
JP JP2020-027618 · Feb 20, 2020 · national
Continuity (1)
Related Publication 20210056336A1 · Feb 25, 2021