IP Library › Granted Patent US 12,614,032
Granted Patent B2
US 12,614,032 · App. 17/973,603 · Granted Apr 28, 2026

Image processing apparatus, image processing method, and storage medium

Inventor: Yusuke Todoroki (Kanagawa, JP)
Assignee: CANON KABUSHIKI KAISHA
G06F40/284G06V30/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,614,032
App. No.
17/973,603
Granted
Apr 28, 2026
Kind
B2
Abstract

An image processing apparatus divides scanned image data including page images obtained by scanning a plurality of documents for each page into image data of each document. The apparatus generates text data by performing character recognition processing for the plurality of page images, sequentially obtains a pair of page images in succession from the plurality of page images and then determines a document delimitation position based on text data of the two page images constituting the pair, and divides the scanned image data at the determined delimitation position. A vector corresponding to tokens is obtained by decomposing the text of each of the two page images constituting the pair is generated and input to a neural network model, the delimitation position is determined by using a score output from the neural network model and represents a possibility value that the two page images constituting the pair belong to different documents.

Claims (38)

1 . An image processing apparatus that divides scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the apparatus comprising:

one or more memories storing instructions; and

one or more processors executing the instructions:

to generate text data by performing character recognition processing for the plurality of page images;

to sequentially obtain a pair of page images in succession from the plurality of page images and then to determine a document delimitation position based on text data of the two page images constituting the pair; and

to divide the scanned image data at the determined delimitation position, wherein, in the determining:

a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;

the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,

in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.

2 . The image processing apparatus according to claim 1 , wherein, in the determining, a portion between the two page images constituting the pair whose score is higher than or equal to a threshold value is determined to be the delimitation position.

3 . The image processing apparatus according to claim 1 , wherein, in the determining, tokens obtained by performing adjustment processing to match tokens obtained by decomposing the text with specifications of the neural network model are converted into the vector.

4 . The image processing apparatus according to claim 3 , wherein the adjustment processing includes reduction processing to reduce the number of tokens so that the tokens can be input to the neural network model.

5 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to truncate part of tokens obtained by decomposing text corresponding to one page in a case where there is an upper limit to the number of tokens that can be input to the neural network model and the total number of tokens obtained by decomposing the text corresponding to one page exceeds the upper limit.

6 . The image processing apparatus according to claim 5 , wherein the processing to truncate part of tokens is processing to extract only tokens corresponding to text in an upper area and a lower area of the page image for the preceding page of the two page images constituting the pair, and for the following page, extract only tokens corresponding to text in an upper area of the page image.

7 . The image processing apparatus according to claim 5 , wherein the processing to truncate part of tokens is processing to extract only tokens corresponding to text in an upper area and a lower area of each of the two page images constituting the pair.

8 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to shorten the decomposition-target text by summarizing text of each of the two page images constituting the pair.

9 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to extract only tokens corresponding to a specific part of speech by performing morphological analysis for text of each of the two page images constituting the pair.

10 . The image processing apparatus according to claim 1 , wherein, in the neural network model, to a natural language processing model having been trained in advance, a unique determination layer is added and for which, fine tuning aiming at determining the delimitation position has been performed.

11 . The image processing apparatus according to claim 10 , wherein the natural language processing model having been trained in advance is BERT (Bidirectional Encoder Representations from Transformers).

12 . The image processing apparatus according to claim 1 , wherein the one or more processors further execute the instructions to obtain the plurality of page images by scanning the plurality of document images.

13 . The image processing apparatus according to claim 1 , wherein the image processing apparatus is a server apparatus.

14 . The image processing apparatus according to claim 1 , wherein the image processing apparatus is a virtual server by cloud computing.

15 . An image processing method of dividing scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the method comprising the steps of:

generating text data by performing character recognition processing for the plurality of page images;

sequentially obtaining a pair of page images in succession from the plurality of page images and then determining a document delimitation position based on text data of the two page images constituting the pair; and

dividing the scanned image data at the determined delimitation position,

wherein, in the determining step:

a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;

the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,

in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.

16 . A non-transitory computer readable storage medium storing a program for causing a computer to perform an image processing method of dividing scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the method comprising the steps of:

generating text data by performing character recognition processing for the plurality of page images;

sequentially obtaining a pair of page images in succession from the plurality of page images and then determining a document delimitation position based on text data of the two page images constituting the pair; and

dividing the scanned image data at the determined delimitation position,

wherein, in the determining step:

a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;

the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,

in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2022
From: TODOROKI, YUSUKE
To: CANON KABUSHIKI KAISHA
Reel/Frame 061931/0274 →
Priority Claims (1)
JP 2021-178618 · Nov 1, 2021 · national
Continuity (1)
Related Publication 20230137350A1 · May 4, 2023
References Cited (22)
US 5813009A · Johnson · 1998 [cited by examiner]
US 11443416B2 · Liao · 2022 [cited by examiner]
US 20150161078A1 · Cholleti · 2015 [cited by examiner]
US 20200125630A1 · Sanghavi · 2020 [cited by examiner]
US 20220171965A1 · Kumar · 2022 [cited by examiner]
US 20230101817A1 · Sinha · 2023 [cited by examiner]
US 20230368557A1 · Kolavennu · 2023 [cited by examiner]
US 20240070436A1 · Xu · 2024 [cited by examiner]
US 20240161522A1 · Khan · 2024 [cited by examiner]
CN 106919545A · 2017 [cited by examiner]
JP 2002024258A · 2002 [cited by examiner]
JP 2002312385A · 2002 [cited by examiner]
JP 4023075B2 · 2007 [cited by examiner]
JP 2011018316A · 2011 [cited by examiner]
JP 2021057710A · 2021 [cited by examiner]
JP 2021086480A · 2021 [cited by examiner]
WO WO2020201746A1 · 2020 [cited by examiner]
Wiedemann, Gregor, and Gerhard Heyer. “Multi-Modal Page Stream Segmentation with Convolutional Neural Networks—Language Resources and Evaluation.” SpringerLink, Springer Netherlands, Sep. 27, 2019 (Year: 2019). [cited by examiner]
Lambert: Layout-Aware Language Modeling for Information Extraction, Garncarek et al (Year: 2021). [cited by examiner]
Delvin, J et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding”, Proceedings of NAACL-HLT 2019, Jun. 2 to 7, 2019, pp. 4171 to 4186 (16 pages). [cited by applicant]
Nakayama, A. et al. “Reducing Prediction Target Data in Resource Construction Sharing Tass Using Active Sampling”, Proceedings of the Twenty-seventh Annual Meeting of the Association for Natural Language Processing. pp.… [cited by applicant]
Office Action issued Sep. 2, 2025, in corresponding Japanese Patent Application No. 2021-178618, with English translation (8 pages). [cited by applicant]