Image processing apparatus, image processing method, and storage medium
An image processing apparatus divides scanned image data including page images obtained by scanning a plurality of documents for each page into image data of each document. The apparatus generates text data by performing character recognition processing for the plurality of page images, sequentially obtains a pair of page images in succession from the plurality of page images and then determines a document delimitation position based on text data of the two page images constituting the pair, and divides the scanned image data at the determined delimitation position. A vector corresponding to tokens is obtained by decomposing the text of each of the two page images constituting the pair is generated and input to a neural network model, the delimitation position is determined by using a score output from the neural network model and represents a possibility value that the two page images constituting the pair belong to different documents.
1 . An image processing apparatus that divides scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the apparatus comprising:
one or more memories storing instructions; and
one or more processors executing the instructions:
to generate text data by performing character recognition processing for the plurality of page images;
to sequentially obtain a pair of page images in succession from the plurality of page images and then to determine a document delimitation position based on text data of the two page images constituting the pair; and
to divide the scanned image data at the determined delimitation position, wherein, in the determining:
a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;
the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,
in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.
2 . The image processing apparatus according to claim 1 , wherein, in the determining, a portion between the two page images constituting the pair whose score is higher than or equal to a threshold value is determined to be the delimitation position.
3 . The image processing apparatus according to claim 1 , wherein, in the determining, tokens obtained by performing adjustment processing to match tokens obtained by decomposing the text with specifications of the neural network model are converted into the vector.
4 . The image processing apparatus according to claim 3 , wherein the adjustment processing includes reduction processing to reduce the number of tokens so that the tokens can be input to the neural network model.
5 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to truncate part of tokens obtained by decomposing text corresponding to one page in a case where there is an upper limit to the number of tokens that can be input to the neural network model and the total number of tokens obtained by decomposing the text corresponding to one page exceeds the upper limit.
6 . The image processing apparatus according to claim 5 , wherein the processing to truncate part of tokens is processing to extract only tokens corresponding to text in an upper area and a lower area of the page image for the preceding page of the two page images constituting the pair, and for the following page, extract only tokens corresponding to text in an upper area of the page image.
7 . The image processing apparatus according to claim 5 , wherein the processing to truncate part of tokens is processing to extract only tokens corresponding to text in an upper area and a lower area of each of the two page images constituting the pair.
8 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to shorten the decomposition-target text by summarizing text of each of the two page images constituting the pair.
9 . The image processing apparatus according to claim 4 , wherein the reduction processing is processing to extract only tokens corresponding to a specific part of speech by performing morphological analysis for text of each of the two page images constituting the pair.
10 . The image processing apparatus according to claim 1 , wherein, in the neural network model, to a natural language processing model having been trained in advance, a unique determination layer is added and for which, fine tuning aiming at determining the delimitation position has been performed.
11 . The image processing apparatus according to claim 10 , wherein the natural language processing model having been trained in advance is BERT (Bidirectional Encoder Representations from Transformers).
12 . The image processing apparatus according to claim 1 , wherein the one or more processors further execute the instructions to obtain the plurality of page images by scanning the plurality of document images.
13 . The image processing apparatus according to claim 1 , wherein the image processing apparatus is a server apparatus.
14 . The image processing apparatus according to claim 1 , wherein the image processing apparatus is a virtual server by cloud computing.
15 . An image processing method of dividing scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the method comprising the steps of:
generating text data by performing character recognition processing for the plurality of page images;
sequentially obtaining a pair of page images in succession from the plurality of page images and then determining a document delimitation position based on text data of the two page images constituting the pair; and
dividing the scanned image data at the determined delimitation position,
wherein, in the determining step:
a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;
the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,
in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.
16 . A non-transitory computer readable storage medium storing a program for causing a computer to perform an image processing method of dividing scanned image data including a plurality of page images obtained by scanning a plurality of documents en bloc for each page into image data of each document, the method comprising the steps of:
generating text data by performing character recognition processing for the plurality of page images;
sequentially obtaining a pair of page images in succession from the plurality of page images and then determining a document delimitation position based on text data of the two page images constituting the pair; and
dividing the scanned image data at the determined delimitation position,
wherein, in the determining step:
a vector corresponding to tokens obtained by decomposing the text of each of the two page images constituting the pair is generated and input the vector to a neural network model;
the delimitation position is determined by using a score output from the neural network model and representing the level of a possibility by a numerical value that the two page images constituting the pair belong to different documents, respectively; and,
in a case when the number of documents of the plurality of documents is known in advance, a portion between the two page images constituting the pair is determined to be the delimitation position for a number of pairs in order from the pair whose output score is the highest, the number being the number of documents of the plurality of documents minus one.