IP Library Granted Patent US 12,664,370
Granted Patent B2
US 12,664,370 · App. 18/074,160 · Granted Jun 23, 2026

Information processing apparatus, information processing method, and storage medium

Inventor: Tomoaki Higo (Kanagawa, JP)
Assignee: Canon Kabushiki Kaisha
G06F40/295G06F40/166G06V10/25G06V30/148G06V30/414
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,370
App. No.
18/074,160
Granted
Jun 23, 2026
Kind
B2
Abstract

The present disclosure relates to a technique of generating a disclosable document image based on a document image including confidential information, without using the confidential information. A document input unit obtains a document image scanned with a scanner, separates the document image into character information and background information, and then outputs them to an extraction unit. The extraction unit performs named entity extraction processing on the obtained character information and background information to extract named entities in the document and attributes thereof, and output an extraction result to a generation unit. The generation unit replaces the named entities in the document image with attribute tags and obtains superimposable ranges to generate attribute tag document data. A management unit registers the received extraction result of the named entities and the attribute tag document data in a database.

Claims (25)

1 . An information processing apparatus comprising:

a controller including a processor and a memory, the controller configured to:

acquire a first document image obtained by scanning a document;

extract character blocks from the first document image;

perform optical character recognition processing on respective extracted character blocks to obtain character strings;

extract a character string corresponding to a named entity of a predetermined attribute from the obtained character strings; and

generate document data including at least attribute information relating to the attribute of the named entity corresponding to the extracted character string and position information of the character block corresponding to the extracted character string,

wherein the document data includes a second document image in which the character block corresponding to the extracted character string in the first document image is replaced by an image indicating the attribute information based on the position information.

2 . The information processing apparatus according to claim 1 , wherein the controller is further configured to: generate a third document image in which the image indicating the attribute information in the second document image is replaced by an arbitrary image.

3 . The information processing apparatus according to claim 2 , wherein the arbitrary image is an image expressing a character string that has the same attribute as the named entity of the extracted character string and that is different from the extracted character string.

4 . The information processing apparatus according to claim 2 , wherein the arbitrary image is an image filled with a predetermined color.

5 . The information processing apparatus according to claim 2 , wherein the arbitrary image is superimposed on a region that includes a region in which the image indicating the attribute information to be replaced by the arbitrary image is superimposed and that does not overlap other objects in the second document image.

6 . The information processing apparatus according to claim 1 , wherein the controller is further configured to: individually save the document data and the extracted character string associating with each other.

7 . The information processing apparatus according to claim 6 , wherein, in the case where similar document data whose similarity to document data to be newly saved is equal to or more than a predetermined threshold and whose attribute information matches that of the document data to be newly saved is saved, only one of the document data to be newly saved and the similar document data is saved.

8 . The information processing apparatus according to claim 6 , wherein the controller is further configured to: reconfigure the first document image based on the document data and the extracted character string associated with the document data.

9 . The information processing apparatus according to claim 1 , wherein the controller is further configured to:

cause a display device to display the first document image, the attribute information, and the extracted character string; and

receive user input and edit the attribute information and the extracted character string based on the user input.

10 . An information processing method comprising:

acquiring a first document image obtained by scanning a document;

extracting character blocks from the first document image;

performing optical character recognition processing on respective extracted character blocks to obtain character strings;

extracting a character string corresponding to a named entity of a predetermined attribute from the obtained character strings; and

generating document data including at least attribute information relating to the attribute of the named entity corresponding to the extracted character string and position information of the character block corresponding to the extracted character string,

wherein the document data includes a second document image in which the character block corresponding to the extracted character string in the first document image is replaced by an image indicating the attribute information based on the position information.