IP Library Granted Patent US 12705749
Granted Patent B2
US 12705749 · App. 18/327,027 · Granted Aug 11, 2026

Image processing apparatus, method and program, and learning apparatus, method and program for extracting image from target image by trained extraction model

Inventor: Satoshi Ihara (Tokyo, JP)
Assignee: FUJIFILM Corporation
G06T7/11G06T3/40G06T7/0012G06T2207/20084G06T2207/30056
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12705749
App. No.
18/327,027
Granted
Aug 11, 2026
Kind
B2
Abstract

A processor is configured to: reduce a target image to derive a reduced image; extract a region of a target structure from the reduced image to derive a reduced structure image including the region of the target structure; extract a corresponding image corresponding to the reduced structure image from the target image; and input the corresponding image and the reduced structure image into an extraction model constructed by machine-learning a neural network to extract a region of the target structure included in the corresponding image from the extraction model.

Claims (40)

1 . An image processing apparatus comprising at least one processor,

wherein the processor is configured to:

acquire a target image;

reduce the target image to derive a reduced image of the target image;

extract a first region from the reduced image to derive a reduced structure image including the first region by a first extraction model without rest of structures of the reduced image;

extract, from the target image, a region which corresponds to the reduced structure image in a corresponding image;

derive an enlarged structure image by enlarging the reduced structure image to be a same size and same resolution as the corresponding image;

input the corresponding image and the enlarged structure image respectively into two channels of an input layer of a second extraction model which is constructed by machine-learning a neural network and includes a plurality of processing layers that perform convolution processing and the input layer which has at least two channels as the second extraction model outputs an extracted image in which a second region of a target structure is included, wherein the extracted image is obtained by extracting the target structure from the corresponding image.

2 . The image processing apparatus according to claim 1 ,

wherein the processor is configured to:

divide the region of the target structure extracted from the reduced image and derive a divided and reduced structure image including each of the divided regions of the target structure;

derive a plurality of divided corresponding images corresponding to the respective divided and reduced structure images from the corresponding image; and

extract the region of the target structure included in the corresponding image in units of the divided corresponding image and the divided and reduced structure image.

3 . A learning apparatus comprising at least one processor,

wherein the processor is configured to:

acquire a reduced structure image including a first region of a target structure, wherein the reduced structure image is derived based on a reduced image of a target image by extracting the first region of the target structure from the reduced image without rest of structures of the reduced image by a first extraction model;

acquire a corresponding image which is extracted from a region of the target image and corresponds to the reduced structure image, wherein an enlarged structure image is to be derived by enlarging the reduced structure image to be a same size and same resolution as the corresponding image; and

construct an extraction model that is configured to extract an extracted image which includes a second region of the target structure from the corresponding image, by machine-learning a neural network using, as supervised training data, a first image including the first region of the target structure extracted from the reduced structure image of the target image including the target structure, a second image corresponding to the corresponding image, and correct answer data representing an extraction result of the target structure from the second image, wherein the extraction model includes an input layer having two channels and a plurality of processing layers that perform convolution processing, the input layer is configured to receive the corresponding image and the enlarged structure.

4 . An image processing method comprising:

acquiring a target image;

reducing the target image to derive a reduced image of the target image;

extracting a first region from the reduced image to derive a reduced structure image including the first region by a first extraction model without rest of structures of the reduced image;

extracting, from the target image, region which corresponds to the reduced structure image in a corresponding image;

deriving an enlarged structure image by enlarging the reduced structure image to be a same size and same resolution as the corresponding image;

inputting the corresponding image and the enlarged structure image respectively into two channels of an input layer of a second extraction model which is constructed by machine-learning a neural network and includes a plurality of processing layers that perform convolution processing and the input layer which has at least two channels as the second extraction model outputs an extracted image in which a second region of a target structure is included, wherein the extracted image is obtained by extracting the target structure from the corresponding image.

5 . A learning method comprising:

acquiring a reduced structure image including a first region of a target structure, wherein the reduced structure image is derived based on a reduced image of a target image by extracting the first region of the target structure from the reduced image without rest of structures of the reduced image by a first extraction model;

acquiring a corresponding image which is extracted from the target image and corresponds to the reduced structure image, wherein an enlarged structure image is to be derived by enlarging the reduced structure image to be a same size and same resolution as the corresponding image; and

constructing an extraction model that is configured to extract an extracted image which includes a second region of the target structure from the corresponding image, by machine-learning a neural network using, as supervised training data, a first image including the first region of the target structure extracted from the reduced structure image of the target image including the target structure, a second image corresponding to the corresponding image, and correct answer data representing an extraction result of the target structure from the second image, wherein the extraction model includes an input layer having two channels and a plurality of processing layers that perform convolution processing, the input layer is configured to receive the corresponding image and the enlarged structure.

6 . A non-transitory computer-readable storage medium that stores an image processing program for causing a computer to execute:

acquiring a target image;

reducing the target image to derive a reduced image of the target image;

extracting a first region from the reduced image to derive a reduced structure image including the first region by a first extraction model without rest of structures of the reduced image;

extracting, from the target image, region which corresponds to the reduced structure image in a corresponding image;

deriving an enlarged structure image by enlarging the reduced structure image to be a same size and same resolution as the corresponding image;

inputting the corresponding image and the enlarged structure image respectively into two channels of an input layer of a second extraction model which is constructed by machine-learning a neural network and includes a plurality of processing layers that perform convolution processing and the input layer which has at least two channels as the second extraction model outputs an extracted image in which a second region of a target structure is included, wherein the extracted image is obtained by extracting the target structure from the corresponding image.

7 . A non-transitory computer-readable storage medium that stores a learning program for causing a computer to execute:

acquiring a reduced structure image including a first region of a target structure, wherein the reduced structure image is derived based on a reduced image of a target image by extracting the first region of the target structure from the reduced image without rest of structures of the reduced image by a first extraction model;

acquiring a corresponding image which is extracted from the target image and corresponds to the reduced structure image, wherein an enlarged structure image is to be derived by enlarging the reduced structure image to be a same size and same resolution as the corresponding image; and

constructing an extraction model that is configured to extract an extracted image which includes a second region of the target structure from the corresponding image, by machine-learning a neural network using, as supervised training data, a first image including the first region of the target structure extracted from the reduced structure image of the target image including the target structure, a second image corresponding to the corresponding image, and correct answer data representing an extraction result of the target structure from the second image, wherein the extraction model includes an input layer having two channels and a plurality of processing layers that perform convolution processing, the input layer is configured to receive the corresponding image and the enlarged structure.