IP Library › Granted Patent US 12,217,387
Granted Patent B2
US 12,217,387 · App. 17/400,142 · Granted Feb 4, 2025

Learning method, learning system, learned model, program, and super resolution image generating device

Inventors: Akira Kudo (Tokyo, JP); Yoshiro Kitamura (Tokyo, JP)
Assignee: FUJIFILM Corporation
G06T3/4053G06N3/045G06N3/088G06T3/40G06T3/4046G06T5/20G06T5/70G06T2200/04G06T2207/10081G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,217,387
App. No.
17/400,142
Filed
Aug 12, 2021
Granted
Feb 4, 2025
Kind
B2
Art Unit
2646
USPC
382/131
Abstract

Provided are a learning method and a learning system of a generative model, a program, a learned model, and a super resolution image generating device that can handle input data of any size and can suppress the amount of calculation at the time of image generation. A learning method according to an embodiment of the present disclosure is a learning method for performing machine learning of a generative model that estimates, from a first image, a second image including higher resolution image information than the first image, the method comprising using a generative adversarial network including a generator which is the generative model and a discriminator which is an identification model that identifies whether provided data is data of a correct image for learning or data derived from an output from the generator and implementing a self-attention mechanism only in a network of the discriminator among the generator and the discriminator.

Claims (54)

1. A learning method for performing machine learning of a generative model that estimates, from a first image, a second image including higher resolution image information than the first image, the method comprising:

using a generative adversarial network including a generator which is the generative model and a discriminator which is an identification model that identifies whether provided data is data of a correct image for learning or data derived from an output from the generator;

using, as learning data, a first learning image including first resolution information having a lower resolution than the second image and a second learning image including second resolution information having a higher resolution than the first learning image, the second learning image being the correct image corresponding to the first learning image;

providing only the first learning image, among the first learning image and the second learning image, to be input to the generator; and

implementing a self-attention mechanism only in a network of the discriminator among the generator and the discriminator, wherein the self-attention mechanism being applied in the network of the discriminator is a part or all of a plurality of convolutional layers of the discriminator.

2. The learning method according to claim 1 ,

wherein each network of the generator and the discriminator is a convolutional neural network.

3. The learning method according to claim 1 ,

wherein the first image is a three-dimensional tomographic image, and

the second image has at least a higher resolution than the first image in a slice thickness direction of the three-dimensional tomographic image.

4. The learning method according to claim 1 ,

wherein the second learning image is an image acquired by using a computed tomography device, and

the first learning image is an image generated by image processing based on the second learning image.

5. The learning method according to claim 4 ,

wherein the image processing of generating the first learning image from the second learning image includes down-sampling processing of the second learning image.

6. The learning method according to claim 5 ,

wherein the image processing of generating the first learning image from the second learning image includes up-sampling processing of an image obtained by the down-sampling processing by performing interpolation processing.

7. The learning method according to claim 4 ,

wherein the image processing of generating the first learning image from the second learning image includes smoothing processing using a Gaussian filter.

8. The learning method according to claim 1 ,

wherein the first learning images and the second learning images in a plurality of types of the learning data used for the machine learning have the same size, respectively.

9. The learning method according to claim 1 ,

wherein the second image is a high-frequency component image indicating information of a high-frequency component, and

the generator estimates a high-frequency component required to increase a resolution of an input image and outputs the high-frequency component image indicating the information of the high-frequency component.

10. The learning method according to claim 9 , further comprising adding the high-frequency component image output from the generator and the image input to the generator,

wherein a virtual second image obtained by the adding is provided to be input to the discriminator.

11. A non-transitory computer-readable recording medium that causes a computer to execute the learning method according to claim 1 in a case in which a command stored in the recording medium is read by the computer.

12. A learned model that is learned by performing the learning method according to claim 1 , the learned model being the generative model that estimates, from the first image, the second image including higher resolution image information than the first image.

13. A super resolution image generating device comprising the generative model which is a learned model that is learned by performing the learning method according to claim 1 ,

wherein the super resolution image generating device generates, from a third image which is input, a fourth image including higher resolution image information than the third image.

14. The super resolution image generating device according to claim 13 ,

wherein an image size of the third image is different from an image size of the first learning image.

15. The super resolution image generating device according to claim 13 , further comprising:

a first interpolation processing unit that performs interpolation processing on the third image to generate an interpolation image; and

a first addition unit that adds the interpolation image and a high-frequency component generated by the generative model,

wherein the interpolation image is input to the generative model, and

the generative model generates the high-frequency component required to increase a resolution of the interpolation image.

16. A learning system for performing machine learning of a generative model that estimates, from a first image, a second image including higher resolution image information than the first image, the system comprising a generative adversarial network including a generator which is the generative model and a discriminator which is an identification model that identifies whether provided data is data of a correct image for learning or data derived from an output from the generator,

wherein a self-attention mechanism is implemented only in a network of the discriminator among the generator and the discriminator, wherein the self-attention mechanism being applied in the network of the discriminator is a part or all of a plurality of convolutional layers of the discriminator,

a first learning image including first resolution information having a lower resolution than the second image and a second learning image including second resolution information having a higher resolution than the first learning image, the second learning image being the correct image corresponding to the first learning image are used as learning data, and

only the first learning image among the first learning image and the second learning image is provided to be input to the generator and learning of the generative adversarial network is performed.

17. The learning system according to claim 16 , further comprising a learning data generating unit that generates the learning data,

wherein the learning data generating unit includes

a fixed-size region cutout unit that cuts out a fixed-size region from an original image including the second resolution information, and

a down-sampling processing unit that performs down-sampling of an image of the fixed-size region cut out by the fixed-size region cutout unit,

the image of the fixed-size region cut out by the fixed-size region cutout unit is used as the second learning image, and

the first learning image is generated by performing down-sampling processing on the second learning image.

18. The learning system according to claim 17 ,

wherein the learning data generating unit further includes

a second interpolation processing unit that performs interpolation processing on an image obtained by the down-sampling processing, and

a smoothing processing unit that performs smoothing by using a Gaussian filter.

19. The learning system according to claim 16 ,

wherein the generator is configured to estimate a high-frequency component required to increase a resolution of an input image and output the high-frequency component image indicating the information of the high-frequency component, and

the learning system further comprises a second addition unit that adds the high-frequency component image output from the generator and the image input to the generator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2021
From: KUDO, AKIRA; KITAMURA, YOSHIRO
To: FUJIFILM CORPORATION
Reel/Frame 057167/0120 →
Priority Claims (1)
JP 2019-036374 · Feb 28, 2019 · national
Continuity (2)
Continuation PCTJP2020007383 · Feb 25, 2020
Related Publication 20210374911A1 · Dec 2, 2021
References Cited (18)
US 20180075581A1 · Shi et al. · 2018 [cited by applicant]
US 20180341836A1 · Lim et al. · 2018 [cited by applicant]
US 20190057488A1 · Li · 2019 [cited by applicant]
US 20210237767A1 · Khoreva · 2021 [cited by examiner]
US 20210374911A1 · Kudo · 2021 [cited by examiner]
EP 3447721 · 2019 [cited by applicant]
“Office Action of Japan Counterpart Application” with English translation thereof, issued on Apr. 26, 2022, p. 1-p. 4. [cited by applicant]
“Search Report of Europe Counterpart Application”, issued on Mar. 23, 2022, pp. 1-8. [cited by applicant]
Christian Ledig et al., “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”, 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 1-14. [cited by applicant]
Junyoung Park et al., “Computed tomography super-resolution using deep convolutional neural network”, Physics in Medicine & Biology, vol. 63, Issue 14, Jul. 2018, pp. 1-13. [cited by applicant]
Zhengchun Liu et al., “TomoGAN: low-dose synchrotron x-ray tomography with generative adversarial networks: discussion”, Journal of the Optical Society of America A, vol. 37, Issue 3, Mar. 2020, pp. 1-17. [cited by applicant]
Khizar Hayat, “Multimedia super-resolution via deep learning: A survey”, Digital Signal Processing, vol. Oct. 2018, pp. 198-217. [cited by applicant]
Ian J. Goodfellow et al., “Generative Adversarial Nets,” arXiv:1406.2661, Jun. 2014, pp. 1-9. [cited by applicant]
Phillip Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks,” CVPR2016, Nov. 2016, pp. 1-17. [cited by applicant]
Han Zhang et al., “Self-Attention Generative Adversarial Networks,” arXiv:1805.08318, May 2018, pp. 1-10. [cited by applicant]
Harsh Nilesh Pathak et al., “Efficient Super Resolution For Large-Scale Images Using Attentional GAN,” IEEE International Conference on Big Data, Dec. 2018, pp. 1777-1786. [cited by applicant]
“International Search Report (Form PCT/ISA/210) of PCT/JP2020/007383,” mailed on Apr. 14, 2020, with English translation thereof, pp. 1-5. [cited by applicant]
“Written Opinion of the International Searching Authority (Form PCT/ISA/237)” of PCT/JP2020/007383, mailed on Apr. 14, 2020, with English translation thereof, pp. 1-7. [cited by applicant]
Cited By (3)
US 12,579,720 US 12,724,323 US 12,737,916