IP Library › Granted Patent US 12,272,032
Granted Patent B2
US 12,272,032 · App. 17/820,795 · Granted Apr 8, 2025

Machine learning-based approaches for synthetic training data generation and image sharpening

Inventors: Devendra K. Jangid (Santa Barbara, CA); John Seokjun Lee (Allen, TX); Hamid R. Sheikh (Allen, TX)
Assignee: Samsung Electronics Co., Ltd.
G06T5/70G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,272,032
App. No.
17/820,795
Granted
Apr 8, 2025
Kind
B2
Abstract

A method includes obtaining an input image that contains blur. The method also includes providing the input image to a trained machine learning model, where the trained machine learning model includes (i) a shallow feature extractor configured to extract one or more feature maps from the input image and (ii) a deep feature extractor configured to extract deep features from the one or more feature maps. The method further includes using the trained machine learning model to generate a sharpened output image. The trained machine learning model is trained using ground truth training images and input training images, where the input training images include versions of the ground truth training images with blur created using demosaic and noise filtering operations.

Claims (76)

1. A method comprising:

obtaining an input image that contains blur;

providing the input image to a trained machine learning model, the trained machine learning model comprising (i) a shallow feature extractor configured to extract one or more feature maps from the input image and (ii) a deep feature extractor configured to extract deep features from the one or more feature maps; and

using the trained machine learning model to generate a sharpened output image;

wherein the trained machine learning model is trained using ground truth training images and input training images, the input training images comprising versions of the ground truth training images with blur created using demosaic and noise filtering operations.

2. The method of claim 1 , further comprising:

training a machine learning model using the ground truth training images and the input training images in order to generate the trained machine learning model.

3. The method of claim 2 , further comprising:

generating the input training images by:

introducing noise into image data used to form the ground truth training images using a synthetic noise model in order to generate noisy images;

performing a demosaic interpolation using the noisy images in order to generate demosaiced images; and

applying a noise filter to the demosaiced images in order to generate the input training images.

4. The method of claim 1 , wherein:

the shallow feature extractor comprises a convolution layer configured to extract the one or more feature maps; and

the deep feature extractor comprises multiple residual convolutional layers and multiple channel attention layers configured to generate one or more improved feature maps, the one or more improved feature maps containing the deep features.

5. The method of claim 1 , wherein:

the deep feature extractor comprises multiple residual groups; and

each residual group comprises multiple residual channel attention blocks (RCABs).

6. The method of claim 1 , wherein the trained machine learning model represents a single-stage machine learning model.

7. The method of claim 1 , wherein the trained machine learning model comprises a raw image enhancement network.

8. An apparatus comprising:

at least one processing device configured to:

obtain an input image that contains blur;

process the input image using a trained machine learning model, the trained machine learning model comprising (i) a shallow feature extractor configured to extract one or more feature maps from the input image and (ii) a deep feature extractor configured to extract deep features from the one or more feature maps; and

use the trained machine learning model to generate a sharpened output image;

wherein the trained machine learning model is trained using ground truth training images and input training images, the input training images comprising versions of the ground truth training images with blur created using demosaic and noise filtering operations.

9. The apparatus of claim 8 , wherein the at least one processing device is further configured to train a machine learning model using the ground truth training images and the input training images in order to generate the trained machine learning model.

10. The apparatus of claim 9 , wherein:

the at least one processing device is further configured to generate the input training images; and

to generate the input training images, the at least one processing device is configured to:

introduce noise into image data used to form the ground truth training images using a synthetic noise model in order to generate noisy images;

perform a demosaic interpolation using the noisy images in order to generate demosaiced images; and

apply a noise filter to the demosaiced images in order to generate the input training images.

11. The apparatus of claim 8 , wherein:

the shallow feature extractor comprises a convolution layer configured to extract the one or more feature maps; and

the deep feature extractor comprises multiple residual convolutional layers and multiple channel attention layers configured to generate one or more improved feature maps, the one or more improved feature maps containing the deep features.

12. The apparatus of claim 8 , wherein:

the deep feature extractor comprises multiple residual groups; and

each residual group comprises multiple residual channel attention blocks (RCABs).

13. The apparatus of claim 8 , wherein the trained machine learning model represents a single-stage machine learning model.

14. The apparatus of claim 8 , wherein the trained machine learning model comprises a raw image enhancement network.

15. A non-transitory computer readable medium containing instructions that when executed cause at least one processor to:

obtain an input image that contains blur;

process the input image using a trained machine learning model, the trained machine learning model comprising (i) a shallow feature extractor configured to extract one or more feature maps from the input image and (ii) a deep feature extractor configured to extract deep features from the one or more feature maps; and

use the trained machine learning model to generate a sharpened output image;

wherein the trained machine learning model is trained using ground truth training images and input training images, the input training images comprising versions of the ground truth training images with blur created using demosaic and noise filtering operations.

16. The non-transitory computer readable medium of claim 15 , wherein the instructions when executed further cause the at least one processor to train a machine learning model using the ground truth training images and the input training images in order to generate the trained machine learning model.

17. The non-transitory computer readable medium of claim 16 , wherein:

the instructions when executed further cause the at least one processor to generate the input training images; and

the instructions that when executed cause the at least one processor to generate the input training images comprise instructions that when executed cause the at least one processor to:

introduce noise into image data used to form the ground truth training images using a synthetic noise model in order to generate noisy images;

perform a demosaic interpolation using the noisy images in order to generate demosaiced images; and

apply a noise filter to the demosaiced images in order to generate the input training images.

18. The non-transitory computer readable medium of claim 15 , wherein:

the shallow feature extractor comprises a convolution layer configured to extract the one or more feature maps; and

the deep feature extractor comprises multiple residual convolutional layers and multiple channel attention layers configured to generate one or more improved feature maps, the one or more improved feature maps containing the deep features.

19. The non-transitory computer readable medium of claim 15 , wherein:

the deep feature extractor comprises multiple residual groups; and

each residual group comprises multiple residual channel attention blocks (RCABs).

20. The non-transitory computer readable medium of claim 15 , wherein the trained machine learning model represents a single-stage machine learning model.

21. A method comprising:

obtaining multiple ground truth training images;

generating multiple input training images using the ground truth training images, wherein generating the input training images comprises performing demosaic and noise filtering operations to add blur to the ground truth training images; and

training a machine learning model to remove blur from images and generate sharpened images using the ground truth training images and the input training images.

22. The method of claim 21 , wherein:

each ground truth training image has a first ISO value; and

each input training image is generated using an associated one of the ground truth training images and has a second ISO value larger than the first ISO value of the associated ground truth training image.

23. The method of claim 21 , wherein performing the demosaic and noise filtering operations comprises:

introducing noise into image data used to form the ground truth training images using a synthetic noise model in order to generate noisy images;

performing a demosaic interpolation using the noisy images in order to generate demosaiced images; and

applying a noise filter to the demosaiced images in order to generate the input training images.

24. The method of claim 23 , wherein the synthetic noise model models read and shot noise.

25. The method of claim 23 , wherein:

the demosaic interpolation converts raw image data into YUV image data; and

the noise filter is applied to the YUV image data in order to generate raw RGB image data.

26. The method of claim 21 , wherein training the machine learning model comprises training the machine learning model based on an L1 Charbonnier with edge loss using an Adam optimizer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2022
From: JANGID, DEVENDRA K.; LEE, JOHN SEOKJUN; SHEIKH, HAMID R.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 060848/0612 →
Continuity (1)
Related Publication 20240062342A1 · Feb 22, 2024
References Cited (14)
US 11756166B2 · Saharia · 2023 [cited by examiner]
US 20180336662A1 · Kimura · 2018 [cited by examiner]
US 20210241421A1 · Pan et al. · 2021 [cited by applicant]
US 20210241429A1 · Pan et al. · 2021 [cited by applicant]
US 20220122235A1 · Liang et al. · 2022 [cited by applicant]
US 20230410259A1 · Melnyk · 2023 [cited by examiner]
CN 113902647A · 2022 [cited by applicant]
Zhang, Deep motion blur removal using noisy/blurry image pairs, Journal of Electronic Imaging, May/Jun. 2021, vol. 30(3) (Year: 2021). [cited by examiner]
Zamir, Learning Enriched Features for Fast Image Restoration and Enhancement, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 45, No. 2, Feb. 2023 (Date of publication Apr. 13, 2022) (Year: 2022). [cited by examiner]
Brown, “Understanding the In-Camera Image Processing Pipeline for Computer Vision,” 2016 IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2016, 354 pages. [cited by applicant]
Zhang et al., “Image Super-Resolution Using Very Deep Residual Channel Attention Networks,” European Conference on Computer Vision, Oct. 2018, 16 pages. [cited by applicant]
Zamir et al., “Multi-Stage Progressive Image Restoration,” IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, 11 pages. [cited by applicant]
Ramanarayanan et al., “MRI Super-Resolution using Laplacian Pyramid Convolutional Neural Networks with Isotropic Undecimated Wavelet Loss,” 42nd Annual International Conference of the IEEE Engineering in Medicine & Biol… [cited by applicant]
Zha et al., “A Lightweight Dense Connected Approach with Attention on Single Image Super-Resolution,” Electronics, vol. 10, Issue 11, Apr. 2021, 14 pages. [cited by applicant]