IP Library Granted Patent US 11,094,043
Granted Patent B2
US 11,094,043 · App. 16/141,843 · Granted Aug 17, 2021

Generation of high dynamic range visual media

Inventors: Nima Khademi Kalantari (San Diego, CA); Ravi Ramamoorthi (Carlsbad, CA)
Assignee: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
G06T5/009G06N3/084G06T5/008G06T5/50G06T7/337G06T2207/10016G06T2207/20208
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,094,043
App. No.
16/141,843
Granted
Aug 17, 2021
Kind
B2
Abstract

Devices, systems and methods for generating high dynamic range images and video from a set of low dynamic range images and video using convolution neural networks (CNNs) are described. One exemplary method for generating high dynamic range visual media includes generating, using a first CNN to merge a first set of images having a first dynamic range, a final image having a second dynamic range that is greater than the first dynamic range. Another exemplary method for generating training data includes generating sets of static and dynamic images having a first dynamic range, generating, based on a weighted sum of the set of static images, a set of ground truth images having a second dynamic range greater than the first dynamic range, and replacing at least one of the set of dynamic images with an image from the set of static images to generate a set of training images.

Claims (40)

1. A method for visual media processing, comprising:

generating a first set of images based on performing an alignment of a set of input images having a first dynamic range, the first set of images having the first dynamic range;

generating a second set of images having a second dynamic range that is greater than the first dynamic range;

aligning the second set of images to generate a third set of images having the second dynamic range; and

generating, using a first convolutional neural network (CNN) and based on at least the third set of images, a final image having the second dynamic range,

wherein the alignment of the set of input images is performed using a plurality of sub-CNNs, wherein each of the plurality of sub-CNNs operates at a distinct resolution, and wherein at least one of the plurality of sub-CNNs is followed by a rectified linear unit (ReLU).

2. The method of claim 1 , wherein each of the first set of images having the first dynamic range have a same exposure or different exposures.

3. The method of claim 1 , wherein performing the alignment of the set of input images is based on an optical flow algorithm.

4. The method of claim 1 , wherein performing the alignment of the set of input images is further based on a plurality of optical flows generated using a corresponding sub-CNN of the plurality of sub-CNNs.

5. The method of claim 4 , further comprising:

generating, using the first CNN with an input comprising the first set of images, a set of estimated images having the second dynamic range; and

minimizing an error between a set of ground truth images and the set of estimated images.

6. The method of claim 1 , wherein a last of the plurality of sub-CNNs is followed by a linear activation function.

7. An apparatus for visual media processing, comprising:

a processor; and

a memory that comprises instructions stored thereupon, wherein the instructions when executed by the processor configure the processor to:

generate a first set of images based on performing an alignment of a set of input images having a first dynamic range, the first set of images having the first dynamic range;

generate a second set of images having a second dynamic range that is greater than the first dynamic range;

align the second set of images to generate a third set of images having the second dynamic range; and

generate, using a first convolutional neural network (CNN) and based on at least the third set of images, a final image having the second dynamic range,

wherein the alignment of the set of input images is performed using a plurality of sub-CNNs, wherein each of the plurality of sub-CNNs operates at a distinct resolution, and wherein at least one of the plurality of sub-CNNs is followed by a rectified linear unit (ReLU).

8. The apparatus of claim 7 , wherein each of the first set of images having the first dynamic range have a same exposure or different exposures.

9. The apparatus of claim 7 , wherein performing the alignment of the set of input images is based on an optical flow algorithm.

10. The apparatus of claim 7 , wherein performing the alignment of the set of input images is further based on a plurality of optical flows generated using a corresponding sub-CNN of the plurality of sub-CNNs.

11. The apparatus of claim 7 , wherein the processor is further configured to:

generate, using the first CNN with an input comprising the first set of images, a set of estimated images having the second dynamic range; and

minimize an error between a set of ground truth images and the set of estimated images.

12. The apparatus of claim 7 , wherein a last of the plurality of sub-CNNs is followed by a linear activation function.

13. A non-transitory computer readable program storage medium having code stored thereon, the code, when executed by a processor, causing the processor to implement a method for visual media processing, the method comprising:

generating a first set of images based on performing an alignment of a set of input images having a first dynamic range, the first set of images having the first dynamic range;

generating a second set of images having a second dynamic range that is greater than the first dynamic range;

aligning the second set of images to generate a third set of images having the second dynamic range; and

generating, using a first convolutional neural network (CNN) and based on at least the third set of images, a final image having the second dynamic range,

wherein the alignment of the set of input images is performed using a plurality of sub-CNNs, wherein each of the plurality of sub-CNNs operates at a distinct resolution, and wherein at least one of the plurality of sub-CNNs is followed by a rectified linear unit (ReLU).

14. The non-transitory computer readable program storage medium of claim 13 , wherein performing the alignment of the set of input images is further based on a plurality of optical flows generated using a corresponding sub-CNN of the plurality of sub-CNNs.

15. The non-transitory computer readable program storage medium of claim 14 , wherein the method further comprises training the first CNN based on:

generating, using the first CNN with an input comprising the first set of images, a set of estimated images having the second dynamic range; and

minimizing an error between a set of ground truth images and the set of estimated images.

16. The non-transitory computer readable program storage medium of claim 13 , wherein each of the first set of images having the first dynamic range have a same exposure or different exposures.

17. The non-transitory computer readable program storage medium of claim 13 , wherein performing the alignment of the set of input images is based on an optical flow algorithm.

Assignments (3)
CONFIRMATORY LICENSE Recorded Dec 6, 2022
From: UNIVERSITY OF CALIFORNIA SAN DEIGO
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 062067/0301 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 25, 2020
From: KALANTARI, NIMA KHADEMI; RAMAMOORTHI, RAVI
To: THE REGENTS OF THE UNIVERSITY OF CALIFORNIA
Reel/Frame 054471/0432 →
CONFIRMATORY LICENSE Recorded Sep 25, 2020
From: UNIVERSITY OF CALIFORNIA, SAN DIEGO
To: NATIONAL SCIENCE FOUNDATION
Reel/Frame 053892/0587 →
Continuity (2)
Provisional Application 62562922 · Sep 25, 2017
Related Publication 20190096046A1 · Mar 28, 2019
Cited By (1)
US 12,452,546