IP Library Granted Patent US 11,503,221
Granted Patent B2
US 11,503,221 · App. 16/837,785 · Granted Nov 15, 2022

System and method for motion warping using multi-exposure frames

Inventors: Anqi Yang (Plano, TX); John W. Glotzbach (Allen, TX); Hamid R. Sheikh (Allen, TX)
Assignee: Samsung Electronics Co., Ltd.
H04N5/2353G06T5/003G06T5/50H04N5/2356H04N19/42G06T2207/10144G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,503,221
App. No.
16/837,785
Filed
Apr 1, 2020
Granted
Nov 15, 2022
Kind
B2
Art Unit
2698
USPC
348/239
Abstract

A method includes obtaining, using at least one image sensor of an electronic device, a first image frame and multiple second image frames of a scene. Each of the second image frames has an exposure time different from an exposure time of the first image frame. The method also includes encoding, using at least one processor, each of the first image frame and the second image frames using a convolutional neural network to generate a corresponding feature map. The method further includes aligning, using the at least one processor, encoded features of the feature map corresponding to the first image frame with encoded features of the feature maps corresponding to the second image frames.

Claims (48)

1. A method comprising:

obtaining, using at least one image sensor of an electronic device, a first image frame and multiple second image frames of a scene, each of the second image frames having an exposure time different from an exposure time of the first image frame;

encoding, using at least one processor, each of the first image frame and the second image frames using a convolutional neural network to generate a corresponding feature map;

generating, using the at least one processor, at least one optical flow map representing pixel-wise differences between at least one pair of image frames among the first and second image frames using at least one optical flow network; and

aligning, using the at least one processor, encoded features of the feature map corresponding to the first image frame with encoded features of the feature maps corresponding to the second image frames by performing warping operations on the encoded features of the feature maps corresponding to the second image frames using the at least one optical flow map.

2. The method of claim 1 , further comprising:

decoding, using the at least one processor, the feature maps having the aligned encoded features using the convolutional neural network to generate a target image frame of the scene.

3. The method of claim 1 , wherein the exposure time of each second image frame is shorter than the exposure time of the first image frame.

4. The method of claim 1 , wherein:

the exposure time of one of the second image frames is longer than the exposure time of the first image frame; and

the exposure time of another of the second image frames is shorter than the exposure time of the first image frame.

5. The method of claim 1 , further comprising:

concatenating the feature maps having the aligned encoded features before decoding the feature maps having the aligned encoded features.

6. The method of claim 1 , wherein the first image frame is used as a reference frame and the second image frames are used as non-reference frames.

7. The method of claim 1 , wherein the convolutional neural network comprises a generative adversarial network.

8. An electronic device comprising:

at least one image sensor; and

at least one processing device configured to:

obtain a first image frame and multiple second image frames of a scene using the at least one image sensor, each of the second image frames having an exposure time different from an exposure time of the first image frame;

encode each of the first image frame and the second image frames using a convolutional neural network to generate a corresponding feature map;

generate at least one optical flow map representing pixel-wise differences between at least one pair of image frames among the first and second image frames using at least one optical flow network; and

align encoded features of the feature map corresponding to the first image frame with encoded features of the feature maps corresponding to the second image frames;

wherein, to align the encoded features, the at least one processing device is configured to perform warping operations on the encoded features of the feature maps corresponding to the second image frames using the at least one optical flow map.

9. The electronic device of claim 8 , wherein the at least one processing device is further configured to:

decode the feature maps having the aligned encoded features using the convolutional neural network to generate a target image frame of the scene.

10. The electronic device of claim 8 , wherein the exposure time of each second image frame is shorter than the exposure time of the first image frame.

11. The electronic device of claim 8 , wherein:

the exposure time of one of the second image frames is longer than the exposure time of the first image frame; and

the exposure time of another of the second image frames is shorter than the exposure time of the first image frame.

12. The electronic device of claim 8 , wherein the at least one processing device is further configured to concatenate the feature maps having the aligned encoded features before decoding the feature maps having the aligned encoded features.

13. The electronic device of claim 8 , wherein the first image frame is used as a reference frame and the second image frames are used as non-reference frames.

14. The electronic device of claim 8 , wherein the convolutional neural network comprises a generative adversarial network.

15. A non-transitory machine-readable medium containing instructions that when executed cause at least one processor of an electronic device to:

obtain, using at least one image sensor of the electronic device, a first image frame and multiple second image frames of a scene, each of the second image frames having an exposure time different from an exposure time of the first image frame;

encode each of the first image frame and the second image frames using a convolutional neural network to generate a corresponding feature map;

generate at least one optical flow map representing pixel-wise differences between at least one pair of image frames among the first and second image frames using at least one optical flow network; and

align encoded features of the feature map corresponding to the first image frame with encoded features of the feature maps corresponding to the second image frames;

wherein the instructions that when executed cause the at least one processor to align the encoded features comprise instructions that when executed cause the at least one processor to perform warping operations on the encoded features of the feature maps corresponding to the second image frames using the at least one optical flow map.

16. The non-transitory machine-readable medium of claim 15 , wherein the instructions when executed further cause the at least one processor to:

decode the feature maps having the aligned encoded features using the convolutional neural network to generate a target image frame of the scene.

17. The non-transitory machine-readable medium of claim 15 , wherein the exposure time of each second image frame is shorter than the exposure time of the first image frame.

18. The non-transitory machine-readable medium of claim 15 , wherein:

the exposure time of one of the second image frames is longer than the exposure time of the first image frame; and

the exposure time of another of the second image frames is shorter than the exposure time of the first image frame.

19. The non-transitory machine-readable medium of claim 15 , wherein the instructions when executed further cause the at least one processor to:

concatenate the feature maps having the aligned encoded features before decoding the feature maps having the aligned encoded features.

20. The non-transitory machine-readable medium of claim 15 , wherein the first image frame is used as a reference frame and the second image frames are used as non-reference frames.

21. The non-transitory machine-readable medium of claim 15 , wherein the convolutional neural network comprises a generative adversarial network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2020
From: YANG, ANQI; GLOTZBACH, JOHN W.; SHEIKH, HAMID R.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 052288/0944 →
Continuity (1)
Related Publication 20210314474A1 · Oct 7, 2021
Cited By (3)
US 12,354,312 US 12,384,409 US 12,586,159