IP Library Granted Patent US 12700064
Granted Patent B2
US 12700064 · App. 18/350,558 · Granted Aug 4, 2026

Multi-stage multi-frame denoising with neural radiance field networks or other machine learning models

Inventors: Yahang Li (Champaign, IL); Nguyen Thang Long Le (Garland, TX); Hamid R. Sheikh (Allen, TX)
Assignee: Samsung Electronics Co., Ltd.
G06T5/50G06T5/70G06T7/30G06T7/579G06T2207/10024G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12700064
App. No.
18/350,558
Granted
Aug 4, 2026
Kind
B2
Abstract

A method includes obtaining, using at least one processing device of an electronic device, raw image frames of a scene. The raw image frames include different sets of raw image frames captured at different viewpoints and different viewing angles relative to the scene. The method also includes performing, using the at least one processing device, blending of each set of raw image frames in order to generate blended image frames of the scene. The method further includes training, using the at least one processing device, a machine learning model using the blended image frames. The machine learning model is trained to generate three-dimensional (3D) information about the scene from viewpoints and viewing angles not captured in the sets of raw image frames.

Claims (82)

1 . A method comprising:

obtaining, using at least one processing device of an electronic device, raw image frames of a scene, wherein the raw image frames include different sets of raw image frames captured at different viewpoints and different viewing angles relative to the scene;

performing, using the at least one processing device, blending of the raw image frames within each set of raw image frames in order to generate multiple blended image frames of the scene; and

training, using the at least one processing device, a machine learning model using the blended image frames, the machine learning model trained to generate three-dimensional (3D) information about the scene from viewpoints and viewing angles not captured in the sets of raw image frames.

2 . The method of claim 1 , wherein performing the blending of the raw image frames within each set of raw image frames comprises, for each set of raw image frames:

performing alignment of at least some of the raw image frames in the set of raw image frames;

performing deghosting of one or more of the raw image frames in the set of raw image frames; and

blending the raw image frames in the set of raw image frames after the alignment and the deghosting.

3 . The method of claim 1 , wherein training the machine learning model comprises:

generating structure from motion using the blended image frames in order to identify 3D feature points within the scene;

identifying poses of one or more imaging sensors that capture the raw image frames based on the 3D feature points within the scene; and

generating rays based on the poses of the one or more imaging sensors.

4 . The method of claim 3 , wherein training the machine learning model further comprises:

using the machine learning model to generate color and density information for points along each ray; and

training the machine learning model to generate the color and density information while minimizing a loss between (i) a rendered image based on the color and density information and (ii) a tone-mapped version of a noisy image.

5 . The method of claim 1 , wherein:

the blended image frames comprise raw blended image frames; and

the machine learning model is trained using the raw blended image frames by, for each set of raw image frames:

identifying a ground truth image, the ground truth image comprising one of: a selected raw image frame from the set of raw image frames or a blended version of multiple selected raw image frames from the set of raw image frames;

generating image data in multiple color channels using the machine learning model;

applying color filter array masking to the image data in the color channels in order to generate masked image data; and

comparing the masked image data and a tone-mapped version of image data from the ground truth image in order to identify a loss associated with the machine learning model.

6 . The method of claim 1 , wherein:

the blended image frames comprise raw blended image frames; and

the machine learning model is trained using the raw blended image frames by, for each set of raw image frames:

selecting one of the raw image frames from the set of raw image frames as a ground truth image;

generating image data in multiple color channels using the machine learning model, the image data having a color filter array form; and

comparing the image data in the color channels and a tone-mapped version of image data from the ground truth image in order to identify a loss associated with the machine learning model.

7 . The method of claim 1 , wherein:

the blended image frames comprise RGB image frames; and

the machine learning model is trained using the RGB image frames.

8 . The method of claim 1 , wherein the machine learning model comprises a neural radiance field network configured to generate color and density information associated with the scene.

9 . An electronic device comprising:

at least one processing device configured to:

obtain raw image frames of a scene, wherein the raw image frames include different sets of raw image frames captured at different viewpoints and different viewing angles relative to the scene;

perform blending of the raw image frames within each set of raw image frames in order to generate multiple blended image frames of the scene; and

train a machine learning model using the blended image frames, the machine learning model trained to generate three-dimensional (3D) information about the scene from viewpoints and viewing angles not captured in the sets of raw image frames.

10 . The electronic device of claim 9 , wherein, to perform the blending of the raw image frames within each set of raw image frames, the at least one processing device is configured, for each set of raw image frames, to:

perform alignment of at least some of the raw image frames in the set of raw image frames;

perform deghosting of one or more of the raw image frames in the set of raw image frames; and

blend the raw image frames in the set of raw image frames after the alignment and the deghosting.

11 . The electronic device of claim 9 , wherein, to train the machine learning model, the at least one processing device is configured to:

generate structure from motion using the blended image frames in order to identify 3D feature points within the scene;

identify poses of one or more imaging sensors that capture the raw image frames based on the 3D feature points within the scene; and

generate rays based on the poses of the one or more imaging sensors.

12 . The electronic device of claim 11 , wherein, to train the machine learning model, the at least one processing device is further configured to:

use the machine learning model to generate color and density information for points along each ray; and

train the machine learning model to generate the color and density information while minimizing a loss between (i) a rendered image based on the color and density information and (ii) a tone-mapped version of a noisy image.

13 . The electronic device of claim 9 , wherein:

the blended image frames comprise raw blended image frames;

the at least one processing device is configured to train the machine learning model using the raw blended image frames; and

to train the machine learning model, the at least one processing device is configured, for each set of raw image frames, to:

identify a ground truth image, the ground truth image comprising one of: a selected raw image frame from the set of raw image frames or a blended version of multiple selected raw image frames from the set of raw image frames;

generate image data in multiple color channels using the machine learning model;

apply color filter array masking to the image data in the color channels in order to generate masked image data; and

compare the masked image data and a tone-mapped version of image data from the ground truth image in order to identify a loss associated with the machine learning model.

14 . The electronic device of claim 9 , wherein:

the blended image frames comprise raw blended image frames;

the at least one processing device is configured to train the machine learning model using the raw blended image frames; and

to train the machine learning model, the at least one processing device is configured, for each set of raw image frames, to:

select one of the raw image frames from the set of raw image frames as a ground truth image;

generate image data in multiple color channels using the machine learning model, the image data having a color filter array form; and

compare the image data in the color channels and a tone-mapped version of image data from the ground truth image in order to identify a loss associated with the machine learning model.

15 . The electronic device of claim 9 , wherein:

the blended image frames comprise RGB image frames; and

the at least one processing device is configured to train the machine learning model using the RGB image frames.

16 . The electronic device of claim 9 , wherein the machine learning model comprises a neural radiance field network configured to generate color and density information associated with the scene.

17 . A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an electronic device to:

obtain raw image frames of a scene, wherein the raw image frames include different sets of raw image frames captured at different viewpoints and different viewing angles relative to the scene;

perform blending of the raw image frames within each set of raw image frames in order to generate multiple blended image frames of the scene; and

train a machine learning model using the blended image frames, the machine learning model trained to generate three-dimensional (3D) information about the scene from viewpoints and viewing angles not captured in the sets of raw image frames.

18 . The non-transitory machine readable medium of claim 17 , wherein the instructions that when executed cause the at least one processor to perform the blending of the raw image frames within each set of raw image frames comprise instructions that when executed cause the at least one processor, for each set of raw image frames, to:

perform alignment of at least some of the raw image frames in the set of raw image frames;

perform deghosting of one or more of the raw image frames in the set of raw image frames; and

blend the raw image frames in the set of raw image frames after the alignment and the deghosting.

19 . The non-transitory machine readable medium of claim 17 , wherein the instructions that when executed cause the at least one processor to train the machine learning model comprise instructions that when executed cause the at least one processor to:

generate structure from motion using the blended image frames in order to identify 3D feature points within the scene;

identify poses of one or more imaging sensors that capture the raw image frames based on the 3D feature points within the scene; and

generate rays based on the poses of the one or more imaging sensors.

20 . The non-transitory machine readable medium of claim 17 , wherein the instructions that when executed cause the at least one processor to train the machine learning model comprise instructions that when executed cause the at least one processor to:

use the machine learning model to generate color and density information for points along each ray; and

train the machine learning model to generate the color and density information while minimizing a loss between (i) a rendered image based on the color and density information and (ii) a tone-mapped version of a noisy image.