IP Library Granted Patent US 12670558
Granted Patent B2
US 12670558 · App. 18/367,845 · Granted Jun 30, 2026

Apparatus for extracting noise from image and method thereof

Inventor: Jae Hoon Cho (Seoul, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION
G06T5/70G06T5/50G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12670558
App. No.
18/367,845
Filed
Sep 13, 2023
Granted
Jun 30, 2026
Kind
B2
Art Unit
2662
USPC
382/100
Abstract

An apparatus for extracting image noise includes an encoder that encodes a first image and a second image in a time-series image to output an encoding feature map. The apparatus also includes a noise information extraction layer that generates a first noise prototype vector to an M-th noise prototype vector of the first image and the second image, and outputs first noise information with reference to at least a part of the first noise prototype vector to the M-th noise prototype vector and the encoding feature map. The apparatus additionally includes a decoder that decodes the first noise information to output second noise information. The apparatus further includes a parameter updater that outputs a loss with reference to the first image, the second image, and the corresponding second noise information, and updates parameters of the encoder, the noise information extraction layer and/or the decoder through back propagation using the loss.

Claims (57)

1 . An apparatus for extracting noise from an image, the apparatus comprising:

an encoder configured to apply an encoding operation to a first image and a second image included in a time-series image to output an encoding feature map of each of the first image and the second image;

a noise information extraction layer configured to apply a noise prototype vector generation operation to the encoding feature map to generate a first noise prototype vector to an M-th noise prototype vector of each of the first image and the second image, and output first noise information of each of the first image and the second image with reference to at least a part of the first noise prototype vector to the M-th noise prototype vector and the encoding feature map;

a decoder configured to apply a decoding operation to the first noise information to output second noise information of each of the first image and the second image; and

a processor,

wherein the processor is configured to:

output a loss with reference to the first image, the second image, and the corresponding second noise information, and update parameters of at least some of the encoder, the noise information extraction layer and the decoder through back propagation using the loss, wherein the loss includes a first loss and a second loss;

generate a first denoised image and a second denoised image by using the first image, the second image, and the second noise information corresponding thereto;

output the first loss with reference to the first denoised image and the second denoised image;

determine a specific noise prototype vector that satisfies a preset similarity condition among the first noise prototype vector to the M-th noise prototype vector with reference to the vector similarity for each pixel; and

output the second loss with reference to an encoding vector for each pixel and the specific noise prototype vector,

wherein the noise information extraction layer is configured to:

apply a first attention mapping operation to an M-th attention mapping operation to the encoding feature map to generate first attention mapping information to M-th attention mapping information of each of the first image and the second image;

generate the first noise prototype vector to the M-th noise prototype vector of each of the first image and the second image with reference to the encoding feature map and the first attention mapping information to the M-th attention mapping information;

generate an integrated noise vector with reference to the first noise prototype vector to the M-th noise prototype vector and vector similarities corresponding thereto, wherein the vector similarity corresponds to a similarity between each of the first noise prototype vector to the M-th noise prototype vector and the encoding feature map;

output the first noise information of each of the first image and the second image with reference to the encoding feature map and the integrated noise vector;

apply the first attention mapping operation to the M-th attention mapping operation to the encoding vector for each pixel corresponding to a plurality of pixels of the encoding feature map to generate first noise attention mapping information to an M-th noise attention mapping information for each pixel corresponding to each encoding vector for each pixel of the first image and the second image;

generate the first noise prototype vector to the M-th noise prototype vector of each of the first image and the second image with reference to the encoding vector for each pixel and the first noise attention mapping information for each pixel to the M-th noise attention mapping information for each pixel;

generate an integrated noise vector for each pixel with reference to the first noise prototype vector to the M-th noise prototype vector and the vector similarity for each pixel corresponding thereto, wherein the vector similarity for each pixel corresponds to a similarity between each of the first noise prototype vector to the M-th noise prototype vector and the encoding vector for each pixel; and

output the first noise information of each of the first image and the second image with reference to the integrated noise vector for each pixel and the encoding vector for each pixel.

2 . The apparatus of claim 1 , wherein the noise information extraction layer is configured to:

perform a channel-wise summation of the integrated noise vector for each pixel and the encoding vector for each pixel to output the first noise information.

3 . The apparatus of claim 1 , wherein:

the loss further includes a third loss, a fourth loss, and a fifth loss, and

the processor is further configured to:

output the third loss with reference to the first denoised image and the second image or to the second denoised image and the first image,

output the fourth loss with reference to a specific denoised image corresponding to a specific image among the first image and the second image, specific noise information corresponding to the specific image, and the specific image, and

output the fifth loss with reference to the first noise prototype vector to the M-th noise prototype vector.

4 . A method of learning image noise extraction, comprising:

outputting an encoding feature map of each of a first image and a second image by applying an encoding operation to a first image and a second image included in a time-series image;

applying a noise prototype vector generation operation to the encoding feature map to generate a first noise prototype vector to an M-th noise prototype vector of each of the first image and the second image;

outputting first noise information of each of the first image and the second image with reference to at least a part of the first noise prototype vector to the M-th noise prototype vector and the encoding feature map;

outputting second noise information of each of the first image and the second image by applying a decoding operation to the first noise information;

outputting a loss with reference to the first image, the second image, and the corresponding second noise information; and

updating parameters of at least some of an encoder, a noise information extraction layer and a decoder through back propagation using the loss, wherein the loss includes a first loss and a second loss,

wherein updating the parameters includes:

generating a first denoised image and a second denoised image by using the first image, the second image, and the second noise information corresponding thereto,

outputting the first loss with reference to the first denoised image and the second denoised image,

determining a specific noise prototype vector that satisfies a preset similarity condition among the first noise prototype vector to the M-th noise prototype vector with reference to the vector similarity for each pixel, and

outputting the second loss with reference to an encoding vector for each pixel and the specific noise prototype vector, and

wherein outputting the first noise information includes:

applying a first attention mapping operation to an M-th attention mapping operation to the encoding feature map to generate first attention mapping information to M-th attention mapping information of each of the first image and the second image,

generating the first noise prototype vector to the M-th noise prototype vector of each of the first image and the second image with reference to the encoding feature map and the first attention mapping information to the M-th attention mapping information,

generating an integrated noise vector with reference to the first noise prototype vector to the M-th noise prototype vector and vector similarities corresponding thereto, wherein the vector similarity corresponds to a similarity between each of the first noise prototype vector to the M-th noise prototype vector and the encoding feature map,

outputting the first noise information of each of the first image and the second image with reference to the encoding feature map and the integrated noise vector,

applying the first attention mapping operation to the M-th attention mapping operation to the encoding vector for each pixel corresponding to a plurality of pixels of the encoding feature map to generate first noise attention mapping information to an M-th noise attention mapping information for each pixel corresponding to each encoding vector for each pixel of the first image and the second image,

generating the first noise prototype vector to the M-th noise prototype vector of each of the first image and the second image with reference to the encoding vector for each pixel and the first noise attention mapping information for each pixel to the M-th noise attention mapping information for each pixel,

generating an integrated noise vector for each pixel with reference to the first noise prototype vector to the M-th noise prototype vector and the vector similarity for each pixel corresponding thereto, wherein the vector similarity for each pixel corresponds to a similarity between each of the first noise prototype vector to the M-th noise prototype vector and the encoding vector for each pixel, and

outputting the first noise information of each of the first image and the second image with reference to the integrated noise vector for each pixel and the encoding vector for each pixel.

5 . The method of claim 4 , wherein outputting the first noise information includes:

performing a channel-wise summation of the integrated noise vector for each pixel and the encoding vector for each pixel to output the first noise information.

6 . The method of claim 4 , wherein:

the loss further includes a third loss, a fourth loss, and a fifth loss, and

updating the parameters further includes:

outputting the third loss with reference to the first denoised image and the second image or to the second denoised image and the first image,

outputting the fourth loss with reference to a specific denoised image corresponding to a specific image among the first image and the second image, specific noise information corresponding to the specific image, and the specific image, and

outputting the fifth loss with reference to the first noise prototype vector to the M-th noise prototype vector.