IP Library › Granted Patent US 12,333,689
Granted Patent B1
US 12,333,689 · App. 17/962,853 · Granted Jun 17, 2025

Methods and apparatus for end-to-end unsupervised multi-document blind image denoising

Inventors: Mehrdad Jabbarzadeh Gangeh (Mountain View, CA); Marcin Plata (Wroclaw, PL); Hamid Reza Motahari-Nezad (Los Altos, CA)
G06T5/70G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,689
App. No.
17/962,853
Granted
Jun 17, 2025
Kind
B1
Abstract

Methods and apparatus for end-to-end unsupervised multi-document blind image denoising is presented. The multi-document blind image denoiser removes various noise types from noisy documents without paired target cleaned documents and preserves the contents for optical character recognition. The end-to-end unsupervised multi-document blind image denoiser integrates a Mixture of Experts with a cycle-consistent GAN as the base network that effectively removes multiple types of noise, including salt & pepper noise, blurred and/or faded text, as well as watermarks from documents at various levels of intensity.

Claims (78)

1. An apparatus comprising:

a processor; and

a memory operatively coupled to the processor, the memory storing instructions to cause the processor to:

receive a noisy image without a paired clean image, the noisy image including at least one noise type of a plurality of noise types;

generate a blind denoised image from the noisy image using at least one denoising layer of a plurality of denoising layers of an end-to-end machine learning model based on the at least one noise type;

generate a blind noised image from the blind denoised image using at least one noising layer of a plurality of noising layers of the end-to-end machine learning model;

calculate a match score based on a comparison between the noisy image and the blind noised image; and

generate a clean image based on the match score and a match threshold.

2. The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to:

classify the noisy image based on an image structure type; and

store the noisy image in a database categorized by the image structure type.

3. The apparatus of claim 1 , wherein each noise type from the plurality of noise types is associated with a denoising layer of the plurality of denoising layers and a noising layer of the plurality of noising layers.

4. The apparatus of claim 1 , wherein the end-to-end machine-learning model includes an unsupervised machine learning model.

5. The apparatus of claim 1 , wherein the end-to-end machine-learning model implements a Mixture of Experts, the Mixture of Experts including:

a plurality of denoising gates, each denoising gate of the plurality of denoising gates configured to activate at least one denoising channel of a plurality of denoising channels of each denoising layer of the plurality of denoising layers;

a plurality of noising gates, each noising gate of the plurality of nosing gates configured to activate at least one noising channel of a plurality of noising channels of each noising layer of the plurality of noising layers; and

an embedder, the memory storing instructions to cause the processor to:

compute, via the embedder, a denoising mixture weight of a plurality of denoising mixture weights for each denoising gate of the plurality of denoising gates based on the noisy image; and

compute, via the embedder, a noising mixture weight of a plurality of noising mixture weights for each noising gate of the plurality of noising gates based on the blind denoised image.

6. The apparatus of claim 5 , wherein the embedder includes a convolutional embedder.

7. The apparatus of claim 5 , wherein the embedder includes a shallow embedder.

8. The apparatus of claim 5 , wherein the memory stores instructions to further cause the processor to:

generate an image pair based on the match score, the image pair including the clean image and the noisy image having the at least one noise type of the plurality of noise types; and

train the embedder based on the image pair including the at least one noise type of the plurality of noise types, to produce a plurality of trained denoising gates and a plurality of trained noising gates.

9. The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to generate a blind score of the blind denoised image.

10. The apparatus of claim 9 , wherein calculating the match score further by comparing the blind score of the blind denoised image and a second blind score of the blind noised image.

11. The apparatus of claim 1 , wherein:

the noisy image including a first noisy image;

the clean image is a first clean image;

the match score is a first match score;

the memory storing instructions to further cause the processor to:

generate a first image pair based on the first match score, the first image pair including the first noisy image and the clean image based on the match score;

receive a second noisy image that includes the at least one noise type of the plurality of noise types;

train a supervised machine learning model based on the first image pair, to produce a trained model; and

execute the trained model using the second image to generate a second clean image.

12. A non-transitory, processor-readable medium storing instructions to cause a processor to:

receive a noisy image without a paired clean image, the noisy image including at least one noise type of a plurality of noise types and an image structure type;

generate a blind denoised image from the noisy image using at least one denoising layer of a plurality of denoising layers of a machine learning model based on the at least one noise type and the image structure type of the noisy image;

generate a blind noised image from the blind denoised image using at least one noising layer of a plurality of noising layers of the machine learning model based on the at least one noise type and the image structure type of the noisy image; and

calculate a match score based on a comparison between the noisy image and the blind noised image to determine a clean image of the noisy image, the clean image having the same image structure type as the noisy image; and

generate an image pair based on the match score, the image pair including the clean image and the noisy image.

13. The non-transitory, processor-readable medium of claim 12 , wherein the non-transitory, processor-readable medium includes instructions to further cause the processor to store the image pair in a database categorized by the image structure type.

14. The non-transitory, processor-readable medium of claim 12 , wherein each noise type from the plurality of noise types is associated with a denoising layer of the plurality of denoising layers and a noising layer of the plurality of noising layers.

15. The non-transitory, processor-readable medium of claim 12 , wherein the machine-learning model includes an unsupervised machine learning model.

16. The non-transitory, processor-readable medium of claim 12 , wherein the non-transitory, processor-readable medium includes a Mixture of Experts, the Mixture of Experts including:

a plurality of denoising gates, each denoising gate of the plurality of denoising gates configured to activate at least one denoising channel of a plurality of denoising channels of each denoising layer;

a plurality of noising gates, each noising gate of the plurality of noising gates configured to activate at least one noising channel of a plurality of noising channels of each noising layer; and

an embedder, the memory storing instructions to cause the processor to:

compute, via the embedder, a denoising mixture weight of a plurality of denoising mixture weights for each denoising gate of the plurality of denoising gates based on the noisy image; and

compute, via the shallow embedder, a noising mixture weight of a plurality of noising mixture weights for each noising gate of the plurality of noising gates based on the blind denoised image.

17. The non-transitory, processor-readable medium of claim 16 , wherein generating the image pair based on the match score includes the non-transitory, processor-readable medium including instructions to further cause processor to:

generate the image pair based on the match score, the image pair including the clean image and the noisy image having the at least one noise type of the plurality of noise types and the same image structure type as the noisy image; and

train the embedder based on the image pair including the at least one noise type of the plurality of noise types and the image structure type, to produce a plurality of trained denoising gates and a plurality of trained noising gates.

18. The non-transitory, processor-readable medium of claim 12 , wherein:

the noisy image including a first noisy image;

the clean image is a first clean image;

the match score is a first match score;

the non-transitory, processor-readable medium includes instructions to further cause processor to:

generate a first image pair based on the first match score, the first image pair including the first noisy image and the clean image based on the match score;

receive a second noisy image, the second noisy image including the at least one noise type of the plurality of noise types;

train a supervised machine learning model based on the first image pair, to produce a trained model; and

execute the trained model using the second image to generate a second clean image.

19. A method, comprising:

receiving a noisy image without a paired clean image, the noisy image including at least one noise type of a plurality of noise types and an image structure type;

generating a blind denoised image from the noisy image using at least one denoising layer of a plurality of denoising layers of a machine learning model based on the at least one noise type and the image structure type of the noisy image;

generating a blind noised image from the blind denoised image using at least one noising layer of a plurality of noising layers of the machine learning model based on the at least one noise type and the image structure type of the noisy image; and

calculating a match score based on a comparison between the noisy image and the blind noised image to determine a clean image of the noisy image; and

generating an image pair based on the match score being above a match threshold, the image pair including the clean image and the noisy image.

20. The method of claim 19 , wherein:

the noisy image including a first noisy image;

the image structure type is a first image structure type;

the at least one noise type is at least a first noise type;

the clean image is a first clean image;

the match score is a first match score;

the image pair is a first image pair the method further including:

receiving a second noisy image, the second noisy image including at least a second noise type of the plurality of noise types and a second image structure type;

training a supervised machine learning model based on the first image pair, to produce a trained model; and

executing the trained model using the second image to generate a second clean image.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: GANGEH, MEHRDAD JABBARZADEH; MOTAHARI-NEZAD, HAMID REZA
To: ERNST & YOUNG U.S. LLP
Reel/Frame 072435/0197 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: EY GDS (CS) POLAND SP. Z O.O.
To: EYGS LLP
Reel/Frame 072435/0245 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: PLATA, MARCIN
To: EY GDS (CS) POLAND SP. Z O.O.
Reel/Frame 072435/0253 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: ERNST & YOUNG U.S. LLP
To: EYGS LLP
Reel/Frame 072435/0345 →
References Cited (21)
US 11508037B2 · Yang · 2022 [cited by examiner]
US 11769239B1 · Zhang · 2023 [cited by examiner]
US 11790492B1 · Ahmad · 2023 [cited by examiner]
US 11842460B1 · Chen · 2023 [cited by examiner]
US 20200074234A1 · Tong · 2020 [cited by examiner]
US 20210104021A1 · Sohn · 2021 [cited by examiner]
US 20210150674A1 · Cai · 2021 [cited by examiner]
US 20220375039A1 · Soh · 2022 [cited by examiner]
US 20230058096A1 · Ferrés · 2023 [cited by examiner]
US 20230103638A1 · Saharia · 2023 [cited by examiner]
US 20230153957A1 · Maleky · 2023 [cited by examiner]
CN 105069794B · 2015 [cited by examiner]
CN 110111266A · 2019 [cited by examiner]
CN 110223254A · 2019 [cited by examiner]
Batson, J. et al., “Noise2Self: Blind Denoising by Self-Supervision,” International Conference on Machine Learning, May 24, 2019, pp. 524-533. [cited by applicant]
Lehtinen, Jaako et al., “Noise2Noise: Learning Image Restoration without Clean Data,” arXiv preprint arXiv:1803.04189, Mar. 12, 2018, 12 pages. [cited by applicant]
Souibgui, M. et al., “DE-GAN: A Conditional Generative Adversarial Network for Document Enhancement,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Sep. 7, 2020, col. 44. No. 3, pp. 1180-1191. [cited by applicant]
Zhang, K. et al., “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising,” IEEE Transactions on Image Processing, Feb. 1, 2017, vol. 26, No. 7, pp. 3142-3155. [cited by applicant]
Gangeh, M. et al., “End-to-End Unsupervised Document Image Blind Denoising,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Virtual Conference (Oct. 10-17, 2021); Ernst & Young et al.; pp. 7888-7897. [cited by applicant]
Krull, A. et al., “Noise2Void—Learning Denoising From Single Noisy Images,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Conference (Jun. 16-20, 2019), pp. 2129-2137. [cited by applicant]
Mao, X. et al., “Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks with Symmetric Skip Connections,” NIPS'16: 30th International Conference on Neural Information Processing Systems, Dec. 5, 2016, … [cited by applicant]
Cited By (3)
US 12,555,204 US 12,657,672 US 12,700,069