IP Library › Granted Patent US 12,555,204
Granted Patent B1
US 12,555,204 · App. 19/204,321 · Granted Feb 17, 2026

Methods and apparatus for end-to-end unsupervised multi-document blind image denoising

Inventors: Mehrdad Jabbarzadeh Gangeh (Mountain View, CA); Marcin Plata (Wroclaw, PL); Hamid Reza Motahari-Nezad (Los Altos, CA)
Assignee: EYGS LLP
G06T5/70G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,555,204
App. No.
19/204,321
Filed
May 9, 2025
Granted
Feb 17, 2026
Kind
B1
Art Unit
2667
USPC
382/100
Abstract

Methods and apparatus for end-to-end unsupervised multi-document blind image denoising is presented. The multi-document blind image denoiser removes various noise types from noisy documents without paired target cleaned documents and preserves the contents for optical character recognition. The end-to-end unsupervised multi-document blind image denoiser integrates a Mixture of Experts with a cycle-consistent GAN as the base network that effectively removes multiple types of noise, including salt & pepper noise, blurred and/or faded text, as well as watermarks from documents at various levels of intensity.

Claims (72)

1 . An apparatus comprising:

a processor; and

a memory operatively coupled to the processor, the memory storing instructions to cause the processor to:

receive a noisy image without a paired clean image, the noisy image being associated with a document type;

input the noisy image to a first machine learning model layer to generate a blind denoised image;

input the blind denoised image to a second machine learning model layer to generate a first classification of the document type;

input the blind denoised image to a third machine learning model layer to generate a blind noised image;

input the blind noised image to a fourth machine learning model layer to generate a second classification of the document type; and

generate a clean image based on a comparison between the first classification of the document type and the second classification of the document type.

2 . The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to:

classify the noisy image based on an image structure type; and

store the noisy image in a database categorized by the image structure type.

3 . The apparatus of claim 1 , wherein the document type includes one of a structured document type, a semi-structured document type, or an unsupervised document type.

4 . The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to:

input the blind denoised image to a fifth machine learning model layer to determine a first noise intensity value; and

input the blind noised image to a sixth machine learning model layer to determine a second noise intensity value, the clean image being generated based further on a comparison between the first noise intensity value and the second noise intensity value.

5 . The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to:

input the blind denoised image to a fifth machine learning model layer to generate a first classification of a noise type; and

input the blind noised image to a sixth machine learning model layer to generate a second classification of the noise type, the clean image being generated based further on a comparison between the first classification of the noise type and the second classification of the noise type.

6 . The apparatus of claim 1 , wherein the memory stores instructions to further cause the processor to:

input the blind denoised image to a fifth machine learning model to classify the blind denoised image as having at least one of salt & pepper (S&P) noise, blurred text, faded text, a shading, or a watermark; and

input the blind noised image to a sixth machine learning model layer to classify the blind noised image as having the at least one of the salt & pepper (S&P) noise, the blurred text, the faded text, the shading, or the watermark, the clean image being generated based further on the classifying the blind denoised image and the classifying the blind noised image.

7 . The apparatus of claim 1 , further including:

a plurality of denoising gates, each denoising gate of the plurality of denoising gates configured to activate at least one denoising channel of a plurality of denoising channels of the first machine learning model layer;

a plurality of noising gates, each noising gate of the plurality of noising gates configured to activate at least one noising channel of a plurality of noising channels of the third machine learning model layer; and

an embedder, the memory storing instructions to cause the processor to:

compute, via the embedder, a denoising mixture weight of a plurality of denoising mixture weights for each denoising gate of the plurality of denoising gates based on the noisy image; and

compute, via the embedder, a noising mixture weight of a plurality of noising mixture weights for each noising gate of the plurality of noising gates based on the blind denoised image.

8 . The apparatus of claim 1 , wherein:

the clean image is a clean instance of the noisy image.

9 . The apparatus of claim 1 , wherein:

an end-to-end machine learning model includes the first machine learning model layer, the second machine learning model layer, the third machine learning model layer, and the fourth machine learning model layer.

10 . A method, comprising:

receiving, at a processor, a noisy image without a paired clean image;

inputting, via the processor, the noisy image to a first machine learning model layer to generate a blind denoised image;

inputting, via the processor, the blind denoised image to a second machine learning model layer to generate a first classification of at least one of an image structure type, a noise type, or a denoising intensity;

inputting, via the processor, the blind denoised image to a third machine learning model layer to generate a blind noised image;

inputting, via the processor, the blind noised image to a fourth machine learning model layer to generate a second classification of the at least one of the image structure type, the noise type, or the denoising intensity; and

generating, via the processor, a clean image based on a comparison between the first classification and the second classification.

11 . The method of claim 10 , wherein the image structure type includes one of a structured document type, a semi-structured document type, or an unsupervised document type.

12 . The method of claim 10 , further comprising:

inputting, via the processor, the blind denoised image to a fifth machine learning model layer to determine a first noise intensity value; and

inputting, via the processor, the blind noised image to a sixth machine learning model layer to determine a second noise intensity value, the clean image being generated based further on a comparison between the first noise intensity value and the second noise intensity value.

13 . The method of claim 10 , further comprising:

inputting, via the processor, the blind denoised image to a fifth machine learning model layer to generate a first classification of a noise type; and

inputting, via the processor, the blind noised image to a sixth machine learning model layer to generate a second classification of the noise type, the clean image being generated based further on a comparison between the first classification of the noise type and the second classification of the noise type.

14 . The method of claim 10 , further comprising:

inputting, via the processor, the blind denoised image to a fifth machine learning model to classify the blind denoised image as having at least one of salt & pepper (S&P) noise, blurred text, faded text, a shading, or a watermark; and

inputting, via the processor, the blind noised image to a sixth machine learning model layer to classify the blind noised image as having the at least one of the salt & pepper (S&P) noise, the blurred text, the faded text, the shading, or the watermark, the clean image being generated based further on the classifying the blind denoised image and the classifying the blind noised image.

15 . The method of claim 10 , wherein:

the inputting the noisy image to the first machine learning model layer includes:

computing, via the processor, a denoising mixture weight for a denoising gate of the first machine learning model layer based on the noisy image, and

activating, via the processor, a denoising channel using the denoising gate that has the denoising mixture weight, to generate the blind denoised image; and

the inputting the blind denoised image to the third machine learning model layer includes:

computing, via the processor, a noising mixture weight for a noising gate of the third machine learning model layer based on the blind denoised image, and activating, via the processor, a noising channel using the noising gate that has the noising mixture weight, to generate the blind noised image.

16 . A non-transitory, processor-readable medium storing instructions that, when executed by a processor, cause the processor to:

receive a noisy image without a paired clean image;

input the noisy image to a first machine learning model layer to generate a blind denoised image;

input the blind denoised image to a second machine learning model layer to generate a first classification of at least one of an image structure type, a noise type, or a denoising intensity;

input the blind denoised image to a third machine learning model layer to generate a blind noised image;

input the blind noised image to a fourth machine learning model layer to generate a second classification of the at least one of the image structure type, the noise type, or the denoising intensity; and

generate a clean image based on a comparison between the first classification and the second classification.

17 . The non-transitory, processor-readable medium of claim 16 , wherein the image structure type includes one of a structured document type, a semi-structured document type, or an unsupervised document type.

18 . The non-transitory, processor-readable medium of claim 16 , further storing instructions to cause the processor to:

input the blind denoised image to a fifth machine learning model layer to determine a first noise intensity value; and

input the blind noised image to a sixth machine learning model layer to determine a second noise intensity value, the clean image being generated based further on a comparison between the first noise intensity value and the second noise intensity value.

19 . The non-transitory, processor-readable medium of claim 16 , further storing instructions to cause the processor to:

input the blind denoised image to a fifth machine learning model layer to generate a first classification of a noise type; and

input the blind noised image to a sixth machine learning model layer to generate a second classification of the noise type, the clean image being generated based further on a comparison between the first classification of the noise type and the second classification of the noise type.

20 . The non-transitory, processor-readable medium of claim 16 , further storing instructions to cause the processor to:

input the blind denoised image to a fifth machine learning model to classify the blind denoised image as having at least one of salt & pepper (S&P) noise, blurred text, faded text, a shading, or a watermark; and

input the blind noised image to a sixth machine learning model layer to classify the blind noised image as having the at least one of the salt & pepper (S&P) noise, the blurred text, the faded text, the shading, or the watermark, the clean image being generated in response to classifying the blind denoised image and classifying the blind noised image.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: GANGEH, MEHRDAD JABBARZADEH; MOTAHARI-NEZAD, HAMID REZA
To: ERNST & YOUNG U.S. LLP
Reel/Frame 072435/0197 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: EY GDS (CS) POLAND SP. Z O.O.
To: EYGS LLP
Reel/Frame 072435/0245 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: PLATA, MARCIN
To: EY GDS (CS) POLAND SP. Z O.O.
Reel/Frame 072435/0253 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 1, 2025
From: ERNST & YOUNG U.S. LLP
To: EYGS LLP
Reel/Frame 072435/0345 →
Continuity (1)
Continuation 17962853 · Oct 10, 2022
References Cited (27)
US 11508037B2 · Yang et al. · 2022 [cited by applicant]
US 11769239B1 · Zhang et al. · 2023 [cited by applicant]
US 11790492B1 · Ahmad et al. · 2023 [cited by applicant]
US 11842460B1 · Chen et al. · 2023 [cited by applicant]
US 12333689B1 · Gangeh · 2025 [cited by examiner]
US 20200074234A1 · Tong et al. · 2020 [cited by applicant]
US 20200175304A1 · Vig · 2020 [cited by examiner]
US 20210104021A1 · Sohn et al. · 2021 [cited by applicant]
US 20210150674A1 · Cai et al. · 2021 [cited by applicant]
US 20220375039A1 · Soh et al. · 2022 [cited by applicant]
US 20220398399A1 · Muffat · 2022 [cited by examiner]
US 20230058096A1 · Ferrés et al. · 2023 [cited by applicant]
US 20230103638A1 · Saharia et al. · 2023 [cited by applicant]
US 20230153957A1 · Maleky et al. · 2023 [cited by applicant]
US 20240037970A1 · Alabi · 2024 [cited by examiner]
CN 105069794B · 2017 [cited by applicant]
CN 110111266A · 2019 [cited by applicant]
CN 110223254A · 2019 [cited by applicant]
Batson, J. et al., “Noise2Self: Blind Denoising by Self-Supervision,” International Conference on Machine Learning, May 24, 2019, pp. 524-533. [cited by applicant]
Gangeh, M. et al., “End-to-End Unsupervised Document Image Blind Denoising,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), Virtual Conference (Oct. 10-17, 2021); Ernst & Young et al.; pp. 7888-7897. [cited by applicant]
Gangeh, M. et al., “End-to-End Unsupervised Document Image Blind Denoising,” arXiv:2105.09437v2 [cs.CV], Oct. 7, 2021. Retrieved from https://arxiv.org/abs/2105.09437; 19 pages. [cited by applicant]
Krull, A. et al., “Noise2Void—Learning Denoising From Single Noisy Images,” 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Conference (Jun. 16-20, 2019), pp. 2129-2137. [cited by applicant]
Lehtinen, Jaako et al., “Noise2Noise: Learning Image Restoration without Clean Data,” arXiv preprint arXiv:1803.04189, Mar. 12, 2018; 12 pages. [cited by applicant]
Mao, X. et al., “Image Restoration Using Very Deep Convolutional Encoder-Decoder Networks with Symmetric Skip Connections,” NIPS'16: 30th International Conference on Neural Information Processing Systems, Dec. 5, 2016, … [cited by applicant]
Souibgui, M. et al., “DE-GAN: A Conditional Generative Adversarial Network for Document Enhancement,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Sep. 7, 2020, col. 44. No. 3, pp. 1180-1191. [cited by applicant]
Wang, X. et al., “Deep Mixture of Experts via Shallow Embedding,” Proceedings of the 35th Uncertainty in Artificial Intelligence Conference, Proceedings of Machine Learning Research (PMLR), Aug. 2020, 115:552-562. [cited by applicant]
Zhang, K. et al., “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising,” IEEE Transactions on Image Processing, Feb. 1, 2017, vol. 26, No. 7, pp. 3142-3155. [cited by applicant]
Cited By (1)
US 12,700,069