IP Library › Granted Patent US 12,597,096
Granted Patent B2
US 12,597,096 · App. 18/188,050 · Granted Apr 7, 2026

Real-time facial restoration and relighting in videos using facial enhancement neural networks

Inventors: Samira Pouyanfar (San Jose, CA); Sunando Sengupta (Berkshire, GB); Eric Chris Wolfgang Sommerlade (Oxford, GB); Anjali S. Parikh (Redmond, WA); Ebey Paulose Abraham (Oxford, GB); Brian Timothy Hawkins (Sammamish, WA); Mahmoud Mohammadi (Redmond, WA)
Assignee: Microsoft Technology Licensing, LLC
G06T5/50G06T5/70G06T5/73G06T7/194G06V40/167G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,096
App. No.
18/188,050
Granted
Apr 7, 2026
Kind
B2
Abstract

The present disclosure relates to an image restoration system that efficiently and accurately produces high-quality images captured under low-light and/or low-quality environmental conditions. To illustrate, when a user is in a low-lit environment and participating in a video stream, the image restoration system enhances the quality of the image by dynamically re-lighting the user's face. Moreover, it significantly enhances the image quality to the extent that other users viewing the video stream are unaware of the poor environmental conditions of the user. In addition, the image restoration system creates and utilizes an image restoration machine-learning model to improve the quality of low-quality images by re-lighting and restoring them in real time. Various implementations combine an autoencoder model with a distortion classifier model to create the image restoration machine-learning model.

Claims (68)

1 . A computer-implemented method comprising:

detecting a face within a digital image that includes the face and an image background;

generating a light-enhanced face image of the face utilizing an image restoration machine-learning model that includes an autoencoder model and a distortion classifier by:

producing a distortion classification by the distortion classifier; and

concatenating the distortion classification with encoded feature vectors generated by an encoder of the autoencoder model;

generating an enhanced digital image by combining the light-enhanced face image with the image background; and

providing the enhanced digital image for display on a computing device.

2 . The computer-implemented method of claim 1 , wherein generating the enhanced digital image includes improving lighting on the face, restoring low-quality facial features to higher-quality facial features, reducing blur, and reducing noise.

3 . The computer-implemented method of claim 1 , wherein generating the light-enhanced face image further comprises:

generating, using the autoencoder model, the encoded feature vectors based on the face within the digital image;

generating, using the distortion classifier, the distortion classification for the face within the digital image; and

generating, using the autoencoder model, the light-enhanced face image from the encoded feature vectors and the distortion classification.

4 . The computer-implemented method of claim 3 , wherein generating the distortion classification includes indicating amounts of noise distortion, blur distortion, exposure distortion, and light distortion in the face detected in the digital image.

5 . The computer-implemented method of claim 1 , wherein the distortion classifier is a lightweight classifier machine learning model that learns distortion types applied to images.

6 . The computer-implemented method of claim 1 , wherein:

detecting the face within the digital image includes:

identifying the face in the digital image;

cropping the face in the digital image to generate a cropped image; and

tracking the face across a set of sequential digital images; and

generating the light-enhanced face image includes:

generating the light-enhanced face image to have image dimensions that match the image dimensions of the cropped image; and

generating a face image mask that separates non-face pixels from face pixels in the light-enhanced face image.

7 . The computer-implemented method of claim 6 , wherein generating the enhanced digital image includes utilizing the face image mask to blend the light-enhanced face image with the image background of the digital image.

8 . The computer-implemented method of claim 1 , wherein:

detecting the face within the digital image includes detecting a colored light shining on the face; and

generating the light-enhanced face image includes removing an effect of the colored light included in the digital image.

9 . The computer-implemented method of claim 1 , wherein generating the light-enhanced face image includes rearranging data in the autoencoder model to maintain lossless spatial dimensionality.

10 . The computer-implemented method of claim 1 , further comprising:

segmenting a body portion from the digital image to generate a body image, the body portion being connected to the face;

generating a body-enhanced image utilizing the image restoration machine-learning model; and

wherein generating the enhanced digital image further comprises combining the body-enhanced image with the light-enhanced face image and the image background.

11 . The computer-implemented method of claim 10 , wherein generating the body-enhanced image further comprises utilizing a scale factor generated for the light-enhanced face image.

12 . The computer-implemented method of claim 1 , further comprising:

tracking the face across a digital video having a set of digital images that includes the digital image;

generating a set of face-enhanced digital images from the set of digital images utilizing the image restoration machine-learning model, wherein the set of face-enhanced digital images includes the enhanced digital image; and

providing the set of face-enhanced digital images for display on the computing device as a face-enhanced digital video.

13 . A system comprising:

an image restoration machine-learning model that includes a distortion classifier and an autoencoder model having an encoder and a generator, wherein the distortion classifier is separate from the autoencoder model;

a processor; and

a computer memory comprising instructions that, when executed by the processor, cause the system to carry out operations comprising:

detecting a face within a digital image that includes the face and an image background;

generating a light-enhanced face image of the face utilizing the image restoration machine-learning model based on:

producing a distortion classification by the distortion classifier;

combining the distortion classification from the distortion classifier with encoded feature vectors output from the encoder of the autoencoder model to generate a combined output; and

using the combined output as an input to the generator of the autoencoder model; and

generating an enhanced digital image by combining the light-enhanced face image with the image background.

14 . The system of claim 13 , wherein the instructions further comprise generating the image restoration machine-learning model by training the distortion classifier and the autoencoder model in parallel to improve accuracy of the generator, wherein the image restoration machine-learning model uses both real digital images and synthetic digital images.

15 . The system of claim 13 , wherein the instructions further comprise generating a synthetic digital image by:

generating a synthetic face;

capturing multiple digital images that each shine light from a different light source on the synthetic face; and

combining the multiple digital images into a combined digital image to generate the synthetic digital image.

16 . The system of claim 13 , wherein the instructions further comprise generating the image restoration machine-learning model by utilizing loss model functions that include pixel-wise loss, feature loss, texture information loss, adversarial loss, classification loss, or segmentation loss.

17 . The system of claim 13 , wherein the distortion classifier predicts degradation types in the digital image and generates the distortion classification to indicate the degradation types in the digital image.

18 . A computer-implemented method comprising:

detecting an object within a digital image that includes an image background;

generating a light-enhanced object image utilizing an object-re-lighting neural network that includes an autoencoder model and a distortion classifier by:

producing a distortion classification by the distortion classifier; and

concatenating the distortion classification with encoded feature vectors generated by an encoder of the autoencoder model;

generating an enhanced digital image having a re-lit object by combining the light-enhanced object image with the image background, wherein the enhanced digital image improves lighting on the object as compared to the digital image and maintains a same or substantially similar image background as the digital image; and

providing the enhanced digital image for display on a computing device.

19 . The computer-implemented method of claim 18 , wherein generating the light-enhanced object image includes:

generating the distortion classification utilizing the distortion classifier of the object-re-lighting neural network;

generating feature vectors utilizing an encoder network of the object-re-lighting neural network; and

generating the light-enhanced object image utilizing a generator of the object-re-lighting neural network from a combination of the distortion classification and the feature vectors.

20 . The computer-implemented method of claim 18 , further comprising:

tracking the object across a set of digital images that includes the digital image;

generating a set of object-enhanced digital images from the set of digital images utilizing the object-re-lighting neural network, wherein the set of object-enhanced digital images includes the enhanced digital image; and

providing the set of object-enhanced digital images for display on the computing device as an object-enhanced digital video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2023
From: POUYANFAR, SAMIRA; SENGUPTA, SUNANDO; SOMMERLADE, ERIC CHRIS WOLFGANG; PARIKH, ANJALI S.; ABRAHAM, EBEY PAULOSE; HAWKINS, BRIAN TIMOTHY; MOHAMMADI, MAHMOUD
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 063063/0180 →
Continuity (1)
Related Publication 20240331094A1 · Oct 3, 2024
References Cited (78)
US 8224039B2 · Ionita · 2012 [cited by examiner]
US 11488288B1 · Heo · 2022 [cited by examiner]
US 11869275B1 · Joseph · 2024 [cited by examiner]
US 11983853B1 · Zhu · 2024 [cited by examiner]
US 20130050395A1 · Paoletti · 2013 [cited by examiner]
US 20210001810A1 · Rivard · 2021 [cited by examiner]
US 20210209344A1 · Guo · 2021 [cited by examiner]
US 20220157012A1 · Ghosh · 2022 [cited by examiner]
US 20230316477A1 · Wu · 2023 [cited by examiner]
US 20230360182A1 · Fanello · 2023 [cited by examiner]
CN 114495201A · 2022 [cited by examiner]
CN 114998124A · 2022 [cited by examiner]
CN 115359525A · 2022 [cited by examiner]
CN 115830325A · 2023 [cited by examiner]
WO WO2021093712A1 · 2021 [cited by examiner]
Zhang, et al., “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising”, In Proceedings of IEEE Transactions on Image Processing vol. 26, Issue 7, Feb. 1, 2017, pp. 3142-3155. [cited by applicant]
Zhang, et al., “Multi-Modality Deep Restoration of Extremely Compressed Face Videos”, In Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 45, Issue 2, Mar. 8, 2022, pp. 2024-2037. [cited by applicant]
Zhang, et al. “Plug-and-Play Image Restoration With Deep Denoiser Prior”, In Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 44, Issue 10, Jun. 14, 2021, pp. 6360-6376. [cited by applicant]
Zhang, et al., “Residual Dense Network for Image Super-Resolution”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern, Jun. 18, 2018, pp. 2472-2481. [cited by applicant]
Zhang, et al., “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 586-595. [cited by applicant]
Zhou, et al., “Towards Robust Blind Face Restoration with Codebook Lookup Transformer”, In Proceedings Advances in Neural Information Processing Systems, Oct. 31, 2022, 13 Pages. [cited by applicant]
Zhu, et al., “Blind Face Restoration via Integrating Face Shape and Generative Priors”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 7652-7661. [cited by applicant]
Zhuang, et al., “Surrogate Gap Minimization Improves Sharpness-Aware Training”, In Proceedings of International Conference on Learning Representations, Jan. 29, 2022, 24 Pages. [cited by applicant]
International Search Report and Written Opinion received for PCT Application No. PCT/US2024/019312, Jun. 24, 2024, 16 pages. [cited by applicant]
Teng, et al., “Blind face restoration via multi-prior collaboration and adaptive feature fusion,” Frontiers in Neurorobotics 16, Article 797231, Feb. 2022, pp. 1-18. [cited by applicant]
Toledo, et al., “Face Reconstruction with Variational Autoencoder and Face Masks,” arXiv preprint arXiv:2112.02139, 2021, Dec. 3, 2021, 12 pages. [cited by applicant]
Wang, et al., “Panini-Net: GAN Prior Based Degradation-Aware Feature Interpolation for Face Restoration,” arXiv preprint: 2203.08444, 2022, Mar. 16, 2022, 9 pages. [cited by applicant]
“Snapdragon Neural Processing Engine SDK”, Retrieved From : https://developer.qualcomm.com/sites/default/files/docs/snpe/, Retrieved on: Sep. 14, 2022, 1 Page. [cited by applicant]
Atoum, et al., “Color-Wise Attention Network for Low-Light Image Enhancement”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 14, 2020, pp. 2130-2139. [cited by applicant]
Chan, et al., “GLEAN: Generative Latent Bank for Large-Factor Image Super-Resolution”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 14240-14249. [cited by applicant]
Chen, et al., “FSRNet: End-to-End Learning Face Super-Resolution with Facial Priors”, Retrieved From : IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, pp. 2492-2501. [cited by applicant]
Chen, et al., “Progressive Semantic-Aware Style Transformation for Blind Face Restoration”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 11891-11900. [cited by applicant]
Chen, et al., “When Vision Transformers Outperform ResNets without Pre-training or Strong Data Augmentations”, In Repository of arXiv:2106.01548v2, Oct. 11, 2021, 19 Pages. [cited by applicant]
Dogan, et al., “Exemplar Guided Face Image Super-Resolution Without Facial Landmarks”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 16, 2019, pp. 1814-1823. [cited by applicant]
Dong, et al., “Deep Wiener Deconvolution: Wiener Meets Deep Learning for Image Deblurring”, In Proceedings of Advances in Neural Information Processing Systems, Dec. 6, 2020, 12 Pages. [cited by applicant]
Dosovitskiy, et al., “An Image Is Worth 16×16 Words: Transformers for Image Recognition at Scale”, In International Conference on Learning Representations, Jun. 3, 2021, 21 Pages. [cited by applicant]
Fan, et al., “FaceFormer: Speech-Driven 3D Facial Animation with Transformers”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 18749-18758. [cited by applicant]
Fu, et al., “An Efficient Hybrid Model for Low-light Image Enhancement in Mobile Devices”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, Jun. 19, 2022, pp. 3056-3065. [cited by applicant]
Gu, et al., “VQFR: Blind Face Restoration with Vector-Quantized Dictionary and Parallel Decoder”, In Proceedings of Computer Vision—ECCV 2022: 17th European Conference, Oct. 23, 2022, 17 Pages. [cited by applicant]
Guo, et al., “Zero-Reference Deep Curve Estimation for Low-Light Image Enhancement”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 1777-1786. [cited by applicant]
He, et al., “GCFSR: a Generative and Controllable Face Super Resolution Method Without Facial and GAN Priors”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 1879-18… [cited by applicant]
He, et al., “Mask r-cnn”, In Proceedings of the IEEE International Conference on Computer Vision, Oct. 22, 2017, pp. 2961-2969. [cited by applicant]
Helou, et al., “Stochastic Frequency Masking to Improve Super-Resolution and Denoising Networks”, In Proceedings of Computer Vision—ECCV: 16th European Conference, Aug. 23, 2020, 18 Pages. [cited by applicant]
Huang, et al., “Neural Compression-Based Feature Learning for Video Restoration”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 5862-5871. [cited by applicant]
Johnson, et al., “Perceptual Losses for Real-Time Style Transfer and Super-Resolution”, In Proceedings of European Conference on Computer Vision, Oct. 11, 2016, pp. 694-711. [cited by applicant]
Karras, et al., “A style-based generator architecture for generative adversarial networks”, In Proceedings of IEEE Conference on Computer Vision and Pattern Recognition, Jun. 15, 2019, pp. 4396-4405. [cited by applicant]
Karras, et al., “Analyzing and Improving the Image Quality of StyleGAN”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 8107-8116. [cited by applicant]
Kupyn, et al., “DeblurGAN: Blind Motion Deblurring Using Conditional Adversarial Networks”, In Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 18, 2018, pp. 8183-8192. [cited by applicant]
Lee, et al., “KNN Local Attention for Image Restoration”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 2129-2139. [cited by applicant]
Li, et al., “Blind Face Restoration via Deep Multi-scale Component Dictionaries”, In Proceedings of Computer Vision 16th European Conference, Aug. 23, 2020, pp. 399-415. [cited by applicant]
Liang, et al., “Swinir: Image restoration using swin transformer”, In Proceedings of 2021 IEEE/CVF International Conference on Computer Vision Workshops, Oct. 11, 2021, pp. 1833-1844. [cited by applicant]
Liu, et al., “Non-Local Recurrent Network for Image Restoration”, In Proceedings of Advances in Neural Information Processing Systems, Dec. 3, 2018, 10 Pages. [cited by applicant]
Matsuoka, et al., “Transformed-Domain Robust Multiple-Exposure Blending With Huber Loss”, In Proceedings of IEEE Access, Nov. 6, 2019, pp. 162282-162296. [cited by applicant]
Mei, et al., “Image Super-Resolution with Non-Local Sparse Attention”, Retrieved From : IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2020, pp. 3516-3525. [cited by applicant]
Menon, et al., “Pulse: Self-Supervised Photo Upsampling via Latent Space Exploration of Generative Models”, In Proceedings of the ieee/cvf Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 2437-2… [cited by applicant]
Milletari, et al., “V-Net: Fully Convolutional Neural Networks for Volumetric Medical Image Segmentation”, In Proceedings of Fourth International Conference on 3D Vision, Oct. 1, 2016, 11 Pages. [cited by applicant]
Nan, et al., “Variational-EM-Based Deep Learning for Noise-Blind Image Deblurring”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 3623-3632. [cited by applicant]
Pang, et al., “Recorrupted-to-Recorrupted: Unsupervised Deep Learning for Image Denoising”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 2043-2052. [cited by applicant]
Parkhi, et al., “Deep Face Recognition”, In Proceedings of the British Machine Vision Conference, Sep. 2015, 12 Pages. [cited by applicant]
Shi, et al., “Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 27, 2016… [cited by applicant]
Sidorov, et al., “Artificial Color Constancy Via Googlenet with Angular Loss Function”, In Journal of Applied Artificial Intelligence, vol. 34, issue 9, Jul. 28, 2020, 14 Pages. [cited by applicant]
Steiner, et al., “How to train your ViT? Data, Augmentation, and Regularization in Vision Transformers”, In Proceedings of arXiv:2106.10270v1, Jun. 18, 2021, 14 Pages. [cited by applicant]
Sudre, et al., “Generalised Dice Overlap as a Deep Learning Loss Function for Highly Unbalanced Segmentations”, In Proceedings of Deep Learning in Medical Image Analysis and Multimodal Learning for Clinical Decision Sup… [cited by applicant]
Tassano, et al., “DVDNET: A Fast Network for Deep Video Denoising”, In Proceedings of IEEE International Conference on Image Processing, Sep. 22, 2019, pp. 1805-1809. [cited by applicant]
Tolstikhin, et al., “MLP-Mixer: An all-MLP Architecture for Vision”, In Proceedings of arXiv:2105.01601v4, Jun. 11, 2021, 16 Pages. [cited by applicant]
Wang, et al., “Panini-net: Gan Prior Based Degradation-Aware Feature Interpolation for Face Restoration”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 36, Issue 3, Jun. 28, 2022, pp. 2576-2584. [cited by applicant]
Wang, et al., “Towards Real-World Blind Face Restoration with Generative Facial Prior”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 9164-9174. [cited by applicant]
Wang, “Uformer: A General U-Shaped Transformer for Image Restoration”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 17662-17672. [cited by applicant]
Yang, et al., “From Fidelity to Perceptual Quality: A Semi-Supervised Approach for Low-Light Image Enhancement”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2020, pp. 3060-… [cited by applicant]
Yang, et al., “GAN Prior Embedded Network for Blind Face Restoration in the Wild”, In Proceedings of IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 20, 2021, pp. 672-681. [cited by applicant]
Yang, et al., “HiFaceGAN: Face Renovation via Collaborative Suppression and Replenishment”, In Proceedings of the 28th ACM International Conference on Multimedia, Oct. 12, 2020, 11 Pages. [cited by applicant]
Yu, et al., “Cascading Convolutional Color Constancy”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, Issue 7, Apr. 3, 2020, pp. 12725-12732. [cited by applicant]
Yu, et al., “Face Super-Resolution Guided by Facial Component Heatmaps”, In Proceedings of the European Conference on Computer Vision, Sep. 8, 2018, 17 Pages. [cited by applicant]
Yu et al., “Path-Restore: Learning Network Path Selection for Image Restoration”, In Proceedings of IEEE Transactions on Pattern Analysis and Machine Intelligence vol. 44, Issue 10, Jul. 13, 2021, pp. 7078-7092. [cited by applicant]
Zameer, et al., “Restormer: Efficient Transformer for High-Resolution Image Restoration”, In Proceedings IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2022, pp. 5718-5729. [cited by applicant]
Zhai, et al., “LiT: Zero-Shot Transfer with Locked-image text Tuning”, In Repository of arXiv:2111.07991v3, Jun. 22, 2022, 26 Pages. [cited by applicant]
International Preliminary Report on Patentability (Chapter I) received for PCT Application No. PCT/US2024/019312, mailed on Oct. 2, 2025, xx pages. [cited by applicant]
Li, et al., “Learning Warped Guidance for Blind Face Restoration”, In Proceedings of the European Conference on Computer Vision, Part XIII, Sep. 2018, 18 Pages. [cited by applicant]