IP Library Granted Patent US 12,387,291
Granted Patent B2
US 12,387,291 · App. 17/718,136 · Granted Aug 12, 2025

Video super-resolution using deep neural networks

Inventors: Maruan Al-Shedivat (Pittsburgh, PA); Yihui He (Pittsburgh, PA); Megan Hardy (Oakland, CA); Andrew Rabinovich (San Francisco, CA)
Assignee: Upwork Inc.
G06T3/4053G06T3/4046
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,291
App. No.
17/718,136
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods and systems for obtaining an input video sequence comprising input video frames; determining i) an input resolution of input video frames and ii) a target output resolution of the plurality of input video frames, wherein the target output resolution is higher than the input resolution; and processing the input video sequence using a neural network to generate an output video sequence, comprising, for each of the plurality of input video frames: processing the input video frame to generate an output video frame having the target output resolution, comprising processing the input video frame using a subnetwork of the neural network corresponding to the input resolution of the plurality of input video frames, the neural network configured to process input video frames having one of a set of possible input resolutions and to generate output video frames having one of a set of possible output resolutions.

Claims (61)

1. A method comprising:

obtaining an input video sequence comprising a plurality of input video frames;

determining i) an input resolution of the plurality of input video frames and ii) a target output resolution of the plurality of input video frames, wherein the target output resolution is higher than the input resolution, wherein the input resolution is dynamically reduced based on a strength of a connection with a destination user device; and

processing the input video sequence using a neural network to generate an output video sequence, comprising, for each of the plurality of input video frames:

processing the input video frame using the neural network to generate an output video frame having the target output resolution, comprising processing the input video frame using a subnetwork of the neural network corresponding to the input resolution of the plurality of input video frames, wherein a resolution of the plurality of input video frames is incrementally upsampled to the target output resolution,

wherein the neural network has been configured through training to process input video frames having one of a predetermined set of possible input resolutions and to generate output video frames having one of a predetermined set of possible output resolutions.

2. The method of claim 1 , wherein the input video sequence is a live video conference video sequence and the output video frames are generated in real time.

3. The method of claim 2 , wherein generating the output video frames in real time comprises generating output frames at a rate of at least 10, 20, 30, or 50 Hz.

4. The method of claim 1 , wherein determining the target output resolution comprises receiving data identifying the target output resolution.

5. The method of claim 1 , wherein:

obtaining an input video sequence comprises obtaining the input video sequence by a cloud system and from a first user device; and

the method further comprises providing the output video sequence to a second user device that is different from the first user device.

6. The method of claim 1 , wherein the neural network has been trained by performing operations comprising:

obtaining a training input video frame;

processing the training input video frame using a trained machine learning model that is configured to detect faces in images;

processing, using an output of the trained machine learning model, the training input video frame to generate an updated training input video frame that isolates a detected face in the training input video frame;

processing the updated training input video frame using the neural network to generate a predicted output video frame; and

updating a plurality of parameters of the neural network according to an error of the predicted output video frame.

7. The method of claim 1 , wherein the neural network comprises one or more gated recurrent units (GRUs) that are configured to maintain information across different input video frames in the input video sequence.

8. The method of claim 1 , wherein the neural network has been trained by performing operations comprising:

processing a training input video frame using the neural network to generate a predicted output video frame;

processing the predicted output video frame using a discriminator neural network to generate a prediction of whether the predicted output video frame was generated by the neural network; and

updating a plurality of parameters of the neural network in order to increase an error of the discriminator neural network.

9. A method comprising:

obtaining an input video sequence comprising a plurality of input video frames having an input resolution, wherein the input resolution is dynamically reduced based on a strength of a connection with a destination user device; and

processing the input video sequence using a neural network to generate an output video sequence, comprising, for each of the plurality of input video frames:

processing the input video frame using the neural network to generate an output video frame having a target output resolution that is higher than the input resolution, wherein a resolution of the plurality of input video frames is incrementally upsampled to the target output resolution,

wherein the neural network has been trained by performing operations comprising:

obtaining a training input video frame;

processing the training input video frame using a trained machine learning model that is configured detect faces in images;

processing, using an output of the trained machine learning model, the training input video frame to generate an updated training input video frame that isolates a detected face in the training input video frame;

processing the updated training input video frame using the neural network to generate a predicted output video frame; and

updating a plurality of parameters of the neural network according to an error of the predicted output video frame.

10. The method of claim 9 , wherein the input video sequence is a live video conference video sequence and the output video frames are generated in real time.

11. The method of claim 10 , wherein generating the output video frames in real time comprises generating output frames at a rate of at least 10, 20, 30, or 50 Hz.

12. The method of claim 9 , wherein:

obtaining an input video sequence comprises obtaining the input video sequence by a cloud system and from a first user device; and

the method further comprises providing the output video sequence to a second user device that is different from the first user device.

13. The method of claim 9 , wherein the neural network comprises one or more gated recurrent units (GRUs) that are configured to maintain information across different input video frames in the input video sequence.

14. The method of claim 9 , wherein the neural network has further been trained by performing operations comprising:

processing a training input video frame using the neural network to generate a predicted output video frame;

processing the predicted output video frame using a discriminator neural network to generate a prediction of whether the predicted output video frame was generated by the neural network; and

updating a plurality of parameters of the neural network in order to increase an error of the discriminator neural network.

15. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one more computers to perform operations comprising:

obtaining an input video sequence comprising a plurality of input video frames;

determining i) an input resolution of the plurality of input video frames and ii) a target output resolution of the plurality of input video frames, wherein the target output resolution is higher than the input resolution, wherein the input resolution is dynamically reduced based on a strength of a connection with a destination user device; and

processing the input video sequence using a neural network to generate an output video sequence, comprising, for each of the plurality of input video frames:

processing the input video frame using the neural network to generate an output video frame having the target output resolution, comprising processing the input video frame using a subnetwork of the neural network corresponding to the input resolution of the plurality of input video frames, wherein a resolution of the plurality of input video frames is incrementally upsampled to the target output resolution,

wherein the neural network has been configured through training to process input video frames having one of a predetermined set of possible input resolutions and to generate output video frames having one of a predetermined set of possible output resolutions.

16. The system of claim 15 , wherein the input video sequence is a live video conference video sequence and the output video frames are generated in real time.

17. The system of claim 16 , wherein generating the output video frames in real time comprises generating output frames at a rate of at least 10, 20, 30, or 50 Hz.

18. The system of claim 15 , wherein determining the target output resolution comprises receiving data identifying the target output resolution.

19. The system of claim 15 , wherein:

obtaining an input video sequence comprises obtaining the input video sequence by a cloud system and from a first user device; and

the operations further comprise providing the output video sequence to a second user device that is different from the first user device.

20. The system of claim 15 , wherein the neural network has been trained by performing operations comprising:

obtaining a training input video frame;

processing the training input video frame using a trained machine learning model that is configured to detect faces in images;

processing, using an output of the trained machine learning model, the training input video frame to generate an updated training input video frame that isolates a detected face in the training input video frame;

processing the updated training input video frame using the neural network to generate a predicted output video frame; and

updating a plurality of parameters of the neural network according to an error of the predicted output video frame.

Assignments (6)
SECURITY INTEREST Recorded Jun 24, 2026
From: UPWORK INC.
To: BANK OF AMERICA, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 075067/0544 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: AL-SHEDIVAT, MARUAN
To: HEADROOM, INC.
Reel/Frame 071628/0437 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: HE, YIHUE
To: HEADROOM, INC.
Reel/Frame 071628/0484 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: HARDY, MEGAN
To: HEADROOM, INC.
Reel/Frame 071628/0604 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2025
From: RABINOVICH, ANDREW
To: HEADROOM, INC.
Reel/Frame 071628/0634 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2024
From: HEADROOM, INC.
To: UPWORK INC.
Reel/Frame 066163/0044 →
Continuity (2)
Provisional Application 63174307 · Apr 13, 2021
Related Publication 20220327663A1 · Oct 13, 2022
References Cited (49)
US 10904476B1 · Adams et al. · 2021 [cited by applicant]
US 11122240B2 · Peters · 2021 [cited by applicant]
US 11158121B1 · Tung · 2021 [cited by examiner]
US 20130101002A1 · Gettings · 2013 [cited by examiner]
US 20180253865A1 · Price · 2018 [cited by applicant]
US 20200334789A1 · Zhang · 2020 [cited by examiner]
US 20200342572A1 · Chen · 2020 [cited by examiner]
US 20200364872A1 · Shelns · 2020 [cited by applicant]
US 20210092462A1 · Cox · 2021 [cited by examiner]
US 20210150278A1 · Dudzik · 2021 [cited by applicant]
US 20210250547A1 · Jiang · 2021 [cited by examiner]
US 20210281867A1 · Golinski · 2021 [cited by applicant]
López-Tapia, Santiago. “Gated Recurrent Networks for Video Super Resolution.”, EUPISCO 2020. 700-704. (Year: 2020). [cited by examiner]
Wenzhe Shi et al. “Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network.”2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR). IEEE, 2016. 1874-18… [cited by examiner]
Caballero et al., “Real-time video Super-Resolution with Spatio-Temporal Networks and Motion Compensation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4778-4787 (abstract on… [cited by applicant]
Donahue et al., “Long-term Recurrent Convolutional Networks for Visual Recognition and Description,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2015, 14 pages. [cited by applicant]
Dong et al., “Image Super-Resolution Using Deep Convolutional Networks,” IEEE transactions on pattern analysis and machine intelligence, 2015, 38(2):295-307. [cited by applicant]
Drulea et al., “Total variation regularization of local-global optical flow,” 2011 14th International IEEE Conference on Intelligent Transportation Systems, 2011, 7 pages. [cited by applicant]
Haris et al., “Deep back-projection networks for super-resolution,” Proceedings of the IEEE conference on computer vision and pattern recognition, Mar. 7, 2018, 10 pages. [cited by applicant]
Huang et al., “Bidirectional recurrent convolutional networks for multi-frame super-resolution,” Advances in Neural Information Processing Systems, 2015, pp. 235-243. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2022/024285, mailed on Jul. 6, 2022, 21 pages. [cited by applicant]
Jo et al., “Deep Video Super-Resolution Network Using Dynamic Upsampling Filters Without Explicit Motion Compensation,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Johnson et al., “DenseCap: Fully convolutional localization networks for dense captioning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, 10 pages. [cited by applicant]
Kappeler et al., “Video super-resolution with convolutional neural networks,” IEEE Transactions on Computational Imaging, 2016, 2(2):109-122 (abstract only). [cited by applicant]
Kim et al., “Deeply-recursive convolutional network for image super-resolution,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, 9 pages. [cited by applicant]
Lai et al., “Deep laplacian pyramid networks for fast and accurate super-resolution,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, pp. 624-632 (abstract only). [cited by applicant]
Ledig et al., “Photo-realistic single image super-resolution using a generative adversarial network,” Proceedings of the IEEE conference on computer vision and pattern recognition, 2017, 19 pages. [cited by applicant]
Lee et al., “Deeply-supervised nets,” Artificial intelligence and statistics, Sep. 25, 2014, 10 pages. [cited by applicant]
Liao et al., “Video super-resolution via deep draft-ensemble learning,” Proceedings of the IEEE International Conference on Computer Vision, 2015, 9 pages. [cited by applicant]
Liu et al., “Robust video super-resolution with learned temporal dynamics,” Proceedings of the IEEE International Conference on Computer Vision, 2017, 9 pages. [cited by applicant]
Liu et al., “Video Super Resolution Based on Deep Learning: A Comprehensive Survey,” Dec. 20, 2020, arXiv:2007.12928v2, 30 pages. [cited by applicant]
Lugmayr et al., “SRFlow: Learning the Super-Resolution Space with Normalizing Flow,” Springer, Aug. 28, 2020, 18 pages. [cited by applicant]
Mao et al., “Deep captioning with multimodal recurrent neural networks (M-RNN),” 2014, arXiv:1412.6632, 17 pages. [cited by applicant]
Nvictia, “How to Reinvent Virtual Collaboration, Video Communications: NVIDIA Blog,” Mar. 17, 2021, retrieved on Jun. 22, 2022, retrieved from URL<https://blogs.nvidia.com/blog/2021/03/17/gtc-maxine-virtual-collaboratio… [cited by applicant]
Sajjadi et al., “Frame-Recurrent video super-resolution,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Mar. 25, 2018, 9 pages. [cited by applicant]
Shi et al., “Convolutional LSTM network: A machine learning approach for precipitation nowcasting,” Advances in Neural Information Processing Systems 28, Sep. 19, 2015, 12 pages. [cited by applicant]
Shi et al., “Real-time single image and video super-resolution using an efficient subpixel convolutional neural network,” Proceedings of the IEEE conference on computer vision and pattern recognition, Sep. 23, 2016, 10 … [cited by applicant]
Tai et al., “Image super-resolution via deep recursive residual network,” Proceedings of the IEEE conference on computer vision and pattern 10 iSeeBetter: Spatio-temporal video super-resolution using recurrent generativ… [cited by applicant]
Tao et al., “Detail Revealing deep video super-resolution,” Proceedings of the IEEE International Conference on Computer Vision, 2017, 9 pages. [cited by applicant]
Venturebeat.com [online], “Headroom launches to combat Zoom fatigue with AI,” Dec. 10, 2021, retrieved on Jun. 22, 2022, retrieved from URL<https://venturebeat.com/2021/12/10/headroom-launches-to-combat-zoom-fatigue-wit… [cited by applicant]
Venugopalan et al., “Translating videos to natural language using deep recurrent neural networks,” Human Language Technologies: The 2015 Annual Conference of the North American Chapter of the ACL, 2015, 11 pages. [cited by applicant]
Wang et al., “ESRGAN: Enhanced super-resolution generative adversarial networks,” Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Sep. 17, 2018, 23 pages. [cited by applicant]
Wang et al., “Multi-Memory Convolutional Neural Network for Video Super-Resolution,” IEEE Transactions on Image Processing, May 1, 2019, 28(5): 2530-2544. [cited by applicant]
Yang et al., “Image super-resolution: Historical overview and future challenges,” Super-resolution imaging, 2010, pp. 20-34 (abstract only). [cited by applicant]
Yu et al., “Video paragraph captioning using hierarchical recurrent neural networks,” Proceedings of the IEEE conference on computer vision and pattern recognition, Apr. 6, 2016, 10 pages. [cited by applicant]
Yuan et al., “Dual Discriminator Generative Adversarial Network for Single Image Super-Resolution,” 2019 12th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics (CISP-BMEI), Oc… [cited by applicant]
Zhang et al., “A Flexible Recurrent Residual Pyramid Network for Video Frame Interpolation,” arxiv.org, Mar. 31, 2020, 18 pages. [cited by applicant]
Zhang et al., “The unreasonable effectiveness of deep features as a perceptual metric,” Proceedings of the IEEE conference on computer vision and pattern recognition, Apr. 10, 2018, 14 pages. [cited by applicant]
Yulin Wang; Glance and Focus: a Dynamic Approach to Reducing Spatial Redundancy in Image Classification ( Year: 2020). [cited by applicant]