IP Library › Granted Patent US 12,250,400
Granted Patent B2
US 12,250,400 · App. 17/670,978 · Granted Mar 11, 2025

Unified space-time interpolation of video information

Inventors: Luming Liang (Redmond, WA); Zhicheng Geng (Austin, TX); Ilya Dmitriyevich Zharkov (Sammamish, WA); Tianyu Ding (Kirkland, WA)
Assignee: Microsoft Technology Licensing, LLC
H04N19/587H04N19/132H04N19/31H04N19/33H04N19/59H04N19/61H04N19/439
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,250,400
App. No.
17/670,978
Granted
Mar 11, 2025
Kind
B2
Abstract

A technique is described herein for temporally and spatially interpolating input video information, to produce output video information having a higher frame rate and a higher resolution compared to that exhibited by the input video information. The technique generates feature information based on plural frames of the input video information. The technique then produces the output video information based on the feature information using an architecture having, in order, a multi-stage encoding operation, a query-generating operation, and a multi-stage decoding operation. Each encoding stage produces an instance of encoder attention information that expresses identified relations across the plural frames of the input video information. Each decoding stage operates on an instance of encoder attention information produced by a corresponding encoding stage. The transformer architecture is compact and is capable of interpolating the input video information in real time.

Claims (53)

1. A method for interpolating video information, comprising:

obtaining input video information having a given first number of plural frames, each frame in the input video information having a given first spatial resolution;

generating feature information based the input video information;

encoding the feature information in a pipeline having plural encoding stages that operate at different respective resolutions, to produce plural instances of encoder attention information and plural instances of encoder output information, each instance of the encoder attention information expressing identified relations across the plural frames of the input video information;

producing a query based on an instance of encoder output information produced by a last encoding stage of the plural encoding stages;

decoding the query in a pipeline having plural decoding stages that operate at different respective resolutions, to produce plural instances of decoder output information, each decoding stage that has a preceding decoding stage receiving an instance of decoding input information produced by the preceding decoding stage, and each particular decoding stage operating on an instance of encoder attention information produced by a particular encoding stage that has a same resolution level as the particular decoding stage; and

producing output video information based on decoder output information produced by a last decoding stage of the plural decoding stages,

said producing including:

performing a reconstruction operation on the decoder output information produced by the last decoding stage, to produce reconstructed information;

interpolating the input video information to produce interpolated video information; and

combining the reconstructed information with the interpolated video information to produce the output video information,

the output video information having a second number of frames that is higher than the first number of frames in the input video information, and having a second spatial resolution that is higher than the first spatial resolution of the input video information.

2. The method of claim 1 , wherein each given stage in the pipeline of encoding stages and the pipeline of decoding stages produces a particular instance of attention information using a transformer-based neural network.

3. The method of claim 2 , wherein the transformer-based neural network performs a first kind of attention operation that involves partitioning first-attention-operation input information into individual windows, and generating window-specific attention information for the individual windows.

4. The method of claim 3 , wherein the transformer-based neural network also performs a second kind of attention operation that involves partitioning second-attention-operation input information into individual windows, and generating window-specific attention information for the individual windows produced by the second kind of attention operation, the individual windows in the second kind of attention operation being shifted relative to the individual windows in the first kind of attention operation.

5. The method of claim 2 , wherein the given stage is a given encoding stage, and wherein the given encoding stage also performs a down-sampling operation on the particular instance of attention information.

6. The method of claim 2 , wherein the given stage is a given encoding stage, and wherein the particular instance of attention information includes key information and value information that is generated based on feature information obtained from the plural frames of the input video information.

7. The method of claim 2 , wherein the given stage is a given decoding stage, and wherein the given decoding stage also performs an up-sampling operation on the particular instance of attention information.

8. The method of claim 1 , wherein said producing a query involves producing encoder output information for at least one added frame that is not present in the input video information.

9. The method of claim 8 , wherein said producing a query produces the encoder output information for said at least one added frame based on the encoder output information produced by the last encoding stage for frames in the input video information that temporally precede and follow the added frame.

10. A computing system for performing an interpolation operation, comprising:

a memory for storing machine-readable instruction;

a processor that performs operations by executing the machine-readable instructions stored in the memory, the operations including:

obtaining input video information having a given first number of plural frames, each frame in the input video information having a given first spatial resolution;

generating feature information based on the input video information; and

producing output video information based on the feature information using a transformer-based encoding operation, followed by a query-generating operation, followed by a transformer-based decoding operation, followed by an interpolating operation, the transformer-based decoding operation being performed based on encoder attention information produced by the transformer-based encoding operation,

the encoder attention information expressing identified relations across the plural frames of the input video information,

the output video information having a second number of frames that is higher than the first number of frames in the input video information, and having a second spatial resolution that is higher than the first spatial resolution of the input video information,

wherein the transformer-based encoding operation includes a pipeline having plural encoding stages that operate at different respective resolutions, and that produce plural instances of encoder attention information and plural instances of encoder output information,

each instance of the encoder attention information expressing identified relations across the plural frames of the input video information,

wherein the query-generating operation produces a query based on an instance of encoder output information produced by a last encoding stage of the plural encoding stages,

wherein the transformer-based decoding operation decodes the query using a pipeline having plural decoding stages that operate at different respective resolutions, and that produce plural instances of decoder output information, each decoding stage that has a preceding decoding stage receiving an instance of decoder input information produced by the preceding decoding stage, and each particular decoding stage operating on an instance of encoder attention information produced by a particular encoding stage that has a same resolution level as the particular decoding stage,

wherein the interpolating operation interpolates the input video information to produce interpolated video information, and wherein the operations further include combining decoder output information produced by a last decoding stage of the plural decoding stages with the interpolated video information to produce the output video information.

11. The computing system of claim 10 , wherein a given encoding stage performs at least one kind of attention operation and a down-sampling operation.

12. The computing system of claim 10 , wherein a given decoding stage performs at least one kind of attention operation and an up-sampling operation.

13. A non-transitory computer-readable storage medium for storing computer-readable instructions, the computer-readable instructions, when executed by one or more hardware processors, performing a method that comprises:

obtaining input video information having a given first number of plural frames, each frame in the input video information having a given first spatial resolution;

generating feature information based on the input video information; and

producing output video information based on the feature information using a transformer-based encoding operation, followed by a query-generating operation, followed by a transformer-based decoding operation, followed by an interpolating operation, the transformer-based decoding operation being performed based on encoder attention information produced by the transformer-based encoding operation,

the encoder attention information expressing identified relations across the plural frames of the input video information,

the output video information having a second number of frames that is higher than the first number of frames in the input video information, and having a second spatial resolution that is higher than the first spatial resolution of the input video information,

wherein the transformer-based encoding operation includes a pipeline having plural encoding stages that operate at different respective resolutions, and that produce plural instances of encoder attention information and plural instances of encoder output information,

each instance of the encoder attention information expressing identified relations across the plural frames of the input video information,

wherein the query-generating operation produces a query based on an instance of encoder output information produced by a last encoding stage of the plural encoding stages,

wherein the transformer-based decoding operation decodes the query using a pipeline having plural decoding stages that operate at different respective resolutions, and that produce plural instances of decoder output information, each decoding stage that has a preceding decoding stage receiving an instance of decoder input information produced by the preceding decoding stage, and each particular decoding stage operating on an instance of encoder attention information produced by a particular encoding stage that has a same resolution level as the particular decoding stage,

wherein the interpolating operation interpolates the input video information to produce interpolated video information, and wherein the operations further include combining decoder output information produced by a last decoding stage of the plural decoding stages with the interpolated video information to produce the output video information.

14. The non-transitory computer-readable storage medium of claim 13 , wherein the pipeline of encoding stages operate using successively lower resolutions, and wherein the pipeline of decoding stages operate using successively higher resolutions.

15. The non-transitory computer-readable storage medium of claim 13 , wherein each given encoding stage in the pipeline of encoding stages performs a first kind of attention operation that involves partitioning first-attention-operation input information into individual windows, and generating window-specific attention information for the individual windows,

each individual window for which the first kind of attention operation is performed being based on feature information obtained from the plural frames of the input video information.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the given encoding stage also performs at a second kind of attention operation that involves partitioning second-attention-operation input information into individual windows, and generating window-specific attention information for individual windows produced by the second kind of attention operation, the individual windows in the second kind of attention operation being shifted relative to the individual windows in the first kind of attention operation,

each individual window for which the second kind of attention operation is performed being based on feature information obtained from the plural frames of the input video information.

17. The non-transitory computer-readable storage medium of claim 13 , further comprising:

performing a reconstruction operation on the decoder output information produced by the last decoding stage of the plural decoding stages, to produce reconstructed information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2022
From: LIANG, LUMING; GENG, ZHICHENG; ZHARKOV, ILYA DMITRIYEVICH; DING, TIANYU
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 059008/0348 →
Continuity (1)
Related Publication 20230262259A1 · Aug 17, 2023
References Cited (93)
US 10733704B2 · Bar-On · 2020 [cited by examiner]
US 11765360B2 · Schroers · 2023 [cited by examiner]
US 20130301933A1 · Salvador · 2013 [cited by examiner]
US 20140177706A1 · Fernandes · 2014 [cited by examiner]
US 20150023611A1 · Salvador · 2015 [cited by examiner]
US 20150104116A1 · Salvador · 2015 [cited by examiner]
US 20150172726A1 · Faramarzi · 2015 [cited by examiner]
US 20190122117A1 · Miyazaki · 2019 [cited by examiner]
US 20220284267A1 · Vitthaladevuni · 2022 [cited by examiner]
US 20230081916A1 · Chee · 2023 [cited by examiner]
CN 110415170A · 2019 [cited by examiner]
CN 112070677A · 2020 [cited by examiner]
CN 112862688B · 2021 [cited by examiner]
CN 113747242A · 2021 [cited by examiner]
CN 114092339A · 2022 [cited by examiner]
WO WO2023081095A1 · 2023 [cited by examiner]
Xiao et al. “Space-time super-resolution for satellite video: A joint framework based on multi-scale spatial-temporal transformer” (Year: 2022). [cited by examiner]
Xiao et al. “Convolutional Hierarchical Attention Network for Query-Focused Video Summarization” (Year: 2020). [cited by examiner]
Pu et al. “ED-ACNN: Novel attention convolutional neural network based on encoder-decoder framework for human traffic prediction” (Year: 2020). [cited by examiner]
Geng, et al., “RSTT: Real-time Spatial Temporal Transformer for Space-Time Video Super-Resolution,” arXiv, e-prints, arXiv:2203.14186v1 [cs.CV], Mar. 27, 2022, 12 pages. [cited by applicant]
Geng, et al., “RSTT: Real-time Spatial Temporal Transformer for Space-Time Video Super-Resolution,” in 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2022, pp. 17420-17430. [cited by applicant]
Bao, et al., “Depth-Aware Video Frame Interpolation,” open access version of paper in Proceeding of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2019, 10 pages. [cited by applicant]
Bao, et al., “MEMC-Net: Motion Estimation and Motion Compensation Driven Neural Network for Video Interpolation and Enhancement,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, No. 3, Mar. 2… [cited by applicant]
Brown, et al., “Language Models are Few-Shot Learners,” arXiv e-prints, arXiv:2005.14165v4 [cs.CL], Jul. 22, 2020, 75 pages. [cited by applicant]
Caballero, et al., “Real-Time Video Super-Resolution with Spatio-Temporal Networks and Motion Compensation,” open access version of paper in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, … [cited by applicant]
Chen, et al., “Pre-Trained Image Processing Transformer,” open access version of paper in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 12 pages. [cited by applicant]
Cheng, et al., “Multiple Video Frame Interpolation via Enhanced Deformable Separable Convolution,” arXiv e-prints, arXiv:2006.08070v2 [cs.CV], Jan. 25, 2021, 18 pages. [cited by applicant]
Cheng, et al., “Video Frame Interpolation via Deformable Separable Convolution,” in Proceedings of the AAAI Conference on Artificial Intelligence, 34(07), 2020, pp. 10607-10614. [cited by applicant]
Claus, et al., “ViDeNN: Deep Blind Video Denoising,” in Proceedings of the 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2019, 10 pages. [cited by applicant]
Dai, et al., “Deformable Convolutional Networks,” open access version of paper in 2017 IEEE International Conference on Computer Vision (ICCV), 2017, 10 pages. [cited by applicant]
Devlin, et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” arXiv e-prints, arXiv:1810.04805v2 [cs.CL], May 24, 2019, 16 pages. [cited by applicant]
Ding, et al., “CDFI: Compression-Driven Network Design for Frame Interpolation,” open access version of paper in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 11 pages. [cited by applicant]
Dosovitskiy, et al., “An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale,” arXiv e-prints, arXiv:2010.11929v2 [cs.CV], Jun. 3, 2021, 22 pages. [cited by applicant]
Gotmare, et al., “A Closer Look at Deep Learning Heuristics: Learning rate restarts, Warmup and Distillation,” arXiv e-prints, arXiv:1810.13243v1 [cs.LG], Oct. 29, 2018, 18 pages. [cited by applicant]
Haris, et al., “Recurrent Back-Projection Network for Video Super-Resolution,” open access version of paper in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 10 pages. [cited by applicant]
Haris, et al., “Space-Time-Aware Multi-Resolution Video Enhancement,” in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 10 pages. [cited by applicant]
Jiang, et al., “Super SloMo: High Quality Estimation of Multiple Intermediate Frames for Video Interpolation,” open access version of paper in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, 9… [cited by applicant]
Jo, et al., “Deep Video Super-Resolution Network Using Dynamic Upsampling Filters Without Explicit Motion Compensation,” open access version of paper in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognitio… [cited by applicant]
Kim, et al., “Deep Video Inpainting,” open access version of paper in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 10 pages. [cited by applicant]
Lee, et al., “AdaCoF: Adaptive Collaboration of Flows for Video Frame Interpolation,” open access version of paper in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 10 pages. [cited by applicant]
Liang, et al., “SwinIR: Image Restoration Using Swin Transformer,” open access version of paper in 2021 IEEE/CVF International Conference on Computer Vision Workshops (ICCVW), Oct. 2021, 12 pages. [cited by applicant]
Liu, et al., “A Bayesian Approach to Adaptive Video Super Resolution,” in CVPR 2011, 2011, pp. 209-216. [cited by applicant]
Liu, et al., “Video Super Resolution Based on Deep Learning: A Comprehensive Survey,” arXiv e-prints, arXiv:2007.12928v2 [cs.CV], Dec. 20, 2020, 30 pages. [cited by applicant]
Liu, et al., “Swin Transformer: Hierarchical Vision Transformer Using Shifted Windows,” open access version of paper in IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, 11 pages. [cited by applicant]
Liu, et al., “Video Frame Synthesis Using Deep Voxel Flow,” open access version of paper in 2017 IEEE International Conference on Computer Vision (ICCV), 2017, 9 pages. [cited by applicant]
Loschilov, et al., “Decoupled Weight Decay Regularization,” available at https://openreview.net/forum?id=Bkg6RiCqY7, open review version of paper in 2019 International Conference on Learning Representations (ICLR), modi… [cited by applicant]
Mahajan, et al., “Moving Gradients: A Path-Based Method for Plausible Image Interpolation,” in SIGGRAPH '09: ACM SIGGRAPH 2009, Article No. 42, Jul. 2009, 12 pages. [cited by applicant]
Meyer, et al., “PhaseNet for Video Frame Interpolation,” open access version of paper in 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, 10 pages. [cited by applicant]
Meyer, et al., “Phase-Based Frame Interpolation for Video,” open access version of paper in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2015, 9 pages. [cited by applicant]
Mudenagudi, et al., “Space-Time Super-Resolution Using Graph-Cut Optimization,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 33, No. 5, May 2011, pp. 995-1008. [cited by applicant]
Niklaus, et al., “Context-Aware Synthesis for Video Frame Interpolation,” open access version of paper in IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, 10 pages. [cited by applicant]
Niklaus, et al., “Softmax Splatting for Video Frame Interpolation,” open access version of paper in IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, 10 pages. [cited by applicant]
Niklaus, et al., “Video Frame Interpolation via Adaptive Convolution,” open access version of paper in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, 10 pages. [cited by applicant]
Niklaus, et al., “Video Frame Interpolation via Adaptive Separable Convolution,” open access version of paper in IEEE International Conference on Computer Vision (ICCV), 2017, 10 pages. [cited by applicant]
Park, et al., “BMBC: Bilateral Motion Estimation with Bilateral Cost Volume for Video Interpolation,” arXiv e-prints, arXiv:2007.12622v1 [cs.CV], Jul. 17, 2020, 16 pages. [cited by applicant]
Ranftl, et al., “Vision Transformers for Dense Prediction,” open access version of paper in IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2021, 10 pages. [cited by applicant]
Rippel, et al., “Learned Video Compression,” open access version of paper in IEEE/CVF International Conference on Computer Vision (ICCV), 2019, 10 pages. [cited by applicant]
Sajjadi, et al., “Frame-Recurrent Video Super-Resolution,” open access version of 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2018, 9 pages. [cited by applicant]
Shechtman, et al., “Increasing Space-Time Resolution in Video,” ECCV 2002, European Conference on Computer Vision, LNCS 2350, Springer-Verlag Berlin Heidelberg, 2002, pp. 753-768. [cited by applicant]
Shechtman, et al., “Space-Time Super-Resolution,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 27, No. 4, Apr. 2005, pp. 531-545. [cited by applicant]
Shi, et al., “Video Frame Interpolation via Generalized Deformable Convolution,” arXiv e-prints, arXiv:2008.10680v3 [cs.CV], Mar. 18, 2021, 13 pages. [cited by applicant]
Strudel, “Segmenter: Transformer for Semantic Segmentation,” arXiv e-prints, arXiv:2105.05633v3 [cs.CV], Sep. 2, 2021, 17 pages. [cited by applicant]
Tag, et al., “In the Eye of the Beholder: The Impact of Frame Rate on Human Eye Blink,” in CHI EA '16: Proceedings of the 2016 CHI Conference Extended Abstracts on Human Factors in Computing Systems, May 2016, pp. 2321-… [cited by applicant]
Tao, et al., “Detail-Revealing Deep Video Super-Resolution,” open access version of paper in 2017 IEEE International Conference on Computer Vision (ICCV), 2017, 9 pages. [cited by applicant]
Tassano, et al., “FastDVDnet: Towards Real-Time Deep Video Denoising Without Flow Estimation,” open access version of paper in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 10 pa… [cited by applicant]
Turletti, et al., “Videoconferencing on the Internet,” in IEEE/ACM Transactions on Networking, vol. 4, No. 3, Jun. 1996, pp. 340-351. [cited by applicant]
Vaswani, et al., “Attention Is All You Need,” in 31st Conference on Neural Information Processing Systems (NIPS 2017), 2017, 11 pages. [cited by applicant]
Wang, et al., “Video Inpainting by Jointly Learning Temporal Structure and Spatial Details,” in Proceedings of the Thirty-Third AAAI Conference on Artificial Intelligence (AAAI-19), vol. 33, 2019, pp. 5232-5239. [cited by applicant]
Wang, et al., “Learning for Video Super-Resolution through HR Optical Flow Estimation,” arXiv e-prints, arXiv:1809.08573v2 [cs.CV], Oct. 25, 2018, 10 pages. [cited by applicant]
Wang, et al., “EDVR: Video Restoration with Enhanced Deformable Convolutional Networks,” open access version of paper in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW), 2019, 10 pa… [cited by applicant]
Wang, et al., “Uformer: A General U-Shaped Transformer for Image Restoration,” arXiv e-prints, arXiv:2106.03106v2 [cs.CV], Nov. 25, 2021, 17 pages. [cited by applicant]
Wu, et al., “Streaming Video over the Internet: Approaches and Directions,” in IEEE Transactions on Circuits and Systems for Video Technology, vol. 11, No. 3, 2001, pp. 282-300. [cited by applicant]
Xiang, et al., “Zooming Slow-Mo: Fast and Accurate One-Stage Space-Time Video Super-Resolution,” open access version of paper in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 10 … [cited by applicant]
Xiang, et al., “Zooming SlowMo: An Efficient One-Stage Framework for Space-Time Video Super-Resolution,” arXiv e-prints, arXiv:2104.07473v1 [cs.CV], Apr. 15, 2021, 14 pages. [cited by applicant]
Xu, et al., “Temporal Modulation Network for Controllable Space-Time Video Super-Resolution,” open access version of paper in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, 10 pag… [cited by applicant]
Xu, et al., “BERT, mBERT, or BiBERT? A Study on Contextualized Embeddings for Neural Machine Translation,” arXiv e-prints, arXiv:2109.04588v1 [cs.CL], Septemer 9, 2021, 13 pages. [cited by applicant]
Xu, et al., “End-to-End Semi-Supervised Object Detection with Soft Teacher,” arXiv e-prints, arXiv:2106.09018v3 [cs.CV], Aug. 6, 2021, 10 pages. [cited by applicant]
Xue, et al., “Video Enhancement with Task-Oriented Flow,” arXiv e-prints, arXiv:1711.09078v3 [cs.CV], Nov. 10, 2019, 20 pages. [cited by applicant]
Yang, et al., “Focal Self-attention for Local-Global Interactions in Vision Transformers,” arXiv e-prints, arXiv:2107.00641v1 [cs.CV], Jul. 1, 2021, 21 pages. [cited by applicant]
Tian, et al., “TDAN: Temporally-Deformable Alignment Network for Video Super-Resolution,” open access version of paper in 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, 10 pages. [cited by applicant]
Yuan, et al., “Zoom-In-to-Check: Boosting Video Interpolation via Instance-level Discrimination,” open access version of paper published in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 201… [cited by applicant]
Zhang, et al., “Image Super-Resolution Using Very Deep Residual Channel Attention Networks,” open access version of paper in 2020 IEEE 15th International Conference on Industrial and Information Systems (ICIIS), 2020, 1… [cited by applicant]
Zheng, et al., “Rethinking Semantic Segmentation from a Sequence-to-Sequence Perspective with Transformers,” open access version of paper in 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Ju… [cited by applicant]
Zhou, et al., “How Video Super-Resolution and Frame Interpolation Mutually Benefit,” in MM '21: Proceedings of the 29th ACM International Conference on Multimedia, Oct. 2021, pp. 5445-5453. [cited by applicant]
Zhu, et al., “Deformable ConvNets v2: More Deformable, Better Results,” open access version of paper in 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, 9 pages. [cited by applicant]
Khandelwal, Renu, “SWIN Transformer: A Unifying Step Between Computer Vision and Natural Language Processing,” available at https://arshren.medium.com/swin-transformer-a-unifying-step-between-computer-vision-and-natural… [cited by applicant]
Alammar, Jay, “The Illustrated Transformer,” available at http://jalammar.github.io/illustrated-transformer/, Github, Jun. 27, 2018, 23 pages. [cited by applicant]
Shaw, et al., “Self-Attention with Relative Position Representations,” arXiv e-prints, arXiv:1803.02155v2 [cs.CL], Apr. 12, 2018, 5 pages. [cited by applicant]
Shi, et al., “Real-Time Single Image and Video Super-Resolution Using an Efficient Sub-Pixel Convolutional Neural Network,” arXiv e-prints, arXiv:1609.05158v2 [cs.CV], Sep. 23, 2016, 10 pages. [cited by applicant]
Zhang, et al., “Cross-Frame Transformer-Based Spatio-Temporal Video Super-Resolution,” in IEEE Transactions on Broadcasting, vol. 68, No. 2, Jun. 2022, pp. 359-369. [cited by applicant]
Huang, et al., “Confidence-Based Global Attention Guided Network for Image Inpainting,” in MultiMedia Modeling, MMM 2021, Lecture Notes in Computer Science, vol. 12572, Springer, Jan. 21, 2021, pp. 200-212. [cited by applicant]
Khan, et al., “Transformers in Vision: A Survey,” arXiv, Cornell University, arXiv:2101.01169v4 [cs.CV], Oct. 3, 2021, 30 pages. [cited by applicant]
PCT Search Report and Written Opinion for PCT/US2022/050038, 16 pages. [cited by applicant]