IP Library › Granted Patent US 12,524,841
Granted Patent B2
US 12,524,841 · App. 17/876,373 · Granted Jan 13, 2026

Techniques for processing videos using temporally-consistent transformer model

Inventors: Yang Zhang (Dubendorf, CH); Mingyang Song (Zurich, CH); Tunc Ozan Aydin (Zurich, CH); Christopher Richard Schroers (Uster, CH)
Assignees: Disney Enterprises, INC.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
G06T5/70G06T3/18G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,841
App. No.
17/876,373
Granted
Jan 13, 2026
Kind
B2
Abstract

Techniques are disclosed for enhancing videos using a machine learning model that is a temporally-consistent transformer model. The machine learning model processes blocks of frames of a video in which the temporally first input video frame of each block of frames is a temporally second to last output video frame of a previous block of frames. After the machine learning model is trained, blocks of video frames, or features extracted from the video frames, can be warped using an optical flow technique and transformed using a wavelet transform technique. The transformed video frames are concatenated along a channel dimension and input into the machine learning model that generates corresponding processed video frames.

Claims (41)

1 . A computer-implemented method for enhancing videos, the method comprising:

receiving, from a video, a first plurality of video frames and a second plurality of video frames;

processing the first plurality of video frames using a machine learning model to generate a first plurality of processed video frames;

replacing a temporally first video frame included in the second plurality of video frames with a temporally second to last video frame included in the first plurality of processed video frames; and

processing the second plurality of video frames using the machine learning model to generate a second plurality of processed video frames.

2 . The computer-implemented method of claim 1 , further comprising performing one or more operations to train the machine learning model using a loss function that penalizes a difference between a temporally last frame of a plurality of processed training video frames generated by the machine learning model and a temporally first frame of a subsequent plurality of processed training video frames generated by the machine learning model.

3 . The computer-implemented method of claim 1 , wherein processing the first plurality of video frames using the machine learning model comprises:

performing one or more transform operations on each video frame included in the first plurality of video frames to generate a plurality of transformed video frames;

concatenating the plurality of transformed video frames along a channel dimension to generate a concatenated set of transformed video frames; and

inputting the concatenated set of transformed video frames into the machine learning model.

4 . The computer-implemented method of claim 3 , wherein the one or more transform operations include at least one of (1) one or more wavelet transform operations or (2) one or more pixel shuffle operations.

5 . The computer-implemented method of claim 1 , further comprising:

generating an optical flow based on the first plurality of video frames; and

warping either the first plurality of video frames or features extracted from the first plurality of video frames based on the optical flow.

6 . The computer-implemented method of claim 1 , further comprising:

adding a plurality of amounts of degradation to a set of video frames to generate a plurality of sets of degraded video frames, wherein each set of degraded video frames included in the plurality of sets of degraded video frames includes a different amount of the degradation; and

training the machine learning model based on the plurality of sets of degraded video frames.

7 . The computer-implemented method of claim 6 , wherein the degradation comprises at least one of noise or blur.

8 . The computer-implemented method of claim 1 , wherein the machine learning model comprises a transformer model.

9 . The computer-implemented method of claim 1 , wherein the machine learning model comprises one or more three-dimensional (3D) convolution layers.

10 . One or more non-transitory computer-readable storage media including instructions that, when executed by one or more processing units, cause the one or more processing units to perform steps for enhancing videos, the steps comprising:

receiving, from a video, a first plurality of video frames and a second plurality of video frames;

processing the first plurality of video frames using a machine learning model to generate a first plurality of processed video frames;

replacing a temporally first video frame included in the second plurality of video frames with a temporally second to last video frame included in the first plurality of processed video frames; and

processing the second plurality of video frames using the machine learning model to generate a second plurality of processed video frames.

11 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the instructions, when executed by the one or more processing units, further cause the one or more processing units to perform the step of:

performing one or more operations to train the machine learning model using a loss function that penalizes a difference between a temporally last frame of a plurality of processed training video frames generated by the machine learning model and a temporally first frame of a subsequent plurality of processed training video frames generated by the machine learning model.

12 . The one or more non-transitory computer-readable storage media of claim 10 , wherein processing the first plurality of video frames using the machine learning model comprises:

performing one or more transform operations on each video frame included in the first plurality of video frames to generate a plurality of transformed video frames;

concatenating the plurality of transformed video frames along a channel dimension to generate a concatenated set of transformed video frames; and

inputting the concatenated set of transformed video frames into the machine learning model.

13 . The one or more non-transitory computer-readable storage media of claim 12 , wherein the one or more transform operations include at least one of one or more wavelet transform operations or one or more pixel shuffle operations.

14 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the instructions, when executed by the one or more processing units, further cause the one or more processing units to perform the steps of:

generating an optical flow based on the first plurality of video frames; and

warping either the first plurality of video frames or features extracted from the first plurality of video frames based on the optical flow.

15 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the instructions, when executed by the one or more processing units, further cause the one or more processing units to perform the steps of:

adding a plurality of amounts of degradation to a set of video frames to generate a plurality of sets of degraded video frames, wherein each set of degraded video frames included in the plurality of sets of degraded video frames includes a different amount of the degradation; and

training the machine learning model based on the plurality of sets of degraded video frames.

16 . The one or more non-transitory computer-readable storage media of claim 15 , wherein the degradation comprises at least one of noise or blur.

17 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the machine learning model comprises a transformer model.

18 . The one or more non-transitory computer-readable storage media of claim 10 , wherein the machine learning model comprises one or more three-dimensional (3D) convolution layers.

Assignments (3)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE RECEIVING PARTY DATA PREVIOUSLY RECORDED AT REEL: 060672 FRAME: 0087. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Dec 12, 2022
From: ZHANG, YANG; SONG, MINGYANG; AYDIN, TUNC OZAN; SCHROERS, CHRISTOPHER RICHARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 062114/0503 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2022
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 062026/0511 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2022
From: ZHANG, YANG; SONG, MINGYANG; AYDIN, TUNC OZAN; SCHROERS, CHRISTOPHER RICHARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
Reel/Frame 060672/0087 →
Continuity (2)
Provisional Application 63316888 · Mar 4, 2022
Related Publication 20230281757A1 · Sep 7, 2023
References Cited (47)
US 11610387B2 · Jouhikainen · 2023 [cited by examiner]
US 12002257B2 · Kandpal · 2024 [cited by examiner]
US 20190347806A1 · Vajapey · 2019 [cited by examiner]
US 20210216778A1 · Ramaswamy · 2021 [cited by examiner]
US 20220051382A1 · Chen · 2022 [cited by examiner]
US 20220058400A1 · Bianconcini · 2022 [cited by examiner]
US 20220083785A1 · Subramanian · 2022 [cited by examiner]
US 20220239925A1 · Topiwala · 2022 [cited by examiner]
US 20220292649A1 · Wang · 2022 [cited by examiner]
US 20230080639A1 · Zoss · 2023 [cited by examiner]
US 20240022760A1 · Li · 2024 [cited by examiner]
WO WO2020015492A1 · 2020 [cited by examiner]
WO WO2023149898A1 · 2023 [cited by examiner]
El-Nouby et al., “Xcit: Cross-Covariance Image Transformers”, 35th Conference on neural information processing systems, 2021, 14 pages. [cited by applicant]
Bertasius et al., “Is Space-Time Attention All You Need for Video Understanding?”, arXiv:2102.05095, Jun. 9, 2021, 13 pages. [cited by applicant]
Buades et al., “A Non-Local Algorithm for Image Denoising”, In 2005 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, IEEE, vol. 2, 2005, pp. 60-65. [cited by applicant]
Cai et al., “Robust Image Denoising Using Kernel Predicting Networks”, Eurographics, 2021, 4 pages. [cited by applicant]
Carion et al., “End-to-End Object Detection with Transformers”, In European Conference on Computer Vision, arxiv:2005.12872, May 28, 2020, pp. 213-229. [cited by applicant]
Chan et al., “BasicVSR++: Improving Video Super-Resolution with Enhanced Propagation and Alignment”, arXiv:2104.13371, Apr. 27, 2021, 12 pages. [cited by applicant]
Chen et al., “Pre-Trained Image Processing Transformer”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 12299-12310. [cited by applicant]
Dai et al., “Dynamic Head: Unifying Object Detection Heads with Attentions”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, arxiv:2106.08322, Jun. 15, 2021, pp. 7373-7382. [cited by applicant]
Dai et al., “CoAtNet: Marrying Convolution and Attention for all Data Sizes”, Advances in Neural Information Processing Systems, 2021, 18 pages. [cited by applicant]
Dosovitskiy et al., “An Image is worth 16x16 Words: Transformers for Image Recognition at Scale”, arXiv:2010.11929, Oct. 22, 2020, 21 pages. [cited by applicant]
Lai et al., “Learning Blind Video Temporal Consistency”, In Proceedings of the European conference on computer vision (ECCV), arxiv: 1808.00449, Aug. 1, 2018, pp. 170-185. [cited by applicant]
Lei et al., “Blind Video Temporal Consistency via Deep Video Prior”, 34th Conference on Neural Information Processing Systems, arxiv:2010.11838, Oct. 22, 2020, 11 pages. [cited by applicant]
Li et al., “Uniformer: Unified Transformer for Efficient Spatiotemporal Representation Learning”, arXiv:2201.04676, Feb. 8, 2022, 19 pages. [cited by applicant]
Liang et al., “VRT: A Video Restoration Transformer”, arxiv:2201.12288, Jun. 15, 2022, 14 pages. [cited by applicant]
Liang et al., “SwinIR: Image Restoration using Swin Transformer”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, arxiv:2108.10257, Aug. 23, 2021, 12 pages. [cited by applicant]
Liu et al., “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows”, arXiv:2103.14030, Aug. 17, 2021, 14 pages. [cited by applicant]
Maggioni et al., “Video Denoising, Deblocking, and Enhancement through Separable 4-D Nonlocal Spatiotemporal Transforms”, IEEE Transactions on Image Processing, vol. 21, No. 9, Sep. 2012, pp. 3952-3966. [cited by applicant]
Maggioni et al., “Efficient Multi-Stage Video Denoising with Recurrent Spatio-Temporal Fusion”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 3466-3475. [cited by applicant]
Perazzi et al., “A Benchmark Dataset and Evaluation Methodology for Video Object Segmentation”, In Proceedings of the IEEE conference on computer vision and pattern recognition, 2016, pp. 724-732. [cited by applicant]
Tassano et al., “Dvdnet: A Fast Network for Deep Video Denoising”, In 2019 IEEE International Conference on Image Processing (ICIP), arxiv:1906.11890, Jun. 4, 2019, 6 pages. [cited by applicant]
Tassano et al., “FastDVDnet: Towards Real-Time Deep Video Denoising Without Flow Estimation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, arxiv:1907.01361, Apr. 29, 2020, 13 pag… [cited by applicant]
Vaksman et al., “Patch Craft: Video Denoising by Deep Modeling and Patch Matching”, arxiv:2103.13767, Oct. 30, 2021, 16 pages. [cited by applicant]
Vaswani et al., “Attention is All You Need”, 34th Conference on Neural Information Processing Systems, arxiv:1706.03762, Dec. 6, 2017, 15 pages. [cited by applicant]
Wang et al., “First Image then Video: A Two-Stage Network for Spatiotemporal Video Denoising”, arXiv:2001.00346, Jan. 22, 2020, 13 pages. [cited by applicant]
Xu et al., “End-to-End Semi-Supervised Object Detection with Soft Teacher”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, arxiv.2106.09018, Aug. 6, 2021, 10 pages. [cited by applicant]
Yang et al., “Focal Self-Attention for Local-Global Interactions in Vision Transformers”, arXiv:2107.00641, Jul. 1, 2021, 21 pages. [cited by applicant]
Yu et al., “Metaformer is Actually What You Need for Vision”, arXiv:2111.11418, Nov. 22, 2021, 14 pages. [cited by applicant]
Yuan et al., “Florence: A New Foundation Model for Computer Vision”, arXiv:2111.11432, Nov. 22, 2021, 17 pages. [cited by applicant]
Yue et al., “Supervised Raw Video Denoising with a Benchmark Dataset on Dynamic Scenes”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, arxiv:2003.14013, Mar. 31, 2020, 10 pages. [cited by applicant]
Zamir et al., “Restormer: Efficient Transformer for High-Resolution Image Restoration”, arXiv:2111.09881, Nov. 18, 2021, 12 pages. [cited by applicant]
Zhai et al., “Scaling Vision Transformers”, arxiv:2106.04560, Jun. 8, 2021, 31 pages. [cited by applicant]
Zhang et al., “Beyond a Gaussian Denoiser: Residual Learning of Deep CNN for Image Denoising”, DOI:10.1109/TIP.2017.2662206, IEEE transactions on Image Processing, vol. 26, No. 7, 2017, 14 pages. [cited by applicant]
Zhang et al., “Learning Deep CNN Denoiser Prior for Image Restoration”, In Proceedings of the IEEE conference on computer vision and pattern recognition, arxiv:1704.03264, Apr. 11, 2017, 10 pages. [cited by applicant]
Zhang et al., “FFDNet: Toward A Fast and Flexible Solution for CNN Based Image Denoising”, IEEE Transactions on Image Processing, vol. 27, No. 9, arxiv:1710.04026, May 22, 2018, 15 pages. [cited by applicant]