IP Library › Granted Patent US 12,309,404
Granted Patent B2
US 12,309,404 · App. 18/349,076 · Granted May 20, 2025

Contextual video compression framework with spatial-temporal cross-covariance transformers

Inventors: Zhenghao Chen (Sydney, AU); Roberto Gerson De Albuquerque Azevedo (Zurich, CH); Christopher Richard Schroers (Uster, CH); Yang Zhang (Dubendorf, CH); Lucas Relic (Zurich, CH)
Assignees: Disney Enterprises, Inc.; ETH Zürich (Eidgenössische Technische Hochschule Zürich)
H04N19/42H04N19/172H04N19/91
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,309,404
App. No.
18/349,076
Granted
May 20, 2025
Kind
B2
Abstract

In some embodiments, a system includes a first component to extract temporal features from a current frame being coded and a previous frame of a video. A second component uses a first transformer to fuse spatial features from the current frame with the temporal features to generate spatio-temporal features as first output. A third component uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output. A fourth component uses a third transformer to reconstruct the current frame based on the first output that is processed using the second output and the temporal features.

Claims (67)

1. A system comprising:

a first component to extract temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame;

a second component that uses a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output;

a third component that uses a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and

a fourth component that uses a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate third output, and wherein the third transformer fuses the temporal features with the third output.

2. The system of claim 1 , wherein the three-dimensional based joint features are:

joint features for a first feature from a first input and a second feature from a second input are generated, wherein the joint features include a temporal channel.

3. The system of claim 2 , wherein at least one of the first transformer, the second transformer and the third transformer:

generates a matrix of attention weights based on the joint features; and

applies the attention weights to the joint features.

4. The system of claim 3 , wherein at least one of the first transformer, the second transformer and the third transformer:

generates a query feature, a key feature, and a value feature from the joint features; and

generates the matrix of attention weights based on the key feature and the value feature; and

applies the attention weights to the value feature.

5. The system of claim 4 , wherein the attention weights are generated by:

combining the query feature and the key feature to generate a combined output; and

applying a function to the combined output to generate the matrix for the attention weights.

6. The system of claim 5 , wherein at least one of the first transformer, the second transformer and the third transformer:

applies the attention weights to the value feature to generate a product feature.

7. The system of claim 6 , wherein at least one of the first transformer, the second transformer and the third transformer:

combines the product feature with the joint features to generate updated joint features.

8. The system of claim 7 , wherein at least one of the first transformer, the second transformer and the third transformer:

applies a gating mechanism on the updated joint features to filter information from the updated joint features to generate a gated joint spatio-temporal features.

9. The system of claim 8 , wherein at least one of the first transformer, the second transformer and the third transformer:

generates a two dimensional spatio-temporal feature from the gated joint spatio-temporal features.

10. The system of claim 1 , wherein the second component:

receives spatial information from the current frame;

combines the spatial information with a first scale of the temporal features to generate first joint spatio-temporal features;

combines, using the first transformer, the first joint spatio-temporal features with a second scale of the temporal features to generate second joint spatio-temporal features; and

combines, using the first transformer, the second joint spatio-temporal features with a third scale of the temporal features to generate third joint spatio-temporal features.

11. The system of claim 10 , wherein the third joint spatio-temporal features are used to generate an encoded bitstream of the current frame.

12. The system of claim 1 , wherein the third component:

receives a first scale of the temporal features;

receives the first output from the second component;

combines the first scale of the temporal features and the first output using the second transformer to generate a spatio-temporal prior; and

entropy codes the spatio-temporal prior to generate a distribution.

13. The system of claim 12 , wherein the distribution is used to encode the first output to an encoded bitstream and decode the encoded bitstream.

14. The system of claim 1 , wherein the fourth component:

receives the first output that is processed using the second output;

combines the first output that is processed using the second output with a first scale of the temporal features to generate first joint spatio-temporal features;

combines, using the third transformer, the first joint spatio-temporal features with a second scale of the temporal features to generate second joint spatio-temporal features; and

combines, using the third transformer, the second joint spatio-temporal features with a third scale of the temporal features to generate third joint spatio-temporal features, wherein the third joint spatio-temporal features are used to reconstruct the current frame.

15. A method comprising:

extracting temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame;

using a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output;

using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and

using a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate a third output, and wherein the third transformer fuses the temporal features with the third output.

16. The method of claim 15 , further comprising:

receiving a first input of a first feature and a second input of a second feature; and

generating joint feature for the first feature and the second feature, wherein the joint features include a temporal channel.

17. The method of claim 16 , wherein at least one of the first transformer, the second transformer and the third transformer:

generates a matrix of attention weights based on the joint features; and

applies the attention weights to the joint features.

18. The method of claim 17 , wherein at least one of the first transformer, the second transformer and the third transformer:

generates a query feature, a key feature, and a value feature from the joint spatio features; and

generates the matrix of attention weights based on the key feature and the value feature; and

applies the attention weights to the value feature.

19. The method of claim 18 , wherein the attention weights are generated by:

combining the query feature and the key feature to generate a combined output; and

applying a function to the combined output to generate the matrix for the attention weights.

20. An apparatus comprising:

one or more computer processors; and

a computer-readable storage medium comprising instructions for controlling the one or more computer processors to be operable for:

extracting temporal features from a current frame being coded and a previous frame of a video, wherein three-dimensional based joint features are determined using the temporal features and spatial features from the current frame;

using a first transformer to receive the three-dimensional based joint features as input and fuse the spatial features from the current frame with the temporal features to generate spatio-temporal features as first output;

using a second transformer to perform entropy coding using the first output and at least a portion of the temporal features to generate a second output, wherein the second transformer is used to fuse the spatio-temporal features with the at least a portion of the temporal features to output fused spatio-temporal features that are entropy encoded to generate the second output; and

using a third transformer to reconstruct the current frame, wherein the first output is processed using the second output to generate a third output and wherein the third transformer fuses the temporal features with the third output.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: DE ALBUQUERQUE AZEVEDO, ROBERTO GERSON; RELIC, LUCAS; CHEN, ZHENGHAO; ZHANG, YANG; SCHROERS, CHRISTOPHER RICHARD
To: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH; ETH ZÜRICH (EIDGENÖSSISCHE TECHNISCHE HOCHSCHULE ZÜRICH)
Reel/Frame 064202/0681 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: THE WALT DISNEY COMPANY (SWITZERLAND) GMBH
To: DISNEY ENTERPRISES, INC.
Reel/Frame 064202/0715 →
Continuity (2)
Provisional Application 63488944 · Mar 7, 2023
Related Publication 20240305801A1 · Sep 12, 2024
References Cited (55)
US 11310509B2 · Topiwala · 2022 [cited by examiner]
US 11582485B1 · Cherian · 2023 [cited by examiner]
US 20170223308A1 · Chen · 2017 [cited by examiner]
US 20220014807A1 · Lin · 2022 [cited by examiner]
US 20220078488A1 · Leleannec · 2022 [cited by examiner]
US 20220092645A1 · Pan · 2022 [cited by examiner]
US 20230090941A1 · Li · 2023 [cited by examiner]
US 20240054757A1 · Guo · 2024 [cited by examiner]
US 20240107088A1 · Kalva · 2024 [cited by examiner]
US 20240167852A1 · Kim · 2024 [cited by examiner]
Abdelaziz Djelouah, et al., “Neural Inter-Frame Compression for Video Coding,” Computer Vision Foundation, 9 pages, 2019. [cited by applicant]
Adam Paszke, “PyTorch: An Imperative Style, High-PerformanceDeep Learning Library,” 33rd Conference on Neural Information Processing Systems (NeurIPS 2019), 12 pages. [cited by applicant]
Alaaeldin El-Nouby, et al., “XCiT: Cross-Covariance Image Transformers,” arXiv:2106.09681v2 [cs.CV] 18 pages, Jun. 2021. [cited by applicant]
Amirhossein Habibian, et al., “Video Compression With Rate-Distortion Autoencoders,” Computer Vision Foundation, 10 pages, 2019. [cited by applicant]
An End-to-End Learning Framework for Video Compression. (n.d.). IEEE Xplore. https://ieeexplore.IEEE.org/document/9072487. [cited by applicant]
Anurag Ranjan, “Optical Flow Estimation using a Spatial Pyramid Network,” arXiv:1611.00850 [cs.CV], Submitted on Nov. 3, 2016, 10 pages. [cited by applicant]
Benjamin Bross, et al., “Overview of the Versatile Video Coding (VVC) Standard and Its Applications,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 10, 29 pages, Oct. 2021. [cited by applicant]
David Minnen, “Joint Autoregressive and Hierarchical Priors for Learned Image Compression,” 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), 10 pages. [cited by applicant]
David S. Taubman, et al., “JPEG2000: Standard for Interactive Imaging,” Roceedings of the IEEE, vol. 90, No. 8, Aug. 2002, 22 pages. [cited by applicant]
Eirikur Agustsson, et al., “Scale-space flow for end-to-end optimized video compression,” Computer Vision Foundation, 2020, 10 pages. [cited by applicant]
Fabian Mentzer, “Practical Full Resolution Learned Lossless Image Compression,” Computer Vision Foundation, 2018, 10 pages. [cited by applicant]
Fabian Mentzer, “VCT: A Video Compression Transformer,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 13 pages. [cited by applicant]
Gary J. Sullivan, “Overview of the High Efficiency Video Coding(HEVC) Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, 20 pages. [cited by applicant]
Gregory K. Wallace, “The JPEG Still Picture Compression Standard,” IEEE Transactions on Consumer Electronics, vol. 38, No. 1, Feb. 1992, 17 pages. [cited by applicant]
Guo Lu, “An End-to-End Learning Framework for Video Compression,” IEEE Transactions on Pattern Analysis and Machine Intelligence, 2020, 17 pages. [cited by applicant]
Guo Lu, “Content Adaptive and Error Propagation Aware Deep Video Compression,” arXiv:2003.11282 [eess.IV], [Submitted on Mar. 25, 2020], 17 pages. [cited by applicant]
Guo Lu, “DVC: An End-to-end Deep Video Compression Framework,” Computer Vision Foundation, 10 pages, 2018. [cited by applicant]
Haiqiang Wang, “MCL-JCV: a JND-Based H.264/AVC Video Quality Assessment Dataset,” IEEE International Conference on Image Processing (ICIP), 2016, 5 pages. [cited by applicant]
https://bellard.org/bpg/, printed from website on on Jul. 7, 2023, 2 pages. [cited by applicant]
https://hevc.hhi.fraunhofer.de/HM-doc/, printed from website on on Jul. 6, 2023, 2 pages. [cited by applicant]
https://jvet.hhi.fraunhofer.de/, printed from website on Jul. 6, 2023, 3 pages. [cited by applicant]
https://pytorch.org/docs/stable/generated/torch.nn.PixelShuffle.html, printed from website on Jul. 7, 2023, 2 pages. [cited by applicant]
https://ultravideo.fi/, printed from website on on Jul. 7, 2023, 3 pages. [cited by applicant]
Jiahao Li, “Deep Contextual Video Compression,” Microsoft Research Asia, 35th Conference on Neural Information Processing Systems (NeurIPS 2021), 12 pages. [cited by applicant]
Jiahao Li, “Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression,” rXiv:2207.05894v1 [eess.IV] Jul. 13, 2022, 17 pages. [cited by applicant]
Jianping Lin, “M-LVC: Multiple Frames Prediction for Learned Video Compression,” Computer Vision Foundation, 2020, 9 pages. [cited by applicant]
Jingyun Liang, et al., “Recurrent Video Restoration Transformer with Guided Deformable Attention,” 36th Conference on Neural Information Processing Systems (NeurIPS 2022), 16 pages. [cited by applicant]
Jingyun Liang, et al., “SwinIR: Image Restoration Using Swin Transformer,” Computer Vision Foundation, 12 pages, 2021. [cited by applicant]
Jingyun Liang, et al., “VRT: A Video Restoration Transformer,” arXiv:2201.12288v2 [cs.CV] Jun. 15, 2022, 14 pages. [cited by applicant]
Johannes Ballé, et al., “Variational Image Compressionwith a Scale Hyperprior,” Xiv:1802.01436v2 [eess.IV] 1, 23 pages, May 2018. [cited by applicant]
Jun Han, et al., Deep Generative Video Compression, “Deep generative video compression,” In Advances in Neural Information Processing Systems, 9287-9298 (2019). [cited by applicant]
Jun Han, et al., “Deep Probabilistic Video Compression,” arXiv preprint arXiv: 1810.02845 (2018), 15 pages. [cited by applicant]
Ming Lu, “Transformer-based Image Compression,” 2022 Data Compression Conference (DCC), 1 page. [cited by applicant]
Mingyang Song, “TempFormer: Temporally Consistent Transformer for Video Denoising,” ECCV 2022: Computer Vision—ECCV 2022, 16 pages. [cited by applicant]
Ren Yang, “Learning for Video Compression with Hierarchical Quality and Recurrent Enhancement,” Computer Vision Foundation, 2020, 10 pages. [cited by applicant]
Syed Waqas Zamir, “Restormer: Efficient Transformer for High-Resolution Image Restoration,” Computer Vision Foundation, arXiv:2111.09881 [cs.CV], Submitted on Nov. 18, 2021, 12 pages. [cited by applicant]
Tianfan Xue, “Video Enhancement with Task-Oriented Flow,” International Journal of Computer Vision (IJCV), 127 (8):1106-1125, 2019. [cited by applicant]
Xihua Sheng, “Temporal Context Mining for Learned VideoCompression,” rXiv:2111.13850v2 [cs.CV] Jan. 30, 2023, 13 pages. [cited by applicant]
Yichen Qian, “Entroformer: a Transformer-Based Entropy Model for Learned Image Compression,” Published as a conference paper at ICLR 2022, 15 pages. [cited by applicant]
Yinhao Zhu, et al., “Transformer-based Transform Coding,” Published as a conference paper at ICLR 2022, 35 pages. [cited by applicant]
Ze Liu, “Swin Transformer: Hierarchical Vision Transformer using Shifted Windows,” Computer Vision Foundation, 2021, 11 pages. [cited by applicant]
Zhengxue Cheng, et al., “Learned Image Compression with Discretized Gaussian Mixture Likelihoods and Attention Modules,” Computer Vision Foundation, 10 pages, 2020. [cited by applicant]
Zhihao Hu, “FVC: A New Framework towards Deep Video Compression in Feature Space,” Computer Vision Foundation, 10 pages, 2021. [cited by applicant]
Zhihao Hu, et al., “Coarse-to-fine Deep Video Coding with Hyperprior-guided Mode Prediction,” Computer Vision Foundation, 10 pages, 2022. [cited by applicant]
Zhihao Hu, et al., Improving Deep Video Compression by Resolution-adaptive Flow Coding, arXiv:2009.05982 [cs. CV], 16 pages, submitted Sep. 13, 2020. [cited by applicant]