IP Library Granted Patent US 12,348,743
Granted Patent B2
US 12,348,743 · App. 18/338,731 · Granted Jul 1, 2025

Method and system for learned video compression

Inventors: Sam Tak Wu Kwong (Kowloon, HK); Haifeng Guo (Kowloon, HK); Shiqi Wang (Kowloon, HK); Dongjie Ye (Kowloon, HK)
Assignee: City University of Hong Kong
H04N19/42H04N19/12H04N19/139H04N19/172H04N19/176H04N19/60H04N19/85
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,348,743
App. No.
18/338,731
Granted
Jul 1, 2025
Kind
B2
Abstract

There is provided a computer-implemented method for learned video compression, which includes processing a current frame (x t ) and previously decoded frame ({circumflex over (x)} t-1 ) of a video data using a motion estimation model to estimate a motion vector (v t ) for every pixel, compressing the motion vector (v t ) and reconstructing the motion vector (v t ) to a reconstructed motion vector ({circumflex over (v)} t ), applying an enhanced context mining (ECM) model to obtain enhanced context ({umlaut over (C)} E ) from the reconstructed motion vector ({circumflex over (v)} t ) and previously decoded frame feature (x̆ t-1 ), compressing the current frame (x t ) with the assistance of the enhanced context ({umlaut over (C)} E ) to obtain a reconstructed frame ({circumflex over (x)}′ t ), and providing the reconstructed frame ({circumflex over (x)}′ t ) to a post-enhancement backend network to obtain a high-resolution frame ({circumflex over (x)} t ).

Claims (49)

1. A computer-implemented method for learned video compression, comprising:

processing a current frame (x t ) and previously decoded frame ({circumflex over (x)} t-1 ) of a video data using a motion estimation model to estimate a motion vector (v t ) for every pixel;

compressing the motion vectors (v t ) and reconstructing the motion vectors (v t ) to reconstructed motion vectors ({circumflex over (v)} t );

applying an enhanced context mining (ECM) model to obtain enhanced context ({umlaut over (C)} E ) from the reconstructed motion vectors ({circumflex over (v)} t ) and previously decoded frame feature (x̆ t-1 ); wherein applying the ECM comprises:

obtaining the motion vectors ({circumflex over (v)} t ) and decoded frame feature (x̆ t-1 ) based on the current input frame (x t ) and previously decoded frame ({circumflex over (x)} t-1 );

warping the motion vectors ({circumflex over (v)} t ) and decoded frame feature (x̆ t-1 ) to obtain a warped feature ({umlaut over (x)} t ); and

processing the warped feature ({umlaut over (x)} t ) using a resblock and convolution layer to obtain a context ( t );

compressing the current frame (x t ) with the assistance of the enhanced context ({umlaut over (C)} E ) to obtain a reconstructed frame ({circumflex over (x)}′ t ); and

providing the reconstructed frame ({circumflex over (x)}′ t ) to a post-enhancement backend network to obtain a high-resolution frame ({circumflex over (x)} t ).

2. The computer-implemented method of claim 1 , wherein the motion estimation model is based on a spatial pyramid network.

3. The computer-implemented method of claim 1 , wherein applying the enhanced context mining (ECM) model comprises utilizing cross-channel interaction and residual learning operation to reduce redundancy across context channels.

4. The computer-implemented method of claim 1 , wherein applying the enhanced context mining (ECM) model further comprises:

incorporating convolution with ReLU on the context ( t ) to obtain residual context (R C ); and

adding the residual context (R C ) to the context ( t ) to obtain the enhanced context ({umlaut over (C)} E ).

5. The computer-implemented method of claim 4 , wherein incorporating convolution with ReLU on the context ( t ) comprises passing the context ( t ) through multiple layers of convolution and ReLU followed by one more convolution layer.

6. The computer-implemented method of claim 1 , wherein the enhanced context mining (ECM) model is designed without batch normalization layers.

7. A computer-implemented method for learned video compression, comprising:

processing a current frame (x t ) and previously decoded frame ({circumflex over (x)} t-1 ) of a video data using a motion estimation model to estimate a motion vector (v t ) for every pixel;

compressing the motion vectors (v t ) and reconstructing the motion vectors (v t ) to reconstructed motion vectors ({circumflex over (v)} t );

applying an enhanced context mining (ECM) model to obtain enhanced context ({umlaut over (C)} E ) from the reconstructed motion vectors ({circumflex over (v)} t ) and previously decoded frame feature (x̆ t-1 );

compressing the current frame (x t ) with the assistance of the enhanced context ({umlaut over (C)} E ) to obtain a reconstructed frame ({circumflex over (x)}′ t ); wherein compressing the current frame (x t ) comprises:

concatenating the input frame (x t ) and the enhanced context ({umlaut over (C)} E ) together;

processing the input frame (x t ) and the enhanced context ({umlaut over (C)} E ) to obtain latent code (y t ) for entropy model; and

transforming the latent code back to pixel space with the assistance of the enhanced context ({umlaut over (C)} E ) to obtain the reconstructed frame ({circumflex over (x)}′ t ); and

providing the reconstructed frame ({circumflex over (x)}′ t ) to a post-enhancement backend network to obtain a high-resolution frame ({circumflex over (x)} t ).

8. The computer-implemented method of claim 1 , wherein the post-enhancement backend network is transformer-based.

9. The computer-implemented method of claim 1 , wherein the post-enhancement backend network comprises multiple transposed gated transformer blocks (TGTBs) and multiple convolution layers.

10. A computer-implemented method for learned video compression, comprising:

processing a current frame (x t ) and previously decoded frame ({circumflex over (x)} t-1 ) of a video data using a motion estimation model to estimate a motion vector (v t ) for every pixel;

compressing the motion vectors (v t ) and reconstructing the motion vectors (v t ) to reconstructed motion vectors ({circumflex over (v)} t );

applying an enhanced context mining (ECM) model to obtain enhanced context ({umlaut over (C)} E ) from the reconstructed motion vectors ({circumflex over (v)} t ) and previously decoded frame feature (x̆ t-1 );

compressing the current frame (x t ) with the assistance of the enhanced context ({umlaut over (C)} E ) to obtain a reconstructed frame ({circumflex over (x)}′ t ); and

providing the reconstructed frame ({circumflex over (x)}′ t ) to a post-enhancement backend network to obtain a high-resolution frame ({circumflex over (x)} t ); wherein providing the reconstructed frame ({circumflex over (x)}′ t ) to the post-enhancement backend network comprises:

applying a convolution layer to the reconstructed frame ({circumflex over (x)}′ t ) to obtain a low-level feature embedding F 0 ∈ H×W×C , where H×W is the spatial height and width, and C denotes the channel numbers;

processing the low-level feature (F 0 ) using one or more transformer blocks to obtain a refined feature (F R ); and

applying a convolution layer to the refined feature (F R ) to obtain residual image R∈ H×W×3 to which the reconstructed frame ({circumflex over (x)}′ t ) is added according to the following equation to obtain {circumflex over (x)} t ={circumflex over (x)}′ t +R, where {circumflex over (x)} t is a high-resolution frame.

11. The computer-implemented method of claim 10 , wherein at least one of the transformer blocks comprises a transposed gated transformer block (TGTB) which is designed without layer normalization.

12. The computer-implemented method of claim 10 , wherein at least one of the transformer blocks is modified to contain a multi-head transposed attention (MTA) and a gated feed-forward network (GFN).

13. The computer-implemented method of claim 12 , wherein the multi-head transposed attention (MTA) comprises calculating self-attention across channels.

14. The computer-implemented method of claim 12 , wherein the multi-head transposed attention (MTA) comprises applying depth-wise convolution.

15. The computer-implemented method of claim 12 , wherein the multi-head transposed attention (MTA) generates, from a feature input X∈ H×W×C , query (Q), key (K) and value (V) projections with the local context, and reshapes the query (Q) to ĤŴ×Ĉ, and key (K) to Ĉ×ĤŴ, to obtain a transposed attention map of size Ĉ×Ĉ .

16. The computer-implemented method of claim 12 ,

wherein the gated feed-forward network (GFN) comprises gating mechanism and depth wise convolutions,

wherein the gating mechanism is achieved as the element-wise product of two parallel paths of transformation layers, one of which is activated with the GELU non-linearity, and

wherein the depth-wise convolution is applied to obtain information from spatially neighboring pixel positions.

17. A system for learned video compression, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for performing or facilitating performing of the computer-implemented method of claim 1 .

18. A non-transitory computer readable medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to execute the computer-implemented method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 21, 2023
From: KWONG, SAM TAK WU; GUO, HAIFENG; WANG, SHIQI; YE, DONGJIE
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 064015/0993 →
Continuity (1)
Related Publication 20240430463A1 · Dec 26, 2024
References Cited (88)
US 10623775B1 · Theis et al. · 2020 [cited by applicant]
US 10977553B2 · Rippel et al. · 2021 [cited by applicant]
US 20230043350A1 · Choe · 2023 [cited by examiner]
US 20240146934A1 · Kwong · 2024 [cited by examiner]
US 20240364878A1 · Yin · 2024 [cited by examiner]
AU 2023272445A1 · 2024 [cited by examiner]
CN 111405283A · 2020 [cited by applicant]
CN 115034959A · 2022 [cited by examiner]
CN 118865048A · 2024 [cited by examiner]
R. Yang, F. Mentzer, L. Van Gool, and R. Timofte, “Learning for video compression with recurrent auto-encoder and recurrent probability model,” IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 3… [cited by applicant]
R. Yang, F. Mentzer, L. V. Gool, and R. Timofte, “Learning for video compression with hierarchical quality and recurrent enhancement,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 20… [cited by applicant]
R. Pourreza and T. Cohen, “Extending neural p-frame codecs for b-frame coding,” in Proceedings of the IEEE International Conference on Computer Vision, 2021, pp. 6680-6689. [cited by applicant]
V. Sze, M. Budagavi, and G. J. Sullivan, “High efficiency video coding (HEVC),” in Integrated Circuit and Systems, Algorithms and Architectures. Springer, 2014, vol. 39, p. 40. [cited by applicant]
Y. Dai, D. Liu, and F. Wu, “A convolutional neural network approach for post-processing in HEVC intra coding,” in MultiMedia Modeling International Conference, 2017, pp. 28-39. [cited by applicant]
Y. Zhang, T. Shen, X. Ji, Y. Zhang, R. Xiong, and Q. Dai, “Residual highway convolutional neural networks for in-loop filtering in HEVC,” IEEE Transactions on Image Processing, vol. 27, No. 8, pp. 3827-3841, 2018. [cited by applicant]
D. Ding, L. Kong, G. Chen, Z. Liu, and Y. Fang, “A switchable deep learning approach for in-loop filtering in video coding,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 7, pp. 1871-1887,… [cited by applicant]
C. Jia, S. Wang, X. Zhang, S. Wang, J. Liu, S. Pu, and S. Ma, “Content-aware convolutional neural network for in-loop filtering in high efficiency video coding,” IEEE Transactions on Image Processing, vol. 28, No. 7, pp… [cited by applicant]
Z. Pan, X. Yi, Y. Zhang, B. Jeon, and S. Kwong, “Efficient in-loop filtering based on enhanced deep convolutional neural networks for HEVC,” IEEE Transactions on Image Processing, vol. 29, pp. 5352-5366, 2020. [cited by applicant]
A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4161-4170. [cited by applicant]
Y. Tian, Y. Zhang, Y. Fu, and C. Xu, “TDAN: Temporally Deformable Alignment Network for Video Super-Resolution,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3357-3366. [cited by applicant]
Y. Wu and K. He, “Group normalization,” in Proceedings of the European Conference on Computer Vision, 2018, pp. 3-19. [cited by applicant]
J. Xu, X. Sun, Z. Zhang, G. Zhao, and J. Lin, “Understanding and improving layer normalization,” Proceedings of Advances in Neural Information Processing Systems, vol. 32, 2019. [cited by applicant]
P. Luo, R. Zhang, J. Ren, Z. Peng, and J. Li, “Switchable normalization for learning-to-normalize deep representation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, No. 2, pp. 712-728, 2019. [cited by applicant]
D. Hendrycks and K. Gimpel, “Gaussian error linear units,” arXiv preprint arXiv:1606.08415, 2016. [cited by applicant]
Z. Wang, E. P. Simoncelli, and A. C. Bovik, “Multiscale structural similarity for image quality assessment,” in Asilomar Conference on Signals, Systems and Computers, vol. 2, 2003, pp. 1398-1402. [cited by applicant]
J. Ballé, V. Laparra, and E. P. Simoncelli, “End-to-end optimization of nonlinear transform codes for perceptual quality,” in Picture Coding Symposium, 2016, pp. 1-5. [cited by applicant]
T. Xue, B. Chen, J. Wu, D. Wei, and W. T. Freeman, “Video enhancement with task-oriented flow,” International Journal of Computer Vision, vol. 127, No. 8, pp. 1106-1125, 2019. [cited by applicant]
F. Bossen et al., “Common test conditions and software reference configurations,” JCTVC-L1100, vol. 12, No. 7, 2013. [cited by applicant]
A. Mercat, M. Viitanen, and J. Vanne, “UVG dataset: 50/120fps 4k sequences for video codec analysis and development,” in Proceedings of the ACM Multimedia Systems Conference, 2020, pp. 297-302. [cited by applicant]
H. Wang, W. Gan, S. Hu, J. Y. Lin, L. Jin, L. Song, P. Wang, I. Katsavounidis, A. Aaron, and C.-C. J. Kuo, “MCL-JCV: A JND-based H. 264/AVC video quality assessment dataset,” in Proceedings of International Conference o… [cited by applicant]
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “PyTorch: An imperative style, high-performance deep learning library,” Proceedings of Advances in Ne… [cited by applicant]
J. Bégaint, F. Racapé, S. Feltman, and A. Pushparaja, “CompressAI: a PyTorch library and evaluation platform for end-to-end compression research,” arXiv preprint arXiv:2011.03029, 2020. [cited by applicant]
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” arXiv preprint arXiv:1711.05101, 2017. [cited by applicant]
G. Bjontegaard, “Calculation of average PSNR differences between RD-curves (VCEG-M33),” in VCEG Meeting (ITU-T SG16 Q. 6), 2001, pp. 2-4. [cited by applicant]
H. Lee, H. Choi, K. Sohn, and D. Min, “Knn local attention for image restoration,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2022, pp. 2139-2149. [cited by applicant]
K. Zhang, W. Zuo, Y. Chen, D. Meng, and L. Zhang, “Beyond a Gaussian denoiser: Residual learning of deep CNN for image denoising,” IEEE transactions on image processing, vol. 26, No. 7, pp. 3142-3155, 2017. [cited by applicant]
L. Torrey and J. Shavlik, “Transfer learning,” in Handbook of research on machine learning applications and trends: algorithms, methods, and techniques. IGI global, 2010, pp. 242-264. [cited by applicant]
G. Lu, C. Cai, X. Zhang, L. Chen, W. Ouyang, D. Xu, and Z. Gao, “Content adaptive and error propagation aware deep video compression,” 2020. [cited by applicant]
T. Wiegand, G. J. Sullivan, G. Bjontegaard, and A. Luthra, “Overview of the H. 264/AVC video coding standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, pp. 560-576, 2003. [cited by applicant]
G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the high efficiency video coding (HEVC) standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668, 2012. [cited by applicant]
B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.-R. Ohm, “Overview of the versatile video coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, … [cited by applicant]
S. Ma, X. Zhang, C. Jia, Z. Zhao, S. Wang, and S. Wang, “Image and video compression with neural networks: A review,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 6, pp. 1683-1698, 2019. [cited by applicant]
K. He, X. Zhang, S. Ren, and J. Sun, “Deep residual learning for image recognition,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2016, pp. 770-778. [cited by applicant]
A. Dosovitskiy, L. Beyer, A. Kolesnikov, D. Weissenborn, X. Zhai, T. Unterthiner, M. Dehghani, M. Minderer, G. Heigold, S. Gelly et al., “An image is worth 16x16 words: Transformers for image recognition at scale,” in I… [cited by applicant]
Z. Liu, Y. Lin, Y. Cao, H. Hu, Y. Wei, Z. Zhang, S. Lin, and B. Guo, “Swin Transformer: Hierarchical vision transformer using shifted windows,” in Proceedings of the IEEE International Conference on Computer Vision, 202… [cited by applicant]
D. Liu, Y. Li, J. Lin, H. Li, and F. Wu, “Deep learning-based video coding: A review and a case study,” ACM Computing Surveys, vol. 53, No. 1, pp. 1-35, 2020. [cited by applicant]
D. Ding, Z. Ma, D. Chen, Q. Chen, Z. Liu, and F. Zhu, “Advances in video compression system using deep neural network: A review and case studies,” Proceedings of the IEEE, vol. 109, No. 9, pp. 1494-1520, 2021. [cited by applicant]
Y. Zhang, S. Kwong, and S. Wang, “Machine learning based video coding optimizations: A survey,” Information Sciences, vol. 506, pp. 395-423, 2020. [cited by applicant]
J. Ballé, V. Laparra, and E. P. Simoncelli, “End-to-end optimized image compression,” in International Conference on Learning Representations, 2017. [cited by applicant]
Z. Cheng, H. Sun, M. Takeuchi, and J. Katto, “Learned image compression with discretized Gaussian mixture likelihoods and attention modules,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognit… [cited by applicant]
G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao, “DVC: An end-to-end deep video compression framework,” in Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 11006-11015. [cited by applicant]
T. Ladune, P. Philippe, W. Hamidouche, L. Zhang, and O. Déforges, “Optical flow and mode selection for learning-based video coding,” in IEEE International Workshop on Multimedia Signal Processing, 2020, pp. 1-6. [cited by applicant]
J. Li, B. Li, and Y. Lu, “Deep contextual video compression,” Proceedings of Advances in Neural Information Processing Systems, vol. 34, pp. 18114-18125, 2021. [cited by applicant]
C. D. Pham, C. Fu, and J. Zhou, “Deep learning based spatial-temporal in-loop filtering for versatile video coding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 1861-1865. [cited by applicant]
F. Zhang, C. Feng, and D. R. Bull, “Enhancing VVC through CNN-based post-processing,” in Proceedings of International Conference on Multimedia and Expo, 2020, pp. 1-6. [cited by applicant]
D. Ma, F. Zhang, and D. R. Bull, “MFRNet: A new CNN architecture for post-processing and in-loop filtering,” IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 378-387, 2020. [cited by applicant]
W. Fedus, B. Zoph, and N. Shazeer, “Switch Transformers: Scaling to Trillion Parameter Models with Simple and Efficient Sparsity,” The Journal of Machine Learning Research, vol. 23, No. 1, pp. 5232-5270, 2022. [cited by applicant]
Y. Liu, M. Ott, N. Goyal, J. Du, M. Joshi, D. Chen, O. Levy, M. Lewis, L. Zettlemoyer, and V. Stoyanov, “RoBERTa: A Robustly Optimized BERT Pretraining Approach,” arXiv preprint arXiv:1907.11692, 2019. [cited by applicant]
A. Vaswani, N. Shazeer, N. Parmar, J. Uszkoreit, L. Jones, A. N. Gomez, Ł. Kaiser, and I. Polosukhin, “Attention is all you need,” Proceedings of Advances in Neural Information Processing Systems, vol. 30, 2017. [cited by applicant]
Y. Mei, Y. Fan, Y. Zhang, J. Yu, Y. Zhou, D. Liu, Y. Fu, T. S. Huang, and H. Shi, “Pyramid attention networks for image restoration,” arXiv preprint arXiv:2004.13824, 2020. [cited by applicant]
S. W. Zamir, A. Arora, S. Khan, M. Hayat, F. S. Khan, and M.-H. Yang, “Restormer: Efficient transformer for high-resolution image restoration,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recogn… [cited by applicant]
G. K. Wallace, “The JPEG still picture compression standard,” IEEE Transactions on Consumer Electronics, vol. 38, No. 1, 1992. [cited by applicant]
C. Christopoulos, A. Skodras, and T. Ebrahimi, “The jpeg2000 still image coding system: an overview,” IEEE Transactions on Consumer Electronics, vol. 46, No. 4, pp. 1103-1127, 2000. [cited by applicant]
F. Bellard, “BPG image format,” 2015, [online available] at https://bellard.org/bpg/. [cited by applicant]
J. Ballé, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston, “Variational image compression with a scale hyperprior,” arXiv preprint arXiv:1802.01436, 2018. [cited by applicant]
D. Minnen, J. Ballé, and G. D. Toderici, “Joint autoregressive and hierarchical priors for learned image compression,” Proceedings of Advances in Neural Information Processing Systems, vol. 31, 2018. [cited by applicant]
J. Lee, S. Cho, and S.-K. Beack, “Context-adaptive entropy model for end-to-end optimized image compression,” arXiv preprint arXiv:1809.10452, 2018. [cited by applicant]
F. Mentzer, E. Agustsson, M. Tschannen, R. Timofte, and L. Van Gool, “Conditional probability models for deep image compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, p… [cited by applicant]
Y. Hu, W. Yang, and J. Liu, “Coarse-to-fine hyper-prior modeling for learned image compression,” in Proceedings of the AAAI Conference on Artificial Intelligence, vol. 34, No. 07, 2020, pp. 11 013-11 020. [cited by applicant]
H. Liu, Y. Zhang, H. Zhang, C. Fan, S. Kwong, C.-C. J. Kuo, and X. Fan, “Deep learning-based picture-wise just noticeable distortion prediction model for image compression,” IEEE Transactions on Image Processing, vol. 2… [cited by applicant]
G. Toderici, S. M. O'Malley, S. J. Hwang, D. Vincent, D. Minnen, S. Baluja, M. Covell, and R. Sukthankar, “Variable rate image compression with recurrent neural networks,” arXiv preprint arXiv:1511.06085, 2015. [cited by applicant]
G. Toderici, D. Vincent, N. Johnston, S. Jin Hwang, D. Minnen, J. Shor, and M. Covell, “Full resolution image compression with recurrent neural networks,” in Proceedings of the IEEE Conference on Computer Vision and Pat… [cited by applicant]
N. Johnston, D. Vincent, D. Minnen, M. Covell, S. Singh, T. Chinen, S. J. Hwang, J. Shor, and G. Toderici, “Improved lossy image compression with priming and spatially adaptive bit rates for recurrent networks,” in Proc… [cited by applicant]
E. Agustsson, M. Tschannen, F. Mentzer, R. Timofte, and L. V. Gool, “Generative adversarial networks for extreme learned image compression,” in Proceedings of the IEEE International Conference on Computer Vision, 2019, … [cited by applicant]
F. Mentzer, G. D. Toderici, M. Tschannen, and E. Agustsson, “High-fidelity generative image compression,” Proceedings of Advances in Neural Information Processing Systems, vol. 33, pp. 11 913-11 924, 2020. [cited by applicant]
T. Zhao, Y. Huang, W. Feng, Y. Xu, and S. Kwong, “Efficient VVC intra prediction based on deep feature fusion and probability estimation,” IEEE Transactions on Multimedia, 2022. [cited by applicant]
L. Zhu, Y. Zhang, N. Li, G. Jiang, and S. Kwong, “Deep learning-based intra mode derivation for versatile video coding,” ACM Transactions on Multimedia Computing, Communications and Applications, vol. 19, No. 2s, pp. 1-… [cited by applicant]
G. Lu, X. Zhang, W. Ouyang, L. Chen, Z. Gao, and D. Xu, “An end-to-end learning framework for video compression,” IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 43, No. 10, pp. 3292-3308, 2020. [cited by applicant]
Z. Hu, Z. Chen, D. Xu, G. Lu, W. Ouyang, and S. Gu, “Improving deep video compression by resolution-adaptive flow coding,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 193-209. [cited by applicant]
Z. Hu, G. Lu, and D. Xu, “FVC: A new framework towards deep video compression in feature space,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 1502-1511. [cited by applicant]
Z. Hu, G. Lu, J. Guo, S. Liu, W. Jiang, and D. Xu, “Coarse-To-Fine Deep Video Coding With Hyperprior-Guided Mode Prediction,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022, pp. 5… [cited by applicant]
E. Agustsson, D. Minnen, N. Johnston, J. Ballé, S. J. Hwang, and G. Toderici, “Scale-space flow for end-to-end optimized video compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recogniti… [cited by applicant]
K. Lin, C. Jia, X. Zhang, S. Wang, S. Ma, and W. Gao, “DMVC: Decomposed motion modeling for learned video compression,” IEEE Transactions on Circuits and Systems for Video Technology, 2022. [cited by applicant]
H. Guo, S. Kwong, C. Jia, and S. Wang, “Enhanced motion compensation for deep video compression,” IEEE Signal Processing Letters, 2023 (submitted). [cited by applicant]
F. Mentzer, G. Toderici, D. Minnen, S.-J. Hwang, S. Caelles, M. Lucic, and E. Agustsson, “VCT: A video compression transformer,” arXiv preprint arXiv:2206.07307, 2022. [cited by applicant]
X. Sheng, J. Li, B. Li, L. Li, D. Liu, and Y. Lu, “Temporal context mining for learned video compression,” IEEE Transactions on Multimedia, 2022. [cited by applicant]
J. Li, B. Li, and Y. Lu, “Hybrid spatial-temporal entropy modelling for neural video compression,” in Proceedings of the ACM International Conference on Multimedia, 2022, pp. 1503-1511. [cited by applicant]
J. Lin, D. Liu, H. Li, and F. Wu, “M-LVC: Multiple frames prediction for learned video compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3546-3554. [cited by applicant]