IP Library Granted Patent US 12,610,089
Granted Patent B2
US 12,610,089 · App. 18/624,518 · Granted Apr 21, 2026

Method and system for learning-based bidirectional video compression

Inventors: Sam Tak Wu Kwong (Kowloon, HK); Haifeng Guo (Kowloon, HK); Shiqi Wang (Kowloon, HK)
Assignee: City University of Hong Kong
H04N19/91G06V10/44G06V10/761H04N19/172H04N19/51
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,610,089
App. No.
18/624,518
Granted
Apr 21, 2026
Kind
B2
Abstract

A computer-implemented method for learning-based bidirectional video compression includes, given a current frame, generating a single reference frame from the current frame and bidirectional frames by using a neural network, estimating a motion between the current frame and the reference frame, obtaining a reconstructed motion by inputting the motion to a motion encoder and decoder, generating a set of temporal contexts based on the reconstructed motion and propagated feature, and compressing the current frame based on the temporal contexts by an inverse channel-wise entropy model which is adapted to reconstruct the relationship among channels such that the channels with less entropy are coded first and the channels with larger entropy are coded with the help of the previously coded channels.

Claims (40)

1 . A computer-implemented method for learning-based bidirectional video compression, comprising steps of:

given a current frame, generating a single reference frame from the current frame and bidirectional frames by using a neural network;

estimating a motion between the current frame and the reference frame;

obtaining a reconstructed motion by inputting the motion to a motion encoder and decoder;

generating a set of temporal contexts based on the reconstructed motion and a propagated feature obtained from coding for a previous frame and stored in a buffer; and

compressing the current frame based on the temporal contexts by an inverse channel-wise entropy model which is adapted to divide latent representation of the current frame to be coded into a plurality of groups (N) along with channels and to reconstruct a relationship among the channels such that the channels with less entropy are coded first and the channels with larger entropy are coded with the help of the previously coded channels.

2 . The computer-implemented method of claim 1 , wherein the step of generating the reference frame comprises:

extracting features from the current frame and the bidirectional frames independently by using the neural network, the bidirectional frames including a previous reconstructed frame and a following reconstructed frame;

obtaining distance values that represent feature similarity between the current frame and the bidirectional frames;

fusing the distance values along with the bidirectional frames to generate the reference frame.

3 . The computer-implemented method of claim 2 , the step of obtaining distance values further comprises conducting normalization subtraction and weighted averaging on the features.

4 . The computer-implemented method of claim 3 , wherein the step of extracting features comprises extracting L layers feature stacks from the current frame, the previous reconstructed frame and the following reconstructed frame.

5 . The computer-implemented method of claim 4 , wherein the step of obtaining distance values comprises, for each distance value:

normalizing the L layers feature stacks from the current frame and the previous or following reconstructed frame, and subtracting the feature stacks in a channel dimension;

scaling each channel and calculating an L2 norm distance; and

averaging across spatial dimensions and all layers.

6 . The computer-implemented method of claim 2 , wherein the distance values comprise a first distance value that represents feature similarity between the current frame and the previous reconstructed frame, and a second distance value that represents feature similarity between the current frame and the following reconstructed frame.

7 . The computer-implemented method of claim 1 , wherein the neural network comprises a Siamese neural network.

8 . The computer-implemented method of claim 1 , wherein the inverse channel-wise entropy model is adapted to predict beginning channel parameters with the help of deep channels such that larger entropy channels have more inputs.

9 . The computer-implemented method of claim 1 , wherein the inverse channel-wise entropy model is adapted to inverse a predicted direction, which predicts the beginning channel parameters with the help of deep channels along with the direction of information accumulation.

10 . The computer-implemented method of claim 1 , wherein the plurality of groups (N) comprise at least a first group (1 st ) of beginning channels and a last group (N th ) of last channels, and the inverse channel wise entropy model is adapted to encode the last group first and take the last group parameters as an input of the next group (N−1 th ) coding.

11 . The computer-implemented method of claim 10 , wherein the element to be coded is divided into the plurality of groups unevenly across the channels.

12 . The computer-implemented method of claim 10 , wherein the first group has the most significant entropy compared with the rest groups, and entropy distribution is reduced from the first group to the last group.

13 . The computer-implemented method of claim 12 , wherein the first group of beginning channels have the most inputs.

14 . The computer-implemented method of claim 12 , wherein the inverse channel-wise entropy model is adapted to encode the first group based on all previous groups.

15 . The computer-implemented method of claim 1 , wherein the inverse channel-wise entropy model is further adapted to obtain reconstructed groups corresponding to the plurality of groups and to merge the reconstructed groups to obtain a reconstructed element.

16 . A system for learning-based bidirectional video compression, comprising:

one or more processors; and

a memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

given a current frame, generating a single reference frame from the current frame and bidirectional frames by using a neural network;

estimating a motion between the current frame and the reference frame;

obtaining a reconstructed motion by inputting the motion to a motion encoder and decoder;

generating a set of temporal contexts based on the reconstructed motion and a propagated feature obtained from coding for a previous frame and stored in a buffer; and

compressing the current frame based on the temporal contexts by an inverse channel-wise entropy model which is adapted to divide latent representation of the current frame to be coded into a plurality of groups (N) along with channels and to reconstruct a relationship among the channels such that the channels with less entropy are coded first and the channels with larger entropy are coded with the help of the previously coded channels.

17 . A non-transitory computer readable medium having instructions stored thereon which, when executed by one or more processors, cause the one or more processors to execute a computer-implemented method for learning-based bidirectional video compression, comprising steps of:

given a current frame, generating a single reference frame from the current frame and bidirectional frames by using a neural network;

estimating a motion between the current frame and the reference frame;

obtaining a reconstructed motion by inputting the motion to a motion encoder and decoder;

generating a set of temporal contexts based on the reconstructed motion and a propagated feature obtained from coding for a previous frame and stored in a buffer; and

compressing the current frame based on the temporal contexts by an inverse channel-wise entropy model which is adapted to divide latent representation of the current frame to be coded into a plurality of groups (N) along with channels and to reconstruct a relationship among the channels such that the channels with less entropy are coded first and the channels with larger entropy are coded with the help of the previously coded channels.

Assignments (1)
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER AND THE TITLE. THE TITLE IS CORRECT. PREVIOUSLY RECORDED AT REEL: 067003 FRAME: 489. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 5, 2024
From: KWONG, SAM TAK WU; GUO, HAIFENG; WANG, SHIQI
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 067025/0674 →
Continuity (1)
Related Publication 20250310574A1 · Oct 2, 2025
References Cited (57)
US 10623775B1 · Theis et al. · 2020 [cited by applicant]
US 10977553B2 · Rippel et al. · 2021 [cited by applicant]
US 20070016405A1 · Mehrotra · 2007 [cited by examiner]
US 20070016412A1 · Mehrotra · 2007 [cited by examiner]
US 20070016414A1 · Mehrotra · 2007 [cited by examiner]
US 20210057084A1 · Takeshima · 2021 [cited by examiner]
US 20230421754A1 · Zhao · 2023 [cited by examiner]
US 20240071039A1 · Eswara · 2024 [cited by examiner]
US 20240129487A1 · Wang · 2024 [cited by examiner]
US 20240146963A1 · Chen · 2024 [cited by examiner]
US 20240185075A1 · Du · 2024 [cited by examiner]
US 20240251098A1 · Chen · 2024 [cited by examiner]
US 20240314357A1 · Wang · 2024 [cited by examiner]
US 20250008094A1 · Deshpande · 2025 [cited by examiner]
US 20250088636A1 · Chen · 2025 [cited by examiner]
CN 111405283 · 2020 [cited by applicant]
V. Sze, M. Budagavi, and G. J. Sullivan, “High efficiency video coding (HEVC),” in Integrated Circuit and Systems, Algorithms and Architectures. Springer, 2014, vol. 39, p. 40. [cited by applicant]
B. Bross, Y.-K. Wang, Y. Ye, S. Liu, J. Chen, G. J. Sullivan, and J.R. Ohm, “Overview of the versatile video coding (VVC) standard and its applications,” IEEE Transactions on Circuits and Systems for Video Technology, v… [cited by applicant]
G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao, “DVC: An end-to-end deep video compression framework,” in Proceedings of the IEEE International Conference on Computer Vision, 2019, pp. 11 006-11 015. [cited by applicant]
T. Ladune, P. Philippe, W. Hamidouche, L. Zhang, and O. D'eforges, “Optical flow and mode selection for learning-based video coding,” in IEEE International Workshop on Multimedia Signal Processing, 2020, pp. 1-6. [cited by applicant]
J. Li, B. Li, and Y. Lu, “Deep contextual video compression,” Proceedings of Advances in Neural Information Processing Systems, vol. 34, pp. 18 114-18 125, 2021. [cited by applicant]
J. Lin, D. Liu, H. Li, and F. Wu, “M-LVC: Multiple frames prediction for learned video compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2020, pp. 3546-3554. [cited by applicant]
R. Yang, F. Mentzer, L. V. Gool, and R. Timofte, “Learning for video compression with hierarchical quality and recurrent enhancement,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 20… [cited by applicant]
M. A. Yilmaz and A. M. Tekalp, “End-to-end rate-distortion optimized learned hierarchical bi-directional video compression,” IEEE Transactions on Image Processing, vol. 31, pp. 974-983, 2021. [cited by applicant]
R. Pourreza and T. Cohen, “Extending neural p-frame codecs for b-frame coding,” in Proceedings of the IEEE International Conference on Computer Vision, 2021, pp. 6680-6689. [cited by applicant]
H. Jiang, D. Sun, V. Jampani, M.-H. Yang, E. Learned-Miller, and J. Kautz, “Super slomo: High quality estimation of multiple intermediate frames for video interpolation,” in Proceedings of the IEEE Conference on Compute… [cited by applicant]
D. Minnen, J. Ball'e, and G. D. Toderici, “Joint autoregressive and hierarchical priors for learned image compression,” Proceedings of Advances in Neural Information Processing Systems, vol. 31, 2018. [cited by applicant]
D. He, Y. Zheng, B. Sun, Y. Wang, and H. Qin, “Checkerboard context model for efficient learned image compression,” in Proceedings of the IEEE International Conference on Computer Vision, 2021, pp. 14 771-14 780. [cited by applicant]
J. Li, B. Li, and Y. Lu, “Hybrid spatial-temporal entropy modelling for neural video compression,” in Proceedings of the ACM International Conference on Multimedia, 2022, pp. 1503-1511. [cited by applicant]
J. Li, B. Li, and Y. Lu, “Neural video compression with diverse contexts,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 22 616-22 626, 2023. [cited by applicant]
D. He, Z. Yang, W. Peng, R. Ma, H. Qin, and Y. Wang, “ELIC: Efficient learned image compression with unevenly grouped space-channel contextual adaptive coding,” in Proceedings of the IEEE Conference on Computer Vision a… [cited by applicant]
D. Minnen and S. Singh, “Channel-wise autoregressive entropy models for learned image compression,” in 2020 IEEE International Conference on Image Processing (ICIP). IEEE, 2020, pp. 3339-3343. [cited by applicant]
J. Ball'e, V. Laparra, and E. P. Simoncelli, “End-to-end optimized image compression,” in International Conference on Learning Representations, 2017. [cited by applicant]
J. Ball'e, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston, “Variational image compression with a scale hyperprior,” in International Conference on Learning Representations, 2018. [cited by applicant]
T. Chen, H. Liu, Z. Ma, Q. Shen, X. Cao, and Y. Wang, “End-to-end learnt image compression via non-local attention optimization and improved context modeling,” IEEE Transactions on Image Processing, vol. 30, pp. 3179-31… [cited by applicant]
Z. Cheng, H. Sun, M. Takeuchi, and J. Katto, “Learned image compression with discretized gaussian mixture likelihoods and attention modules,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognit… [cited by applicant]
Z. Pan, X. Yi, Y. Zhang, B. Jeon, and S. Kwong, “Efficient in-loop filtering based on enhanced deep convolutional neural networks for hevc,” IEEE Transactions on Image Processing, vol. 29, pp. 5352-5366, 2020. [cited by applicant]
A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2017, pp. 4161-4170. [cited by applicant]
Z. Hu, Z. Chen, D. Xu, G. Lu, W. Ouyang, and S. Gu, “Improving deep video compression by resolution-adaptive flow coding,” in Proceedings of the European Conference on Computer Vision, 2020, pp. 193-209. [cited by applicant]
H. Guo, S. Kwong, C. Jia, and S. Wang, “Enhanced motion compensation for deep video compression,” IEEE Signal Processing Letters, vol. 30, pp. 673-677, 2023. [cited by applicant]
Z. Guo, R. Feng, Z. Zhang, X. Jin, and Z. Chen, “Learning cross-scale weighted prediction for efficient neural video compression,” IEEE Transactions on Image Processing, vol. 32, pp. 3567-3579, 2023. [cited by applicant]
D. Jin, J. Lei, B. Peng, Z. Pan, L. Li, and N. Ling, “Learned video compression with efficient temporal context learning,” IEEE Transactions on Image Processing, vol. 32, pp. 3188-3198, 2023. [cited by applicant]
E. Agustsson, D. Minnen, N. Johnston, J. Balle, S. J. Hwang, and G. Toderici, “Scale-space flow for end-to-end optimized video compression,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recogniti… [cited by applicant]
X. Sheng, J. Li, B. Li, L. Li, D. Liu, and Y. Lu, “Temporal context mining for learned video compression,” IEEE Transactions on Multimedia, 2022. [cited by applicant]
H. Guo, S. Kwong, D. Ye, and S. Wang, “Enhanced context mining and filtering for learned video compression,” IEEE Transactions on Multimedia, pp. 1-13, 2023. [cited by applicant]
H. Wang and Z. Chen, “Exploring long- and short-range temporal information for learned video compression,” IEEE Transactions on Image Processing, vol. 33, pp. 780-792, 2024. [cited by applicant]
D. Alexandre, H.-M. Hang, and W.-H. Peng, “Hierarchical B-frame Video Coding Using Two-Layer CANF without Motion Coding,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2023, pp. 10 24… [cited by applicant]
R. Zhang, P. Isola, A. A. Efros, E. Shechtman, and O. Wang, “The unreasonable effectiveness of deep features as a perceptual metric,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 201… [cited by applicant]
Z. Hu, G. Lu, and D. Xu, “Fvc: A new framework towards deep video compression in feature space,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2021, pp. 1502-1511. [cited by applicant]
F. Bossen et al., “Common test conditions and software reference configurations,” JCTVC-L1100, vol. 12, No. 7, 2013. [cited by applicant]
A. Mercat, M. Viitanen, and J. Vanne, “UVG dataset: 50/120fps 4k sequences for video codec analysis and development,” in Proceedings of the ACM Multimedia Systems Conference, 2020, pp. 297-302. [cited by applicant]
H. Wang, W. Gan, S. Hu, J. Y. Lin, L. Jin, L. Song, P. Wang, I. Katsavounidis, A. Aaron, and C.-C. J. Kuo, “MCL-JCV: A JNDbased H. 264/AVC video quality assessment dataset,” in Proceedings of International Conference on… [cited by applicant]
I. Loshchilov and F. Hutter, “Decoupled weight decay regularization,” International Conference on Learning Representations, 2018. [cited by applicant]
F. L. Vadim Seregin, Jie Chen and K. Zhang, “JVET AHG report: ECM software development (AHG6),” in JVET-AA0006, 2022. [cited by applicant]
G. Bjontegaard, “Calculation of Average PSNR Differences between RD-curves (VCEG-M33),” in VCEG Meeting (ITU-T SG16 Q. 6), 2001, pp. 2-4. [cited by applicant]
G. Lu, C. Cai, X. Zhang, L. Chen, W. Ouyang, D. Xu, and Z. Gao, “Content adaptive and error propagation aware deep video compression,” 2020. [cited by applicant]
R. Yang, F. Mentzer, L. Van Gool, and R. Timofte, “Learning for video compression with recurrent auto-encoder and recurrent probability model,” IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 3… [cited by applicant]