IP Library Granted Patent US 12,542,908
Granted Patent B2
US 12,542,908 · App. 18/051,572 · Granted Feb 3, 2026

System and method for facilitating machine-learning based media compression

Inventors: Sam Tak Wu Kwong (Kowloon, HK); Rongqun Lin (Kowloon, HK); Shiqi Wang (Kowloon, HK)
Assignee: City University of Hong Kong
H04N19/139H04N19/13H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,542,908
App. No.
18/051,572
Granted
Feb 3, 2026
Kind
B2
Abstract

A computer-implemented method for facilitating machine-learning based media (e.g., video) compression. The method includes receiving a motion data set associated with motion-related difference between a first image and a second image, and processing the motion data set using a neural network to determine a plurality of motion data subsets. The method also includes processing the plurality of motion data subsets using one or more features associated with the first image to obtain a plurality of motion-warped feature data sets each associated with a respective motion data subset; and processing the plurality of motion-warped feature data sets to facilitate generation of context data for facilitating conditional coding based compression of the second image.

Claims (67)

1 . A computer-implemented method for facilitating machine-learning based media compression, comprising:

(a) receiving a decoded motion data set associated with motion-related difference between a first image and a second image;

(b) processing the decoded motion data set using a neural network to determine a plurality of motion data subsets at decoder side, wherein the neural network comprises convolution layers and residual blocks to generate motion diversities, and the motion diversities are added to the decoded motion data set to generate the plurality of motion data subsets;

(c) processing the plurality of motion data subsets using one or more features associated with the first image to obtain a plurality of motion-warped feature data sets with temporal information, each of the plurality of motion-warped feature data sets being associated with a respective motion data subset; and

(d) processing the plurality of motion-warped feature data sets to facilitate generation of context data for facilitating conditional coding based compression of the second image, wherein the step (d) comprises:

(d1) processing the plurality of motion-warped feature data sets by a hypotheses attention module (HAM) to determine a plurality of attention based weights each associated with a respective one of the plurality of motion-warped feature data sets, and applying each of the plurality of attention based weights to a respective one of the plurality of motion-warped feature data sets to obtain a plurality of attention-weighted motion-warped feature data sets; and

(d2) processing the plurality of attention-weighted motion-warped feature data sets using a neural network to obtain the context data.

2 . The computer-implemented method of claim 1 , wherein the motion data set comprises a motion vector.

3 . The computer-implemented method of claim 2 , wherein the motion vector comprises a reconstructed motion vector m t .

4 . The computer-implemented method of claim 3 , further comprising:

determining the reconstructed motion vector m t .

5 . The computer-implemented method of claim 4 , wherein determining the reconstructed motion vector m t comprises:

processing the first image and the second image to determine a first motion vector m t associated with motion-related difference between the first image and the second image; and

performing a motion compression operation on the first motion vector m t to determine the reconstructed motion vector m t ;

wherein the reconstructed motion vector m t corresponds to a second motion vector.

6 . The computer-implemented method of claim 5 , wherein the processing includes processing the first image and the second image using a spatial pyramid network (SpyNet) to determine the first motion vector m t .

7 . The computer-implemented method of claim 5 , wherein the motion compression operation comprises:

encoding the first motion vector based on a hyper-prior based entropy model to obtain motion data bitstream; and

decoding the motion data bitstream to obtain the reconstructed motion vector m t .

8 . The computer-implemented method of claim 2 , wherein (b) comprises:

(b1) generating a first motion matrix M ini based on the motion vector;

(b2) processing the first motion matrix M ini using a neural network to determine a motion diversity function F div ;

(b3) generating a second motion matrix M final based on the first motion matrix M ini and the motion diversity function F div ; and

(b4) processing the second motion matrix M final to obtain the plurality of motion data subsets.

9 . The computer-implemented method of claim 8 , wherein (b1) comprises:

duplicating and concatenating the motion vector to obtain the first motion matrix M ini .

10 . The computer-implemented method of claim 8 , wherein (b3) comprises:

generating a second motion matrix M final based on M final =M ini +F div (M ini ).

11 . The computer-implemented method of claim 8 , wherein (b4) comprises:

splitting the second motion matrix M final into the plurality of motion data subsets.

12 . The computer-implemented method of claim 1 , wherein each of the plurality of motion data subsets corresponds to a respective motion.

13 . The computer-implemented method of claim 1 , wherein (c) comprises:

(c1) extracting the one or more features from the first image; and

(c2) warping each of the plurality of motion data subsets with the one or more extracted features associated with the first image to obtain the plurality of motion-warped feature data sets.

14 . The computer-implemented method of claim 13 , wherein (c1) comprises:

processing the first image using a neural network to extract the one or more features.

15 . The computer-implemented method of claim 1 , wherein the hypotheses attention module comprises an attention based neural network which comprises a squeeze-and-excitation layer and a multi-scale neural network.

16 . The computer-implemented method of claim 15 , wherein (d1) comprises:

generating a feature matrix M fs based on concatenating the plurality of motion-warped feature data sets;

processing the feature matrix M fs using the squeeze-and-excitation layer to obtain a re-calibrated feature matrix M fs ; and

processing the re-calibrated feature matrix M fs using the multi-scale neural network to determine a weight matrix W;

wherein the weight matrix W includes the plurality of attention based weights.

17 . The computer-implemented method of claim 1 , wherein the context data is arranged to be applied to an entropy model to facilitate conditional coding based compression of the second image.

18 . The computer-implemented method of claim 17 , wherein the entropy model is an auto-aggressive entropy model.

19 . The computer-implemented method of claim 1 , further comprising:

(e) performing a conditional compression operation on the second image to obtain a compressed second image.

20 . The computer-implemented method of claim 19 , wherein the conditional compression operation comprises:

encoding the second image based on an auto-regressive entropy model to obtain a bitstream; and

decoding the bitstream to obtain the compressed second image.

21 . The computer-implemented method of claim 1 , wherein the first image is a compressed reference image.

22 . The computer-implemented method of claim 1 , wherein the first image and the second image correspond to consecutive frames of a video, and wherein the first image corresponds to a frame immediately before the second image.

23 . A system for facilitating machine-learning based media compression, comprising:

one or more processors; and

memory storing one or more programs configured to be executed by the one or more processors, the one or more programs including instructions for:

(a) receiving a decoded motion data set associated with motion-related difference between a first image of a video and a second image of the video;

(b) processing the decoded motion data set using a neural network to determine a plurality of motion data subsets at decoder side, wherein the neural network comprises convolution layers and residual blocks to generate motion diversities, and the motion diversities are added to the decoded motion data set to generate the plurality of motion data subsets;

(c) processing the plurality of motion data subsets using one or more features associated with the first image to obtain a plurality of motion-warped feature data sets with temporal information, each of the plurality of motion-warped feature data sets being associated with a respective motion data subset; and

(d) processing the plurality of motion-warped feature data sets to facilitate generation of context data for facilitating conditional coding based compression of the second image, wherein the step (d) comprises:

(d1) processing the plurality of motion-warped feature data sets by a hypotheses attention module (HAM) to determine a plurality of attention based weights each associated with a respective one of the plurality of motion-warped feature data sets, and applying each of the plurality of attention based weights to a respective one of the plurality of motion-warped feature data sets to obtain a plurality of attention-weighted motion-warped feature data sets; and

(d2) processing the plurality of attention-weighted motion-warped feature data sets using a neural network to obtain the context data.

24 . A non-transitory computer-readable storage medium storing one or more programs configured to be executed by one or more processors, the one or more programs including instructions for performing or facilitating performing of a method for facilitating machine-learning based media compression, the method comprising:

(a) receiving a decoded motion data set associated with motion-related difference between a first image of a video and a second image of the video;

(b) processing the decoded motion data set using a neural network to determine a plurality of motion data subsets at decoder side, wherein the neural network comprises convolution layers and residual blocks to generate motion diversities, and the motion diversities are added to the decoded motion data set to generate the plurality of motion data subsets;

(c) processing the plurality of motion data subsets using one or more features associated with the first image to obtain a plurality of motion-warped feature data sets with temporal information, each of the plurality of motion-warped feature data sets being associated with a respective motion data subset; and

(d) processing the plurality of motion-warped feature data sets to facilitate generation of context data for facilitating conditional coding based compression of the second image, wherein the step (d) comprises:

(d1) processing the plurality of motion-warped feature data sets by a hypotheses attention module (HAM) to determine a plurality of attention based weights each associated with a respective one of the plurality of motion-warped feature data sets, and applying each of the plurality of attention based weights to a respective one of the plurality of motion-warped feature data sets to obtain a plurality of attention-weighted motion-warped feature data sets; and

(d2) processing the plurality of attention-weighted motion-warped feature data sets using a neural network to obtain the context data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2022
From: KWONG, SAM TAK WU; LIN, RONGQUN; WANG, SHIQI
To: CITY UNIVERSITY OF HONG KONG
Reel/Frame 061828/0545 →
Continuity (1)
Related Publication 20240146934A1 · May 2, 2024
References Cited (55)
US 10332001B2 · Rippel et al. · 2019 [cited by applicant]
US 10623775B1 · Theis et al. · 2020 [cited by applicant]
US 10977553B2 · Rippel et al. · 2021 [cited by applicant]
US 20220394240A1 · Zhang · 2022 [cited by examiner]
CN 111405283 · 2022 [cited by applicant]
G. J. Sullivan, J.-R. Ohm, W.-J. Han, and T. Wiegand, “Overview of the High Efficiency Video Coding (HEVC) Standard,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, pp. 1649-1668, 2012. [cited by applicant]
J. Zhang, C. Jia, M. Lei, S. Wang, S. Ma, and W. Gao, “Recent development of AVS Video Coding Standard: AVS3,” in 2019 Picture Coding Symposium (PCS), 2019, pp. 1-5. [cited by applicant]
B. Bross, J. Chen, J.-R. Ohm, G. J. Sullivan, and Y.-K. Wang, “Developments in international video coding standardization after AVC, with an overview of Versatile Video Coding (VVC),” Proceedings of the IEEE, vol. 109, … [cited by applicant]
S. Ma, X. Zhang, C. Jia, Z. Zhao, S. Wang, and S. Wang, “Image and video compression with neural networks: A review,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 30, No. 6, pp. 1683-1698, 2019. [cited by applicant]
G. Toderici, S. M. O'Malley, S. J. Hwang, D. Vincent, D. Minnen, S. Baluja, M. Covell, and R. Sukthankar, “Variable rate image compression with recurrent neural networks,” arXiv preprint arXiv:1511.06085, 2015. [cited by applicant]
G. Toderici, D. Vincent, N. Johnston, S. Jin Hwang, D. Minnen, J. Shor, and M. Covell, “Full resolution image compression with recurrent neural networks,” in Proceedings of the IEEE conference on Computer Vision and Pat… [cited by applicant]
N. Johnston, D. Vincent, D. Minnen, M. Covell, S. Singh, T. Chinen, S. J. Hwang, J. Shor, and G. Toderici, “Improved lossy image compression with priming and spatially adaptive bit rates for recurrent networks,” in Proc… [cited by applicant]
J. Balle, V. Laparra, and E. P. Simoncelli, “End-to-end optimized image compression,” arXiv preprint arXiv:1611.01704, 2016. [cited by applicant]
J. Balle, D. Minnen, S. Singh, S. J. Hwang, and N. Johnston, “Vari-ational image compression with a scale hyperprior,” arXiv preprint arXiv:1802.01436, 2018. [cited by applicant]
D. Minnen, J. Balle, and G. D. Toderici, “Joint autoregressive and hierarchical priors for learned image compression,” Advances in neural information processing systems, vol. 31, 2018. [cited by applicant]
Z. Cheng, H. Sun, M. Takeuchi, and J. Katto, “Learned image compression with discretized gaussian mixture likelihoods and attention modules,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Reco… [cited by applicant]
D. He, Y. Zheng, B. Sun, Y. Wang, and H. Qin, “Checkerboard context model for efficient learned image compression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 14771-1… [cited by applicant]
G. Lu, W. Ouyang, D. Xu, X. Zhang, C. Cai, and Z. Gao, “DVC: An end-to-end deep video compression framework,” in Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2019, pp. 11006-11015. [cited by applicant]
A. Habibian, T. v. Rozendaal, J. M. Tomczak, and T. S. Cohen, “Video compression with rate-distortion autoencoders,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 7033-7042. [cited by applicant]
A. Djelouah, J. Campos, S. Schaub-Meyer, and C. Schroers, “Neural inter-frame compression for video coding,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 6421-6429. [cited by applicant]
O. Rippel, S. Nair, C. Lew, S. Branson, A. G. Anderson, and L. Bourdev, “Learned video compression,” in Proceedings of the IEEE/CVF International Conference on Computer Vision, 2019, pp. 3454-3463. [cited by applicant]
E. Agustsson, D. Minnen, N. Johnston, J. Balle, S. J. Hwang, and G. Toderici, “Scale-space flow for end-to-end optimized video compression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recog… [cited by applicant]
G. Lu, C. Cai, X. Zhang, L. Chen, W. Ouyang, D. Xu, and Z. Gao, “Content adaptive and error propagation aware deep video compression,” 2020. [cited by applicant]
Z. Hu, Z. Chen, D. Xu, G. Lu, W. Ouyang, and S. Gu, “Improving deep video compression by resolution-adaptive flow coding,” in European Conference on Computer Vision. Springer, 2020, pp. 193-209. [cited by applicant]
J. Lin, D. Liu, H. Li, and F. Wu, “M-LVC: Multiple frames prediction for learned video compression,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 3546-3554. [cited by applicant]
R. Yang, F. Mentzer, L. V. Gool, and R. Timofte, “Learning for video compression with hierarchical quality and recurrent enhancement,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition… [cited by applicant]
R. Yang, F. Mentzer, L. Van Gool, and R. Timofte, “Learning for video compression with recurrent auto-encoder and recurrent probability model,” IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 3… [cited by applicant]
J. Li, B. Li, and Y. Lu, “Deep contextual video compression,” Advances in Neural Information Processing Systems, vol. 34, 2021. [cited by applicant]
Z. Hu, G. Lu, and D. Xu, “FVC: A new framework towards deep video compression in feature space,” in Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2021, pp. 1502-1511. [cited by applicant]
S. T. H. Shah and X. Xuezhi, “Traditional and modern strategies for optical flow: an investigation,” SN Applied Sciences, vol. 3, No. 3, pp. 1-14, 2021. [cited by applicant]
B. Girod, “Why b-pictures work: A theory of multi-hypothesis motioncompensated prediction,” in Proceedings 1998 International Conference on Image Processing. ICIP98 (Cat. No. 98CB36269), vol. 2. IEEE, 1998, pp. 213-217. [cited by applicant]
B. Girod, “Efficiency analysis of multihypothesis motion-compensated prediction for video coding,” IEEE Transactions on Image Processing, vol. 9, No. 2, pp. 173-183, 2000. [cited by applicant]
M. Flierl, T. Wiegand, and B. Girod, “Rate-constrained multihypothesis prediction for motion-compensated video compression,” IEEE Transactions on Circuits and Systems for Video Technology, vol. 12, No. 11, pp. 957-969, … [cited by applicant]
M. T. Orchard and G. J. Sullivan, “Overlapped block motion compensation: An estimation-theoretic approach,” IEEE Transactions on Image Processing, vol. 3, No. 5, pp. 693-699, 1994. [cited by applicant]
G. Sullivan, “Multi-hypothesis motion compensation for low bit-rate video coding,” in 1993 IEEE International Conference on Acoustics, Speech, and Signal Processing, vol. 5. IEEE, 1993, pp. 437-440. [cited by applicant]
Z. Wang, S. Wang, X. Zhang, S. Wang, and S. Ma, “Multi-hypothesis prediction based on implicit motion vector derivation for video coding,” in 2018 25th IEEE International Conference on Image Processing (ICIP). IEEE, 201… [cited by applicant]
G. K. Wallace, “The JPEG still picture compression standard,” IEEE transactions on consumer electronics, vol. 38, No. 1, pp. xviii-xxxiv, 1992. [cited by applicant]
D. S. Taubman and M. W. Marcellin, “JPEG2000: Standard for interactive imaging,” Proceedings of the IEEE, vol. 90, No. 8, pp. 1336-1357, 2002. [cited by applicant]
F. Bellard, “BPG image format,” 2015. [Online]. Available: https://bellard.org/bpg/. [cited by applicant]
C.-Y. Wu, N. Singhal, and P. Krahenbuhl, “Video compression through image interpolation,” in Proceedings of the European conference on computer vision (ECCV), 2018, pp. 416-431. [cited by applicant]
J. Pessoa, H. Aidos, P. Tomas, and M. A. Figueiredo, “End-to-end learning of video compression using spatio-temporal autoencoders,” in 2020 IEEE Workshop on Signal Processing Systems (SiPS). IEEE, 2020, pp. 1-6. [cited by applicant]
M. Liou, “Overview of the px64 kbit/s video coding standard,” Communications of the ACM, vol. 34, No. 4, pp. 59-63, 1991. [cited by applicant]
S.-W. Wu and A. Gersho, “Joint estimation of forward and backward motion vectors for interpolative prediction of video,” IEEE Transactions on Image Processing, vol. 3, No. 5, pp. 684-687, 1994. [cited by applicant]
M. Flierl, T. Wiegand, and B. Girod, “A locally optimal design algorithm for block-based multi-hypothesis motion-compensated prediction,” in Proceedings DCC'98 Data Compression Conference (Cat. No. 98TB100225). IEEE, 19… [cited by applicant]
W.-Y. Kung, C.-S. Kim, and C.-C. Kuo, “Multi-hypothesis motion compensated prediction (mhmcp) for error-resilient visual communication,” in Proceedings of 2004 International Symposium on Intelligent Multimedia, Video an… [cited by applicant]
W.-Y. Kung, C.-S. Kim, and C.-C. J. Kuo, “Analysis of multi-hypothesis motion compensated prediction for robust video transmission,” in 2004 IEEE International Symposium on Circuits and Systems (ISCAS), vol. 3. IEEE, 20… [cited by applicant]
A. Ranjan and M. J. Black, “Optical flow estimation using a spatial pyramid network,” in 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017, pp. 2720-2729. [cited by applicant]
K. C. Chan, X. Wang, K. Yu, C. Dong, and C. C. Loy, “Understanding deformable alignment in video super-resolution,” arXiv preprint arXiv:2009.07265, vol. 4, No. 3, p. 4, 2020. [cited by applicant]
T. Xue, B. Chen, J. Wu, D. Wei, and W. T. Freeman, “Video enhancement with task-oriented flow,” International Journal of Computer Vision, vol. 127, No. 8, pp. 1106-1125, 2019. [cited by applicant]
A. Mercat, M. Viitanen, and J. Vanne, “UVG dataset: 50/120fps 4k sequences for video codec analysis and development,” in Proceedings of the 11th ACM Multimedia Systems Conference, 2020, pp. 297-302. [cited by applicant]
A. Paszke, S. Gross, F. Massa, A. Lerer, J. Bradbury, G. Chanan, T. Killeen, Z. Lin, N. Gimelshein, L. Antiga et al., “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural informatio… [cited by applicant]
J. Begaint, F. Racap' e, S. Feltman, and A. Pushparaja, “Compressai: a pytorch library and evaluation platform for end-to-end compression research,” arXiv preprint arXiv:2011.03029, 2020. [cited by applicant]
D. P. Kingma and J. Ba, “Adam: A method for stochastic optimization,” arXiv preprint arXiv:1412.6980, 2014. [cited by applicant]
R. Yang, L. Van Gool, and R. Timofte, “OpenDVC: An open source implementation of the DVC video compression method,” arXiv preprint arXiv:2006.15862, 2020. [cited by applicant]
G. Bjøntegaard, “Calculation of average PSNR differences between RDcurves (VCEG-M33),” in VCEG Meeting (ITU-T SG16 Q. 6), 2001, pp. 2-4. [cited by applicant]