IP Library › Granted Patent US 12,501,050
Granted Patent B2
US 12,501,050 · App. 18/457,079 · Granted Dec 16, 2025

Efficient warping-based neural video codec

Inventors: Ties Jehan Van Rozendaal (Amsterdam, NL); Hoang Cong Minh Le (Santee, CA); Tushar Singhal (San Diego, CA); Amir Said (San Diego, CA); Krishna Buska (San Diego, CA); Guillaume Konrad Sautiere (Amsterdam, NL); Anjuman Raha (San Diego, CA); Auke Joris Wiggers (Amsterdam, NL); Frank Steven Mayer (San Diego, CA); Liang Zhang (San Diego, CA); Abhijit Khobare (San Diego, CA); Muralidhar Reddy Akula (San Diego, CA)
Assignee: QUALCOMM Incorporated
H04N19/137H04N19/159H04N19/176H04N19/192
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,501,050
App. No.
18/457,079
Granted
Dec 16, 2025
Kind
B2
Abstract

An example computing device may include memory and one or more processors. The one or more processors may be configured to parallel entropy decode encoded video data from a received bitstream to generate entropy decoded data. The one or more processors may be configured to predict a motion vector based on the entropy decoded data. The one or more processors may be configured to decode a motion vector residual from the entropy decoded data. The one or more processors may be configured to add the motion vector residual and motion vector. The one or more processors may be configured to warp previous reconstructed video data with an overlapped block-based warp function using the motion vector to generate predicted current video data. The one or more processors may be configured to sum the predicted current video data with a residual block to generate current reconstructed video data.

Claims (62)

1 . A device for coding video data, the device comprising:

memory for storing the video data, the video data comprising previous reconstructed video data and current reconstructed video data; and

one or more processors configured to:

parallel entropy decode encoded video data from a received bitstream to generate entropy decoded data;

predict a block-based motion vector based on a flow extrapolator applied to the entropy decoded data to generate a predicted motion vector;

decode a motion vector residual, with a decoder of a flow auto-encoder, from the entropy decoded data;

add the motion vector residual to the predicted motion vector to generate the block-based motion vector;

warp the previous reconstructed video data with an overlapped block-based warp function using the block-based motion vector to generate predicted current video data; and

sum the predicted current video data with a residual block, wherein the residual block is output from a decoder of a residual auto-encoder, to generate the current reconstructed video data.

2 . The device of claim 1 , wherein the motion vector residual, is a pixel-based motion vector residual.

3 . The device of claim 2 , wherein the flow auto-encoder is quantization-aware trained.

4 . The device of claim 1 , wherein the overlapped block-based warp function is configured to:

warp a block of the previous reconstructed video data a plurality of times using a respective motion vector of a respective surrounding block to generate warping results; and

average the warping results using a decay.

5 . The device of claim 1 , wherein as part of parallel entropy decoding the encoded video data, the one or more processors are configured to parallel entropy decode the encoded video data with at least one graphics processing unit.

6 . The device of claim 1 , wherein as part of warping the previous reconstructed video data, the one or more processors are configured to block-based frame interpolation (FINT) warp the previous reconstructed video data.

7 . The device of claim 6 , wherein as part of block-based FINT warping the previous reconstructed video data, the one or more processors are configured to use a FINT kernel to FINT warp the previous reconstructed video data.

8 . The device of claim 7 , wherein the FINT kernel is implemented in a neural network signal processor.

9 . The device of claim 1 , wherein the encoded video data represents YUV420 video data and the current reconstructed video data comprises YUV420 video data.

10 . The device of claim 1 , wherein the one or more processors are further configured to quantize at least a portion of the entropy decoded data.

11 . The device of claim 10 , wherein as part of quantizing the at least a portion of the entropy decoded data, the one or more processors are configured to quantize at least one of a latent, a mean, or a scale.

12 . The device of claim 10 , wherein as part of quantizing the at least a portion of the entropy decoded data, the one or more processors are configured to quantize the at least a portion of the entropy decoded data using int8.

13 . The device of claim 1 , wherein the one or more processors are configured to apply the flow extrapolator to the entropy decoded data to generate extrapolated flow.

14 . The device of claim 13 , wherein the one or more processors are configured to preform additive flow prediction using the extrapolated flow.

15 . The device of claim 1 , wherein the encoded video data comprises luma data.

16 . A method of decoding video data, the method comprising:

parallel entropy decoding encoded video data from a received bitstream to generate entropy decoded data;

predicting a block-based motion vector based on a flow extrapolator applied to the entropy decoded data to generate a predicted motion vector;

decoding a motion vector residual, with a decoder of a flow auto-encoder, from the entropy decoded data;

adding the motion vector residual and the predicted motion vector to generate the block-based motion vector;

warping previous reconstructed video data with an overlapped block-based warp function using the block-based motion vector to generate predicted current video data; and

summing the predicted current video data with a residual block, wherein the residual block is output from a decoder of a residual auto-encoder, to generate current reconstructed video data.

17 . The method of claim 16 , wherein decoding the motion vector residual comprises decoding a pixel-based motion vector residual using a neural network model.

18 . The method of claim 17 , wherein the neural network model is quantization-aware trained.

19 . The method of claim 16 , wherein warping the previous reconstructed video data with the overlapped block-based warp function comprises:

warping a block of the previous reconstructed video data a plurality of times using a respective motion vector of a respective surrounding block to generate warping results; and

averaging the warping results using a decay.

20 . The method of claim 16 , wherein parallel entropy decoding the encoded video data comprises parallel entropy decoding the encoded video data with at least one graphics processing unit.

21 . The method of claim 16 , wherein warping the previous reconstructed video data comprises block-based frame interpolation (FINT) warping the previous reconstructed video data.

22 . The method of claim 21 , wherein block-based FINT warping the previous reconstructed video data comprises using a FINT kernel to FINT warp the previous reconstructed video data.

23 . The method of claim 22 , wherein the FINT kernel is implemented in a neural network signal processor.

24 . The method of claim 16 , wherein the encoded video data represents YUV420 video data and the current reconstructed video comprises YUV420 video data.

25 . The method of claim 16 , further comprising quantizing at least a portion of the entropy decoded data.

26 . The method of claim 25 , wherein quantizing the at least a portion of the entropy decoded data comprises quantizing at least one of a latent, a mean, or a scale.

27 . The method of claim 25 , wherein the quantizing the at least a portion of the entropy decoded data comprises quantizing the at least a portion of the entropy decoded data using int8.

28 . The method of claim 16 , further comprising applying the flow extrapolator to the entropy decoded data to generate extrapolated flow.

29 . The method of claim 28 , further comprising preforming additive flow prediction using the extrapolated flow.

30 . The method of claim 16 , wherein the encoded video data comprises luma data.

31 . A device for coding video data, the device comprising:

means for parallel entropy decoding encoded video data from a received bitstream to generate entropy decoded data;

means for predicting a block-based motion vector based on a flow extrapolator applied to the entropy decoded data to generate a predicted motion vector;

means for decoding a motion vector residual, with a decoder of a flow auto-encoder, from the entropy decoded data;

means for adding the motion vector residual and the predicted motion vector to generate the block-based motion vector;

means for warping previous reconstructed video data with an overlapped block-based warp function using the block-based motion vector to generate predicted current video data; and

means for summing the predicted current video data with a residual block, wherein the residual block is output from a decoder of a residual auto-encoder, to generate current reconstructed video data.

32 . A non-transitory, computer-readable storage medium storing instructions that, when executed, cause one or more processors to:

parallel entropy decode encoded video data from a received bitstream to generate entropy decoded data;

predict a block-based motion vector based on a flow extrapolator applied to the entropy decoded data to generate a predicted motion vector;

decode a motion vector residual, with a decoder of a flow auto-encoder, from the entropy decoded data;

add the motion vector residual to the predicted motion vector to generate the block-based motion vector;

warp previous reconstructed video data with an overlapped block-based warp function using the block-based motion vector to generate predicted current video data; and

sum the predicted current video data with a residual block, wherein the residual block is output from a decoder of a residual auto-encoder, to generate current reconstructed video data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2025
From: VAN ROZENDAAL, TIES JEHAN; LE, HOANG CONG MINH; SINGHAL, TUSHAR; SAID, AMIR; BUSKA, KRISHNA; SAUTIERE, GUILLAUME KONRAD; RAHA, ANJUMAN; WIGGERS, AUKE JORIS; MAYER, FRANK STEVEN; ZHANG, LIANG; KHOBARE, ABHIJIT; AKULA, MURALIDHAR REDDY
To: QUALCOMM INCORPORATED
Reel/Frame 070963/0908 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 21, 2023
From: VAN ROZENDAAL, TIES JEHAN; LE, HOANG CONG MINH; SINGHAL, TUSHAR; SAID, AMIR; BUSKA, KRISHNA; SAUTIERE, GUILLAUME KONRAD; RAHA, ANJUMAN; WIGGERS, AUKE JORIS; MAYER, FRANK STEVEN; ZHANG, LIANG; KHOBARE, ABHIJIT; AKULA, MURALIDHAR REDDY
To: QUALCOMM INCORPORATED
Reel/Frame 065638/0337 →
Continuity (3)
Provisional Application 63497411 · Apr 20, 2023
Provisional Application 63489306 · Mar 9, 2023
Related Publication 20240305785A1 · Sep 12, 2024
References Cited (104)
US 6421467B1 · Mitra · 2002 [cited by applicant]
US 11374952B1 · Coskun et al. · 2022 [cited by applicant]
US 11532155B1 · Garimella et al. · 2022 [cited by applicant]
US 11599972B1 · Xu · 2023 [cited by examiner]
US 11825090B1 · Said · 2023 [cited by applicant]
US 11869221B2 · Johnston et al. · 2024 [cited by applicant]
US 20030108099A1 · Nagumo et al. · 2003 [cited by applicant]
US 20070092122A1 · Xiao · 2007 [cited by examiner]
US 20100074356A1 · Ashikhmin · 2010 [cited by applicant]
US 20100215101A1 · Jeon · 2010 [cited by examiner]
US 20120232913A1 · Terriberry et al. · 2012 [cited by applicant]
US 20130094831A1 · Suzuki · 2013 [cited by applicant]
US 20150179166A1 · Nagao · 2015 [cited by applicant]
US 20160219261A1 · Chen · 2016 [cited by examiner]
US 20170064302A1 · Na et al. · 2017 [cited by applicant]
US 20190265955A1 · Wolf · 2019 [cited by applicant]
US 20200027247A1 · Minnen et al. · 2020 [cited by applicant]
US 20200099954A1 · Hemmer et al. · 2020 [cited by applicant]
US 20200275130A1 · Bokov et al. · 2020 [cited by applicant]
US 20200351509A1 · Lee et al. · 2020 [cited by applicant]
US 20200356835A1 · Robinson et al. · 2020 [cited by applicant]
US 20210075937A1 · Hall · 2021 [cited by examiner]
US 20210125380A1 · Lee · 2021 [cited by examiner]
US 20210232872A1 · Ries et al. · 2021 [cited by applicant]
US 20210287780A1 · Korani et al. · 2021 [cited by applicant]
US 20220272372A1 · Dinh et al. · 2022 [cited by applicant]
US 20220375030A1 · Choi · 2022 [cited by examiner]
US 20230185953A1 · Weggenmann et al. · 2023 [cited by applicant]
US 20230245317A1 · Morard et al. · 2023 [cited by applicant]
US 20230262222A1 · Said · 2023 [cited by applicant]
US 20230262267A1 · Said et al. · 2023 [cited by applicant]
US 20240015318A1 · Pourreza · 2024 [cited by examiner]
US 20240016456A1 · De Zambotti et al. · 2024 [cited by applicant]
US 20240104786A1 · Johnston et al. · 2024 [cited by applicant]
US 20240121392A1 · Said · 2024 [cited by applicant]
Haojie et al. (“Hao”) (“Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 8, Aug. 1, 2021 (Aug. 1, 2… [cited by examiner]
Agustsson E., et al., “Scale-Space Flow for End-to-End Optimized Video Compression”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 13, 2020 (Jun. 13, 2020), pp. 8500-8509, XP0338… [cited by applicant]
Akiyo., et al., Xiph.org Video Test Media [derf's collection], Xiph.org, pp. 1-18. [cited by applicant]
Aminabadi R.Y., et al., “DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale”, arXiv:2207.00032v1 [cs.LG], Jun. 30, 2022, 13 Pages. [cited by applicant]
Balle., J,, et al., “Integer Networks for Data Compression with Latent-variable Models”, Published as a conference paper at ICLR 2019, pp. 1-10. [cited by applicant]
Balle J., et al., “Variational Image Compression with A Scale Hyperprior”, arXiv:1802.01436v2 [eess.IV], May 1, 2018, XP055632204, pp. 1-23. [cited by applicant]
Bengio Y., et al., “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation”, Computer Science and Operations Research Department, Montreal University, Aug. 15, 2013, pp. 1-12, arXiv p… [cited by applicant]
Bossen F., et al., “VTM Software Manual”, JVET-Software Manual, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG5, Date saved: Oct. 26, 2023, pp. 1-64. [cited by applicant]
Bossen F., “Common Test Conditions and Software Reference Configurations”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 12th Meeting: Geneva, CH Jan. 14-23, 2013, JCTVC… [cited by applicant]
Bross B., et al.,“Developments in International Video Coding Standardization After AVC, with an Overview of Versatile Video Coding (VVC)”, Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1463-1493. [cited by applicant]
Brummer B., et al., “End-to-end Optimized Image Compression with Competition of Prior Distributions”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 17, 2021, XP0911007… [cited by applicant]
Cao J., et al., “Extreme Learning Machine with Affine Transformation Inputs in an Activation Function”, IEEE Transactions on Neural Networks and Learning Systems, vol. 30, No. 7, Jul. 2019, pp. 2093-2107. [cited by applicant]
Dimitriadis S., et al., “Revealing Cross-Frequency Causal Interactions during a Mental Arithmetic Task through Symbolic Transfer Entropy: A Novel Vector-Quantization Approach”, IEEE Transactions on Neural Systems and Re… [cited by applicant]
Ding D., et al., “Advances in Video Compression System Using Deep Neural Network: A Review and Case Studies”, Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1494-1520. [cited by applicant]
Gabrie M., et al., “Entropy and Mutual Information in Models of Deep Neural Networks”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada, 2018, pp. 1-11. [cited by applicant]
Galpin F., et al., “Entropy Coding Improvement for Low-complexity Compressive Auto-Encoders”, arXiv:2303.05962v1 [eess.IV] Mar. 10, 2023, InterDigital, Inc. Rennes, France, 10 Pages. [cited by applicant]
Golinski A., et al., “Feedback Recurrent Autoencoder for Video Compression”, arXiv:2004.04342v1 [ cs.LG], Apr. 9, 2020, pp. 1-29. [cited by applicant]
Habibian A., et al., “Video Compression with Rate-Distortion Autoencoders”, arXiv:1908.05717v2 [eess.IV], Nov. 13, 2019, Cornell University Library, 201 Olin Library Cornell University, Ithaca, NY, 14853, Aug. 14, 2019,… [cited by applicant]
He D., et al., “Post-Training Quantization for Cross-Platform Learned Image Compression”, arXiv:2202.07513v2 [eess.IV], Nov. 30, 2022, 25 Pages. [cited by applicant]
Hong W., et al., “Efficient Neural Image Decoding via Fixed-Point Inference”, IEEE, Transactions on Circuits and Systems for Video Technology, vol. 31, No. 9, Sep. 2021, pp. 3618-3630. [cited by applicant]
ITU-T: HSTP-VID-WPOM Working Practices Using Objective Metrics for Evaluation of Video Coding Efficiency Experiments, Technical Paper, Telecommunication Standardization Sector of ITU, Jul. 3, 2020, pp. 13. [cited by applicant]
ITU-T H.265: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, High Efficiency Video Coding, The International Telecommunication Union, Jun. 2019, 696 Pages. [cited by applicant]
ITU-T H.266: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, Versatile Video Coding, The International Telecommunication Union, Aug. 2020, 516 pages. [cited by applicant]
Kaeli D., et al., “Heterogeneous Computing with OpenCL 2.0 Third Edition”, Elsevier Science, Burlington, 2015, pp. 1-330. [cited by applicant]
Kingma D.P., et al., “An Introduction to Variational Autoencoders”, Foundations and Trends in Machine Learning, arXiv:1906.02691v3 [cs.LG] Dec. 11, 2019, 89 pgs. [cited by applicant]
Kingma D.P., et al., “Auto-Encoding Variational Bayes”, ICLR 2014 Conference, Dec. 2013, pp. 1-14, arXiv preprint arXiv:1312.6114v10 [stat.ML] May 1, 2014. [cited by applicant]
Koyuncu E., et al., “Device Interoperability for Learned Image Compression with Weights and Activations Quantization”, arXiv:2212.01330v1 [eess.IV], Dec. 2, 2022, IEEE, pp. 1-5. [cited by applicant]
Le H., et al., “MobileCodec: Neural Inter-Frame Video Compression on Mobile Devices”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jul. 18, 2022, 8 Pages, XP091273560. [cited by applicant]
Li J., et al., “Deep Contextual Video Compression”, arXiv:2109.15047v2 [eess.IV], Dec. 14, 2021, 35th Conference on Neural Information Processing Systems, Sydney, Australia, 2021, 19 Pages. [cited by applicant]
Li J., et al., “Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression”, arXiv:2207.05894v1 [eess.IV], Jul. 13, 2022, 17 Pages. [cited by applicant]
Li J., et al., “Neural Video Compression with Diverse Contexts”, arXiv:2302.14402v3 [eess.IV], Mar. 14, 2023, 19 Pages. [cited by applicant]
Lu G., et al., “DVC: An End-To-End Deep Video Compression Framework”, arXiv:1812.00101v3 [eess.IV], Apr. 7, 2019, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019, XP033686979, … [cited by applicant]
Ma S., et al., “Image and Video Compression with Neural Networks: A Review”, arxiv.org, IEEE Transactions on Circuits and Systems for Video Technology, Cornell University Library, 201 Olin Library Cornell University Ith… [cited by applicant]
Mentzer F., et al., “M2T: Masking Transformers Twice for Faster Decoding”, arXiv:2304.07313v1 [eess.IV], Apr. 14, 2023, 13 Pages. [cited by applicant]
Mercat A., et al., “UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development”, Ultra Video Group, Tampere University, Finland, 2020, 6 Pages. [cited by applicant]
Minnen D., et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression”, 32nd Conference on Neural Information Processing Systems, Montreal, Canada, 2018, 10 Pages. [cited by applicant]
Nagel M., et al., “A White Paper on Neural Network Quantization”, arXiv:2106.08295v1 [cs.LG], Jun. 15, 2021 Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 15, 2021, 27 pages, XP08… [cited by applicant]
Nogaki S., et al., “An Overlapped block Motion Compensation for High Quality Motion Picture Coding”, Proceedings of the International Symposium on Circuits and Systems, San Diego, May 10-13, 1992; [Proceedings of the In… [cited by applicant]
Orchard M.T., et al., “Overlapped Block Motion Compensation: an Estimation-theoretic Approach”, IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US, vol. 3, No. 5, Sep. 1, 1994, pp. 693-699, X… [cited by applicant]
Pearlman W.A., et al., “Digital Signal Compression: Principles and Practice”, Cambridge University Press, 2011, 439 pages. [cited by applicant]
Pourreza R., et al., “Boosting Neural Video Codecs by Exploiting Hierarchical Redundancy”, arXiv:2208.04303v2 [eess.IV], Sep. 16, 2022, Qualcomm Technologies, pp. 1-11. [cited by applicant]
Qualcomm: “World's First Software-Based Neural Video Decoder Running HD Format in Real-Time on a Commercial Smartphone [Video] Qualcomm AI Research demonstrates 1280x704 HD video being decoded real-time at 30+ frames pe… [cited by applicant]
Rippel O., et al., “ELF-VC: Efficient Learned Flexible-Rate Video Coding”, arXiv:2104.14335v1 [eess.IV] Apr. 29, 2021, 14 Pages. [cited by applicant]
Rippel O., et al., “Learned Video Compression”, arXiv:1811.06981v1 [eess.IV], Nov. 16, 2018, 2019 IEEE/CVF International Conference on Computer Vision(ICCV), IEEE, Oct. 27, 2019 (Oct. 27, 2019), pp. 3453-3462, XP0337237… [cited by applicant]
Rozendaal T.V., et al., “Instance-Adaptive Video Compression: Improving Neural Codecs by Training on the Test Set”, arXiv:2111.10302v2 [eess.IV] Jun. 23, 2023, Qualcomm Technologies, pp. 1-29. [cited by applicant]
Rozendaal T.V., et al., “Overfitting for Fun and Profit: Instance-adaptive Data Compression”, arXiv:2101.08687v2 [ cs.LG], Jun. 1, 2021, Published as a Conference Paper at ICLR 2021, pp. 1-18. [cited by applicant]
Said A., “Arithmetic Coding”, Chapter 5, 2003, pp. 101-152. [cited by applicant]
Said A., et al., “Compressed Data Organization for High Throughput Parallel Entropy Coding”, LG Electronics Mobile Research, San Jose, CA, USA, Sep. 2015, 9 Pages. [cited by applicant]
Said A., et al., “Optimized Learned Entropy Coding Parameters for Practical Neural-Based Image and Video Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 20… [cited by applicant]
Said A., “Introduction to Arithmetic Coding—Theory and Practice”, Technical Report, Apr. 21, 2004, pp. 1-67, Retrieved from the Internet: URL: http://www.hpl.hp.com/techreports/2004/HPL-2004-76.pdf. [cited by applicant]
Shi J., et al., “Rate-Distortion Optimized Post-Training Quantization for Learned Image Compression”, arXiv: 2211.02854v3 [eess.IV], Oct. 9, 2023, IEEE, pp. 1-16. [cited by applicant]
Shi Y., et al., “AlphaVC: High-Performance and Efficient Learned Video Compression”, Huawei Technologies, Beijing China, 16 Pages. [cited by applicant]
Siddegowda S., et al., “Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)”, arXiv:2201.08442v1 [cs.LG] Jan. 20, 2022, pp. 1-39. [cited by applicant]
Sullivan G.J., et al., “Overview of the High Efficiency Video Coding (HEVC) Standard”, IEEE Transactions on Circuits and Systems for Video Technology, IEEE Service Center, Piscataway, NJ, US, vol. 22, No. 12, Dec. 1, 20… [cited by applicant]
Sun H., et al., “End-to-End Learned Image Compression with Fixed Point Weight Quantization”, arXiv:2007.04684v1 [eess.IV], Jul. 9, 2020, 5 Pages. [cited by applicant]
Sun H., et al., “End-to-End Learned Image Compression with Quantized Weights and Activations”, arXiv:2111.09348v1 [eess.IV], Nov. 17, 2021, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, pp. 1-14. [cited by applicant]
Sun H., et al., “Learned Image Compression with Fixed-point Arithmetic”, Picture Coding Symposium (PCS), IEEE, 2021, 5 Pages. [cited by applicant]
Sun H., et al., “Q-LIC: Quantizing Learned Image Compression with Channel Splitting”, arXiv:2205.14510v1 [eess.IV], May 28, 2022, pp. 1-9. [cited by applicant]
Theis L., et al., “Lossy Image Compression with Compressive Autoencoders”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY14853, Mar. 1, 2017 (Mar. 1, 2017), XP080753545, pp. 1-19. [cited by applicant]
Tkalcic M., et al., “Colour Spaces—Perceptual, Historical and Applicational Background”, EUROCON Ljubljana, Slovenia, IEEE, 2003, pp. 304-308. [cited by applicant]
Tomar S., “Converting Video Formats with FFmpeg”, Linux Journal, Apr. 28, 2006, pp. 1-4. [cited by applicant]
Wang H., et al., “MCL-JCV: A JND-Based H.264/AVC Video Quality Assessment Dataset”, 2016 IEEE International Conference on Image Processing (ICIP), 2016, pp. 1509-1513. [cited by applicant]
Wiegand T., et al., “Overview of the H.264 / AVC Video Coding Standard”, IEEE Transactions On Circuits and Systems for Video Technology, Jul. 2003, pp. 1-19. [cited by applicant]
Xue T., et al., “Video Enhancement with Task-Oriented Flow”, International Journal of Computer Vision, arXiv:1711.09078v3 [cs.CV] Nov. 10, 2019, pp. 1-20. [cited by applicant]
Zhao J., et al., “A Universal Encoder Rate Distortion Optimization Framework for Learned Compression”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , IEEE, Jun. 19, 2021, pp. 188… [cited by applicant]
Zhu Y., et al., “Transformer-Based Transform Coding”, Sep. 29, 2021, XP093005391, the whole document, pp. 1-35. [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/018215—ISA/EPO—May 17, 2024, 13 Pages. [cited by applicant]
Liu H., et al., “Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 8, Aug. 1, 2021, pp. 3182-3196, X… [cited by applicant]
Zhao S., et al., “Global Matching with Overlapping Attention for Optical Flow Estimation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 18, 2022, pp. 17571-17580, XP034194379, s… [cited by applicant]
Cited By (1)
US 12,744,944