IP Library Granted Patent US 12,470,715
Granted Patent B2
US 12,470,715 · App. 18/534,073 · Granted Nov 11, 2025

Neural-network media compression using quantized entropy coding distribution parameters

Inventor: Amir Said (San Diego, CA)
Assignee: QUALCOMM Incorporated
H04N19/13H04N19/124H04N19/134H04N19/136H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,470,715
App. No.
18/534,073
Granted
Nov 11, 2025
Kind
B2
Abstract

A media coder performs entropy coding techniques for media data coded using neural-based techniques. A media coder is configured to determine a probability distribution function parameter for a data element of a data stream coded by a neural-based media compression technique, wherein the probability distribution function parameter is a function of a standard deviation of a probability distribution function of the data stream, determine a code vector based on the probability distribution function parameter, and entropy code the data element using the code vector.

Claims (44)

1 . A method of coding media data, the method comprising:

determining a probability distribution function parameter for a data element of a data stream coded by a neural-based media compression technique, wherein the probability distribution function parameter is a function of a standard deviation of a probability distribution function of the data stream and wherein the probability distribution function parameter is based on a distribution of the data stream optimized for quantization,

determining a code vector based on the probability distribution function parameter; and

entropy coding the data element using the code vector.

2 . The method of claim 1 , further comprising:

quantizing the probability distribution function parameter prior to determining the code vector.

3 . The method of claim 1 , wherein the probability distribution function parameter is u, the standard deviation of the probability distribution function is σ, a minimum standard deviation is σ min , a maximum standard deviation is σ max , and wherein the relationship between u and σ is defined as: σ=T u→σ (u)=T λ→σ (T u→λ (u)), where T λ→σ (λ)=exp(λ ln(σ max /σ min )+ln(σ min )), where functions T are defined according to an algorithm that measures coding redundancy or defined by solving an ordinary differential equation.

4 . The method of claim 1 , further comprising:

generating the data element using the neural-based compression technique; and

quantizing the data element to create a quantized data element,

wherein entropy coding the data element using the code vector comprises entropy encoding the quantized data element using the code vector.

5 . The method of claim 4 , wherein generating the data element using the neural-based compression technique comprises:

processing an image or video picture using an image analysis neural network to generate the data element.

6 . The method of claim 5 , further comprising:

capturing the image or video picture using a camera.

7 . The method of claim 1 , wherein entropy coding the data element using the code vector comprises entropy decoding an encoded data element using the code vector to create a quantized data element, the method further comprising:

dequantizing the quantized data element to create a reconstructed data element.

8 . The method of claim 7 , further comprising:

processing the reconstructed data element using an image synthesis neural network to reconstruct an image or video picture.

9 . The method of claim 8 , further comprising:

displaying the image or the video picture.

10 . An apparatus configured to code media data, the apparatus comprising:

a memory; and

one or more processors in communication with the memory, the one or more processors configured to:

determine a probability distribution function parameter for a data element of a data stream coded by a neural-based media compression technique, wherein the probability distribution function parameter is a function of a standard deviation of a probability distribution function of the data stream and wherein the probability distribution function parameter is based on a distribution of the data stream optimized for quantization;

determine a code vector based on the probability distribution function parameter; and

entropy code the data element using the code vector.

11 . The apparatus of claim 10 , wherein the one or more processors are further configured to:

quantize the probability distribution function parameter prior to determining the code vector.

12 . The apparatus of claim 10 , wherein the probability distribution function parameter is u, the standard deviation of the probability distribution function is σ, a minimum standard deviation is σ min , a maximum standard deviation is σ max , and wherein the relationship between u and σ is defined as: σ=T u→σ (u)=T λ→σ (T u→λ (u)), where T λ→σ (λ)=exp(λ ln(σ max /σ min )+ln(σ min )), where functions T are defined according to an algorithm that measures coding redundancy or defined by solving an ordinary differential equation.

13 . The apparatus of claim 10 , wherein the one or more processors are further configured to:

generate the data element using the neural-based compression technique; and

quantize the data element to create a quantized data element,

wherein entropy coding the data element using the code vector comprises entropy encoding the quantized data element using the code vector.

14 . The apparatus of claim 13 , wherein to generate the data element using the neural-based compression technique, the one or more processors are further configured to:

process an image or video picture using an image analysis neural network to generate the data element.

15 . The apparatus of claim 14 , further comprising:

a camera configured to capture the image or video picture.

16 . The apparatus of claim 10 , wherein entropy coding the data element using the code vector comprises entropy decoding an encoded data element using the code vector to create a quantized data element, and wherein the one or more processors are further configured to:

dequantize the quantized data element to create a reconstructed data element.

17 . The apparatus of claim 16 , wherein the one or more processors are further configured to:

process the reconstructed data element using an image synthesis neural network to reconstruct an image or video picture.

18 . The apparatus of claim 17 , further comprising:

a display configured to display the image or video picture.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2023
From: SAID, AMIR
To: QUALCOMM INCORPORATED
Reel/Frame 065815/0849 →
Continuity (3)
Continuation 17814426 · Jul 22, 2022
Provisional Application 63267857 · Feb 11, 2022
Related Publication 20240121392A1 · Apr 11, 2024
References Cited (102)
US 6421467B1 · Mitra · 2002 [cited by examiner]
US 11374952B1 · Coskun · 2022 [cited by examiner]
US 11532155B1 · Garimella · 2022 [cited by examiner]
US 11599972B1 · Xu et al. · 2023 [cited by applicant]
US 11825090B1 · Said · 2023 [cited by applicant]
US 11869221B2 · Johnston · 2024 [cited by examiner]
US 11876969B2 · Said · 2024 [cited by applicant]
US 20030108099A1 · Nagumo · 2003 [cited by examiner]
US 20070092122A1 · Xiao et al. · 2007 [cited by applicant]
US 20100074356A1 · Ashikhmin · 2010 [cited by examiner]
US 20100215101A1 · Jeon et al. · 2010 [cited by applicant]
US 20120232913A1 · Terriberry · 2012 [cited by examiner]
US 20130094831A1 · Suzuki · 2013 [cited by examiner]
US 20150179166A1 · Nagao · 2015 [cited by examiner]
US 20160219261A1 · Chen et al. · 2016 [cited by applicant]
US 20170064302A1 · Na · 2017 [cited by examiner]
US 20190265955A1 · Wolf · 2019 [cited by examiner]
US 20200027247A1 · Minnen · 2020 [cited by examiner]
US 20200099954A1 · Hemmer · 2020 [cited by examiner]
US 20200275130A1 · Bokov et al. · 2020 [cited by applicant]
US 20200351509A1 · Lee · 2020 [cited by examiner]
US 20200356835A1 · Robinson · 2020 [cited by examiner]
US 20210075937A1 · Hall · 2021 [cited by applicant]
US 20210125380A1 · Lee et al. · 2021 [cited by applicant]
US 20210232872A1 · Ries · 2021 [cited by examiner]
US 20210287780A1 · Korani · 2021 [cited by examiner]
US 20220272372A1 · Dinh et al. · 2022 [cited by applicant]
US 20220375030A1 · Choi et al. · 2022 [cited by applicant]
US 20230185953A1 · Weggenmann · 2023 [cited by examiner]
US 20230245317A1 · Morard · 2023 [cited by examiner]
US 20230262267A1 · Said et al. · 2023 [cited by applicant]
US 20240016456A1 · de Zambotti · 2024 [cited by examiner]
US 20240104786A1 · Johnston et al. · 2024 [cited by applicant]
Agustsson E., et al., “Scale-Space Flow for End-to-End Optimized Video Compression”, 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 13, 2020 (Jun. 13, 2020), pp. 8500-8509, XP0338… [cited by applicant]
Akiyo., et al., Xiph.org Video Test Media [derf's collection], Xiph.org, pp. 1-18. [cited by applicant]
Aminabadi R.Y., et al., “DeepSpeed Inference: Enabling Efficient Inference of Transformer Models at Unprecedented Scale”, arXiv:2207.00032v1 [cs.LG], Jun. 30, 2022, 13 Pages. [cited by applicant]
Balle., J,, et al., “Integer Networks for Data Compression with Latent-variable Models”, Published as a conference paper at ICLR 2019, pp. 1-10. [cited by applicant]
Balle J., et al., “Variational Image Compression with A Scale Hyperprior”, arXiv: 1802.01436v2 [eess.IV], May 1, 2018, XP055632204, pp. 1-23, Section 2. [cited by applicant]
Bengio Y., et al., “Estimating or Propagating Gradients Through Stochastic Neurons for Conditional Computation”, Computer Science and Operations Research Department, Montreal University, Aug. 15, 2013, pp. 1-12, arXiv p… [cited by applicant]
Bossen F., et al., “VTM Software Manual”, JVET-Software Manual, Joint Video Experts Team (JVET) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG5, Date saved: Oct. 26, 2023, pp. 1-64. [cited by applicant]
Bossen F., “Common Test Conditions and Software Reference Configurations”, Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11, 12th Meeting: Geneva, CH Jan. 14-23, 2013, JCTVC… [cited by applicant]
Bross B., et al.,“Developments in International Video Coding Standardization After AVC, with an Overview of Versatile Video Coding (VVC)”, Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1463-1493. [cited by applicant]
Brummer B., et al., “End-to-end Optimized Image Compression with Competition of Prior Distributions”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 17, 2021, XP0911007… [cited by applicant]
Cao J., et al., “Extreme Learning Machine with Affine Transformation Inputs in an Activation Function”, IEEE Transactions on Neural Networks and Learning Systems, vol. 30, No. 7, Jul. 2019, pp. 2093-2107. [cited by applicant]
Co-pending U.S. Appl. No. 18/457,079, filed Aug. 28, 2023. [cited by applicant]
Dimitriadis S., et al., “Revealing Cross-Frequency Causal Interactions during a Mental Arithmetic Task through Symbolic Transfer Entropy: A Novel Vector-Quantization Approach”, IEEE Transactions on Neural Systems and Re… [cited by applicant]
Ding D., et al., “Advances in Video Compression System Using Deep Neural Network: A Review and Case Studies”, Proceedings of the IEEE, vol. 109, No. 9, Sep. 2021, pp. 1494-1520. [cited by applicant]
Gabrie M., et al., “Entropy and Mutual Information in Models of Deep Neural Networks”, 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montréal, Canada, 2018, pp. 1-11. [cited by applicant]
Galpin F., et al., “Entropy Coding Improvement for Low-complexity Compressive Auto-Encoders”, arXiv:2303.05962v1 [eess.IV] Mar. 10, 2023, InterDigital, Inc. Rennes, France, 10 Pages. [cited by applicant]
Golinski A., et al., “Feedback Recurrent Autoencoder for Video Compression”, arXiv:2004.04342v1 [ cs.LG], 15th Asian Conference on Computer Vision, Kyoto, Japan, Nov. 30, 2020-Dec. 4, 2020, Revised Selected Papers, Part… [cited by applicant]
Habibian A., et al., “Video Compression with Rate-Distortion Autoencoders”, arXiv: 1908.05717v2 [eess.IV], Nov. 13, 2019, Cornell University Library, 201 Olin Library Cornell University, Ithaca, NY, 14853, Aug. 14, 2019… [cited by applicant]
He D., et al., “Post-Training Quantization for Cross-Platform Learned Image Compression”, arXiv:2202.07513v2 [eess.IV], Nov. 30, 2022, 25 Pages. [cited by applicant]
Hong W., et al., “Efficient Neural Image Decoding via Fixed-Point Inference”, IEEE, Transactions on Circuits and Systems for Video Technology, vol. 31, No. 9, Sep. 2021, pp. 3618-3630. [cited by applicant]
International Search Report and Written Opinion—PCT/US2023/060543—ISA/EPO—Mar. 24, 2023. [cited by applicant]
ITU-T: HSTP-VID-WPOM Working Practices Using Objective Metrics for Evaluation of Video Coding Efficiency Experiments, Technical Paper, Telecommunication Standardization Sector of ITU, Jul. 3, 2020, pp. 13. [cited by applicant]
ITU-T H.265: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, High Efficiency Video Coding, The International Telecommunication Union, Jun. 2019, 696 Pages. [cited by applicant]
ITU-T H.266: “Series H: Audiovisual and Multimedia Systems Infrastructure of Audiovisual Services—Coding of Moving Video”, Versatile Video Coding, The International Telecommunication Union, Aug. 2020, 516 pages. [cited by applicant]
Kaeli D., et al., “Heterogeneous Computing with OpenCL 2.0 Third Edition”, Elsevier Science, Burlington, 2015, pp. 1-330. [cited by applicant]
Kingma D.P., et al., “An Introduction to Variational Autoencoders”, Foundations and Trends in Machine Learning, arXiv:1906.02691v3 [cs.LG] Dec. 11, 2019, 89 pgs. [cited by applicant]
Kingma D.P., et al., “Auto-Encoding Variational Bayes”, ICLR 2014 Conference, Dec. 2013, pp. 1-14, arXiv preprint arXiv:1312.6114v10 [stat.ML] May 1, 2014. [cited by applicant]
Koyuncu E., et al., “Device Interoperability for Learned Image Compression with Weights and Activations Quantization”, arXiv:2212.01330v1 [eess.IV], Dec. 2, 2022, IEEE, pp. 1-5. [cited by applicant]
Le H., et al., “MobileCodec: Neural Inter-Frame Video Compression on Mobile Devices”, arxiv.org, Cornell University Library, 201 OLIN Library Cornell University Ithaca, NY 14853, Jul. 18, 2022, 8 Pages, XP091273560, Sec… [cited by applicant]
Li J., et al., “Deep Contextual Video Compression”, arXiv:2109.15047v2 [eess.IV], Dec. 14, 2021, 35th Conference on Neural Information Processing Systems, Sydney, Australia, 2021, 19 Pages. [cited by applicant]
Li J., et al., “Hybrid Spatial-Temporal Entropy Modelling for Neural Video Compression”, arXiv:2207.05894v1 [eess.IV], Jul. 13, 2022, 17 Pages. [cited by applicant]
Li J., et al., “Neural Video Compression with Diverse Contexts”, arXiv:2302.14402v3 [eess.IV], Mar. 14, 2023, 19 Pages. [cited by applicant]
Lu G., et al., “DVC: An End-To-End Deep Video Compression Framework”, arXiv:1812.00101v3 [eess.IV], Apr. 7, 2019, IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 15, 2019, XP033686979, … [cited by applicant]
Ma S., et al., “Image and Video Compression with Neural Networks: A Review”, arxiv.org, IEEE Transactions on Circuits and Systems for Video Technology, Cornell University Library, 201 OLIN Library Cornell University Ith… [cited by applicant]
Mentzer F., et al., “M2T: Masking Transformers Twice for Faster Decoding”, arXiv:2304.07313v1 [eess.IV], Apr. 14, 2023, 13 Pages. [cited by applicant]
Mercat A., et al., “UVG Dataset: 50/120fps 4K Sequences for Video Codec Analysis and Development”, Ultra Video Group, Tampere University, Finland, 2020, 6 Pages. [cited by applicant]
Minnen D., et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression”, 32nd Conference on Neural Information Processing Systems, Montreal, Canada, 2018, 10 Pages. [cited by applicant]
Nagel M., et al., “A White Paper on Neural Network Quantization”, arXiv:2106.08295v1 [cs.LG], Jun. 15, 2021 Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 15, 2021, 27 Pages, XP08… [cited by applicant]
Nogaki S., et al., “An Overlapped block Motion Compensation for High Quality Motion Picture Coding”, Proceedings of the International Symposium on Circuits and Systems, San Diego, May 10-13, 1992; [Proceedings of the In… [cited by applicant]
Orchard M.T., et al., “Overlapped Block Motion Compensation: an Estimation-theoretic Approach”, IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US, vol. 3, No. 5, Sep. 1, 1994, pp. 693-699, X… [cited by applicant]
Pearlman W.A., et al., “Digital Signal Compression: Principles and Practice”, Cambridge University Press, 2011, 439 pages. [cited by applicant]
Pourreza R., et al., “Boosting Neural Video Codecs by Exploiting Hierarchical Redundancy”, arXiv:2208.04303v2 [eess.IV], Sep. 16, 2022, Qualcomm Technologies, pp. 1-11. [cited by applicant]
Qualcomm: “World's First Software-Based Neural Video Decoder Running HD Format in Real-Time on a Commercial Smartphone [Video] Qualcomm AI Research demonstrates 1280x704 HD video being decoded real-time at 30+ frames pe… [cited by applicant]
Rippel O., et al., “ELF-VC: Efficient Learned Flexible-Rate Video Coding”, arXiv:2104.14335v1 [eess.IV] Apr. 29, 2021, 14 Pages. [cited by applicant]
Rippel O., et al., “Learned Video Compression”, arXiv:1811.06981v1 [eess.IV], Nov. 16, 2018, 2019 IEEE/CVF International Conference on Computer Vision(ICCV), IEEE, Oct. 27, 2019 (Oct. 27, 2019), pp. 3453-3462, XP0337237… [cited by applicant]
Rozendaal T.V., et al., “Instance-Adaptive Video Compression: Improving Neural Codecs by Training on the Test Set”, arXiv:2111.10302v2 [eess.IV] Jun. 23, 2023, Qualcomm Technologies, pp. 1-29. [cited by applicant]
Rozendaal T.V., et al., “Overfitting for Fun and Profit: Instance-adaptive Data Compression”, arXiv:2101.08687v2 [ cs.LG], Jun. 1, 2021, Published as a Conference Paper at ICLR 2021, pp. 1-18. [cited by applicant]
Said A., “Arithmetic Coding”, Chapter 5, 2003, pp. 101-152. [cited by applicant]
Said A., et al., “Compressed Data Organization for High Throughput Parallel Entropy Coding”, LG Electronics Mobile Research, San Jose, CA, USA, Sep. 2015, 9 Pages. [cited by applicant]
Said A., et al., “Optimized Learned Entropy Coding Parameters for Practical Neural-Based Image and Video Compression”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 20… [cited by applicant]
Said A., “Introduction to Arithmetic Coding—Theory and Practice”, Technical Report, Apr. 21, 2004, pp. 1-67, Retrieved from the Internet: URL: http://www.hpl.hp.com/techreports/2004/HPL-2004-76.pdf. [cited by applicant]
Shi J., et al., “Rate-Distortion Optimized Post-Training Quantization for Learned Image Compression”, arXiv: 2211.02854v3 [eess.IV], Oct. 9, 2023, IEEE, pp. 1-16. [cited by applicant]
Shi Y., et al., “AlphaVC: High-Performance and Efficient Learned Video Compression”, Huawei Technologies, Beijing China, 16 Pages. [cited by applicant]
Siddegowda S., et al., “Neural Network Quantization with AI Model Efficiency Toolkit (AIMET)”, arXiv:2201.08442v1 [cs.LG] Jan. 20, 2022, pp. 1-39. [cited by applicant]
Sullivan G.J., et al., “Overview of the High Efficiency Video Coding (HEVC) Standard”, IEEE Transactions on Circuits and Systems for Video Technology, IEEE Service Center, Piscataway, NJ, US, vol. 22, No. 12, Dec. 1, 20… [cited by applicant]
Sun H., et al., “End-to-End Learned Image Compression with Fixed Point Weight Quantization”, arXiv:2007.04684v1 [eess.IV], Jul. 9, 2020, 5 Pages. [cited by applicant]
Sun H., et al., “End-to-End Learned Image Compression with Quantized Weights and Activations”, arXiv:2111.09348v1 [eess.IV], Nov. 17, 2021, Journal of Latex Class Files, vol. 14, No. 8, Aug. 2015, pp. 1-14. [cited by applicant]
Sun H., et al., “Learned Image Compression with Fixed-point Arithmetic”, Picture Coding Symposium (PCS), IEEE, 2021, 5 Pages. [cited by applicant]
Sun H., et al., “Q-LIC: Quantizing Learned Image Compression with Channel Splitting”, arXiv:2205.14510v1 [eess.IV], May 28, 2022, pp. 1-9. [cited by applicant]
Theis L., et al., “Lossy Image Compression with Compressive Autoencoders”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY14853, Mar. 1, 2017 (Mar. 1, 2017), XP080753545, pp. 1-19, … [cited by applicant]
Tkalcic M., et al., “Colour Spaces—Perceptual, Historical and Applicational Background”, Eurocon Ljubljana, Slovenia, IEEE, 2003, pp. 304-308. [cited by applicant]
Tomar S., “Converting Video Formats with FFmpeg”, Linux Journal, Apr. 28, 2006, pp. 1-4. [cited by applicant]
Wang H., et al., “MCL-JCV: A JND-Based H.264/AVC Video Quality Assessment Dataset”, 2016 IEEE International Conference on Image Processing (ICIP), 2016, pp. 1509-1513. [cited by applicant]
Wiegand T., et al., “Overview of the H.264 / AVC Video Coding Standard”, IEEE Transactions on Circuits and Systems for Video Technology, Jul. 2003, pp. 1-19. [cited by applicant]
Xue T., et al., “Video Enhancement with Task-Oriented Flow”, International Journal of Computer Vision, arXiv:1711.09078v3 [cs.CV] Nov. 10, 2019, pp. 1-20. [cited by applicant]
Zhao J., et al., “A Universal Encoder Rate Distortion Optimization Framework for Learned Compression”, 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops (CVPRW) , IEEE, Jun. 19, 2021, pp. 188… [cited by applicant]
Zhu Y., et al., “Transformer-Based Transform Coding”, Sep. 29, 2021, XP093005391, the whole document, pp. 1-35. [cited by applicant]
Liu H., et al., “Neural Video Coding Using Multiscale Motion Compensation and Spatiotemporal Context Model”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 8, Aug. 1, 2021, pp. 3182-3196, X… [cited by applicant]
Zhao S., et al., “Global Matching with Overlapping Attention for Optical Flow Estimation”, 2022 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), IEEE, Jun. 18, 2022, pp. 17571-17580, XP034194379, s… [cited by applicant]