IP Library › Granted Patent US 12,432,389
Granted Patent B2
US 12,432,389 · App. 18/563,734 · Granted Sep 30, 2025

Video compression using optical flow

Inventors: George Dan Toderici (Mountain View, CA); Eirikur Thor Agustsson (Zurich, CH); Fabian Julius Mentzer (Zurich, CH); David Charles Minnen (Mountain View, CA); Johannes Balle (San Francisco, CA); Nicholas Johnston (San Jose, CA)
Assignee: Google LLC
H04N19/91G06T3/18G06T5/70H04N19/124H04N19/137G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,432,389
App. No.
18/563,734
Granted
Sep 30, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for compressing video data. In one aspect, a method comprises: receiving a video sequence of frames; generating, using a flow prediction network, an optical flow between two sequential frames, wherein the two sequential frames comprise a first frame and a second frame that is subsequent the first frame; generating from the optical flow, using a first autoencoder neural network: a predicted optical flow between the first frame and the second frame; and warping a reconstruction of the first frame according to the predicted optical flow and subsequently applying a blurring operation to obtain an initial predicted reconstruction of the second frame.

Claims (76)

1. A method of compressing video performed by a data processing apparatus, comprising:

receiving a video sequence of frames;

generating, using a flow prediction network, an optical flow between two sequential frames, wherein the two sequential frames comprise a first frame and a second frame that is subsequent the first frame;

generating from the optical flow, using a first autoencoder neural network:

a predicted optical flow between the first frame and the second frame; and

a confidence mask that comprises a plurality of confidence values characterizing uncertainty in the predicted optical flow between the first frame and the second frame;

warping a reconstruction of the first frame according to the predicted optical flow and subsequently applying a blurring operation according the confidence mask to obtain an initial predicted reconstruction of the second frame;

generating, using a second autoencoder neural network, a prediction of a residual that is a difference between the second frame and the initial predicted reconstruction of the second frame;

combining the initial predicted reconstruction of the second frame and the prediction of the residual to obtain a predicted second frame;

wherein:

each of the first and second autoencoder neural networks respectively comprise an encoder network and a generator network; and

the generator network of the second autoencoder neural network is a component of a generative adversarial neural network (GANN).

2. The method of claim 1 , wherein:

the first frame and the second frame are subsequent to a third frame, and wherein the third frame is an initial frame in the video sequence; and

further comprising, prior to processing the second and third frames:

generating from the third frame, using a third autoencoder neural network, a predicted reconstruction of the third frame;

generating, using the flow prediction network, an optical flow between third frame and the first frame;

generating from the optical flow, using the first autoencoder neural network:

a predicted optical flow between the third frame and the first frame; and

a confidence mask;

warping the reconstruction of the third frame according to the predicted optical flow and subsequently applying a blurring operation according the confidence mask to obtain an initial predicted reconstruction of the first frame;

generating, using the second autoencoder neural network, a prediction of a residual that is a difference between the first frame and the initial predicted reconstruction of the first frame; and

combining the initial predicted reconstruction of the first frame and the prediction of the residual to obtain a predicted first frame;

wherein:

the third autoencoder neural network comprises an encoder network and a generator network;

the third generator network of the third autoencoder neural network is a component of a generative adversarial neural network (GANN).

3. The method of claim 1 , further comprising:

encoding, using the second autoencoder neural network, a residual to obtain a residual latent;

obtaining, using the third encoder neural network, a free latent by encoding the initial prediction of the second frame; and

concatenating the free latent and the residual latent;

wherein generating, using the second autoencoder neural network, the prediction of the residual comprises generating the predicted residual by the second autoencoder neural network using the concatenation of the free latent and the residual latent.

4. The method of claim 3 , further comprising entropy encoding a quantization of the residual latent, wherein the entropy encoded quantization of the residual latent is included in compressed video data representing the video.

5. The method of claim 3 , wherein encoding the residual to obtain the residual latent comprises:

processing the residual using the encoder neural network of the second autoencoder neural network to generate the residual latent.

6. The method of claim 3 , wherein obtaining the free latent by encoding the initial prediction of the second frame comprises:

processing the initial prediction of the second frame using an encoder neural network to generate the free latent.

7. The method of claim 3 , wherein generating the prediction of the residual comprises:

processing the concatenation of the free latent and the residual latent using the generator neural network of the second autoencoder neural network to generate the prediction of the residual.

8. The method of claim 1 , wherein combining the initial predicted reconstruction of the second frame and the prediction of the residual to obtain the predicted second frame comprises:

generating the predicted second frame by summing the initial predicted reconstruction of the second frame and the prediction of the residual.

9. The method of claim 1 , wherein generating the predicted optical flow between the first frame and the second frame comprises:

processing the optical flow generated by the flow prediction network using the encoder network of the first autoencoder network to generate a flow latent representing the optical flow; and

processing a quantization of the flow latent using the generator neural network of the first autoencoder neural network to generate the predicted optical flow.

10. The method of claim 9 , further comprising entropy encoding the quantization of the flow latent, wherein the entropy encoded quantization of the flow latent is included in compressed video data representing the video.

11. The method of claim 1 , wherein the first and second autoencoder neural networks have been trained on a set of training videos to optimize an objective function that includes an adversarial loss.

12. The method of claim 11 , wherein for one or more video frames of each training video, the adversarial loss is based on a discriminator score, wherein the discriminator score is generated by operations comprising:

generating an input to a discriminator neural network, wherein the input comprises a reconstruction of the video frame that is generated using the first and second autoencoder neural networks; and

providing the input to the discriminator neural network, wherein the discriminator neural network is configured to:

receive an input comprising an input video frame; and

process the input to generate an output discriminator score defining a likelihood that the video frame was generated using the first and second autoencoder neural networks.

13. A non-transitory computer storage medium encoded with a computer program, the program comprising instructions that when executed by data processing apparatus cause the data processing apparatus to perform operations for compressing video, the operations comprising:

receiving a video sequence of frames;

generating, using a flow prediction network, an optical flow between two sequential frames, wherein the two sequential frames comprise a first frame and a second frame that is subsequent the first frame;

generating from the optical flow, using a first autoencoder neural network:

a predicted optical flow between the first frame and the second frame; and

a confidence mask that comprises a plurality of confidence values characterizing uncertainty in the predicted optical flow between the first frame and the second frame;

warping a reconstruction of the first frame according to the predicted optical flow and subsequently applying a blurring operation according the confidence mask to obtain an initial predicted reconstruction of the second frame;

generating, using a second autoencoder neural network, a prediction of a residual that is a difference between the second frame and the initial predicted reconstruction of the second frame;

combining the initial predicted reconstruction of the second frame and the prediction of the residual to obtain a predicted second frame;

wherein:

each of the first and second autoencoder neural networks respectively comprise an encoder network and a generator network; and

the generator network of the second autoencoder neural network is a component of a generative adversarial neural network (GANN).

14. A system, comprising:

a data processing apparatus; and

a computer storage medium encoded with a computer program, the program comprising instructions that when executed by the data processing apparatus cause the data processing apparatus to perform operations for compressing video, the operations comprising:

receiving a video sequence of frames;

generating, using a flow prediction network, an optical flow between two sequential frames, wherein the two sequential frames comprise a first frame and a second frame that is subsequent the first frame;

generating from the optical flow, using a first autoencoder neural network:

a predicted optical flow between the first frame and the second frame; and

a confidence mask that comprises a plurality of confidence values characterizing uncertainty in the predicted optical flow between the first frame and the second frame;

warping a reconstruction of the first frame according to the predicted optical flow and subsequently applying a blurring operation according the confidence mask to obtain an initial predicted reconstruction of the second frame;

generating, using a second autoencoder neural network, a prediction of a residual that is a difference between the second frame and the initial predicted reconstruction of the second frame;

combining the initial predicted reconstruction of the second frame and the prediction of the residual to obtain a predicted second frame;

wherein:

each of the first and second autoencoder neural networks respectively comprise an encoder network and a generator network; and

the generator network of the second autoencoder neural network is a component of a generative adversarial neural network (GANN).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 30, 2023
From: TODERICI, GEORGE DAN; AGUSTSSON, EIRIKUR THOR; MENTZER, FABIAN JULIUS; MINNEN, DAVID CHARLES; BALLE, JOHANNES; JOHNSTON, NICHOLAS
To: GOOGLE LLC
Reel/Frame 065714/0621 →
Continuity (2)
Provisional Application 63218853 · Jul 6, 2021
Related Publication 20240223817A1 · Jul 4, 2024
References Cited (73)
US 1399531A · Walter · 1921 [cited by examiner]
US 20100194741A1 · Finocchio · 2010 [cited by examiner]
US 20150365696A1 · Garud · 2015 [cited by examiner]
US 20160292826A1 · Beall · 2016 [cited by examiner]
US 20170124433A1 · Chandraker · 2017 [cited by examiner]
US 20190164296A1 · Chikkerur · 2019 [cited by examiner]
US 20200053388A1 · Schroers · 2020 [cited by examiner]
US 20210044804A1 · Lu et al. · 2021 [cited by applicant]
US 20210073589A1 · Orhon · 2021 [cited by examiner]
US 20220183208A1 · Sibley · 2022 [cited by examiner]
US 20240161312A1 · Jeong · 2024 [cited by examiner]
Agustsson et al., “Generative Adversarial Networks for Extreme Learned Image Compression,” Presented at Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, Oct. 27-Nov. 2, 2019,… [cited by applicant]
Agustsson et al., “Scale-Space Flow for End-to-End Optimized Video Compression,” Presented at 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seattle, Washington, Jun. 13-19, 2020, pp. 8503-8… [cited by applicant]
Ahmed et al., “Discrete Cosine Transform,” IEEE Transactions on Computers, Jan. 1974, C-23(1):90-93. [cited by applicant]
Aitken et al., “Checkerboard artifact free sub-pixel convolution: A note on sub-pixel convolution, resize convolution and convolution resize,” CoRR, Submitted on Jul. 10, 2017, arXiv:1707.02937v1, 16 pages. [cited by applicant]
Ballé et al., “Variational image compression with a scale hyperprior,” Presented at the International Conference on Learning Representations (ICLR), Vancouver, BC, Canada, Apr. 30-May 3, 2018, 47 pages. [cited by applicant]
Bankoski et al., “Technical overview of VP8, an open source video codec for the web,” Presented at 2011 IEEE International Conference on Multimedia and Expo, Barcelona, Spain, Jul. 11-15, 2011, pp. 1-6. [cited by applicant]
Bellard, “BPG Image format,” Apr. 21, 2018, retrieved on Dec. 10, 2024, retrieved from URL <https://bellard.org/bpg/>, 2 pages. [cited by applicant]
Bhardwaj et al., “An unsupervised information-theoretic perceptual quality metric,” CoRR, Submitted on Jan. 10, 2021, arXiv:2006.06752v3, 19 pages. [cited by applicant]
Blau et al., “Rethinking Lossy Compression: The Rate-Distortion-Perception Tradeoff,” CoRR, Submitted on Jul. 30, 2019, arXiv:1901.07821v4, 21 pages. [cited by applicant]
Blau et al., “The Perception-Distortion Tradeoff,” Presented at Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Salt Lake City, Utah, Jun. 18-12, 2018, pp. 6228-6237. [cited by applicant]
Chen et al., “An Overview of Core Coding Tools in the AV1 Video Codec,” Presented at 2018 Picture Coding Symposium (PCS), San Francisco, CA, USA, Jun. 24-27, 2018, pp. 41-45. [cited by applicant]
cloud.google.com [online], “AI Platform Data Labeling Service pricing,” upon information and belief, available no later than Jul. 6, 2021, retrieved on Dec. 13, 2024, retrieved from URL <https://cloud.google.com/ai-plat… [cited by applicant]
developers.google.com [online], “VideoCategories,” Available on or before Jun. 4, 2020, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20200604194442/https://developers.google.com/youtube/v3/docs… [cited by applicant]
Djelouah et al., “Neural Inter-Frame Compression for Video Coding,” Presented at Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul Korea, Oct. 27-Nov. 2, 2019, pp. 6421-6429. [cited by applicant]
github.com [online], “Netflix/vmaf—Video Multi-Method Assessment Fusion,” Feb. 22, 2017, retrieved on Dec. 13, 2024, retrieved from URL <https://github.com/Netflix/vmaf/>, 4 pages. [cited by applicant]
Golinski et al., “Feedback Recurrent Autoencoder for Video Compression,” Presented at the Proceedings of the Asian Conference on Computer Vision (ACCV), Kyoto, Japan, Nov. 30-Dec. 4, 2020; Revised Selected Papers, Part … [cited by applicant]
Goodfellow et al., “Generative Adversarial Nets,” Presented at the Annual Conference on Neural Information Processing Systems, Montreal, Quebec, Canada, Dec. 8-13, 2024, 27:2672-2680. [cited by applicant]
Goyal, “Theoretical foundations of transform coding,” IEEE Signal Processing Magazine, Sep. 2001, 18(5):9-21. [cited by applicant]
Guo et al., “Dvc: An end-to-end deep video compression framework,” Presented at the Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Long Beach, California, Jun. 15-20, 2019, pp. 11006-… [cited by applicant]
Habibian et al., “Video Compression With Rate-Distortion Autoencoders,” Presented at Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Seoul, Korea, Oct. 24-Nov. 2, 2019, pp. 7033-7042. [cited by applicant]
Heusel et al., “GANs Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium,” Presented at the Thirty-First Annual Conference on Neural Information Processing Systems (NIPS), Long Beach, California… [cited by applicant]
International Preliminary Report on Patentability in International Appln. No. PCT/US2022/036111, mailed on Jan. 18, 2024, 13 pages. [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/US2022/036111, mailed on Oct. 19, 2022, 17 pages. [cited by applicant]
itu.int [online], “H.264 : Advanced video coding for generic audiovisual services,” Available on or before Jul. 3, 2021, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20210703024340/https://www.… [cited by applicant]
itu.int [online], “H.265 : High efficiency video coding ,” Available on or before Oct. 29, 2020, via Internet Archive: Wayback Machine URL <https://web.archive.org/web/20201029050752/https://www.itu.int/rec/T-REC-H.265/… [cited by applicant]
Jonschkowski et al., “What Matters in Unsupervised Optical Flow,” CoRR, Submitted on Aug. 14, 2020, arXiv:2006.04902v2, 16 pages. [cited by applicant]
Krizhevsky et al., “ImageNet Classification with Deep Convolutional Neural Networks,” Presented at the Twenty-Sixth Annual Conference on Neural Information Processing Systems (NIPS), Lake Tahoe, Nevada, Dec. 3-8, 2012, … [cited by applicant]
Ladune et al., “Optical Flow and Mode Selection for Learning-based Video Coding,” CoRR, Submitted on Aug. 6, 2020, arXiv:2008.02580v1, 6 pages. [cited by applicant]
Li et al., “Deep Contextual Video Compression,” Presented at the Thirty-Fifth Annual Conference on Neural Information Processing Systems (NIPS), Virtual, Dec. 6-14, 2021, 34:18114-18125. [cited by applicant]
Liu et al., “Conditional Entropy Coding for Efficient Video Compression,” CoRR, Submitted on Aug. 20, 2020, arXiv:2008.09180v1, 23 pages. [cited by applicant]
Liu et al., “Neural Video Compression using Spatio-Temporal Priors,” CoRR, Submitted on Feb. 21, 2019, arXiv:1902.07383v2, 5 pages. [cited by applicant]
Mentzer et al., “High-Fidelity Generative Image Compression,” CoRR, Submitted on Jul. 10, 2020, arXiv:2006.09965v2, 18 pages. [cited by applicant]
Mentzer et al., “Towards Generative Video Compression,” CoRR, Submitted on Jul. 26, 2021, arXiv:2107.12038v1, 19 pages. [cited by applicant]
Mercat et al., “UVG dataset: 50/120fps 4K sequences for video codec analysis and development,” Presented at the MMSys '20: Proceedings of the 11th ACM Multimedia Systems Conference, Istanbul, Turkey, Jun. 8-11, 2020, pp… [cited by applicant]
Minnen et al., “Channel-Wise Autoregressive Entropy Models for Learned Image Compression,” CoRR, Submitted on Jul. 17, 2020, arXiv:2007.08739v1, 16 pages. [cited by applicant]
Minnen et al., “Joint Autoregressive and Hierarchical Priors for Learned Image Compression,” Presented at the 32nd Conference on Neural Information Processing Systems (NeurIPS 2018), Montreal, Canada, Dec. 2-8, 2018, 31… [cited by applicant]
Mirza et al., “Conditional Generative Adversarial Nets,” CoRR, Submitted on Nov. 6, 2014, arXiv:1411.1784v1, 7 pages. [cited by applicant]
Mukherjee et al., “The latest open-source video codec VP9—An overview and preliminary results,” Presented at the 2013 Picture Coding Symposium (PCS), San Jose, California, Dec. 8-11, 2013, pp. 390-393. [cited by applicant]
Nehab et al., “A Fresh Look at Generalized Sampling,” Now Foundations and Trends, 2012, 87 pages. [cited by applicant]
Odena et al., “Deconvolution and Checkerboard Artifacts,” Distill, Oct. 17, 2016, 1(10):e3, 10 pages. [cited by applicant]
Platt et al., “Constrained Differential Optimization for Neural Networks,” California Institute of Technology, Jan. 1988, 11 pages. [cited by applicant]
Rippel et al., “ELF-VC: Efficient Learned Flexible-Rate Video Coding,” CoRR, Submitted on Apr. 29, 2021, arXiv:2104.14335v1, 14 pages. [cited by applicant]
Rippel et al., “Learned Video Compression,” Presented at the Proceedings of the IEEE/CVF International Conference on Computer Vision, Seoul, Korea, Oct. 27-Nov. 2, 2019; IEEE Xplore, Feb. 27, 2020, pp. 3454-3463. [cited by applicant]
Rippel et al., “Real-time adaptive image compression,” CoRR, Submitted on May 16, 2017, arXiv:1705.05823v1, 16 pages. [cited by applicant]
Santurkar et al., “Generative compression,” CoRR, Submitted on Jun. 4, 2017, arXiv:1703.01467v2, 10 pages. [cited by applicant]
Shulman et al., “Regularization of discontinuous flow fields,” [1989] Proceedings Workshop on Visual Motion, Mar. 20-22, 1989, pp. 81-86. [cited by applicant]
Theis et al., “A coding theorem for the rate-distortion-perception function,” CoRR, Submitted on Apr. 28, 2021, arXiv:2104.13662v1, 5 pages. [cited by applicant]
Theis et al., “On the advantages of stochastic encoders,” CoRR, Submitted on Apr. 29, 2021, arXiv:2102.09270v2, 8 pages. [cited by applicant]
Tschannen et al., “Deep generative models for distributionpreserving lossy compression,” CoRR, Submitted on Oct. 28, 2018, arXiv:1805.11057v2, 27 pages. [cited by applicant]
Van Rozendaal et al., “Lossy compression with distortion constrained optimization,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, 2020, pp. 166-167. [cited by applicant]
Wang et al., “MCL-JCV: a JND-based H.264/AVC video quality assessment dataset,” 2016 IEEE International Conference on Image Processing (ICIP), Sep. 25-28, 2016, pp. 1509-1513. [cited by applicant]
Wang et al., “Multiscale structural similarity for image quality assessment,” Proceedings of the 37th IEEE Asilomar Conference on Signals, Systems and Computers, Pacific Grove, CA, Nov. 9-12, 2003. [cited by applicant]
Wang et al., “Understanding Convolution for Semantic Segmentation,” Presented as the 2018 IEEE Winter Conference on Applications of Computer Vision (WACV), Lake Tahoe, Nevada, Mar. 12-15, 2018, pp. 1451-1460. [cited by applicant]
Wang et al., “Video-to-video synthesis,” Advances in Neural Information Processing Systems (NeurIPS), 2018, 13 pages. [cited by applicant]
Wu et al., “Video Compression through Image Interpolation,” Presented at Proceedings of the European Conference on Computer Vision (ECCV), Germany, Munich, Sep. 8-14, 2018, pp. 416-431. [cited by applicant]
Yang et al., “An introduction to neural data compression,” CoRR, Submitted on Feb. 14, 2022, arXiv:2202.06533v1, 90 pages. [cited by applicant]
Yang et al., “Learning for video compression with recurrent auto-encoder and recurrent probability model,” CoRR, Submitted on Dec. 6, 2020, arXiv:2006.13560v4, 17 pages. [cited by applicant]
Yang, “Hierarchical Autoregressive Modeling for Neural Video Compression,” CoRR, Submitted on May 4, 2021, arXiv:2010.10258v2, 15 pages. [cited by applicant]
Zhang et al., “The unreasonable effectiveness of deep features as a perceptual metric,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2018, pp. 586-595. [cited by applicant]
Zhang, “Making convolutional networks shift-invariant again,” CoRR, Submitted on Jun. 9, 2019, arXiv:1904.11486v2, 17 pages. [cited by applicant]
Office Action issued in Japanese Appln. No. 2024-500559, mailed on Feb. 25, 2025, 10 pages (with machine translation). [cited by applicant]
Pourreza et al., “Extending Neural P-frame Codecs for B-frame Coding,” CoRR, Submitted on Mar. 30, 2021, arXiv:2104.00531v1, 16 pages. [cited by applicant]