IP Library › Granted Patent US 12,177,462
Granted Patent B2
US 12,177,462 · App. 18/337,331 · Granted Dec 24, 2024

Video compression using deep generative models

Inventors: Amirhossein Habibian (Amsterdam, NL); Ties Jehan Van Rozendaal (Amsterdam, NL); Taco Sebastiaan Cohen (Amsterdam, NL)
Assignee: QUALCOMM INCORPORATED
H04N19/20G06F18/21G06N3/044G06N3/045G06N3/047G06N3/084G06V10/764G06V10/82G06V20/46H04N19/124H04N19/14H04N19/179H04N19/186H04N19/46H04N23/90
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,177,462
App. No.
18/337,331
Granted
Dec 24, 2024
Kind
B2
Abstract

Certain aspects of the present disclosure are directed to methods and apparatus for compressing video content using deep generative models. One example method generally includes receiving video content for compression. The received video content is generally encoded into a latent code space through an encoder, which may be implemented by a first artificial neural network. A compressed version of the encoded video content is generally generated through a trained probabilistic model, which may be implemented by a second artificial neural network, and output for transmission.

Claims (68)

1. A method for compressing video, comprising:

receiving video content for compression;

encoding the received video content into a plurality of codes in a latent code space through an encoder implemented by a first artificial neural network, the encoding being based, at least in part, on semantic information derived from the received video content and that includes information indicative of a plurality of segments of the received video content, wherein:

each respective segment of the plurality of segments has different characteristics from other segments of the plurality of segments, and

each respective segment of the plurality of segments is associated with a respective code of the plurality of codes in the latent code space;

generating, based on the plurality of codes associated with the plurality of segments, a compressed version of the encoded video content through a probabilistic model implemented by a second artificial neural network; and

outputting the compressed version of the encoded video content for transmission.

2. The method of claim 1 , wherein the information about the received video content comprises a content mask indicative of an amount of lossy compression to use in compressing semantically different areas of the received video content.

3. The method of claim 2 , wherein the content mask comprises a binary mask distinguishing a first segment and a second segment generated based on a machine learning model trained based on foreground and background content identified in a plurality of training videos.

4. The method of claim 3 , wherein encoding the received video content into the latent code space comprises:

quantizing foreground content using a first amount of compression loss; and

quantizing background content using a second amount of compression loss, wherein the first amount of compression loss is less than the second amount of compression loss.

5. The method of claim 2 , further comprising generating the content mask using a recurrent convolutional neural network, wherein the content mask distinguishes between at least semantically different content in the received video content.

6. The method of claim 5 , further comprising generating information identifying the semantically different content in the received video content based on processing the received video content through a mask generating convolutional neural network.

7. The method of claim 1 , wherein the information about the received video content comprises data from a fixed environment from which the video content was captured.

8. The method of claim 7 , wherein:

the received video content comprises a plurality of video clips captured within the fixed environment, and

encoding the received video content into the latent code space comprises encoding fixed content in the plurality of video clips to a same code in the latent code space.

9. The method of claim 8 , wherein the plurality of video clips captured within the fixed environment comprises video clips of a fixed ambient scene captured by a camera located in a fixed location.

10. The method of claim 8 , wherein the plurality of video clips captured within the fixed environment comprises video clips captured from a fixed vantage point on a moving platform.

11. The method of claim 1 , wherein:

the video content comprises a plurality of channels,

the plurality of channels comprise one or more additional data channels in addition to one or more luminance channels in video content captured by a first camera, and

the information about the received video content comprises correlations between modalities in the plurality of channels.

12. The method of claim 11 , wherein the one or more additional data channels comprise one or more color channels and a depth information channel.

13. The method of claim 11 , wherein the one or more additional data channels comprise one or more channels capturing data within a range of visible wavelengths and one or more channels capturing data outside of the range of visible wavelengths.

14. The method of claim 11 , wherein the video content comprises videos captured of a subject from different perspectives, wherein the videos are captured by the first camera and one or more second cameras.

15. The method of claim 1 , wherein the probabilistic model comprises an auto-regressive model of a probability distribution over four-dimensional tensors, the probability distribution illustrating a likelihood that different codes can be used to compress the encoded video content.

16. The method of claim 15 , wherein the probabilistic model generates data based on a four-dimensional tensor, wherein dimensions of the four-dimensional tensor comprise time, a channel, and spatial dimensions of the received video content.

17. A method for decompressing video, comprising:

receiving a compressed version of an encoded video content, the encoded video content having been encoded based, at least in part, on information about source video content, wherein the information comprises semantic information derived from the received video content and that includes information indicative of a plurality of segments of the source video content, wherein:

each respective segment of the plurality of segments has different characteristics from other segments of the plurality of segments, and

each respective segment of the plurality of segments is associated with a respective code of the plurality of codes in a latent code space;

decompressing the compressed version of the encoded video content into a code in the latent code space through a probabilistic model implemented by a first artificial neural network;

decompressing the code in the latent code space into a reconstruction of the encoded video content through a decoder implemented by a second artificial neural network; and

outputting the reconstruction of the encoded video content.

18. The method of claim 17 , wherein the information about the source video content comprises a content mask indicative of an amount of lossy compression used in compressing semantically different areas of the source video content.

19. The method of claim 18 , wherein the reconstruction of the encoded video content comprises first content reconstructed with a first amount of compression loss and second content reconstructed with a second amount of compression loss, wherein the first amount of compression loss is less than the second amount of compression loss.

20. The method of claim 17 , wherein the information about the source video content comprises data from a fixed environment from which the video content was captured.

21. The method of claim 20 , wherein the encoded video content comprises a plurality of video clips captured within the fixed environment such that fixed content in the plurality of video clips is encoded to a same code in the latent code space.

22. The method of claim 21 , wherein the plurality of video clips captured within the fixed environment comprise video clips of a fixed ambient scene captured by a camera located in a fixed location.

23. The method of claim 21 , wherein the plurality of video clips captured within the fixed environment comprise video clips captured from a fixed vantage point on a moving platform.

24. The method of claim 17 , wherein:

the video content comprises a plurality of channels;

the plurality of channels comprise one or more additional data channels in addition to one or more luminance channels in video content captured by a first camera; and

the information about the source video content comprises correlations between modalities in the plurality of channels.

25. The method of claim 24 , wherein the one or more additional channels comprise one or more color channels and a depth information channel.

26. The method of claim 24 , wherein the one or more additional channels comprise one or more channels capturing data within a range of visible wavelengths and one or more channels capturing data outside of the range of visible wavelengths.

27. The method of claim 24 , wherein the video content comprises videos captured of a subject from different perspectives by the first camera and one or more additional cameras.

28. The method of claim 17 , wherein the probabilistic model comprises an auto-regressive model of a probability distribution over four-dimensional tensors, the probability distribution illustrating a likelihood that different codes can be used to decompress the compressed version of the encoded video content into the code in the latent space, and wherein dimensions of each four-dimensional tensor comprise time, a channel, and spatial dimensions of the source video content.

29. A system for compressing video, comprising:

at least one processor configured to:

receive video content for compression,

encode the received video content into a latent code space through an encoder implemented by a first artificial neural network, the encoding being based, at least in part, on information about the received video content, wherein the information comprises semantic information derived from the received video content and that includes information indicative of a plurality of segments of the received video content, wherein:

each respective segment of the plurality of segments has different characteristics from other segments of the plurality of segments, and

each respective segment of the plurality of segments is associated with a respective code of the plurality of codes in the latent code space,

generate a compressed version of the encoded video content through a probabilistic model implemented by a second artificial neural network, and

output the compressed version of the encoded video content for transmission; and

a memory coupled to the at least one processor.

30. A system for decompressing video, comprising:

at least one processor configured to:

receive a compressed version of an encoded video content, the encoded video content having been encoded based, at least in part, on information about source video content, wherein the information comprises semantic information derived from the received video content and that includes information indicative of a plurality of segments of the source video content, wherein:

each respective segment of the plurality of segments has different characteristics from other segments of the plurality of segments, and

each respective segment of the plurality of segments is associated with a respective code of the plurality of codes in a latent code space,

decompress the compressed version of the encoded video content into a code in the latent code space through a probabilistic model implemented by a first artificial neural network,

decompress the code in the latent code space into a reconstruction of the encoded video content through a decoder implemented by a second artificial neural network, and

output the reconstruction of the encoded video content; and

a memory coupled to the at least one processor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2023
From: HABIBIAN, AMIRHOSSEIN; VAN ROZENDAAL, TIES JEHAN; COHEN, TACO SEBASTIAAN
To: QUALCOMM INCORPORATED
Reel/Frame 064633/0736 →
Continuity (3)
Continuation 16826221 · Mar 21, 2020
Provisional Application 62821778 · Mar 21, 2019
Related Publication 20230336754A1 · Oct 19, 2023