IP Library Granted Patent US 12,192,518
Granted Patent B2
US 12,192,518 · App. 17/443,913 · Granted Jan 7, 2025

Video encoding by providing geometric proxies

Inventors: Michael Hemmer (San Francisco, CA); Ameesh Makadia (New York, NY)
Assignee: GOOGLE LLC
H04N19/597H04N19/186H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,192,518
App. No.
17/443,913
Granted
Jan 7, 2025
Kind
B2
Abstract

Compressing a frame of video includes receiving a frame of a video, identifying a three dimensional (3D) object in the frame, matching the 3D object to a stored 3D object, compressing the frame of the video using a color prediction scheme based on the 3D object and the stored 3D object, and storing the compressed frame with metadata, the metadata identifying the 3D object, indicating a position of the 3D object in the frame of the video and indicating an orientation of the 3D object in the frame of the video.

Claims (145)

1. A method comprising:

receiving a frame of a video;

identifying a three-dimensional (3D) object in the frame;

matching the 3D object to a stored 3D object;

generating a first 3D object proxy based on the stored 3D object;

transforming the first 3D object proxy based on the 3D object identified in the frame;

compressing the frame of the video using a color prediction scheme based on the transformed first 3D object proxy and a transformed second 3D object proxy based on the stored 3D object; and

storing the compressed frame with metadata, the metadata identifying the 3D object, indicating a position of the 3D object in the frame of the video and indicating an orientation of the 3D object in the frame of the video.

2. The method of claim 1 , wherein the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

generating the second 3D object proxy based on the stored 3D object;

identifying the 3D object in a key frame of the video;

transforming the second 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating residuals for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

3. The method of claim 1 , wherein the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

generating the second 3D object proxy based on the stored 3D object;

identifying the 3D object in a key frame of the video;

transforming the second 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy; and

generating residuals for the 3D object based on the color attributes for the transformed first 3D object proxy and default color attributes for the transformed second 3D object proxy.

4. The method of claim 1 , wherein the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an auto encoder, wherein the transforming the first 3D object proxy includes transforming the encoded first 3D object proxy based on the 3D object identified in the frame;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating residuals for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

5. The method of claim 1 , wherein the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an auto encoder, wherein the transforming the first 3D object proxy includes transforming the encoded first 3D object proxy based on the 3D object identified in the frame;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy; and

generating residuals for the 3D object based on the color attributes for the transformed first 3D object proxy and default color attributes for the transformed second 3D object proxy.

6. The method of claim 1 , further comprising:

before storing the 3D object:

identifying at least one 3D object of interest associated with the video;

determining a plurality of mesh attributes associated with the 3D object of interest;

determining a position associated with the 3D object of interest;

determining an orientation associated with the 3D object of interest;

determining a plurality of color attributes associated with the 3D object of interest; and

reducing a number of variables associated with the mesh attributes for the 3D object of interest using an autoencoder.

7. The method of claim 1 , wherein compressing the frame of the video includes determining position coordinates of the 3D object relative to an origin coordinate of a background 3D object in a key frame.

8. The method of claim 1 , wherein

the stored 3D object includes default color attributes, and

the color prediction scheme uses the default color attributes.

9. The method of claim 1 , further comprising:

identifying at least one 3D object of interest associated with the video;

generating at least one stored 3D object based on the at least one 3D object of interest, each of the at least one stored 3D object being defined by a mesh including a collection of points connected by faces, each point storing at least one attribute, the at least one attribute including a position coordinate for the respective point; and

storing the at least one stored 3D object in association with the video.

10. A method comprising:

receiving a frame of a video;

identifying a three-dimensional (3D) object in the frame;

matching the 3D object to a stored 3D object;

generating a 3D object proxy based on the stored 3D object;

transforming the 3D object proxy based on the 3D object identified in the frame;

mapping color attributes from the 3D object to the transformed 3D object proxy;

compressing the frame of the video using a color prediction scheme based on the transformed 3D object proxy, a second transformed 3D object proxy and the stored 3D object; and

storing the frame of the video.

11. The method of claim 10 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

generating the second 3D object proxy based on the stored 3D object;

identifying the 3D object in a key frame of the video;

transforming the second 3D object proxy as the second transformed 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

12. The method of claim 10 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an autoencoder;

transforming the encoded first 3D object proxy based on metadata associated with the 3D object;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy as the second transformed 3D object proxy based on metadata associated with the 3D object identified in the key frame;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

13. The method of claim 10 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an autoencoder;

transforming the encoded first 3D object proxy based on metadata associated with the 3D object;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy as the second transformed 3D object proxy based on metadata associated with the 3D object identified in the key frame; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and default attributes for the transformed second 3D object proxy.

14. The method of claim 10 , further comprising:

receiving at least one latent representation for a 3D shape; and

using a machine trained generative modeling technique to:

determine a plurality of mesh attributes associated with the 3D shape;

determining a position associated with the 3D shape;

determining an orientation associated with the 3D shape; and

determining a plurality of color attributes associated with the 3D shape; and

storing the 3D shape as the stored 3D object.

15. The method of claim 10 , further comprising:

receiving a neural network used by an encoder of an autoencoder to reduce a number of variables associated with mesh attributes, position, orientation and color attributes for at least one 3D object of interest;

generating points associated with a mesh for the at least one 3D object of interest using the neural network in an encoder of the autoencoder, the generation of the points including generating position attributes, orientation attributes and color attributes; and

storing the at least one 3D object of interest as the stored 3D object.

16. A method comprising:

receiving a frame of a video;

identifying a three-dimensional (3D) object in the frame;

matching the 3D object to a stored 3D object;

generating a 3D object proxy based on the stored 3D object;

encoding the 3D object proxy using an autoencoder;

transforming the encoded 3D object proxy as a first transformed 3D object proxy based on metadata associated with the 3D object;

compressing the frame of the video using a color prediction scheme based on the transformed 3D object proxy, a second transformed 3D object proxy, and the stored 3D object; and

storing the frame of the video.

17. The method of claim 16 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

generating the second 3D object proxy based on the stored 3D object;

identifying the 3D object in a key frame of the video;

transforming the second 3D object proxy as the second transformed 3D object proxy based on the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

18. The method of claim 16 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an autoencoder;

transforming the encoded first 3D object proxy based on metadata associated with the 3D object;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy as the second transformed 3D object proxy based on metadata associated with the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy;

mapping color attributes from the 3D object identified in the key frame to the transformed second 3D object proxy; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and the color attributes for the transformed second 3D object proxy.

19. The method of claim 16 , wherein the 3D object proxy is a first 3D object proxy and the compressing of the frame of the video using the color prediction scheme based on the 3D object and the stored 3D object includes:

encoding the first 3D object proxy using an autoencoder;

transforming the encoded first 3D object proxy based on metadata associated with the 3D object;

generating the second 3D object proxy based on the stored 3D object;

encoding the second 3D object proxy using an autoencoder;

identifying the 3D object in a key frame of the video;

transforming the encoded second 3D object proxy as the second transformed 3D object proxy based on metadata associated with the 3D object identified in the key frame;

mapping color attributes from the 3D object to the transformed first 3D object proxy; and

generating color attributes for the 3D object based on the color attributes for the transformed first 3D object proxy and default attributes for the transformed second 3D object proxy.

20. The method of claim 16 , further comprising:

receiving at least one latent representation for a 3D shape; and

using a machine trained generative modeling technique to:

determine a plurality of mesh attributes associated with the 3D shape;

determining a position associated with the 3D shape;

determining an orientation associated with the 3D shape; and

determining a plurality of color attributes associated with the 3D shape; and

storing the 3D shape as the stored 3D object.

21. The method of claim 16 , further comprising:

receiving a neural network used by an encoder of an autoencoder to reduce a number of variables associated with mesh attributes, position, orientation and color attributes for at least one 3D object of interest;

generating points associated with a mesh for the at least one 3D object of interest using the neural network in an encoder of the autoencoder, the generation of the points including generating position attributes, orientation attributes and color attributes; and

storing the at least one 3D object of interest as the stored 3D object.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 4, 2021
From: HEMMER, MICHAEL; MAKADIA, AMEESH
To: GOOGLE LLC
Reel/Frame 057072/0866 →
Continuity (2)
Division 16143165 · Sep 26, 2018
Related Publication 20210360286A1 · Nov 18, 2021
References Cited (43)
US 5818463A · Tao et al. · 1998 [cited by applicant]
US 5880737A · Griffin et al. · 1999 [cited by applicant]
US 6518965B2 · Dye et al. · 2003 [cited by applicant]
US 6593925B1 · Hakura et al. · 2003 [cited by applicant]
US 8942283B2 · Pace · 2015 [cited by applicant]
US 9106977B2 · Pace · 2015 [cited by applicant]
US 10192115B1 · Sheffield et al. · 2019 [cited by applicant]
US 20050063596A1 · Yomdin et al. · 2005 [cited by applicant]
US 20050141762A1 · Zhao et al. · 2005 [cited by applicant]
US 20100303150A1 · Hsiung et al. · 2010 [cited by applicant]
US 20110032987A1 · Lee et al. · 2011 [cited by applicant]
US 20110052045A1 · Kameyama · 2011 [cited by examiner]
US 20110058609A1 · Chaudhury et al. · 2011 [cited by applicant]
US 20120170809A1 · Picazo · 2012 [cited by applicant]
US 20140198182A1 · Ward et al. · 2014 [cited by applicant]
US 20170344850A1 · Kobori · 2017 [cited by examiner]
US 20180173826A1 · White · 2018 [cited by examiner]
US 20190362551A1 · Sheffield et al. · 2019 [cited by applicant]
US 20200145661A1 · Jeon et al. · 2020 [cited by applicant]
CN 101622874A · 2010 [cited by applicant]
EP 1434171A2 · 2004 [cited by applicant]
WO 2015177162A1 · 2015 [cited by applicant]
Thanou et al., “Graph-Based Compression of Dynamic 3D Point Cloud Sequences”, IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US vol. 25, No. 4, Apr. 1, 2016; 14 pages (Year: 2016). [cited by examiner]
Q. Tan, L. Gao, Y.-K. Lai and S. Xia, “Variational Autoencoders for Deforming 3D Mesh Models,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, USA, 2018, pp. 5841-5850, doi: 10.1… [cited by examiner]
U.S. Appl. No. 13/394,486, filed Jun. 1, 2010, Picazo. [cited by applicant]
U.S. Appl. No. 12/805,455, filed Jul. 30, 2010, Lee, et al. [cited by applicant]
U.S. Appl. No. 12/896,051, filed Oct. 1, 2010, Kameyama. [cited by applicant]
U.S. Appl. No. 15/591,690, filed May 10, 2017, Kobori, et al. [cited by applicant]
U.S. Appl. No. 15/990,429, filed May 25, 2018, Sheffield, et al. [cited by applicant]
International Search Report and Written Opinion for Application No. PCT/US2019/051566, mailed on Mar. 3, 2020, 17 pages. [cited by applicant]
Kundu et al.; “3D-RCNN: Instance-Level 3D Object Reconstuction Via Render-And-Compare”; 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018; pp. 3559-3568. [cited by applicant]
Song et al.; “IM2PANO3D: Extrapolating 360 Structure and Semantics Beyond the Field of View”; 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 18, 2018, 18 pages. [cited by applicant]
Thanou et al.; “Graph-Based Compression of Dynamic 3D Point Cloud Sequences”; IEEE Transactions on Image Processing, IEEE Service Center, Piscataway, NJ, US, vol. 25, No. 4, Apr. 1, 2016; 14 pages. [cited by applicant]
Yang et al.; “Content-Based 3-D Model Retrieval: A Survey”; IEEE Transactions on Systems, Man, and Cybernetics: Part C: Applications and Reviews, IEEE Service Center, Piscataway, NJ, US; vol. 37, No. 6; Nov. 1, 2007; pp… [cited by applicant]
Fathi, Alireza , et al., “Semantic Instance Segmentation via Deep Metric Learning”, Arxiv.org (https://arxiv.org/abs/1703.10277), Mar. 30, 2017, 9 pages. [cited by applicant]
Goodfellow, Ian J., et al., “Generative Adversarial Nets”, Arxiv.org (https://arxiv.org/abs/1406.2661), Jun. 10, 2014, 9 pages. [cited by applicant]
Kingman, Diederik P., et al., “Auto-Encoding Variational Bayes”, Arxiv.org (https://arxiv.org/abs/1312.6114), May 1, 2014, 14 pages. [cited by applicant]
Litany, Or , et al., Arxiv.org (https://arxiv.org/abs/1712.00268), Apr. 3, 2018, 10 pages. [cited by applicant]
Liu, Shikun , et al., “Learning a Hierarchical Latent-Variable Model of 3D Shapes”, Arxiv.org (https://arxiv.org/pdf/1712.00268.pdf), Aug. 4, 2018, 10 pages. [cited by applicant]
Muller, Karsten , et al., “3D Video Formats and Coding Methods”, 2010 IEEE International Conference On Image Processing, ICIP, Sep. 26, 2010, pp. 2389-2392. [cited by applicant]
Tan, Qingyang , et al., “Variational Autoencoders for Deforming 3D Mesh Models”, Arxiv.org (https://arxiv.org/abs/1709.04307), Mar. 29, 2018, 10 pages. [cited by applicant]
Babu, et al., “Object-based Surveillance Video Compression using Foreground Motion Compensation”, 2006 9th International Conference on Control, Automation, Robotics and Vision, pp. 1-6, Jul. 16, 2007. [cited by applicant]
Han, “Research on Adaptive Extraction of Video Objects and Video Compression Coding”, China Doctoral Dissertations Full-text Database (Information Science and Technology), No. 06, pp. 1136-1140, Jun. 15, 2005. [cited by applicant]
Cited By (1)
US 12,555,272