IP Library Granted Patent US 12,243,145
Granted Patent B2
US 12,243,145 · App. 17/927,101 · Granted Mar 4, 2025

Re-timing objects in video via layered neural rendering

Inventors: Forrester H. Cole (Cambridge, MA); Erika Lu (Lexington, MA); Tali Dekel (Arlington, MA); William T. Freeman (Acton, MA); David Henry Salesin (Saualito, CA); Michael Rubinstein (Natick, MA)
Assignee: GOOGLE LLC
G06T13/80G06V10/454G06V10/82G06V20/46G06V20/49G11B27/005G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,145
App. No.
17/927,101
Granted
Mar 4, 2025
Kind
B2
Abstract

A computer-implemented method for decomposing videos into multiple layers ( 212, 213 ) that can be re-combined with modified relative timings includes obtaining video data including a plurality of image frames ( 201 ) depicting one or more objects. For each of the plurality of frames, the computer-implemented method includes generating one or more object maps descriptive of a respective location of at least one object of the one or more objects within the image frame. For each of the plurality of frames, the computer-implemented method includes inputting the image frame and the one or more object maps into a machine-learned layer Tenderer model. ( 220 ) For each of the plurality of frames, the computer-implemented method includes receiving, as output from the machine-learned layer Tenderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps. The object layers include image data illustrative of the at least one object and one or more trace effects at least partially attributable to the at least one object such that the one or more object layers and the background layer can be re-combined with modified relative timings.

Claims (52)

1. A computer-implemented method for decomposing videos into multiple layers that can be individually retimed and re-combined with modified relative timings, the computer-implemented method comprising:

obtaining, by a computing system comprising one or more computing devices, video data, the video data comprising a plurality of image frames depicting one or more objects; and

for each of the plurality of image frames:

generating, by the computing system, one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame;

inputting, by the computing system, the image frame and the one or more object maps into a machine-learned layer renderer model, comprising iteratively individually inputting each of the one or more object maps into the machine-learned layer renderer model;

receiving, by the computing system as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps, wherein each of the one or more object layers comprises image data illustrative of the at least one object and one or more trace effects at least partially attributable to the at least one object; and

generating, by the computing system, a retimed video by:

retiming at least one of the background layer or the one or more object layers; and

re-combining the one or more retimed lavers.

2. The computer-implemented method of claim 1 , wherein inputting, by the computing system, the image frame and the one or more object maps into the machine-learned layer renderer model comprises iteratively individually receiving, as output from the machine-learned layer renderer model and by the computing system, each of the one or more object layers respective to the one or more object maps.

3. The computer-implemented method of claim 1 , wherein the background layer and the one or more object layers comprise one or more color channels and an opacity matte.

4. The computer-implemented method of claim 1 , wherein the machine-learned layer renderer model comprises a neural network.

5. The computer-implemented method of claim 1 , wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.

6. The computer-implemented method of claim 5 , wherein the training was performed on downsampled video and then upsampled.

7. The computer-implemented method of claim 1 , wherein the one or more object maps comprise one or more texture maps.

8. The computer-implemented method of claim 1 , wherein the one or more object maps comprise one or more re-sampled texture maps.

9. The computer-implemented method of claim 8 , wherein obtaining, by the computing system, one or more object maps comprises:

obtaining, by the computing system, one or more UV maps, each of the UV maps indicative of the at least one object of the one or more objects depicted within the one or more frames;

obtaining, by the computing system, a background deep texture map and one or more object deep texture maps; and

resampling, by the computing system, the one or more object deep texture maps based at least in part on the one or more UV maps.

10. The computer-implemented method of claim 9 , wherein generating, by the computing system, the one or more UV maps comprises:

identifying, by the computing system, one or more keypoints; and

obtaining, by the computing system, one or more UV maps based on the one or more keypoints.

11. The computer-implemented method of claim 1 , further comprising:

transferring, by the computing system, high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.

12. A computing system configured to decompose video data into a plurality of layers, the computing system comprising:

one or more processors; and

one or more non-transitory computer-readable media that store instructions that, when executed by the one or more processors, cause the computing system to perform operations, the operations comprising:

obtaining video data, the video data comprising a plurality of image frames depicting one or more objects; and

for each of the plurality of image frames:

generating one or more object maps, wherein each of the one or more object maps is descriptive of a respective location of at least one object of the one or more objects within the image frame;

inputting the image frame and the one or more object maps into a machine-learned layer renderer model, comprising iteratively individually inputting each of the one or more object maps into the machine-learned layer renderer model;

receiving, as output from the machine-learned layer renderer model, a background layer illustrative of a background of the video data and one or more object layers respectively associated with one of the one or more object maps, wherein each of the one or more object layers comprises image data illustrative of the at least one object and one or more trace effects at least partially attributable to the at least one object; and

generating, by the computing system, a retimed video by:

retiming at least one of the background layer or the one or more object layers; and

re-combining the one or more retimed layers.

13. The computing system of claim 12 , wherein inputting the image frame and the one or more object maps into the machine-learned layer renderer model comprises iteratively individually receiving, as output from the machine-learned layer renderer model and by the computing system, each of the one or more object layers respective to the one or more object maps.

14. The computing system of claim 12 , wherein the background layer and the one or more object layers comprise one or more color channels and an opacity matte.

15. The computing system of claim 12 , wherein the machine-learned layer renderer model comprises a neural network.

16. The computing system of claim 12 , wherein the machine-learned layer renderer model has been trained based at least in part on a reconstruction loss, a mask loss, and a regularization loss.

17. The computing system of claim 16 , wherein the training was performed on downsampled video and then upsampled.

18. The computing system of claim 12 , wherein the one or more object maps comprise one or more texture maps.

19. The computing system of claim 12 , wherein obtaining one or more object maps comprises:

obtaining one or more UV maps, each of the UV maps indicative of the at least one object of the one or more objects depicted within the one or more frames;

obtaining a background deep texture map and one or more object deep texture maps; and

resampling the one or more object deep texture maps based at least in part on the one or more UV maps.

20. The computing system of claim 19 , wherein obtaining the one or more UV maps comprises:

identifying one or more keypoints; and

generating one or more UV maps based on the one or more keypoints.

21. The computing system of claim 12 , wherein the instructions further comprise:

transferring high resolution details of the video data in a post processing step subsequent to receiving the background layer and the one or more object layers.

22. The computing system of claim 12 , wherein the one or more trace effects comprise at least one of: shadows, reflections, splashes, or motion of loose clothing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 6, 2023
From: COLE, FORRESTER H.; LU, ERIKA; DEKEL, TALI; FREEMAN, WILLIAM T.; SALESIN, DAVID HENRY; RUBINSTEIN, MICHAEL
To: GOOGLE LLC
Reel/Frame 062296/0386 →
Continuity (1)
Related Publication 20230206955A1 · Jun 29, 2023
References Cited (71)
US 20040160453A1 · Horton · 2004 [cited by examiner]
US 20080034292A1 · Brunner · 2008 [cited by examiner]
US 20180075590A1 · Yamasaki · 2018 [cited by examiner]
US 20180158230A1 · Yan · 2018 [cited by examiner]
US 20190228264A1 · Huang · 2019 [cited by examiner]
US 20190289321A1 · Liu · 2019 [cited by examiner]
US 20200074642A1 · Wilson et al. · 2020 [cited by applicant]
US 20210004962A1 · Tsai · 2021 [cited by examiner]
WO WO2001035641 · 2001 [cited by applicant]
WO WO2003027766 · 2003 [cited by applicant]
Lu, “Layered Neural Rendering for Retiming People in Video”, SIGGRAPH Asia 2020, https://www.youtube.com/watch?v=KAVCHRImucw, Sep. 10, 2020, retrieved on Dec. 12, 2023, 21 pages. [cited by applicant]
Aberman et al., “Deep Video-Based Performance Cloning”, arXiv:1808.06847vl, Aug. 21, 2018, 10 pages. [cited by applicant]
Agarwala et al., “Panoramic Video Textures”, https://grail.cs.washington.edu/projects/panovidtex/panovidtex.pdf, retrieved on Dec. 12, 2022, 8 pages. [cited by applicant]
Alayrac et al., “Controllable Attention for Structured Layered Video Decomposition”, arXiv:1910.11306v1, Oct. 24, 2019, 14 pages. [cited by applicant]
Alayrac et al., “The Visual Centrifuge: Model-Free Layered Video Representations”, arXiv:1812.01461v2, Apr. 4, 2019, 12 pages. [cited by applicant]
Bai et al., “Video SnapCut: Robust Video Object Cutout Using Localized Classifiers”, Association for Computing Machinery—Transactions on Graphics, vol. 28, Issue 3, Aug. 2009, Article No. 70, pp. 1-11. [cited by applicant]
Barnes et al., “Video Tapestries with Continuous Temporal Zoom”, Association for Computing Machinery—Transactions on Graphics, vol. 29, Issue 4, Jul. 2010, pp. 1-9. [cited by applicant]
Bennett et al., “Computational Time-Lapse Video”, Association for Computing Machinery—Transactions on Graphics, vol. 26, Issue 3, Jul. 2007, 6 pages. [cited by applicant]
Castro et al., “Let's Dance: Learning from Online Dance Videos”, arXiv:1801.07388v1, Jan. 23, 2018, 10 pages. [cited by applicant]
Chan et al., “Everybody Dance Now”, 2019 Conference on Computer Vision and Pattern Recognition, Long Beach, California, United States, Jun. 15-19, 2019, pp. 5933-5942. [cited by applicant]
Chuang et al., “Animating Pictures with Stochastic Motion Textures”, Association for Computing Machinery—Transactions on Graphics, vol. 24, Issue 3, Jul. 2005, 8 pages. [cited by applicant]
Chuang et al., “Video Matting of Complex Scenes”, SIGGRAPH '02: 29th Annual Conference on Computer Graphics and Interactive Techniques, San Antonio, Texas, United States, Jul. 23-26, 2002, 6 pages. [cited by applicant]
Davis et al., “Visual Rhythm and Beat”, Association for Computing Machinery—Transactions on Graphics, vol. 37, Issue 4, Aug. 2018, pp. 1-11. [cited by applicant]
Fang et al., “RMPE: Regional Multi-Person Pose Estimation”, arXiv:1612.00137v5, Feb. 4, 2018, 10 pages. [cited by applicant]
Fradet et al, “Semi-Automatic Motion Segmentation with Motion Layer Mosaics”, 10th European Conference on Computer Vision, Marseille, France, Oct. 12-18, 2008, pp. 210-223. [cited by applicant]
Gafni et al., “Vid2Game: Controllable Characters Extracted from Real-World Videos”, arXiv:1904.08379v1, Apr. 17, 2019, 14 pages. [cited by applicant]
Gandelsman et al., ““Double-DIP”: Unsupervised Image Decomposition via Coupled Deep-Image-Priors”, arXiv:1812.00467v2, Dec. 5, 2018, 10 pages. [cited by applicant]
Goldman et al., “Video Object Annotation, Navigation, and Composition”, 21st Annual Association for Computing Machinery Symposium on User Interface Software and Technology, Monterey, California, United States, Oct. 19-2… [cited by applicant]
Grundmann et al., “Auto-Directed Video Stabilization with Robust LI Optimal Camera Paths”, 2011 Conference on Computer Vision and Pattern Recognition, Colorado Springs, Colorado, United States, Jun. 20-25, 2011, pp. 225… [cited by applicant]
Guler et al., “DensePose: Dense Human Pose Estimation in The Wild”, arXiv:1802.00434v1, Feb. 1, 2018, 12 pages. [cited by applicant]
Hou et al., “Context-Aware Image Matting for Simultaneous Foreground and Alpha Estimation”, arXiv:1909.09725v2, Oct. 2, 2019, 10 pages. [cited by applicant]
International Preliminary Report of Patentability for PCT/US2020/034296, mailed on Dec. 1, 2022, 8 pages. [cited by applicant]
International Search Report and Written Opinion for PCT/US2020/034296, mailed on Feb. 11, 2021, 13 pages. [cited by applicant]
Isola et al., “Image-to-Image Translation with Conditional Adversarial Networks”, arXiv:1611.07004v3, Nov. 26, 2018, 17 pages. [cited by applicant]
Jojic et al., “Learning Flexible Sprites in Video Layers”, Institute of Electrical and Electronics Engineers Computer Society Conference on Computer Vision and Pattern Recognition, Dec. 8-14, 2001, Kauai, Hawaii, United… [cited by applicant]
Joshi et al., “Real-Time Hyperlapse Creation via Optimal Frame Selection”, Association for Computing Machinery—Transactions on Graphics, vol. 34, Issue 4, Aug. 2015, 9 pages. [cited by applicant]
Kim et al., “Deep Video Portraits”, arXiv:1805.11714vl, May 29, 2018, 14 pages. [cited by applicant]
Kingma et al., “Adam: A Method for Stochastic Optimization”, arXiv:1412.6980v9, Jan. 30, 2017, 15 pages. [cited by applicant]
Kumar et al., “Learning Layered Motion Segmentations of Video”, Tenth International Conference on Computer Vision, San Diego, California, United States, Jun. 20-25, 2005, 8 pages. [cited by applicant]
Lan et al., “FFNet: Video Fast-Forwarding via Reinforcement Learning”, arXiv:1805.02792v1, May 8, 2018, 10 pages. [cited by applicant]
Lee et al., “MetaPix: Few-Shot Video Retargeting”, arXiv:1910.04742v2, Mar. 24, 2020, 13 pages. [cited by applicant]
Li et al., “Video Object Cut and Paste”, SIGGRAPH '05: Association for Computing Machinery, Jul. 1, 2005, pp. 595-600. [cited by applicant]
Liu et al., “Neural Rendering and Reenactment of Human Actor Videos”, arXiv:1809.03658v3, May 9, 2019, 14 pages. [cited by applicant]
Loper et al., “SMPL: A Skinned Multi-Person Linear Model”, Association for Computing Machinery—Transactions on Graphics, vol. 34, Issue 6, Nov. 2015, 16 pages. [cited by applicant]
Martin-Brualla et al., “LookinGood: Enhancing Performance Capture with Real-time Neural Re-Rendering”, arXiv:1811.05029v1, Nov. 12, 2018, 14 pages. [cited by applicant]
McCann et al., “Physics-Based Motion Retiming”, 2006 Association for Computing Machinery's Special Interest Group on Computer Graphics and Interactive Techniques Symposium on Computer Animation, Vienna, Austria, Sep. 2-… [cited by applicant]
Meshry et al., “Neural Rerendering in the Wild”, arXiv:1904.04290v1, Apr. 8, 2019, 16 pages. [cited by applicant]
Nandoriya et al., “Video Reflection Removal Through Spatio-Temporal Optimization”, 2017 International Conference on Computer Vision, Venice, Italy, Oct. 22-29, 2017, pp. 2430-2438. [cited by applicant]
Newson et al., “Video Inpainting of Complex Scenes”, arXiv:1503.05528v2, Jun. 8, 2015, 27 pages. [cited by applicant]
Poleg et al., “EgoSampling: Fast-Forward and Stereo for Egocentric Videos”, 2015 Conference on Computer Vision and Pattern Recognition, Boston, Massachusetts, United States, Jun. 7-12, 2015, pp. 4768-4776. [cited by applicant]
Porter et al., “Compositing Digital Images”, ACM SIGGRAPH Computer Graphics, vol. 18, Issue Jul. 3, 1984, pp. 253-259. [cited by applicant]
Pritch et al., “Nonchronological Video Synopsis and Indexing”, Transactions on Pattern Analysis and Machine Intelligence, vol. 30, No. 11, Nov. 2008, pp. 1971-1984. [cited by applicant]
Rublee et al., “ORB: An Efficient Alternative to SIFT or SURF”, 2011 International Conference on Computer Vision, Washington, District of Columbia, United States, Nov. 6-13, 2011, pp. 2564-2571. [cited by applicant]
Silva et al., “A Weighted Sparse Sampling and Smoothing Frame Transition Approach for Semantic Fast-Forward First-Person Videos”, 2018 Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, United … [cited by applicant]
Shin et al., “3D Scene Reconstruction with Multi-layer Depth and Epipolar Transformers”, arXiv:1902.06729v2, Aug. 27, 2019, 28 pages. [cited by applicant]
Sitzmann et al., “Deep Voxels: Learning Persistent 3D Feature Embeddings”, arXiv:1812.01024v2, Apr. 11, 2019, 10 pages. [cited by applicant]
Srinivasan et al., “Pushing the Boundaries of View Extrapolation with Multiplane Images”, arXiv:1905.00413v1, May 1, 2019, 12 pages. [cited by applicant]
Thies et al., “Deferred Neural Rendering: Image Synthesis Using Neural Textures”, arXiv:1904.12356v1, Apr. 28, 2019, 12 pages. [cited by applicant]
Tomasi et al., “Correlation Optimized Warping and Dynamic Time Warping as Preprocessing Methods for Chromatographic Data”, Journal of Chemometrics, vol. 18, Issue 5, May 2004, pp. 231-241. [cited by applicant]
Ulyanov et al., “Deep Image Prior”, 2018 Conference on Computer Vision and Pattern Recognition, Salt Lake City, Utah, United States, Jun. 18-23, 2018, pp. 9446-9454. [cited by applicant]
Wang et al., “Interactive Video Cutout”, SIGGRAPH '05: Association for Computing Machinery—Transactions on Graphics, vol. 24, Issue 3, Jul. 2005, pp. 585-594. [cited by applicant]
Wang et al., “Representing Moving Images with Layers”, Transactions on Image Processing, vol. 3, No. 5, Sep. 1, 1994, pp. 625-638. [cited by applicant]
Wexler et al., “Space-Time Completion of Video”, Transactions on Pattern Analysis and Machine Intelligence, vol. 29, No. 3, Mar. 2007, pp. 463-476. [cited by applicant]
Xiu et al., “Pose Flow: Efficient Online Pose Tracking”, arXiv:1802.00977v2, Jul. 2, 2018, 12 pages. [cited by applicant]
Xu et al., “Deep Image Matting”, arXiv:1703.03872v3, Apr. 11, 2017, 10 pages. [cited by applicant]
Xue et al., “A Computational Approach for Obstruction-Free Photography”, Association for Computing Machinery—Transactions on Graphics, vol. 34, Issue 4, Aug. 2015, pp. 1-11. [cited by applicant]
Yoo et al., “Motion Retiming by Using Bilateral Time Control Surfaces”, vol. 47, Apr. 2015, pp. 59-67. [cited by applicant]
Zhou et al., “Dance, Dance Generation: Motion Transfer for Internet Videos”, arXiv:1904.00129v1, Mar. 30, 2019, 9 pages. [cited by applicant]
Zhou et al., “Stereo Magnification: Learning View Synthesis using Multiplane Images”, arXiv:1805.09817vl, May 24, 2018, 12 pages. [cited by applicant]
Zhou et al., “Time-Mapping Using Space-Time Saliency”, 2014 Conference on Computer Vision and Pattern Recognition, Columbus, Ohio, United States, Jun. 23-28, 2014, pp. 3358-3365. [cited by applicant]
Zitnick et al., “High-Quality Video View Interpolation Using a Layered Representation”, Association for Computing Machinery—Transactions on Graphics, vol. 23, Issue 3, Aug. 2004, pp. 600-608. [cited by applicant]
Cited By (1)
US 12,567,162