IP Library Granted Patent US 12,265,574
Granted Patent B2
US 12,265,574 · App. 18/391,262 · Granted Apr 1, 2025

Using interpolation to generate a video from static images

Inventors: Janne Kontkanen (San Francisco, CA); Jamie Aspinall (Mountain View, CA); Dominik Kaeser (New York City, NY); Navin Sarma (Palo Alto, CA); Brian Curless (Seattle, WA); David Salesin (Sausalito, CA)
Assignee: Google LLC
G06F16/739G06F16/75G06F16/7867G06N20/00G06T7/20H04N5/2628
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,265,574
App. No.
18/391,262
Granted
Apr 1, 2025
Kind
B2
Abstract

A media application selects, from a collection of images associated with a user account, candidate pairs of images, where each pair includes a first static image and a second static image from the user account. The media application applies a filter to select a particular pair of images from the candidate pairs of images. The media application generates, using an image interpolator, one or more intermediate images based on the particular pair of images. The media application generates a video that includes three or more frames arranged in a sequence, where a first frame of the sequence is the first static image, a last frame of the sequence is the second static image, and each of the one or more intermediate images is a corresponding intermediate frame of the sequence between the first frame and the last frame.

Claims (37)

1. A computer-implemented method comprising:

selecting, from a collection of images associated with a user account, candidate pairs of images, wherein each candidate pair includes a first static image of a scene and a second static image of the scene from the user account;

applying a filter to select a particular pair of images from the candidate pairs of images based on the particular pair of images failing to meet a threshold similarity;

generating, using an image interpolator, one or more intermediate images based on the particular pair of images;

providing the first static image as input to a depth machine-learning model, the depth machine-learning model outputting a three-dimensional representation of the scene; and

generating a video that includes three or more frames arranged in a sequence, wherein a first frame of the sequence is the first static image, a last frame of the sequence is the second static image, and each of the one or more intermediate images is a corresponding intermediate frame of the sequence between the first frame and the last frame, wherein the video includes the three-dimensional representation of the scene.

2. The method of claim 1 , wherein the filter further selects the particular pair of images based on the particular pair of images including a subject that a user associated with the user account expressed a preference for including in the video.

3. The method of claim 1 , wherein the video includes at least one camera effect selected from a group of zooming, panning, and combinations thereof.

4. The method of claim 1 , wherein the depth machine-learning model includes a convolutional neural network that outputs a low-resolution version of the three-dimensional representation and iteratively outputs improvements to the low-resolution version of the three-dimensional representation.

5. The method of claim 1 , wherein the depth machine-learning model is trained using training data that includes depth maps.

6. The method of claim 1 , wherein the first static image is in an x-y plane and the three-dimensional representation includes z-axis coordinates of objects in the three-dimensional representation.

7. The method of claim 1 , wherein the image interpolator includes an interpolation machine-learning model that receives the first static image and the second static image as input and that outputs the one or more intermediate images.

8. A system comprising:

one or more processors; and

a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

selecting, from a collection of images associated with a user account, candidate pairs of images, wherein each candidate pair includes a first static image of a scene and a second static image of the scene from the user account;

applying a filter to select a particular pair of images from the candidate pairs of images based on the particular pair of images failing to meet a threshold similarity;

generating, using an image interpolator, one or more intermediate images based on the particular pair of images;

providing the first static image as input to a depth machine-learning model, the depth machine-learning model outputting a three-dimensional representation of the scene; and

generating a video that includes three or more frames arranged in a sequence, wherein a first frame of the sequence is the first static image, a last frame of the sequence is the second static image, and each of the one or more intermediate images is a corresponding intermediate frame of the sequence between the first frame and the last frame, wherein the video includes the three-dimensional representation of the scene.

9. The system of claim 8 , wherein the filter selects the particular pair of images based on the particular pair of images exceeding a motion threshold.

10. The system of claim 8 , wherein the video includes at least one camera effect selected from a group of zooming, panning, and combinations thereof.

11. The system of claim 8 , wherein the depth machine-learning model includes a convolutional neural network that outputs a low-resolution version of the three-dimensional representation and iteratively outputs improvements to the low-resolution version of the three-dimensional representation.

12. The system of claim 8 , wherein the depth machine-learning model is trained using training data that includes depth maps.

13. The system of claim 8 , wherein the first static image is in an x-y plane and the three-dimensional representation includes z-axis coordinates of objects in the three-dimensional representation.

14. The system of claim 8 , wherein the image interpolator includes an interpolation machine-learning model that receives the first static image and the second static image as input and that outputs the one or more intermediate images.

15. A non-transitory computer-readable medium with instructions stored thereon that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

selecting, from a collection of images associated with a user account, candidate pairs of images, wherein each candidate pair includes a first static image of a scene and a second static image of the scene from the user account;

applying a filter to select a particular pair of images from the candidate pairs of images based on the particular pair of images failing to meet a threshold similarity;

generating, using an image interpolator, one or more intermediate images based on the particular pair of images;

providing the first static image as input to a depth machine-learning model, the depth machine-learning model outputting a three-dimensional representation of the scene; and

generating a video that includes three or more frames arranged in a sequence, wherein a first frame of the sequence is the first static image, a last frame of the sequence is the second static image, and each of the one or more intermediate images is a corresponding intermediate frame of the sequence between the first frame and the last frame, wherein the video includes the three-dimensional representation of the scene.

16. The computer-readable medium of claim 15 , wherein the filter selects the particular pair of images based on the particular pair of images falling below a motion threshold.

17. The computer-readable medium of claim 15 , wherein the video includes at least one camera effect selected from a group of zooming, panning, and combinations thereof.

18. The computer-readable medium of claim 15 , wherein the depth machine-learning model includes a convolutional neural network that outputs a low-resolution version of the three-dimensional representation and iteratively outputs improvements to the low-resolution version of the three-dimensional representation.

19. The computer-readable medium of claim 15 , wherein the depth machine-learning model is trained using training data that includes depth maps.

20. The computer-readable medium of claim 15 , wherein the first static image is in an x-y plane and the three-dimensional representation includes z-axis coordinates of objects in the three-dimensional representation.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 20, 2023
From: KONTKANEN, JANNE; ASPINALL, JAMIE; KAESER, DOMINIK; SARMA, NAVIN; SALESIN, DAVID; CURLESS, BRIAN
To: GOOGLE LLC
Reel/Frame 065925/0131 →
Continuity (3)
Continuation 17566462 · Dec 30, 2021
Provisional Application 63190234 · May 18, 2021
Related Publication 20240126810A1 · Apr 18, 2024
References Cited (27)
US 10489897B2 · Staranowicz · 2019 [cited by examiner]
US 10958869B1 · Chi · 2021 [cited by examiner]
US 11288543B1 · Tovchigrechko · 2022 [cited by examiner]
US 20030025810A1 · Pilu et al. · 2003 [cited by applicant]
US 20100058397A1 · Rogers · 2010 [cited by examiner]
US 20120311623A1 · Davis · 2012 [cited by examiner]
US 20180315174A1 · Staranowicz et al. · 2018 [cited by applicant]
US 20190197667A1 · Paluri · 2019 [cited by examiner]
US 20200356827A1 · Dinerstein et al. · 2020 [cited by applicant]
JP 2003141559 · 2003 [cited by applicant]
JP 2021010099 · 2021 [cited by applicant]
KR 20200057844 · 2020 [cited by applicant]
KR 1020200130105 · 2020 [cited by applicant]
KR 1020210018182 · 2021 [cited by applicant]
WO 2021025717 · 2021 [cited by applicant]
EPO, Partial International Search for International Patent Application No. PCT/US2021/065764, Apr. 4, 2022, 8 pages. [cited by applicant]
EPO, Written Opinion for International Patent Application No. PCT/US2021/065764, Jul. 13, 2022, 13 pages. [cited by applicant]
EPO, International Search Report for International Patent Application No. PCT/US2021/065764, Jul. 13, 2022, 6 pages. [cited by applicant]
Hua, et al., “Content based photograph slide show with incidental music”, ISCAS '03, vol. 2, May 25, 2003, pp. IL648-ll651. [cited by applicant]
Hua, et al., “Photo2Video—A System for Automatically Converting Photographic Series Into Video”, IEEE Transactions on Circuits and Systems for Video Technology, IEEE Service Center, Piscataway, NJ, vol. 16, No. 7, Jul. … [cited by applicant]
Liao, et al., “Depth Map Design and Depth-based Effects With a Single Image”, Graphics Interface. Vol. 2; retrieved from Internet: https://www.researchgate.net/profile/Jingtang-Liao/publication/317044981_Depth_Map_Desig… [cited by applicant]
USPTO, Non-final Office Action for U.S. Appl. No. 17/566,462, Jun. 5, 2023, 18 pages. [cited by applicant]
USPTO, Notice of Allowance for U.S. Appl. No. 17/566,462, Sep. 25, 2023, 6 pages. [cited by applicant]
Van Amersfoort, et al., “Frame interpolation with multi-scale deep loss functions and generative adversarial networks”, arXiv preprint arXiv:1711.06045, 2017, 17 pages. [cited by applicant]
JPO, Office Action (with English translation) for Japanese Patent Application No. 2023-528278, May 21, 2024, 10 pages. [cited by applicant]
KIPO, Office Action (with English translation) for Korean Patent Application No. 10-2023-7010192, Oct. 7, 2024, 13 pages. [cited by applicant]
JPO, Notice of Allowance (with English translation) for Japanese Patent Application No. 2023-528278, Oct. 29, 2024, 5 pages. [cited by applicant]