IP Library › Granted Patent US 12,744,972
Granted Patent B2
US 12,744,972 · App. 18/901,340 · Granted Sep 22, 2026

Generating corrected simulations using video generation models

Inventors: Praneet Dutta (Sunnyvale, CA); Medhini Narasimhan (London, GB); Timo Immanuel Denk (Zürich, CH); Ishaan Malhi (Sunnyvale, CA)
Assignee: GDM Holding LLC
H04N21/816G06T7/20G06V20/70G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,972
App. No.
18/901,340
Granted
Sep 22, 2026
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for generating corrected simulations using video generation models. One of the methods includes obtaining an input video including a sequence of frames depicting a state transition of an environment that does not meet a state transition criterion; generating, based on a frame from a sequence of frames that depicts an incorrect end state of the state transition, a synthetic ending frame depicting a corrected end state of the state transition; and processing an input including one or more key frames from sequence of frames of the input video and the synthetic ending frame depicting the corrected end state of the state transition using a video generation model to generate an output video depicting a synthetic state transition of the environment that meets the state transition criterion.

Claims (49)

1 . A method, comprising:

obtaining an input video comprising a sequence of frames depicting a state transition of an environment that does not meet a state transition criterion;

generating, based on a frame from a sequence of frames that depicts an incorrect end state of the state transition, a synthetic ending frame depicting a corrected end state of the state transition; and

processing an input comprising one or more key frames from sequence of frames of the input video and the synthetic ending frame depicting the corrected end state of the state transition using a video generation model to generate an output video depicting a synthetic state transition of the environment that meets the state transition criterion.

2 . The method of claim 1 , comprising:

processing the sequence of the frames of the input video using a visual language model to generate a respective state annotation for each frame in the sequence of the frames of the input video; and

determining the one or more key frames from the sequence of the frames of the input video based on the respective state annotations for the sequence of the frames of the input video.

3 . The method of claim 1 , wherein the sequence of the frames of the input video comprises a transition frame, wherein the state transition of the environment depicted in the transition frame and frames in the input video before the transition frame meets the state transition criterion, and the state transition of the environment depicted in frames in the input video after the transition frame does not meet the state transition criterion.

4 . The method of claim 3 , wherein the input to the video generation model comprises the transition frame.

5 . The method of claim 3 , wherein the input to the video generation model comprises one or more frames that precede the transition frame.

6 . The method of claim 1 , comprising:

processing an input comprising at least the frame that depicts the incorrect end state of the state transition using an image editing model to generate the synthetic ending frame depicting the corrected end state of the state transition.

7 . The method of claim 1 , wherein the one or more key frames comprise a starting frame depicting the environment before the state transition happens, and the method comprises:

obtaining a set of points on the starting frame;

obtaining a target trajectory for the set of points associated with a target condition of the environment; and

processing the input comprising the starting frame, the synthetic ending frame, and the target trajectory for the set of points using the video generation model to generate the output video that meets the state transition criterion and is conditioned on the starting frame, the synthetic ending frame, and the target trajectory for the set of points, wherein a first frame of the output video is the starting frame, a last frame of the output video is the synthetic ending frame, and locations for the set of points in at least some frames of the output video approximately follow the target trajectory.

8 . The method of claim 1 , further comprising:

generating control data for controlling one or more objects in the environment that causes the one or more objects to follow respective trajectories for each of the one or more objects depicted in the output video.

9 . The method of claim 1 , wherein the state transition of the environment comprises a landing or a takeoff of an aircraft.

10 . The method of claim 1 , further comprising:

obtaining a set of points on an object in the environment on a starting frame of the output video;

processing the output video using a point tracking model to generate trajectories for the set of points in the output video; and

generating an evaluation result for the output video based on the trajectories for the set of points in the output video.

11 . The method of claim 10 , wherein generating the evaluation result for the output video comprises:

determining that at least one trajectory of the trajectories for the set of points in the output video is discontinuous; and

in response to determining that at least one trajectory of the trajectories for the set of points in the output video is discontinuous, generating the evaluation result for the output video indicating that the output video has an error.

12 . The method of claim 10 , wherein generating the evaluation result for the output video comprises:

determining a difference value between the trajectories for the set of points in the output video and reference trajectories for the set of points generated by a simulation engine that is based on one or more laws of physics; and

determining whether the trajectories for the set of points in the output video meet the one or more laws of physics based on whether the difference value is less than a threshold.

13 . A system comprising one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

obtaining an input video comprising a sequence of frames depicting a state transition of an environment that does not meet a state transition criterion;

generating, based on a frame from a sequence of frames that depicts an incorrect end state of the state transition, a synthetic ending frame depicting a corrected end state of the state transition; and

processing an input comprising one or more key frames from sequence of frames of the input video and the synthetic ending frame depicting the corrected end state of the state transition using a video generation model to generate an output video depicting a synthetic state transition of the environment that meets the state transition criterion.

14 . The system of claim 13 , wherein the operations comprise:

processing the sequence of the frames of the input video using a visual language model to generate a respective state annotation for each frame in the sequence of the frames of the input video; and

determining the one or more key frames from the sequence of the frames of the input video based on the respective state annotations for the sequence of the frames of the input video.

15 . The system of claim 13 , wherein the sequence of the frames of the input video comprises a transition frame, wherein the state transition of the environment depicted in the transition frame and frames in the input video before the transition frame meets the state transition criterion, and the state transition of the environment depicted in frames in the input video after the transition frame does not meet the state transition criterion.

16 . The system of claim 15 , wherein the input to the video generation model comprises the transition frame.

17 . The system of claim 15 , wherein the input to the video generation model comprises one or more frames that precede the transition frame.

18 . The system of claim 13 , wherein the operations comprise:

processing an input comprising at least the frame that depicts the incorrect end state of the state transition using an image editing model to generate the synthetic ending frame depicting the corrected end state of the state transition.

19 . The system of claim 13 , wherein the one or more key frames comprise a starting frame depicting the environment before the state transition happens, and the operations comprise:

obtaining a set of points on the starting frame;

obtaining a target trajectory for the set of points associated with a target condition of the environment; and

processing the input comprising the starting frame, the synthetic ending frame, and the target trajectory for the set of points using the video generation model to generate the output video that meets the state transition criterion and is conditioned on the starting frame, the synthetic ending frame, and the target trajectory for the set of points, wherein a first frame of the output video is the starting frame, a last frame of the output video is the synthetic ending frame, and locations for the set of points in at least some frames of the output video approximately follow the target trajectory.

20 . One or more non-transitory storage media encoded with instructions that when executed by a computing device cause the computing device to perform operations comprising:

obtaining an input video comprising a sequence of frames depicting a state transition of an environment that does not meet a state transition criterion;

generating, based on a frame from a sequence of frames that depicts an incorrect end state of the state transition, a synthetic ending frame depicting a corrected end state of the state transition; and

processing an input comprising one or more key frames from sequence of frames of the input video and the synthetic ending frame depicting the corrected end state of the state transition using a video generation model to generate an output video depicting a synthetic state transition of the environment that meets the state transition criterion.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2025
From: DEEPMIND TECHNOLOGIES LIMITED
To: GDM HOLDING LLC
Reel/Frame 072361/0231 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 30, 2025
From: DUTTA, PRANEET; NARASIMHAN, MEDHINI GULGANJALLI; DENK, TIMO IMMANUEL; MALHI, ISHAAN
To: DEEPMIND TECHNOLOGIES LIMITED
Reel/Frame 071564/0335 →
Continuity (1)
Related Publication 20260095632A1 · Apr 2, 2026
References Cited (27)
US 11189189B2 · Rosolio · 2021 [cited by examiner]
US 11869236B1 · Callari · 2024 [cited by applicant]
US 12488569B2 · Evans · 2025 [cited by examiner]
US 20170294135A1 · Lechner · 2017 [cited by examiner]
US 20210141986A1 · Ganille · 2021 [cited by examiner]
US 20240135618A1 · Zhang et al. · 2024 [cited by applicant]
US 20250119624A1 · Oh · 2025 [cited by examiner]
CN 112529895A · 2021 [cited by applicant]
Adiwardana et al., “Towards a Human-like Open-Domain Chatbot,” CoRR, Submitted on Feb. 27, 2020, arXiv:2001.09977v3, 38 pages. [cited by applicant]
Avrahami et al., “Blended diffusion for text-driven editing of natural images,” Paper, Presented at Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, New Orleans, LA, USA, Jun. 18-24, 20… [cited by applicant]
Brown et al., “Language Models are Few-Shot Learners,” CoRR, Submitted on Jul. 22, 2020, arXiv:2005.14165v4, 75 pages. [cited by applicant]
Doersch et al. “TAPIR: Tracking Any Point with Per-Frame Initialization and Temporal Refinement,” Paper, Presented at Proceedings of the IEEE/CVF International Conference on Computer Vision, Paris, France, Oct. 1-6, 202… [cited by applicant]
Donahue, “Label-Conditioned Next-Frame Video Generation with Neural Flows,” CoRR, Submitted on Oct. 15, 2019, arXiv:1910.11106v1, 9 pages. [cited by applicant]
aviationsafetymagazine.com [online], “Fixing Your Bounce,” Jan. 27, 2019, retrieved on Oct. 23, 2024, retrieved from URL <https://www.aviationsafetymagazine.com/features/fixing-your-bounce/>, 7 pages. [cited by applicant]
Hoffmann et al., “Training Compute-Optimal Large Language Models,” CoRR, Submitted on Mar. 29, 2022, arXiv:2203.15556v1, 36 pages. [cited by applicant]
Huang et al., “VBench: Comprehensive Benchmark Suite for Video Generative Models,” CoRR, Submitted on Nov. 29, 2023, arXiv:2311.17982v1, 28 pages. [cited by applicant]
boldmethod.com [online], “9 Approach To Landing Problems, And How To Recover From Each One,” Apr. 13, 2021, retrieved on Oct. 23, 2024, retrieved from URL <https://www.boldmethod.com/blog/lists/2021/04/nine-approach-to-… [cited by applicant]
Ling et al., “EditGAN: High-Precision Semantic Image Editing,” Paper, Presented at the 35th Conference on Neural Information Processing Systems (NeurIPS 2021), Virtual Conference, Dec. 6-14, 2021; Advances in Neural Inf… [cited by applicant]
MuJoCo.org [online], “MuJoCo, Advanced physics simulation,” 2021, retrieved on Sep. 20, 2024, retrieved from URL <https://mujoco.org/>, 3 pages. [cited by applicant]
Rae et al., “Scaling Language Models: Methods, Analysis & Insights from Training [cited by applicant]
Raffel et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” CoRR, Submitted on Oct. 24, 2019, arXiv:1910.10683v2, 53 pages. [cited by applicant]
Tulyakov et al., “MoCoGAN: Decomposing Motion and Content for Video Generation,” CoRR, Submitted on Dec. 14, 2017, arXiv:1707.04933v2, 13 pages. [cited by applicant]
Wikipedia.org [Online], Microsoft Flight Simulator (2020 video game), Mar. 11, 2024, retrieved on Sep. 20, 2024, retrieved from URL <https://en.wikipedia.org/wiki/Microsoft_Flight_Simulator_(2020_video game)>, 30 pages. [cited by applicant]
Yu et al., “Recurrent Deconvolutional Generative Adversarial Networks with Application to Text Guided Video Generation,” CoRR, Submitted on Aug. 13, 2020, arXiv:2008.05856v1, 11 pages. [cited by applicant]
Zeng et al., “Make Pixels Dance: High-Dynamic Video Generation,” CoRR, Submitted on Nov. 18, 2023, arXiv:2311.10982v1, 11 pages. [cited by applicant]
Extended Search Report in European Appln. No. 25200473.4, mailed on Feb. 10, 2026, 9 pages. [cited by applicant]
Innamorati et al., “Neural Re-Simulation for Generating Bounces in Single Images,” 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 27, 2019, pp. 8718-8727. [cited by applicant]