IP Library Granted Patent US 12,608,621
Granted Patent B2
US 12,608,621 · App. 18/065,672 · Granted Apr 21, 2026

Generating artificial video with changed domain

Inventors: Akhil Perincherry (Dearborn, MI); Arpita Chand (Dearborn, MI)
Assignee: Ford Global Technologies, LLC
G06N3/09G06N3/094G06V10/7715G06V10/776G06V10/7788G06V10/806G06V10/82G06V20/58G06N3/045G06N3/047G06N3/08G06N3/088G06T9/002H04N5/265H04N19/139H04N19/172H04N19/182H04N19/42
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,608,621
App. No.
18/065,672
Granted
Apr 21, 2026
Kind
B2
Abstract

A computer includes a processor and a memory, and the memory stores instructions executable by the processor to receive an input video of a scene and audio data associated with the input video, the input video being in a first domain; execute an encoder to map the input video and the audio data to a latent vector in a lower-dimensional latent space; and execute a generator to generate an output video of the scene from the latent vector, the output video being in a second domain. The encoder and the generator are trained to maintain temporal consistency between the input video and the output video by using the audio data.

Claims (40)

1 . A computer comprising a processor and a memory, the memory storing instructions executable by the processor to:

receive an input video of a scene and audio data associated with the input video, the input video being in a first domain;

execute an encoder to map the input video and the audio data to a latent vector in a lower-dimensional latent space; and

execute a generator to generate an output video of the scene from the latent vector, the output video being in a second domain;

wherein the encoder and the generator are trained to maintain temporal consistency between the input video and the output video by using the audio data.

2 . The computer of claim 1 , wherein the encoder and the generator are supervised by a discriminator during training.

3 . The computer of claim 2 , wherein the discriminator supervises the training of the encoder and the generator by testing a consistency of the output video with the audio data.

4 . The computer of claim 3 , wherein, while training the encoder and the generator, the discriminator uses a correlation between the output video and the audio data to test the consistency of the output video with the audio data.

5 . The computer of claim 4 , wherein, while training the encoder and the generator, the discriminator receives the correlation from a correlation module, the correlation module being pretrained with contrastive learning.

6 . The computer of claim 2 , wherein the discriminator supervises the training of the encoder and the generator by testing a consistency of the output video with the second domain.

7 . The computer of claim 6 , wherein the instructions further include instructions to determine an adversarial loss based on an output of the discriminator and to update the encoder and the generator based on the adversarial loss.

8 . The computer of claim 1 , wherein the first domain and the second domain are mutually exclusive environmental conditions of the scene.

9 . The computer of claim 8 , wherein the environmental conditions are one of a lighting condition or a weather condition.

10 . The computer of claim 1 , wherein the first domain and the second domain are mutually exclusive visual rendering characteristics of the input video and output video.

11 . The computer of claim 10 , wherein the visual rendering characteristics are one of a resolution, a color representation scheme, or simulatedness.

12 . The computer of claim 1 , wherein

the instructions further include instructions to extract visual features from the input video; and

executing the encoder is based on the visual features.

13 . The computer of claim 12 , wherein

the instructions further include instructions to extract audio features from the audio data and to fuse the visual features and the audio features; and

executing the encoder is based on the fusion of the visual features and the audio features.

14 . The computer of claim 1 , wherein

the instructions further include instructions to extract audio features from the audio data; and

executing the encoder is based on the audio features.

15 . The computer of claim 1 , wherein the encoder is trained to include semantic content of the input video in the latent vector and to exclude domain data of the input video from the latent vector.

16 . The computer of claim 1 , wherein

the encoder is a first encoder;

the generator is a first generator;

the latent vector is a first latent vector; and

training the first encoder and the first generator includes executing a second encoder to map the output video and the audio data to a second latent vector in the lower-dimensional latent space, and executing a second generator to generate a test video of the scene from the second latent vector in the first domain.

17 . The computer of claim 16 , wherein training the first encoder and the first generator includes updating the first encoder and the first generator based on a difference between the test video and the input video.

18 . The computer of claim 1 , wherein

the instructions further include instructions to train a machine-learning model on training data; and

the training data includes the output video.

19 . The computer of claim 18 , wherein the machine-learning model is an object-recognition model.

20 . A method comprising:

receiving an input video of a scene and audio data associated with the input video, the input video being in a first domain;

executing an encoder to map the input video and the audio data to a latent vector in a lower-dimensional latent space; and

executing a generator to generate an output video of the scene from the latent vector, the output video being in a second domain;

wherein the encoder and the generator are trained to maintain temporal consistency between the input video and the output video by using the audio data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2022
From: PERINCHERRY, AKHIL; CHAND, ARPITA
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 062082/0477 →
Continuity (1)
Related Publication 20240202533A1 · Jun 20, 2024
References Cited (7)
US 10885111B2 · Chaudhury et al. · 2021 [cited by applicant]
US 20200134426A1 · Ketz · 2020 [cited by examiner]
US 20200234725A1 · Garbacea · 2020 [cited by examiner]
US 20200242507A1 · Gan et al. · 2020 [cited by applicant]
CN 113935418A · 2022 [cited by applicant]
CN 114419204A · 2022 [cited by applicant]
Aldausari, N, et al., “Video Generative Adversarial Networks: a Review,” ACM Comput. Surv., vol. 55, No. 2, Article 30, Jan. 2022, 10 pages. [cited by applicant]