IP Library › Granted Patent US 12,670,362
Granted Patent B2
US 12,670,362 · App. 17/491,226 · Granted Jun 30, 2026

Video synthesis within a messaging system

Inventors: Menglei Chai (Los Angeles, CA); Kyle Olszewski (Los Angeles, CA); Jian Ren (Marina Del Ray, CA); Yu Tian (Piscataway, NJ); Sergey Tulyakov (Marina del Rey, CA)
Assignee: Snap Inc,.
G06N3/045G06F18/214G06N3/08G06T7/20G06T2207/10016G06T2207/20081G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,670,362
App. No.
17/491,226
Filed
Sep 30, 2021
Granted
Jun 30, 2026
Kind
B2
Art Unit
2127
USPC
706/16
Abstract

Aspects of the present disclosure involve a system comprising a computer-readable storage medium storing a program and method for video synthesis. The program and method provide for accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator; generating an updated GAN based on the primary GAN, by performing operations comprising identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, training the motion generator based on the input data, and adjusting weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator; and generating a synthesized video based on the primary GAN and the input data.

Claims (45)

1 . A video synthesis method comprising:

accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator;

generating an updated GAN based on the primary GAN, by performing operations comprising

identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, and

training the motion generator based on the input data using a contrastive image discriminator with contrastive loss to constrain generated images to possess similar quality and content, the contrastive loss comprising positive pairs of augmented frames from a same video that share same content and negative pairs of augmented images from different videos; and

generating a synthesized video based on the primary GAN and the input data,

wherein the pre-trained image generator is trained on a first dataset of a first domain while learning motion generator parameters using an second dataset of a second domain, the first domain and the second domain corresponding to different types of captured subjects, and the first dataset and the second dataset corresponding to different types of data.

2 . The video synthesis method of claim 1 , wherein generating the updated GAN based on the primary GAN further comprises:

adjusting weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator.

3 . The video synthesis method of claim 1 , wherein the motion domain dataset corresponds to a motion trajectory vector that is used to synthesize each individual data frame corresponding to the generated synthesized video.

4 . The video synthesis method of claim 1 , wherein the pre-trained image generator is configured to receive the initial latent code and output from the motion generator, to generate the synthesized video.

5 . The video synthesis method of claim 1 , wherein the pre-trained image generator is pre-trained with a primary dataset comprising at least one of real images or a content dataset.

6 . The video synthesis method of claim 5 , wherein a generator corresponding to the motion generator and the pre-trained image generator is trained with a secondary dataset that is different than the primary dataset.

7 . The video synthesis method of claim 1 , wherein the motion generator is configured to receive the initial latent code to predict consecutive latent codes.

8 . The video synthesis method of claim 1 , wherein the motion generator is implemented with two long short-term memory neural networks.

9 . The method of claim 1 , wherein the motion generator learns motion patterns from the second domain to synthesize temporally consistent video frames in the first domain, thereby enabling generation of realistic videos of the first domain using motion extracted from the second domain.

10 . The method of claim 9 , wherein the first domain corresponds to captured subjects of animal faces,

wherein the second domain corresponds to captured subjects of human facial expressions, and

wherein that the pre-trained image generator is trained on the video dataset of the animal faces while learning motion generator parameters using the image dataset of the human facial expressions.

11 . A system comprising:

at least one processor; and

a memory storing instructions that, when executed by the at least one processor, configure the system to perform operations comprising:

accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator;

generating an updated GAN based on the primary GAN, by performing operations comprising

identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, and

training the motion generator based on the input data using a contrastive image discriminator with contrastive loss to constrain generated images to possess similar quality and content, the contrastive loss comprising positive pairs of augmented frames from a same video that share same content and negative pairs of augmented images from different videos; and

generating a synthesized video based on the primary GAN and the input data,

wherein the pre-trained image generator is trained on a first dataset of a first domain while learning motion generator parameters using an second dataset of a second domain, the first domain and the second domain corresponding to different types of captured subjects, and the first dataset and the second dataset corresponding to different types of data.

12 . The system of claim 11 , wherein generating the updated GAN based on the primary GAN further comprises:

adjust weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator.

13 . The system of claim 11 , wherein the motion domain dataset corresponds to a motion trajectory vector that is used to synthesize each individual data frame corresponding to the generated synthesized video.

14 . The system of claim 11 , wherein the pre-trained image generator is configured to receive the initial latent code and output from the motion generator, to generate the synthesized video.

15 . The system of claim 11 , wherein the pre-trained image generator is pre-trained with a primary dataset comprising at least one of real images or a content dataset.

16 . The system of claim 15 , wherein a generator corresponding to the motion generator and the pre-trained image generator is trained with a secondary dataset that is different than the primary dataset.

17 . The system of claim 11 , wherein the motion generator is configured to receive the initial latent code to predict consecutive latent codes.

18 . The system of claim 11 , wherein the motion generator is implemented with two long short-term memory neural networks.

19 . A non-transitory computer-readable storage medium, the computer-readable storage medium including instructions that when executed by a computer, cause the computer to perform operations comprising:

accessing a primary generative adversarial network (GAN) comprising a pre-trained image generator, a motion generator comprising a plurality of neural networks, and a video discriminator;

generating an updated GAN based on the primary GAN, by performing operations comprising

identifying input data of the updated GAN, the input data comprising an initial latent code and a motion domain dataset, and

training the motion generator based on the input data using a contrastive image discriminator with contrastive loss to constrain generated images to possess similar quality and content, the contrastive loss comprising positive pairs of augmented frames from a same video that share same content and negative pairs of augmented images from different videos; and

generating a synthesized video based on the primary GAN and the input data,

wherein the pre-trained image generator is trained on a first dataset of a first domain while learning motion generator parameters using an second dataset of a second domain, the first domain and the second domain corresponding to different types of captured subjects, and the first dataset and the second dataset corresponding to different types of data.

20 . The computer-readable storage medium of claim 19 , wherein generating the updated GAN based on the primary GAN further comprises:

adjust weights of the plurality of neural networks of the primary GAN based on an output of the video discriminator.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2025
From: CHAI, MENGLEI; OLSZEWSKI, KYLE; REN, JIAN; TIAN, YU; TULYAKOV, SERGEY
To: SNAP INC.
Reel/Frame 070142/0155 →
Continuity (2)
Provisional Application 63198151 · Sep 30, 2020
Related Publication 20220101104A1 · Mar 31, 2022
References Cited (27)
US 10546230B2 · Kurata · 2020 [cited by examiner]
US 11188790B1 · Kumar · 2021 [cited by examiner]
US 20180314716A1 · Kim et al. · 2018 [cited by applicant]
US 20200204822A1 · Liu · 2020 [cited by examiner]
CN 108460812 · 2018 [cited by applicant]
CN 102682420A · 2019 [cited by examiner]
CN 110210386 · 2019 [cited by applicant]
CN 110568442 · 2019 [cited by applicant]
CN 111325817 · 2020 [cited by applicant]
CN 116261745A · 2023 [cited by applicant]
KR 102084782B1 · 2020 [cited by examiner]
WO WO2022072725A1 · 2022 [cited by applicant]
Sun, Ximeng, Huijuan Xu, and Kate Saenko. “Twostreamvan: Improving motion modeling in video generation.” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2020. (Year: 2020). [cited by examiner]
Kang Tae Won, KR-102084782-B1 (English Translation), Mar. 4, 2020 (Year: 2020). [cited by examiner]
Abdullah Hamdi and Bernard Ghanem. IAN: Combining Generative Adversarial Networks for Imaginative Face Generation. Arxiv.org. arXiv: 1904.07916. (Apr. 16, 2019). [cited by examiner]
Dongxu Wei, Xiaowei Xu, Haibin Shen, and Kejie Huang. GAC-GAN: A General Method for Appearance-Controllable Human Video Motion Transfer. Archiv.org. arXiv:1911.10672. (Feb. 9, 2020). [cited by examiner]
“International Application Serial No. PCT/US2021/053012, International Search Report mailed Jan. 21, 2022”, 4 pgs. [cited by applicant]
“International Application Serial No. PCT/US2021/053012, Written Opinion mailed Jan. 21, 2022”, 7 pgs. [cited by applicant]
Clark, Aidan, “Adversarial Video Generation on Complex Datasets”, arXiv:1907.06571v2, arxiv.org, Cornell University Library, 201Olin Library Cornell University Ithaca, NY 14853, (2019), 21 pgs. [cited by applicant]
Karras, Tero, “Training Generative Adversarial Networks with Limited Data”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, (Jun. 11, 2020), 32 pgs. [cited by applicant]
Tulyakov, Sergey, “MoCoGAN: Decomposing Motion and Content for Video Generation”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, Ny 14853, (Jul. 17, 2017), 11 pgs. [cited by applicant]
“Chinese Application Serial No. 202180066923.5, Office Action mailed Jun. 25, 2025”, w/ English Translation, 21 pgs. [cited by applicant]
“International Application Serial No. PCT/US2021/053012, International Preliminary Report on Patentability mailed Apr. 13, 2023”, 9 pgs. [cited by applicant]
“Chinese Application Serial No. 202180066923.5, Response filed Oct. 20, 2025 to Office Action mailed Jun. 25, 2025”, W/ English Claims, 17 pgs. [cited by applicant]
“Korean Application Serial No. 10-2023-7014448, Notice of Preliminary Rejection mailed Jan. 20, 2026”, w/ English translation, 12 pgs. [cited by applicant]
Cao, Yangjie, “Review of computer vision based on generative adversarial networks”, Journal of Image and Graphics, vol. 23, No. 10, (Nov. 7, 2018), 1433-1449. [cited by applicant]
Tulyakov, Sergey, “MoCoGAN: Decomposing Motion and Content for Video Generation”, CVPR, (Dec. 14, 2017), 13 pgs. [cited by applicant]