IP Library Granted Patent US 12,299,796
Granted Patent B2
US 12,299,796 · App. 18/067,642 · Granted May 13, 2025

Generation of story videos corresponding to user input using generative models

Inventors: Bingchen Liu (Los Angeles, CA); Kin Chung Wong (Los Angeles, CA); Peilin Li (Los Angeles, CA); Lexin Tang (Los Angeles, CA)
Assignee: LEMON INC.
G06T13/80G06F3/167G06F40/30G06T7/50G10L13/02G06T2200/24G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,299,796
App. No.
18/067,642
Granted
May 13, 2025
Kind
B2
Abstract

The present disclosure provides systems and methods for video generation corresponding to a user input. Given a user input, a story video with content relevant to the user input can be generated. One aspect includes a computing system comprising a processor and memory. The processor can be configured to execute a program using portions of the memory to receive the user input, generate a story text based on the user input, generate a plurality of story images based on the story text, and output a story including the story text and a story video having content corresponding to the story text, wherein the story video includes the plurality of story images. Additionally or alternatively, the story video can include audio data and a plurality of generated animated videos, each animated video corresponding to a story image in the plurality of story images.

Claims (48)

1. A computing system for video generation corresponding to a user input, the computing system comprising:

a processor and memory, the processor configured to execute a program using portions of the memory to:

receive the user input;

generate a story text based on the user input, wherein the story text comprises a plurality of sentences generated by:

generating a sentence based on the user input; and

iteratively generating successive sentences using one or more of previously generated sentences and the user input;

generate a plurality of story images based on the story text, each of the plurality of story images corresponding to one of the plurality of sentences; and

output a story, the story including the story text and a story video having content corresponding to the story text, wherein the story video comprises a plurality of animated videos, each animated video corresponding to a respective story image of the plurality of story images.

2. The computing system of claim 1 , wherein the story video is generated by:

for each of the plurality of animated videos:

performing monocular depth estimation on the respective story image corresponding to the animated video;

generating the animated video by animating the corresponding respective story image using information from the monocular depth estimation.

3. The computing system of claim 1 , wherein the processor is further configured to provide audio data, and wherein the story includes the audio data.

4. The computing system of claim 3 , wherein the audio data is provided based on a selection by a user.

5. The computing system of claim 3 , wherein the audio data is provided based on the story text.

6. The computing system of claim 1 , wherein the user input comprises an input sentence.

7. The computing system of claim 1 , wherein the user input includes an artistic style descriptor, and the plurality of story images is generated based on the artistic style descriptor.

8. The computing system of claim 1 , wherein the plurality of sentences is generated using a sequence-to-sequence transformer model.

9. The computing system of claim 1 , wherein the plurality of story images is generated using a generative diffusion model.

10. A method for video generation corresponding to a user input, the method comprising:

receiving the user input;

generating a story text based on the user input, wherein the story text comprises a plurality of sentences generated by:

generating a sentence based on the user input; and

iteratively generating successive sentences using one or more of previously generated sentences and the user input;

generating a plurality of story images based on the story text, each of the plurality of story images corresponding to one of the plurality of sentences; and

outputting a story including the story text and a story video, wherein the story video comprises a plurality of animated videos, each animated video corresponding to a respective story image of the plurality of story images.

11. The method of claim 10 , wherein the story video is generated by:

for each of the plurality of animated videos:

performing monocular depth estimation on the respective story image corresponding to the animated video;

generating the animated video by animating the corresponding respective story image using information from the monocular depth estimation.

12. The method of claim 10 , further comprising providing audio data based on the story text using a sentiment analysis model, wherein the story includes the audio data.

13. The method of claim 10 , wherein the plurality of sentences is generated using a sequence-to-sequence transformer model.

14. The method of claim 10 , wherein the plurality of story images is generated using a generative diffusion model and a language-text matching model.

15. A computing system for video generation corresponding to a user input, the computing system comprising:

a processor and memory, the processor configured to execute a program using portions of the memory to:

receive the user input, wherein the user input includes one or more words;

generate a story text based on the user input, wherein the story text includes a plurality of sentences having contextual coherence, and wherein the plurality of sentences is generated by:

generating a sentence based on the user input; and

iteratively generating successive sentences using one or more of previously generated sentences and the user input;

generate a plurality of story images, wherein each of the plurality of story images corresponds to a sentence in the plurality of sentences;

select audio data based on the story text; and

output a story including the story text, the selected audio data, and a story video having content corresponding to the story text, wherein the story video comprises a plurality of animated videos, each animated video corresponding to a respective story image of the plurality of story images.

16. The computing system of claim 15 , wherein the story video is generated by:

for each of the plurality of animated videos:

performing monocular depth estimation on the respective story image corresponding to the animated video;

generating the animated video by animating the corresponding respective story image using information from the monocular depth estimation.

17. The computing system of claim 15 , wherein the audio data is selected based on the story text using a sentiment analysis model.

18. The computing system of claim 15 , wherein the plurality of story images is generated using a generative diffusion model and a language-text matching model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: WONG, KIN CHUNG; LIU, BINGCHEN; LI, PEILIN; TANG, LEXIN
To: BYTEDANCE INC.
Reel/Frame 065184/0530 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 11, 2023
From: BYTEDANCE INC.
To: LEMON INC.
Reel/Frame 065184/0920 →
Continuity (1)
Related Publication 20230118966A1 · Apr 20, 2023
References Cited (35)
US 11886955B2 · David · 2024 [cited by applicant]
US 20020039199A1 · Nose et al. · 2002 [cited by applicant]
US 20120044244A1 · Huang et al. · 2012 [cited by applicant]
US 20190130221A1 · Bose et al. · 2019 [cited by applicant]
US 20200036861A1 · Gondek et al. · 2020 [cited by applicant]
US 20200065342A1 · Panuganty · 2020 [cited by examiner]
US 20200134089A1 · Sankaran · 2020 [cited by examiner]
US 20200168252A1 · Bhuruth · 2020 [cited by examiner]
US 20200272695A1 · Dogan et al. · 2020 [cited by applicant]
US 20200342328A1 · Revaud et al. · 2020 [cited by applicant]
US 20220084204A1 · Li et al. · 2022 [cited by applicant]
US 20220121702A1 · Kale et al. · 2022 [cited by applicant]
US 20220277218A1 · Fan et al. · 2022 [cited by applicant]
US 20230237725A1 · Zoss et al. · 2023 [cited by applicant]
US 20230267315A1 · Kingma et al. · 2023 [cited by applicant]
US 20230368337A1 · Karras et al. · 2023 [cited by applicant]
US 20230377099A1 · Kreis · 2023 [cited by examiner]
US 20230377226A1 · Saharia et al. · 2023 [cited by applicant]
US 20230410270A1 · Fujita · 2023 [cited by applicant]
US 20240005604A1 · Kreis et al. · 2024 [cited by applicant]
US 20240037810A1 · Gong et al. · 2024 [cited by applicant]
US 20240087179A1 · Min et al. · 2024 [cited by applicant]
US 20240111894A1 · Kreis et al. · 2024 [cited by applicant]
US 20240112088A1 · Yu et al. · 2024 [cited by applicant]
US 20240115954A1 · Olson et al. · 2024 [cited by applicant]
US 20240153093A1 · Xu et al. · 2024 [cited by applicant]
EP 3620935A1 · 2020 [cited by applicant]
WO 2023239358A1 · 2023 [cited by applicant]
Zbinden, R., “Implementing and Experimenting with Diffusion Models for Text-to-Image Generation,” Master Thesis, École Polytechnique Fédérale de Lausanne, Machine Learning and Optimization Lab (MLO), Sep. 23, 2022, 63 p… [cited by applicant]
Xu, S., “LIP-Diffusion-LM: Apply Diffusion Model on Image Captioning,” arXiv Computer Vision and Pattern Recognition (cs.CV), Oct. 10, 2022, 10 pages. [cited by applicant]
Kim, G. et al., “DiffusionCLIP: Text-Guided Diffusion Models for Robust Image Manipulation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR 2022), Jun. 19, 2022, New Orleans, Lou… [cited by applicant]
Nichol, A. et al., “GLIDE: Towards Photorealistic Image Generation and Editing with Text-Guided Diffusion Models,” Proceedings of the 39th International Conference on Machine Learning (ICML 2022), Jul. 17, 2022, Baltimo… [cited by applicant]
Ramesh, A. et al., “Hierarchical Text-Conditional Image Generation with CLIP Latents,” arXiv Computer Vision and Pattern Recognition (cs.CV), Apr. 13, 2022, 27 pages. [cited by applicant]
United States Patent and Trademark Office, Office Action Issued in U.S. Appl. No. 18/052,865, filed Aug. 22, 2024, 19 pages. [cited by applicant]
United States Patent and Trademark Office, Office Action Issued in U.S. Appl. No. 18/052,865, filed Feb. 26, 2025, 18 pages. [cited by applicant]
Cited By (2)
US 12,494,004 US 12,700,145