IP Library › Granted Patent US 12,562,194
Granted Patent B2
US 12,562,194 · App. 18/775,954 · Granted Feb 24, 2026

Method, apparatus, device and medium for generating a video

Inventors: Yan Zeng (Beijing, CN); Guoqiang Wei (Beijing, CN); Hang Li (Beijing, CN)
Assignee: Beijing Youzhuju Network Technology Co., Ltd.
G11B27/036G06F40/40G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,562,194
App. No.
18/775,954
Granted
Feb 24, 2026
Kind
B2
Abstract

Provided are a method, apparatus, device and medium for generating a video. In one method, a plurality of images for respectively describing a plurality of target images in a target video are received. A text for describing a content of the target video is received. The target video is generated based on the plurality of images and the text according to a generation model. With exemplary implementations of the present disclosure, the plurality of images received can serve as guiding data to determine a development direction of a story in the video, which contributes to the generation of a richer and more realistic dynamic video.

Claims (67)

1 . A method for generating a video, comprising:

receiving a first plurality of images for respectively describing a plurality of target images in a target video, wherein the first plurality of images comprises a first image for describing a first head image in the plurality of target images, and a second image for describing a first tail image in the plurality of target images;

receiving a text for describing a content of the target video; and

generating the target video based on the first plurality of images and the text according to a generation model, wherein the generating the target video comprises:

determining a feature for generating the target video, a first image feature of the first image being placed at a first position of the feature for generating the target video and a second image feature of the second image being placed at a tail position of the feature for generating the target video; and

generating the target video based on the feature for generating the target video and the text.

2 . The method according to claim 1 , wherein the second image is received via at least any of:

an import control for importing an address of the second image;

a draw control for drawing the second image; or

an edit control for editing the first image to generate the second image.

3 . The method according to claim 1 , further comprising:

providing a second plurality of images in the target video;

in response to receiving an interaction for a third image in the second plurality of images, determining a second head image for describing a further target video subsequent to the target video;

acquiring a further text for describing a content of the further target video, and a fourth image for describing a second tail image of the further target video; and

generating the further target video based on the third image, the fourth image and the further text according to the generation model.

4 . The method according to claim 1 , wherein the images and the text are determined by:

receiving a plurality of images for respectively describing a plurality of keyframes in a video comprising a plurality of video clips;

receiving a long text for describing a content of the video; and

determining the images and the text based on the plurality of images and the long text according to a machine learning model.

5 . The method according to claim 1 , wherein the images and the text are determined by:

receiving a summary text for describing a video comprising a plurality of video clips;

determining a plurality of texts for respectively describing contents of the plurality of video clips in the video based on the summary text according to a machine learning model;

generating a target image based on a target text in the plurality of texts according to the machine learning model; and

taking the target image as the image and the target text as the text.

6 . The method according to claim 1 , further comprising:

determining prompt information associated with a further target video subsequent to the target video based on the target video according to a machine learning model, wherein the prompt information comprises at least one image and a further text; and

generating the further target video based on the at least one image and the further text according to the generation model.

7 . The method according to claim 1 , wherein generating the target video further comprises:

receiving an influencing factor for specifying a degree of influence of the second image;

dividing, based on the influencing factor, a plurality of steps for calling a diffusion model in the generation model into a first stage and a second stage;

determining, in the first stage and the second stage, a reconstruction feature of the target video based on the first image and the second image according to the diffusion model; and

generating the target video based on the reconstruction feature by using a decoder model in the generation model.

8 . The method according to claim 7 , wherein determining the reconstruction feature of the target video comprises:

determining, in the first stage, an intermediate reconstruction feature associated with the target video based on the first image and the second image according to the diffusion model; and

determining, in the second stage, the reconstruction feature based on the first image and the intermediate reconstruction feature according to the diffusion model.

9 . The method according to claim 1 , wherein the generation model is determined by:

determining a first reference image and a second reference image from a plurality of reference images in a reference video;

receiving a reference text for describing the reference video; and

acquiring the generation model based on the first reference image, the second reference image and the reference text, the generation model being configured to generate the target video based on the text and any of the first image and the second image.

10 . The method according to claim 9 , wherein the first reference image is located at a head of the reference video, and the second reference image is located within a predetermined range at a tail of the reference video.

11 . The method according to claim 9 , wherein the generation model comprises an encoder model and a diffusion model, and acquiring the generation model based on the first reference image, the second reference image and the reference text comprises:

determining a first reference feature of the reference video by using the encoder model, the first reference feature comprising a plurality of reference image features of the plurality of reference images;

determining a second reference feature of the reference video by using the encoder model, the second reference feature comprising a first reference image feature of the first reference image and a second reference image feature of the second reference image; and

determining the diffusion model based on the first reference feature, the second reference feature, and the reference text.

12 . The method according to claim 11 , wherein a first position of the first reference image feature in the second reference feature corresponds to a position of the first reference image in the reference video, and a second position of the second reference image feature in the second reference feature corresponds to a position of the second reference image in the reference video.

13 . The method according to claim 12 , wherein a dimension of the second reference feature is equal to a dimension of the first reference feature, and features at positions other than the first position and the second position in the second reference feature are set to be empty.

14 . The method according to claim 13 , further comprising: setting the second reference image feature to be empty according to a predetermined condition.

15 . The method according to claim 11 , wherein determining the diffusion model based on the first reference feature, the second reference feature, and the reference text comprises:

performing noising processing for the first reference feature and the second reference feature respectively to generate a first noise reference feature and a second noise reference feature;

connecting the first noise reference feature and the second noise reference feature to generate a noise reference feature of the reference video;

determining a reconstruction feature of the reference video based on the noise reference feature and the reference text by using the diffusion model; and

updating the diffusion model based on a difference between the reconstruction feature and the reference feature.

16 . The method according to claim 1 , wherein the first head image is a first image frame in the target video, and the first tail image is a last image frame in the target video.

17 . An electronic device, comprising:

at least one processing unit; and

at least one memory coupled to the at least one processing unit and storing instructions for execution by the at least one processing unit, wherein the instructions, when executed by the at least one processing unit, cause the electronic device to perform acts for generating a video, the acts comprising:

receiving a first plurality of images for respectively describing a plurality of target images in a target video to be generated;

receiving a text for describing a content of the target video to be generated; and

generating the target video based on the first plurality of images and the text according to a generation model, wherein the generating the target video comprises:

determining a feature for generating the target video, a first image feature of the first image is placed at a first position of the feature for generating the target video and a second image feature of the second image is placed at a tail position of the feature for generating the target video; and

generating the target video based on the feature for generating the target video and the text.

18 . A non-transitory computer readable storage medium having a computer program stored thereon, wherein the computer program, when executed by a processor, causes the processor to implement acts for generating a video, the acts comprising:

receiving a first plurality of images for respectively describing a plurality of target images in a target video;

receiving a text for describing a content of the target video; and

generating the target video based on the first plurality of images and the text according to a generation model, wherein the generating the target video comprises:

determining a feature for generating the target video, a first image feature of the first image is placed at a first position of the feature for generating the target video and a second image feature of the second image is placed at a tail position of the feature for generating the target video; and

generating the target video based on the feature for generating the target video and the text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 20, 2026
From: ZENG, YAN; WEI, GUOQIANG; LI, HANG
To: BEIJING YOUZHUJU NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073520/0069 →
Priority Claims (1)
CN 202311543966.5 · Nov 17, 2023 · national
Continuity (1)
Related Publication 20250166667A1 · May 22, 2025
References Cited (48)
US 10218954B2 · Lakhani · 2019 [cited by examiner]
US 11277556B2 · Katou · 2022 [cited by examiner]
US 11797780B1 · Finegan · 2023 [cited by examiner]
US 11941885B2 · Balannik · 2024 [cited by examiner]
US 12010371B2 · Kikuchi · 2024 [cited by examiner]
US 12412370B2 · Hou · 2025 [cited by examiner]
US 20190325084A1 · Peng · 2019 [cited by examiner]
US 20230103947A1 · Shields · 2023 [cited by examiner]
US 20230118966A1 · Liu et al. · 2023 [cited by applicant]
US 20230260284A1 · Balannik · 2023 [cited by applicant]
US 20230377324A1 · Kim · 2023 [cited by examiner]
US 20240171807A1 · Fletcher · 2024 [cited by examiner]
US 20240420404A1 · Kasap · 2024 [cited by examiner]
US 20250229183A1 · Iwasaki · 2025 [cited by examiner]
CN 103650002A · 2014 [cited by applicant]
CN 112818955A · 2021 [cited by applicant]
CN 113051420A · 2021 [cited by applicant]
CN 113784171A · 2021 [cited by applicant]
CN 115186133A · 2022 [cited by applicant]
CN 116233491A · 2023 [cited by applicant]
CN 116320216A · 2023 [cited by applicant]
CN 116363563A · 2023 [cited by applicant]
CN 116740204A · 2023 [cited by applicant]
CN 116916112A · 2023 [cited by applicant]
CN 116939320A · 2023 [cited by applicant]
JP 2021033961A · 2021 [cited by applicant]
JP 2023062173A · 2023 [cited by applicant]
JP 2023095832A · 2023 [cited by applicant]
KR 1020180065498A · 2018 [cited by applicant]
KR 1020200032614A · 2020 [cited by applicant]
WO 2020150688A1 · 2020 [cited by applicant]
WO 2021164326A1 · 2021 [cited by applicant]
WO 2022221080A1 · 2022 [cited by applicant]
China National Intellectual Property Administration, Office Action Issued in Application No. 202311543966.5, Jul. 31, 2024, 9 pages. [cited by applicant]
Wang, Z. et al., “A Review of Text-to-Visual Speech Synthesis,” Journal of Computer Research and Development, vol. 43, No. 1, Jan. 28, 2006, 8 pages. Submitted with English translation of abstract. [cited by applicant]
Xu, Z. et al., “Moving target detection of the video images,” Computer Era, vol. 2006, No. 8, Aug. 25, 2006, 3 pages. Submitted with English translation of abstract. [cited by applicant]
China National Intellectual Property Administration, Notice of Allowance Issued in Application No. 202311543966.5, Nov. 6, 2024, 6 pages. [cited by applicant]
Japan Patent Office, Office Action Issued in Application No. 2024114361, Nov. 19, 2024, 8 pages. [cited by applicant]
Denton, R. et al., “Unsupervised Learning of Disentangled Representations from Video,” arXiv:1705.10915v1, May 31, 2017, 13 pages. [cited by applicant]
Tulyakov, S. et al., “MoCoGaN: Decomposing Motion and Content for Video Generation,” arXiv:1707.04993v2, Dec. 14, 2017, 14 pages. [cited by applicant]
China National Intellectual Property Administration, Notice of Grant of Patent Right for invention from Chinese patent application No. 202311543966.5 mailed on Nov. 6, 2024, 6 pages. [cited by applicant]
European Patent Office, Extended European Search Report for European Application No. 24189153.0, mailed Dec. 12, 2024, 8 Pages. [cited by applicant]
Lin H., et al., “VideoDirectorGPT: Consistent Multi-scene Video Generation via LLM-guided Planning”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 26, 2023, XP09162386… [cited by applicant]
Japan Patent Office, Notification of Grant Issued in Application No. 2024-114361, Mar. 25, 2025, 6 pages. [cited by applicant]
Dorkenwald, M. et al., “Stochastic Image-to-Video Synthesis using cINNs,” 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 17, 2021, 18 pages. [cited by applicant]
Hong, W. et al., “CogVideo: Large-scale Pretraining for Text-to-Video Generation via Transformers,” Available Online at https://arxiv.org/abs/2205.15868, May 29, 2022, 15 pages. [cited by applicant]
Korean Intellectual Property Office, Office Action Issued in Application No. 10-2024-0094655, Apr. 30, 2025, 12 pages. [cited by applicant]
China National Intellectual Property Administration, Office Action Issued in Application No. 202510020700.5, Sep. 27, 2025, 15 pages. [cited by applicant]