IP Library Granted Patent US 12,524,940
Granted Patent B2
US 12,524,940 · App. 18/622,479 · Granted Jan 13, 2026

Method, apparatus, device and storage medium for video generation

Inventors: Aini Tang (Beijing, CN); Tianqi Zhang (Beijing, CN); Qizhi Zhang (Beijing, CN); Huimin Zhou (Beijing, CN); Hanqi Zheng (Beijing, CN); Haohua Zhong (Beijing, CN); Haoran Zhang (Beijing, CN); Gen Li (Beijing, CN)
Assignee: Beijing Zitiao Network Technology Co., Ltd.
G06T11/60G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,524,940
App. No.
18/622,479
Granted
Jan 13, 2026
Kind
B2
Abstract

The disclosure provides a method, an apparatus, a device and a storage medium for video generation. The method comprises: obtaining first text information used to describe a video effect requirement; obtaining at least one multimedia material; and generating a target video based on the first text information and the at least one multimedia material. The at least one multimedia material is presented in the target video. A video effect of the target video meets the video effect requirement described in the first text information. The target video is used to present a combination of at least one video segment. The at least one video segment is formed respectively based on respective video-image materials in the at least one multimedia material. The respective video-image materials comprise a video material and/or an image material.

Claims (82)

1 . A method of video generation, comprising:

displaying an interface configured to display a plurality of multimedia materials and to enable a selection of at least one multimedia material from the plurality of multimedia materials;

in response to selecting at least two multimedia materials, displaying an input box on the same interface via which the at least two multimedia materials are selected, wherein the input box is configured to receive text information indicating a requirement to be met by a target video to be generated;

extracting a feature label from the requirement indicated by first text information received via the input box;

comparing the feature label of the first text information with templates in a template library and determining a first video editing template that matches the feature label of the first text information;

extracting feature labels from the at least two multimedia materials;

comparing the feature labels of the at least two multimedia materials with the templates in the template library and determining a second video editing template that matches the feature labels of the at least two multimedia materials;

displaying the first video editing template and the second video editing template on a preview page;

generating a new video editing template based on the first video editing template and the second video editing template; and

generating the target video based on the new video editing template, the first text information received via the input box, and the at least two multimedia materials, wherein the target video is a new video, wherein the target video meets the requirement indicated by the first text information, the target video presents a combination of at least one video segment, the at least one video segment is formed respectively based on respective video-image materials in the at least two multimedia materials, and the respective video-image materials comprise at least one of video materials or image materials.

2 . The method of claim 1 , wherein the generating the target video comprises:

generating a video editing draft based on the first text information and the at least two multimedia materials, wherein the video editing draft comprises the at least two multimedia materials and editing information, the editing information is used to indicate an editing operation for the at least two multimedia materials, and the editing operation is at least used to edit the respective video-image materials in the at least two multimedia materials respectively into the at least one video segment, a video editing effect and/or the at least one multimedia material corresponding to the editing operation meet the video effect requirement described in the first text information; and

generating the target video based on the video editing draft.

3 . The method of claim 2 , wherein the generating a video editing draft based on the first text information and the at least two multimedia materials comprises:

determining at least one video editing template based on the first text information and the at least two multimedia materials, wherein the editing effect of the at least one video editing template meets the video effect requirement described in the first text information; and

applying an editing operation indicated by a target video editing template from the at least one video editing template to the at least two multimedia materials, to generate the video editing draft.

4 . The method of claim 3 , wherein after the determining at least one video editing template based on the first text information and the at least two multimedia materials, and the method further comprises:

selecting a third video editing template from the at least one video editing template and presenting the third video editing template on a preview page for a video editing effect, so that the preview page is used to preview the video effect obtained by importing the at least one multimedia material into the third video editing template, and the preview page being configured with an update recommendation control; and

in response to a trigger operation for the update recommendation control, selecting a fourth video editing template in the at least one video editing template, and replacing the third video editing template presented on the preview page with the fourth video editing template, so that the preview page is used to preview the video effect obtained by importing the at least one multimedia material into the fourth video editing template.

5 . The method of claim 3 , wherein after the determining at least one video editing template based on the first text information and the two multimedia materials, the method further comprises:

displaying, on a preview page, a fifth video editing template in the at least one video editing template;

obtaining adjusted text information, in response to a text adjustment operation on the preview page for the first text information;

determining a second video editing template set based on the adjusted text information and the at least one multimedia material; and

replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second video editing template set.

6 . The method of claim 5 , wherein before the determining a second video editing template set based on the adjusted text information and the at least one multimedia material, the method further comprises:

receiving a material adjustment operation for the at least one multimedia material to obtain an adjusted multimedia material; and

wherein the determining a second video editing template set based on the adjusted text information and the at least one multimedia material comprises determining the second video editing template set based on the adjusted text information and the adjusted multimedia material.

7 . The method of claim 1 , wherein obtaining the at least two multimedia materials comprises:

determining a matching first multimedia material from a user material set based on an analysis result for the first text information; or

generating a second multimedia material based on the analysis result for the first text information.

8 . The method of claim 1 , wherein the method further comprises:

displaying the input box in response to an importing operation for the at least two multimedia materials.

9 . The method of claim 1 , wherein the method further comprises:

displaying at least one video label, wherein the video label is configured to characterize the requirement of the target video to be generated; and

obtaining the first text information based on an operation of adding a target video label in the at least one video label to the input box.

10 . A non-transitory computer-readable storage medium, wherein the computer-readable storage medium stores an instruction that, when executed on a terminal device, causes the terminal device to implement operations comprising:

displaying an interface configured to display a plurality of multimedia materials and to enable a selection of at least one multimedia material from the plurality of multimedia materials;

in response to selecting at least two multimedia materials, displaying an input box on the same interface via which the at least two multimedia materials are selected, wherein the input box is configured to receive text information indicating a requirement to be met by a target video to be generated;

extracting a feature label from the requirement indicated by first text information received via the input box;

comparing the feature label of the first text information with templates in a template library and determining a first video editing template that matches the feature label of the first text information;

extracting feature labels from the at least two multimedia materials;

comparing the feature labels of the at least two multimedia materials with the templates in the template library and determining a second video editing template that matches the feature labels of the at least two multimedia materials;

displaying the first video editing template and the second video editing template on a preview page;

generating a new video editing template based on the first video editing template and the second video editing template; and

generating the target video based on the new video editing template, the first text information received via the input box, and the at least two multimedia materials, wherein the target video is a new video, wherein the target video meets the requirement indicated by the first text information, the target video presents a combination of at least one video segment, the at least one video segment is formed respectively based on respective video-image materials in the at least two multimedia materials, and the respective video-image materials comprise at least one of video materials or image materials.

11 . The non-transitory computer-readable storage medium of claim 10 , wherein the generating the target video comprises:

generating a video editing draft based on the first text information and the at least two multimedia materials, wherein the video editing draft comprises the at least two multimedia materials and editing information, the editing information is used to indicate an editing operation for the at least two multimedia materials, and the editing operation is at least used to edit the respective video-image materials in the at least two multimedia materials respectively into the at least one video segment, a video editing effect and/or the at least one multimedia material corresponding to the editing operation meet the video effect requirement described in the first text information; and

generating the target video based on the video editing draft.

12 . The non-transitory computer-readable storage medium of claim 11 , wherein the generating a video editing draft based on the first text information and the at least two multimedia materials comprises:

determining at least one video editing template based on the first text information and the at least two multimedia materials, wherein the editing effect of the at least one video editing template meets the video effect requirement described in the first text information; and

applying an editing operation indicated by a target video editing template from the at least one video editing template to the at least two multimedia materials, to generate the video editing draft.

13 . The non-transitory computer-readable storage medium of claim 12 , wherein after the determining at least one video editing template based on the first text information and the at least two multimedia materials, and the operations further comprise:

selecting a third video editing template from the at least one video editing template and presenting the third video editing template on a preview page for a video editing effect, so that the preview page is used to preview the video effect obtained by importing the at least one multimedia material into the third video editing template, and the preview page being configured with an update recommendation control; and

in response to a trigger operation for the update recommendation control, selecting a fourth video editing template in the at least one video editing template, and replacing the third video editing template presented on the preview page with the fourth video editing template, so that the preview page is used to preview the video effect obtained by importing the at least one multimedia material into the fourth video editing template.

14 . The non-transitory computer-readable storage medium of claim 12 , wherein after the determining at least one video editing template based on the first text information and the at least two multimedia materials, the operations further comprise:

displaying, on a preview page, a fifth video editing template in the at least one video editing template;

obtaining adjusted text information, in response to a text adjustment operation on the preview page for the first text information;

determining a second video editing template set based on the adjusted text information and the at least one multimedia material; and

replacing the fifth video editing template displayed on the preview page with a sixth video editing template in the second video editing template set.

15 . The non-transitory computer-readable storage medium of claim 10 , wherein obtaining the at least two multimedia materials comprises:

determining a matching first multimedia material from a user material set based on an analysis result for the first text information; or

generating a second multimedia material based on the analysis result for the first text information.

16 . The non-transitory computer-readable storage medium of claim 10 , wherein the operations further comprise:

displaying the input box in response to an importing operation for the at least two multimedia materials.

17 . The non-transitory computer-readable storage medium of claim 10 , wherein the operations further comprise:

displaying at least one video label, wherein the video label is configured to characterize the requirement of the target video to be generated; and

obtaining the first text information based on an operation of adding a target video label in the at least one video label to the input box.

18 . A device for video processing, comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program, when executed by the processor, implement operations comprising:

displaying an interface configured to display a plurality of multimedia materials and to enable a selection of at least one multimedia material from the plurality of multimedia materials;

in response to selecting at least two multimedia materials, displaying an input box on the same interface via which the at least two multimedia materials are selected, wherein the input box is configured to receive text information indicating a requirement to be met by a target video to be generated;

extracting a feature label from the requirement indicated by first text information received via the input box;

comparing the feature label of the first text information with templates in a template library and determining a first video editing template that matches the feature label of the first text information;

extracting feature labels from the at least two multimedia materials;

comparing the feature labels of the at least two multimedia materials with the templates in the template library and determining a second video editing template that matches the feature labels of the at least two multimedia materials;

displaying the first video editing template and the second video editing template on a preview page;

generating a new video editing template based on the first video editing template and the second video editing template; and

generating the target video based on the new video editing template, the first text information received via the input box, and the at least two multimedia materials, wherein the target video is a new video, wherein the target video meets the requirement indicated by the first text information, the target video presents a combination of at least one video segment, the at least one video segment is formed respectively based on respective video-image materials in the at least two multimedia materials, and the respective video-image materials comprise at least one of video materials or image materials.

19 . The device of claim 18 , the operations further comprising:

displaying at least one video label, wherein the video label is configured to characterize the requirement of the target video to be generated; and

obtaining the first text information based on an operation of adding a target video label in the at least one video label to the input box.

20 . The device of claim 18 , the operations further comprising:

displaying the new video editing template on a preview page, wherein the preview page is configured to enable a preview of the target video by importing the at least two multimedia materials into a video editing template.

Assignments (7)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: TANG, AINI
To: SHENZHEN LEMON TECHNOLOGY CO., LTD.
Reel/Frame 073227/0169 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: ZHANG, TIANQI; ZHANG, HAORAN
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073227/0277 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: ZHANG, QIZHI; ZHOU, HUIMIN; ZHENG, HANQI; ZHONG, HAOHUA
To: LEMON TECHNOLOGY (SHENZHEN) CO., LTD.
Reel/Frame 073227/0418 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: LI, GEN
To: BEIJING OCEAN ENGINE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073227/0495 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: SHENZHEN LEMON TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073227/0590 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: LEMON TECHNOLOGY (SHENZHEN) CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073227/0645 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2025
From: BEIJING OCEAN ENGINE NETWORK TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073227/0750 →
Priority Claims (1)
CN 202310446304.X · Apr 23, 2023 · national
Continuity (2)
Continuation PCTCN2023136857 · Dec 6, 2023
Related Publication 20240355024A1 · Oct 24, 2024
References Cited (20)
US 10049477B1 · Kokemohr · 2018 [cited by examiner]
US 20070083851A1 · Huang et al. · 2007 [cited by applicant]
US 20180096708A1 · Choi et al. · 2018 [cited by applicant]
US 20190295532A1 · Ammedick · 2019 [cited by examiner]
US 20200411053A1 · Han · 2020 [cited by examiner]
US 20220188352A1 · Wu · 2022 [cited by applicant]
US 20230162502A1 · Patel · 2023 [cited by examiner]
US 20230282240A1 · Wiersema · 2023 [cited by examiner]
US 20240107127A1 · Chen et al. · 2024 [cited by applicant]
CN 110996017A · 2020 [cited by applicant]
CN 112579826A · 2021 [cited by examiner]
CN 112637675A · 2021 [cited by applicant]
CN 113518160A · 2021 [cited by applicant]
CN 115442539A · 2022 [cited by applicant]
WO WO2022088783A1 · 2022 [cited by applicant]
Unschooler, “Meitu Video Editing Tutorial”, downloaded @https://www.youtube.com/watch?v=6VUoS7AUjEs, uploaded on Jun. 30, 2020) (Year: 2020). [cited by examiner]
Stackoverflow forum, downloaded @https://stackoverflow.com/questions/43575859/any-way-to-use-image-with-select-list, available online on Apr. 23, 2017 (Year: 2017). [cited by examiner]
Xiohoo, Xiohoo Quick Tips—Video Editing on Meitu: Transitions Tool, downloaded @ Xiohoo Quick Tips—Video Editing on Meitu: TransitionsTool (Year: 2023). [cited by examiner]
International Patent Application No. PCT/CN2023/136857; Int'l Search Report; dated Feb. 27, 2024; 2 pages. [cited by applicant]
European Patent Application No. 23866709.1; Extended Search Report; dated Jan. 17, 2025; 10 pages. [cited by applicant]