IP Library › Granted Patent US 12,586,610
Granted Patent B2
US 12,586,610 · App. 18/573,097 · Granted Mar 24, 2026

Method, apparatus, device, storage medium and program product for video generation

Inventors: Xinwei Li (Beijing, CN); Jiajin Cao (Beijing, CN)
Assignee: Beijing Zitiao Network Technology Co., Ltd.
G11B27/036G06T11/60G11B27/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,610
App. No.
18/573,097
Granted
Mar 24, 2026
Kind
B2
Abstract

Embodiments of the present disclosure relates to a method, apparatus, device, storage medium, and program product for video generation. The method comprises: generating initial multimedia data based on received text data; obtaining a target editing template in response to an editing template obtaining request; applying the editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data; and generating target video based on the target multimedia data. Embodiments of the present disclosure generates video by directly applying the editing operation in the obtained editing template to the multimedia data, without the need for users to manually clip the video. This can not only reduce the time cost of video production, but also improve the quality of video production.

Claims (63)

1 . A method for video generation comprising:

generating initial multimedia data based on received text data, the initial multimedia data comprising spoken speech of the text data and a video image matching the text data, the initial multimedia data comprising at least one multimedia segment, the at least one multimedia segment respectively corresponding to at least one text segment divided from the text data, a target multimedia segment in the at least one multimedia segment corresponding to a target text segment in the at least one text segment, the target multimedia segment comprising a target video segment and a target speech segment, the target video segment comprising a video image matching the target text segment, and the target speech segment comprises a spoken speech matching the target text segment;

displaying a video editing area comprising a template control, wherein the template control is used to indicate editing the initial multimedia data using existing templates;

displaying a mask area in response to a triggering operation of the template control;

displaying at least one template theme control on the mask area;

in response to a triggering operation for a template theme control of the at least one template theme control, determining an editing template corresponding to the triggering operation as the target editing template;

obtaining the target editing template;

applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data; and

generating a target video based on the target multimedia data.

2 . The method of claim 1 , wherein the video image comprises subtitle text matching the target text segment.

3 . The method of claim 1 , wherein the editing operation indicated by the target editing template comprises a video synthesis operation; and

wherein applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data comprises:

synthesizing a video segment in the target editing template with the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data.

4 . The method of claim 3 , wherein synthesizing a video segment in the target editing template with the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data comprises:

loading the video segment in the target editing template to a predetermined position of the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data, wherein the predetermined position comprises a position before the first frame of media data of the initial multimedia data and/or a position after the last frame of media data of the initial multimedia data.

5 . The method of claim 1 , wherein the editing operation indicated by the target editing template comprises a transition setting operation; and

wherein applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data comprises:

adding a transition effect to the multimedia segment in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

6 . The method of claim 1 , wherein the editing operation indicated by the target editing template comprises a virtual object addition operation; and

wherein applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data comprises:

adding a virtual object in the target editing template to a predetermined position of the initial multimedia data based on the virtual object addition operation to obtain the target multimedia data.

7 . The method of claim 1 , wherein the editing operation indicated by the target editing template comprises a background audio addition operation; and

wherein applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data comprises:

mixing a background audio in the target editing template with the spoken speech in the initial multimedia data based on the background audio addition operation to obtain target multimedia data.

8 . The method of claim 1 , wherein the editing operation indicated by the target editing template comprises a keyword extraction operation; and

wherein applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data comprises:

extracting a keyword from the at least one target text segment; and

adding the keyword to the target multimedia segment corresponding to the target text segment.

9 . The method of claim 8 , wherein adding the keyword to the target multimedia segment corresponding to the target text segment comprises:

obtaining key text information matching the keyword; and

adding the keyword and the key text information to the target multimedia segment corresponding to the target text segment.

10 . An electronic device comprising:

one or more processors; and

a memory storing one or more programs thereon,

the one or more programs, when executed by the one or more processors, causing the electronic device to:

generate initial multimedia data based on received text data, the initial multimedia data comprising spoken speech of the text data and a video image matching the text data, the initial multimedia data comprising at least one multimedia segment, the at least one multimedia segment respectively corresponding to at least one text segment divided from the text data, a target multimedia segment in the at least one multimedia segment corresponding to a target text segment in the at least one text segment, the target multimedia segment comprising a target video segment and a target speech segment, the target video segment comprising a video image matching the target text segment, and the target speech segment comprises a spoken speech matching the target text segment;

display a video editing area comprising a template control, wherein the template control is used to indicate editing the initial multimedia data using existing templates;

display a mask area in response to a triggering operation of the template control;

display at least one template theme control on the mask area;

in response to a triggering operation for a template theme control of the at least one template theme control, determine an editing template corresponding to the triggering operation as the target editing template;

obtain the target editing template;

apply an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data; and

generate a target video based on the target multimedia data.

11 . The electronic device of claim 10 , wherein the video image comprises subtitle text matching the target text segment.

12 . The electronic device of claim 10 , wherein the editing operation indicated by the target editing template comprises a video synthesis operation; and

wherein the one or more programs, when executed by the one or more processors, cause the electronic device to apply an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data by:

synthesizing a video segment in the target editing template with the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data.

13 . The electronic device of claim 12 , wherein the one or more programs, when executed by the one or more processors, cause the electronic device to synthesize a video segment in the target editing template with the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data by:

loading the video segment in the target editing template to a predetermined position of the multimedia segments in the initial multimedia data based on the video synthesis operation to obtain the target multimedia data, wherein the predetermined position comprises a position before the first frame of media data of the initial multimedia data and/or a position after the last frame of media data of the initial multimedia data.

14 . The electronic device of claim 10 , wherein the editing operation indicated by the target editing template comprises a transition setting operation; and

wherein the one or more programs, when executed by the one or more processors, cause the electronic device to apply an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data by:

adding a transition effect to the multimedia segment in the initial multimedia data based on the transition setting operation to obtain the target multimedia data.

15 . The electronic device of claim 10 , wherein the editing operation indicated by the target editing template comprises a virtual object addition operation; and

wherein the one or more programs, when executed by the one or more processors, cause the electronic device to apply an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data by:

adding a virtual object in the target editing template to a predetermined position of the initial multimedia data based on the virtual object addition operation to obtain the target multimedia data.

16 . A non-transitory computer-readable storage medium storing a computer program thereon, the program, when executed by a processor, implementing:

generating initial multimedia data based on received text data, the initial multimedia data comprising spoken speech of the text data and a video image matching the text data, the initial multimedia data comprising at least one multimedia segment, the at least one multimedia segment respectively corresponding to at least one text segment divided from the text data, a target multimedia segment in the at least one multimedia segment corresponding to a target text segment in the at least one text segment, the target multimedia segment comprising a target video segment and a target speech segment, the target video segment comprising a video image matching the target text segment, and the target speech segment comprises a spoken speech matching the target text segment;

displaying a video editing area comprising a template control, wherein the template control is used to indicate editing the initial multimedia data using existing templates;

displaying a mask area in response to a triggering operation of the template control;

displaying at least one template theme control on the mask area;

in response to a triggering operation for a template theme control of the at least one template theme control, determining an editing template corresponding to the triggering operation as the target editing template;

applying an editing operation indicated by the target editing template to the initial multimedia data to obtain target multimedia data; and

generating a target video based on the target multimedia data.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2026
From: LI, XINWEI
To: SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.
Reel/Frame 073360/0033 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 5, 2026
From: CAO, JIAJIN; SHANGHAI SUIXUNTONG ELECTRONIC TECHNOLOGY CO., LTD.
To: BEIJING ZITIAO NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 073360/0152 →
Priority Claims (1)
CN 202210508063.2 · May 10, 2022 · national
Continuity (1)
Related Publication 20240296871A1 · Sep 5, 2024
References Cited (24)
US 11244488B2 · Watanabe · 2022 [cited by examiner]
US 11722727B2 · Li · 2023 [cited by examiner]
US 12154598B1 · Warnick · 2024 [cited by examiner]
US 20180143741A1 · Uriostegui · 2018 [cited by examiner]
US 20220044026A1 · Huang · 2022 [cited by applicant]
US 20220130427A1 · Allibhai · 2022 [cited by examiner]
CN 109756751A · 2019 [cited by applicant]
CN 110121103A · 2019 [cited by applicant]
CN 111243632A · 2020 [cited by applicant]
CN 111460183A · 2020 [cited by applicant]
CN 112449231A · 2021 [cited by applicant]
CN 112579826A · 2021 [cited by applicant]
CN 110572722B · 2021 [cited by applicant]
CN 112738623A · 2021 [cited by applicant]
CN 113452941A · 2021 [cited by applicant]
CN 113473182A · 2021 [cited by applicant]
CN 114339399A · 2022 [cited by applicant]
JP 2021033367A · 2021 [cited by applicant]
JP 2021069117A · 2021 [cited by applicant]
Communication pursuant to Rules 70(2) and 70a(2) EPC for European Application No. 23802924.3, mailed Oct. 22, 2024, 1 page. [cited by applicant]
Extended European Search Report for European Application No. 23802924.3, mailed Oct. 2, 2024, 8 pages. [cited by applicant]
Notice of Reasons for Refusal issued in JP Appl. No. 2023-578709 dated Jan. 21, 2025, English translation (14 pages). [cited by applicant]
“Home TV series movies variety shows”, Video production, Retrieved from the link: “https://v.youku.com/video?vid=XNTEzNjgzMTUyNA%3D%3D”, 2025, pp. 1-2. [cited by applicant]
Office action received from Chinese patent application No. 202210508063.2 mailed on Jun. 17, 2025, 16 pages (8 pages English Translation and 8 pages Original Copy). [cited by applicant]