IP Library Granted Patent US 12682932
Granted Patent B2
US 12682932 · App. 18/652,942 · Granted Jul 14, 2026

Systems and methods for automatically generating a video production

Inventors: Jackson Grant (Brisbane, AU); Bhautik Jitendra Joshi (Sydney, AU); Cheng Chen (Melbourne, AU); Aditya Sangram Singh Rana (Vienna, AT); Emilio Tylson Baixauli (Vienna, AT)
Assignee: Canva Pty Ltd
G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682932
App. No.
18/652,942
Granted
Jul 14, 2026
Kind
B2
Abstract

Described herein is a computer implemented method for automatically generating a video production. The method includes determining a production description and a set of media items. A media item description corresponding to each media item is generated, and a prompt based on the production description and the media item descriptions is generated. The method further includes generating, using the prompt, cohesion information that includes a caption corresponding to each media item and automatically generating the video production based on the set of media items and the cohesion information. The video production is generated to include a set of one more scenes; each scene corresponds to a particular media item; and each scene is generated so the caption corresponding to the particular media item that the scene corresponds to is displayed during playback of the scene.

Claims (71)

1 . A computer implemented method for automatically generating a video production, the method comprising:

determining a production description, the production description including text that generally describes the video production that is to be generated;

determining a set of media items, the set of media items including one or more media items that are to be included in the video production that is to be generated;

generating, by a computer processing unit, a set of media item descriptions, the set of media item descriptions including a media item description corresponding to each media item;

generating a prompt based on the production description, the set of media item descriptions and an ordering component, wherein the prompt is a text string and includes order determination text;

generating, using the prompt, cohesion information, the cohesion information including a set of captions and scene order information, the set of captions including a caption corresponding to each media item; and

automatically generating the video production based on the set of media items and the cohesion information, wherein:

the video production is generated to include a set of one or more scenes ordered based on the scene order information;

each scene corresponds to a particular media item; and

each scene is generated so the caption corresponding to the particular media item that the scene corresponds to is displayed during playback of the scene.

2 . The computer implemented method of claim 1 , wherein the set of media items includes one or more video-type media items.

3 . The computer implemented method of claim 1 , wherein the set of media items includes one or more image-type media items.

4 . The computer implemented method of claim 1 , wherein the set of media item descriptions is generated using a first machine learning model.

5 . The computer implemented method of claim 4 , wherein:

the set of media items includes a first media item that is a video-type media item; and

generating the set of media item descriptions includes generating a first media item description corresponding to the first media item by:

determining a representative image of the first media item; and

processing the representative image of the first media item using the first machine learning model.

6 . The computer implemented method of claim 4 , wherein

the set of media items includes a second media item that is an image-type media item; and

generating the set of media item descriptions includes generating a second media item description corresponding to the second media item by processing the second media item using the first machine learning model.

7 . The computer implemented method of claim 1 , wherein:

the prompt is generated based on prompt generation data that defines one or more set text components and one or more placeholder components; and

the prompt is generated to include:

the one or more set text components; and

replacement text in place of each placeholder component.

8 . The computer implemented method of claim 7 , wherein:

the prompt generation data includes a first placeholder component; and

generating the prompt includes generating first replacement text to be used in place of the first placeholder component, the first replacement text being based on the set of media item descriptions.

9 . The computer implemented method of claim 7 , wherein:

the prompt generation data includes a second placeholder component; and

generating the prompt includes generating second replacement text to be used in place of the second placeholder component, the second replacement text being based on the production description.

10 . The computer implemented method of claim 7 , wherein:

the prompt generation data includes a third placeholder component; and

generating the prompt includes generating third replacement text in place of the third placeholder component, the third replacement text being based on a number of media items in the set of media items.

11 . The computer implemented method of claim 7 , wherein the one or more set text components include one or more of: an output format text component; a tone component; and a form component.

12 . The computer implemented method of claim 1 , wherein the cohesion information is generated by using the prompt as input to a second machine learning model.

13 . The computer implemented method of claim 12 , wherein the second machine learning model is a large language model.

14 . The computer implemented method of claim 1 , wherein:

the set of media items includes a third media item and the cohesion information includes a third caption corresponding to the third media item; and

automatically generating the video production includes:

generating a first scene based on a third media item;

generating a first element that includes text based on the third caption; and

associating the first element with the first scene.

15 . A computer processing system comprising:

one or more computer processing units; and

non-transitory computer-readable medium storing instructions which, when executed by the one or more computer processing units, cause the one or more computer processing units to perform a method comprising:

determining a production description, the production description including text that generally describes the video production that is to be generated;

determining a set of media items, the set of media items including one or more media items that are to be included in the video production that is to be generated;

generating, by a computer processing unit, a set of media item descriptions, the set of media item descriptions including a media item description corresponding to each media item;

generating a prompt based on the production description, the set of media item descriptions and an ordering component, wherein the prompt is a text string and includes order determination text;

generating, using the prompt, cohesion information, the cohesion information including a set of captions and scene order information, the set of captions including a caption corresponding to each media item; and

automatically generating the video production based on the set of media items and the cohesion information, wherein:

the video production is generated to include a set of one or more scenes ordered based on the scene order information;

each scene corresponds to a particular media item; and

each scene is generated so the caption corresponding to the particular media item that the scene corresponds to is displayed during playback of the scene.

16 . The computer processing system of claim 15 , wherein:

the set of media items includes a first media item that is a video-type media item; and

generating the set of media item descriptions includes generating a first media item description corresponding to the first media item by:

determining a representative image of the first media item; and

processing the representative image of the first media item using a first machine learning model.

17 . Non-transitory storage storing instructions executable by one or more computer processing units to cause the one or more computer processing units to perform a method comprising:

determining a production description, the production description including text that generally describes the video production that is to be generated;

determining a set of media items, the set of media items including one or more media items that are to be included in the video production that is to be generated;

generating, by a computer processing unit, a set of media item descriptions, the set of media item descriptions including a media item description corresponding to each media item;

generating a prompt based on the production description, the set of media item descriptions and an ordering component, wherein the prompt is a text string and includes order determination text;

generating, using the prompt, cohesion information, the cohesion information including a set of captions and scene order information, the set of captions including a caption corresponding to each media item; and

automatically generating the video production based on the set of media items and the cohesion information, wherein:

the video production is generated to include a set of one or more scenes ordered based on the scene order information;

each scene corresponds to a particular media item; and

each scene is generated so the caption corresponding to the particular media item that the scene corresponds to is displayed during playback of the scene.