IP Library Granted Patent US 12,587,717
Granted Patent B2
US 12,587,717 · App. 18/668,847 · Granted Mar 24, 2026

Facilitating video generation

Inventors: Philip Martin Meier (San Diego, CA); Csaba Matyas Petre (San Diego, CA)
Assignee: 10Z, LLC.
H04N21/816H04N21/23418H04N21/23424H04N21/4318H04N21/8456
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,587,717
App. No.
18/668,847
Granted
Mar 24, 2026
Kind
B2
Abstract

Features described herein generally relate to content production. Particularly, the present disclosure relates to facilitating video generation. Using machine-learning models, a storyboard can be generated from an inspirational video, video attributes can be determined for the storyboard, editing scores and actions can be determined for candidate videos, candidate videos can be edited based on the editing scores and actions, and the edited candidate videos can be combined to generate a video.

Claims (55)

1 . A method for facilitating video generation, the method comprising:

identifying a first video;

segmenting the first video into a plurality of video segments;

automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;

automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;

identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;

determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and

editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;

combining the second videos to generate a third video; and

outputting the third video.

2 . The method of claim 1 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.

3 . The method of claim 1 , wherein the plurality of video attributes is automatically determined.

4 . The method of claim 1 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing scores.

5 . The method of claim 4 , wherein the other machine-learning model used to determine the plurality of editing scores or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.

6 . The method of claim 1 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing actions.

7 . The method of claim 6 , wherein the other machine-learning model used to determine the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.

8 . A system for facilitating video generation comprising:

one or more processors; and

one or more memories, the one or more memories storing instructions which, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

identifying a first video;

segmenting the first video into a plurality of video segments;

automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;

automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;

identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;

determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and

editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;

combining the second videos to generate a third video; and

outputting the third video.

9 . The system of claim 8 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.

10 . The system of claim 8 , wherein the plurality of video attributes is automatically determined.

11 . The system of claim 8 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing scores.

12 . The system of claim 11 , wherein the other machine-learning model used to determine the plurality of editing scores or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.

13 . The system of claim 8 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing actions.

14 . The system of claim 13 , wherein the other machine-learning model used to determine the plurality of editing actions or the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.

15 . One or more non-transitory computer-readable media storing computer-readable instructions that, when executed by one or more processors, cause the one or more processors to perform operations including:

identifying a first video;

segmenting the first video into a plurality of video segments;

automatically generating a plurality of storyboard panels based on the plurality of video segments using a machine-learning model, each storyboard panel of the plurality of storyboard panels representing at least one video segment of the plurality of video segments;

automatically determining a plurality of video attributes by processing each video segment of the plurality of video segments, each video attribute of the plurality of video attributes corresponding to a storyboard panel of the plurality of storyboard panels;

identifying a plurality of second videos, each second video of the plurality of second videos corresponding to at least one video attribute of the plurality of video attributes;

determining a plurality of editing scores or a plurality of editing actions for the plurality of second videos by processing the plurality of second videos using another machine-learning model; and

editing the plurality of second videos based on the plurality of editing scores or the plurality of editing actions;

combining the second videos to generate a third video; and

outputting the third video.

16 . The one or more non-transitory computer-readable media of claim 15 , wherein each storyboard panel of the plurality of storyboard panels includes a textual label describing at least one of content depicted in a video segment of the plurality of video segments and a video attribute of the plurality of video attributes.

17 . The one or more non-transitory computer-readable media of claim 15 , wherein the plurality of video attributes is automatically determined using the other machine-learning model.

18 . The one or more non-transitory computer-readable media of claim 15 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing scores.

19 . The one or more non-transitory computer-readable media of claim 18 , wherein the plurality of editing scores or the plurality of editing actions includes

the plurality of editing actions.

20 . The one or more non-transitory computer-readable media of claim 19 , wherein the other machine-learning model used to determine the plurality of editing scores and used to determine the plurality of editing actions includes an encoder and decoder, a computational fabric, or a combination thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2024
From: MEIER, PHILIP MARTIN; PETRE, CSABA MATYAS
To: 10Z, LLC.
Reel/Frame 067784/0132 →
Continuity (4)
Continuation PCTUS2022051101 · Nov 28, 2022
Provisional Application 63281700 · Nov 21, 2021
Provisional Application 63319303 · Mar 12, 2022
Related Publication 20240314406A1 · Sep 19, 2024
References Cited (11)
US 9992556B1 · Price · 2018 [cited by examiner]
US 10771867B1 · Chemolosov · 2020 [cited by examiner]
US 20190130192A1 · Kauffmann · 2019 [cited by examiner]
US 20190370554A1 · Meier · 2019 [cited by examiner]
US 20210304468A1 · Doggett · 2021 [cited by examiner]
US 20210390317A1 · Kim · 2021 [cited by examiner]
US 20250203178A1 · Mishra · 2025 [cited by examiner]
EP 2044591B1 · 2014 [cited by applicant]
Apostolidis, E. et al. “Video Summarization Using Deep Neural Networks: A Survey”, Proceedings of the IEEE, vol. 109, No. 11, Nov. 2, 2021. [cited by applicant]
International Search Report and Written Opinion for PCT Patent Application No. PCT/US2022/051101, dated Feb. 24, 2023. [cited by applicant]
Sah, S et al., “Semantic Text Summarization of Long Videos”, 2017 IEEE Winter Conference on Applications of Computer Vision, Mar. 2017. [cited by applicant]