IP Library Granted Patent US 12,725,635
Granted Patent B2
US 12,725,635 · App. 18/446,673 · Granted Sep 1, 2026

Systems and methods for automating video editing

Inventors: Geneviève Patterson (Nottingham, NH); Hannah Wensel (Los Angeles, CA)
Assignee: Visual Supply Company
G11B27/031G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,635
App. No.
18/446,673
Granted
Sep 1, 2026
Kind
B2
Abstract

Provided are systems and methods for automatic video processing that employ machine learning models to process input video and understand user video content in a semantic and cultural context. This recognition enables the processing system to recognize interesting temporal events, and build narrative video sequences automatically, for example, by linking or interleaving temporal events or other content with film-based categorizations. In further embodiments, the implementation of the processing system is adapted to mobile computing platforms which can be distributed as an “app” within various app stores. In various example, the mobile apps turn everyday users into professional videographers. In further embodiments, music selection and dialog based editing can likewise be automated via machine learning models to create dynamic and professional quality video segments.

Claims (38)

1 . A video processing system, comprising:

a video processing component, executed by at least one processor, configured to perform operations comprising:

accepting a first user sourced video input generated by the first user;

executing a first machine learning process to perform operations comprising:

analyzing the user sourced video input and decomposing the user sourced video input into video segments comprising the user sourced video input, the decomposing into the video segments being based, at least in part, on determining importance of sections and content within the user sourced video input, the video segments comprising trimmed clip segments having variable durations;

transforming the video segments into a semantic embedding space comprising feature vectors of the video segments in a multi-dimensional space;

classifying the transformed video segments into at least one of cinematic categories or spatial cinematic layout categories based on determining similarity between classified video segments having associated cinematic categories and spatial cinematic layout categories and the transformed video segments;

in response to selection of an automatic editing function, editing automatically at least a first or second video segment of the first user sourced video input, the edit execution including at least one of altering the duration of the first or second video segment and introducing at least one visual effect into the first or second video segment based, at least in part, on the cinematic categories or the spatial cinematic layout categories and a machine learning model trained or fine-tuned on approved edits, the edits including altering the duration and introducing at least one visual effect to one or more video segments; and

generating a rough-cut video output including a sequence of video including edited versions of the first or second video segment and at least some of the plurality of video segments from the first user sourced video.

2 . The system of claim 1 , wherein the operations further comprise:

automatically identifying using a second machine learning process a narrative goal based on analysis of the first user sourced video input; and

defining a new sequencing of the first user sourced video to convey the narrative goal based on a third machine learning algorithm trained on film-based categorizations of video segments, the film-based categorizations including at least cinematic style.

3 . The system of claim 1 , wherein the video processing component includes at least a first neural network configured to transform the first user sourced video input into a semantic embedding space.

4 . The system of claim 3 , wherein the first neural network comprises a convolutional neural network.

5 . The system of claim 3 , wherein the first neural network is configured to classify user video into visual concept categories.

6 . The system of claim 4 , wherein the video processing component further comprises a second neural network configured to determine a narrative goal associated with the first user sourced video input or the sequence of video to be displayed.

7 . The system of claim 6 , wherein the second neural network comprises a long term short term memory recurrent network.

8 . The system of claim 1 , further comprising a second neural network configured to classify visual beats within user sourced video.

9 . The system of claim 8 , wherein the operations further comprise automatically selecting at least one soundtrack for the first user sourced video input.

10 . The system of claim 1 , wherein the semantic embedding space comprises respective numerical representation of the respective video segments in a multiple dimensioned space, and the numerical values from the sematic embedding space are input into a neural network to output a matching film idiom from the neural network.

11 . A computer implemented method for automatic video processing, the method comprising:

accepting, by at least one processor, a first user sourced video input generated by the first user;

analyzing, by the at least one processor, the user sourced video input and decomposing the user sourced video input into video segments comprising the user sourced video input, the decomposing being based, at least in part, on determining importance of sections and content within the user sourced video input, the video segments comprising trimmed clip segments having variable durations;

transforming, by the at least one processor, the video segments into a semantic embedding space comprising feature vectors of the video segments in a multi-dimensional space;

classifying, by the at least one processor, the transformed video segments into at least one of contextual cinematic categories or spatial cinematic layout categories based on determining similarity between classified video segments having associated cinematic categories and spatial cinematic layout categories and the transformed video segments;

editing, by the at least one processor, automatically at least a first or second video segment of the first user sourced video input in response to selection of an automatic editing function, the editing including at least one of altering the duration of the first or second video segment and introducing at least one visual effect into the first or second video segment based at least in part on the cinematic categories or the spatial cinematic layout categories and a machine learning model trained on reviewed or approved edits, the edits including altering the duration and introducing at least one visual effect to one or more video segments; and

generating, by the at least one processor, a rough-cut video output including a sequence of video including edited versions of first or second video segment and at least some of the plurality of video segments from the first user sourced video.

12 . The method of claim 11 , wherein the method further comprises:

automatically identifying, by the at least one processor, a narrative goal using a second machine learning process based on analysis of the first user sourced video input; and

defining, by the at least one processor, a new sequence of the first user sourced video input to convey the narrative goal based on a third machine learning algorithm trained on film-based categorizations of video segments, the film-based categorizations including at least cinematic style.

13 . The method of claim 11 , wherein the method further comprises executing at least a first neural network configured to transform the first user sourced video input into a semantic embedding space.

14 . The method of claim 13 , wherein the first neural network comprises a convolutional neural network.

15 . The method of claim 14 , wherein the method further comprises determining, by a second neural network, a narrative goal associated with the first user sourced video or the sequence of video to be displayed.

16 . The method of claim 15 , wherein the second neural network comprises a long term short term memory recurrent network.

17 . The method of claim 13 , wherein the method further comprises classifying user video into visual concept categories with the first neural network.

18 . The method of claim 11 , wherein the method further comprises classifying, by a third neural network, visual beats within the first user sourced video.

19 . The method of claim 18 , wherein the method further comprises automatically selecting at least one soundtrack for the user sourced video.

20 . The method of claim 11 , wherein the semantic embedding space comprises respective numerical representations of the respective video segments in a multiple dimensioned space, and the method includes processing at least some of the numerical values from the sematic embedding space as input into a neural network to output a matching film idiom from the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: PATTERSON, GENEVIÈVE; WENSEL, HANNAH
To: VISUAL SUPPLY COMPANY
Reel/Frame 065399/0600 →
Continuity (3)
Continuation 17189865 · Mar 2, 2021
Provisional Application 62983923 · Mar 2, 2020
Related Publication 20230386520A1 · Nov 30, 2023
References Cited (115)
US 8577210B2 · Kashima · 2013 [cited by applicant]
US 9146942B1 · Hodges et al. · 2015 [cited by applicant]
US 10062415B2 · Eppolito et al. · 2018 [cited by applicant]
US 10129573B1 · Sahasrabudhe · 2018 [cited by examiner]
US 10419773B1 · Wei · 2019 [cited by examiner]
US 10750245B1 · Zeiler et al. · 2020 [cited by applicant]
US 11170389B2 · Wang et al. · 2021 [cited by applicant]
US 11769528B2 · Patterson et al. · 2023 [cited by applicant]
US 20070074115A1 · Patten et al. · 2007 [cited by applicant]
US 20070192107A1 · Sitomer et al. · 2007 [cited by applicant]
US 20120057852A1 · Devleeschouwer et al. · 2012 [cited by applicant]
US 20130031479A1 · Flower · 2013 [cited by applicant]
US 20130195429A1 · Fay et al. · 2013 [cited by applicant]
US 20130343729A1 · Rav-Acha et al. · 2013 [cited by applicant]
US 20140133834A1 · Shannon · 2014 [cited by applicant]
US 20140255009A1 · Svendsen et al. · 2014 [cited by applicant]
US 20150208023A1 · Boyle et al. · 2015 [cited by applicant]
US 20160092561A1 · Liu et al. · 2016 [cited by applicant]
US 20160148407A1 · Hodges et al. · 2016 [cited by applicant]
US 20170124400A1 · Yehezkel Rohekar · 2017 [cited by examiner]
US 20170150235A1 · Mei et al. · 2017 [cited by applicant]
US 20170238055A1 · Chang et al. · 2017 [cited by applicant]
US 20180115706A1 · Kang et al. · 2018 [cited by applicant]
US 20190012697A1 · Nemani · 2019 [cited by examiner]
US 20190208124A1 · Newman et al. · 2019 [cited by applicant]
US 20190303403A1 · More et al. · 2019 [cited by applicant]
US 20190364211A1 · Chun · 2019 [cited by applicant]
US 20200021718A1 · Barbu et al. · 2020 [cited by applicant]
US 20200184278A1 · Zadeh et al. · 2020 [cited by applicant]
US 20200228692A1 · Wakamatsu · 2020 [cited by examiner]
US 20210117685A1 · Sureshkumar et al. · 2021 [cited by applicant]
US 20210272599A1 · Patterson et al. · 2021 [cited by applicant]
International Search Report and Written Opinion mailed May 21, 2021, in connection with International Application No. PCT/US2021/020424. [cited by applicant]
International Preliminary Report on Patentability dated Sep. 15, 2022, in connection with International Application No. PCT/US2021/020424. [cited by applicant]
Extended European Search Report dated Feb. 23, 2024, in connection with European Application No. 21763998.8. [cited by applicant]
[No Author Listed], Explore Computer Vision APIs. WWDC 2020. 9 pages. https://developer.apple.com/videos/play/wwdc2020/10673/ [Last accessed Sep. 13, 2022]. [cited by applicant]
[No Author Listed], 50 Must-Know Stats About Video Marketing. Insivia Agile Marketing 2016. 12 pages. https://www.insivia.com/50-must-know-stats-about-video-marketing-2016/ [Last accessed May 5, 2021]. [cited by applicant]
[No Author Listed], Mobile Fact Sheet. Pew Research Center. Apr. 2021. 6 pages. http://www.pewinternet.org/fact-sheet/mobile/ [Last accessed May 5, 2021]. [cited by applicant]
[No Author Listed], NSF Seed Fund. Phase I. 2018. 3 pages. bootcamp.https://seedfund.nsf.gov/resources/awardees/phase-1/bootcamp/ [Last accessed May 5, 2021]. [cited by applicant]
[No Author Listed], The State of Traditional TV: Updated with Q2 2017 Data. Dec. 2017. 13. 15 pages. https://www.marketingcharts.com/featured-24817 [Last accessed May 5, 2021]. [cited by applicant]
Abu-El-Haija et al., Youtube-8m: A large-scale video classification benchmark. arXiv:1609.08675v1 [cs.CV]. 2016 Sep. 27. 10 pages. [cited by applicant]
Bilen et al., Dynamic image networks for action recognition. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2016; 3034-42. [cited by applicant]
Buolamwini et al., Gender Shades: Intersectional Accuracy Disparities in Commercial Gender Classification. Proceedings of Machine Learning Research. 2018;81:1-15. [cited by applicant]
Chan et al., Listen, Attend and Spell. arXiv:1508.01211v2 [cs.CL] Aug. 20, 2015. 16 pages. [cited by applicant]
Christensen et al., Know your customers' jobs to be done. Harvard Business review 94.9 (2016) 54:14 pages. [cited by applicant]
Davis et al., Visual Rhythm and Beat. ACM Trans. Graph. 37.4 (2018) 122-1. [cited by applicant]
Deng et al., ImageNet: A Large-Scale Hierarchical Image Database. ReaserchGate. CVPR. IEEE. Jun. 2009. 9 Pages. [cited by applicant]
Donahue et al., Long-term recurrent convolutional networks for visual recognition and description. Proceedings of the IEEE conference on computer vision and pattern recognition. 2015. 2625-2634. [cited by applicant]
Fouhey et al., From Lifestyle Vlogs to Everyday Interactions. arXiv:1712.02310 [cs.CV]. Dec. 6, 2017. 11 pages. [cited by applicant]
Gallello, My favorite usability testing question. Medium. Nov. 7, 2017. 2 pages. https://medium.com/@cgallello/my-favorite-usability-test-question-62fbe3aa9373 [Last accessed May 5, 2021]. [cited by applicant]
Gandhi et al. Multi-clip video editing from a single viewpoint. Proceedings of the 11 [cited by applicant]
Gandhi et al., A computational framework for vertical video editing. 4 [cited by applicant]
Goyal et al., The “something something” video database for learning and evaluating visual common sense. arXiv:1706.04261v2 [cs.CV] Jun. 15, 2017. 21 pages. [cited by applicant]
Gu et al., AVA: A video dataset of spatio-temporally localized atomic visual actions. arXiv:1705.08421v4 [cs.CV] Apr. 30, 2018. 15 pages. [cited by applicant]
Haimson et al., What makes live events engaging on Facebook Live, Periscope, and Snapchat. Proceedings of the 2017 CHI Conference on Human Factors in Computing Systems. ACM. 2017; 48-60. [cited by applicant]
He et al., Mask R-CNN. arXiv:1703.06870v3 [cs.CV]. Jan. 24, 2018. 12 pages. [cited by applicant]
Johansson-Sköldberg et al., Design thinking: past, present, and possible futures. Creativity and innovation management. 22.2 (2013): 121-146. [cited by applicant]
Karayev et al., Recognizing image style. arXiv preprint arXiv:1311.3715v3:23 Jul. 23, 2014. 20 pages. [cited by applicant]
Kot et al., Image and Video Source Class Identification. Digital Image Forensics. Springer, New York, NY, 2013. Jul. 31, 2012: 157-178. [cited by applicant]
Leake et al. Computational video editing for dialogue-driven scenes. ACM Transactions on Graphic (TOG). 2017; 36(130). [cited by applicant]
Liao et al., Audeosynth: music-driven video montage. ACM Transactions on Graphics (TOG) 34.4. (2015): 1-10. [cited by applicant]
Lin et al., Chinese Story Generation Using Conditional Generative Adversarial Network. 2020 International Conference on Artificial Intelligence in Information and Communication (ICAIIC), Fukuoka, Japan. 2020: 457-62. [cited by applicant]
Lin et al., Microsoft COCO: Common Objects in Context. arXiv:1405.0312v3 [cs.CV] Feb. 21, 2015. 15 pages. [cited by applicant]
Lino et al., Computational model of film editing for interactive storytelling. International Conference on Interactive Digital Storytelling. Springer. 3022; 305-308. [cited by applicant]
Lu et al. Story-driven summarization for egocentric video. Computer Vision and Pattern Recognition (CVPR). 2013 IEEE Conference on. IEEE. 2013; 2714-2721. [cited by applicant]
Manyika et al., Independent Work: Choice, Necessity, and the Gig Economy. McKinsey Global Institute. Oct. 2016. 148 pages. [cited by applicant]
Marshall, By 2020, 75% of Mobile Traffic will be Video [Cisco Study]. Feb. 2016. 6 pages. https://tubularinsights.com/2020-mobile-video-traffic [Last accessed May 5, 2021]. [cited by applicant]
Marszalek, et al., Actions in Context. IEEE Conference on Computer Vision & Patterson Recognition. Jun. 2009; 2929-2936. [cited by applicant]
Mccue, Top 10 Video Marketing Trends and Statistics Roundup 2017. Sep. 22, 2017. www.forbes.com/sites/tjmccue/2017/09/22/top-10-vieo-marketing-trends-and-statistics-roundup-2017 [Last Accessed May 5, 2021]. [cited by applicant]
Mccue. Video Marketing In 2018 Continues To Explode As Way To Reach Customers. Jun. 22, 2018. www.forbes.com/sites/tjmccue/2018/06/22/video-marketing-2018-trends-continues-to-explode-as-the-way-to-reach-customers [Last … [cited by applicant]
Merabti et al. A virtual director inspired by real directors. AAA! Workshop on Intelligent Cinematography and Editing. 2014; 27:28-36. [cited by applicant]
Merabti et al. A virtual director using hidden markov models. Computer Graphics Forum. Wiley Online Library. 2016; 35(8):51-67. [cited by applicant]
Mithun et al., Learning Joint Embedding with Multimodal Cues for Cross-Modal Video-Text Retrieval. ICMR. Jun. 11, 2018. 9 pages. [cited by applicant]
Monfort et al., Moments in Time Dataset: One Million Videos For Event Understanding. ArXIV:1801.03150v3 [cs.CV] Feb. 16, 2019. 8 pages. [cited by applicant]
Murch, In the blink of an eye: A perspective on film editing. Silman-James Press, 2001; 20 pages. [cited by applicant]
Naderiparizi et al., Glimpse: A programmable early-discard camera architecture for continuous mobile vision. Proceedings of the 15 [cited by applicant]
Nagarajan et al., Attributes as operators: factorizing unseen attribute-object compositions. ECCV. 2018. 17 pages. [cited by applicant]
Ngamkan et al., Building Models for Mobile Video Understanding. ICLR 2019. 4 pages. [cited by applicant]
Nguyen et al., The open world of micro-videos. arXiv:1603.094392v2 [cs.CV] Apr. 1, 2016. 17 pages. [cited by applicant]
Nieto et al., Systematic Exploration of Computational Music Structure Research. ISMIR. Aug. 1, 2016. 7 pages. [cited by applicant]
Owens, Measure your product's perceived usability with one simple number. Dec. 2016. 5 pages. https://medium.theuxblog.com/measure-your-products-usability-with-one-simple-number-3ecef1cb757e [Last accessed May 5, 2021]. [cited by applicant]
Pan et al., Recurrent residual module for fast inference in videos. Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018; 1536-1545. [cited by applicant]
Patterson et al., Coco attributes: Attributes for people, animals, and objects. European Conference on Computer Vision. Springer. 2016; 85-100. [cited by applicant]
Patterson et al., Sun attribute database: Discovering, annotating, and recognizing scene attributes. Brown university. 2012. 8 pages. [cited by applicant]
Patterson et al., The SUN attribute Database: Beyond Categories for Deeper Scene Understanding. Int J. Comput Vis. Jan. 2014. 18;108: 59-81. [cited by applicant]
Patterson et al., Tropel: Crowdsourcing detectors with minimal training. HCOMP. 2015; 150-159. [cited by applicant]
Pavel et al., Vidcrit: Video-based asynchronous video review. Proceedings of the 29th Annual Symposium on User Interface Software and Technology. ACM. 2016: 517-528. [cited by applicant]
Pennington et al., GloVe: Global Vectors for Word Representation. EMNLP. Oct. 2014. 1532-1543. [cited by applicant]
Ringer et al., Deep unsupervised multi-view detection of video game stream highlights. arXiv:1807.09715v1. Jul. 25, 2018. 6 pages. [cited by applicant]
Ronchi et al., Benchmarking and error diagnosis in multi-instance pose estimation IEEE on Computer Vision and Pattern Recognition. 2018. 4510-4520. [cited by applicant]
Salazar, Diary Studies: Understanding Long-Term User Behavior and Experiences. NN/g Nielsen Norman Group. Jun. 5, 2016. 6 pages. [cited by applicant]
Sandler et al., MobileNetV2: Inverted Residuals and Linear Bottlenecks. IEEE 2018. 4510-4520. [cited by applicant]
Schalkwyk, Google https://ai.googleblog.com/2019/03/an-all-neural-on-device-speech.html [Last accessed May 5, 2021]. [cited by applicant]
Simonyan et al., Two-stream convolutional networks for action recognition in videos. Advances in neural information processing systems. 2014; 568-576. [cited by applicant]
Spangler, Cord-Cutting Explodes: 22 Million U.S. Adults Will Have Canceled Cable, Satellite TV by End of 2017. Variety. Sep. 2017. https://variety.com/2017/biz/news/cord-cutting-2017-estimtes-cancel-cable-satellite-tv-1… [cited by applicant]
Suris et al., Cross-Modal Embeddings For Video and Audio Retrieval. arXiv:1801.02200v1 [cs.IR]. Jan. 7, 2018. 6 pages. [cited by applicant]
Swanson, Instagram and Snapchat are Most Popular Social Networks for Teens; Black Teens are Most Active on Social Media, Messaging Apps. Associated Press-NORC. Apr. 9, 2015. 18 pages. [cited by applicant]
Tankovska. Instagram-Statistics & Facts. Statista. Jun. 4, 2021. 8 pages. [cited by applicant]
Tran et al., Learning spatiotemporal features with 3d convolutional networks. Computer Vision (ICCV). 2015 IEEE International Conference on. IEEE. 2015: 4489-4497. [cited by applicant]
Truong et al., Quickcut: An interactive tool for editing narrated video. UIST. 2016: 497-507. [cited by applicant]
Turner, How many smartphones are in the world? Bank My Cell. May 2021. 16 pages. https://www.bankmycell.com/blog/how-many-phones-are-in-the-world [Last accessed May 5, 2021]. [cited by applicant]
Zazelenchuk, Data collection for usability research. May 5, 2008. 5 pages. https://www.userfocus.co.uk/articles.dataloggingtools.html [Last accessed May 5, 2021]. [cited by applicant]
Zhang, The Importance of Cameras in the Smartphone War. Samsung. PetaPixel. Feb. 12, 2015. 11 pages. [cited by applicant]
Zhou et al., Places: A 10 million image database for scene recognition. IEEE transactions on pattern analysis and machine intelligence. 40.6 (2017): 1452-1464. [cited by applicant]
“U.S. Appl. No. 17/189,865, Examiner Interview Summary mailed Sep. 19, 2022”, 2 pgs. [cited by applicant]
“U.S. Appl. No. 17/189,865, Final Office Action mailed Dec. 23, 2022”, 14 pgs. [cited by applicant]
“U.S. Appl. No. 17/189,865, Non Final Office Action mailed Mar. 31, 2022”, 12 pgs. [cited by applicant]
“U.S. Appl. No. 17/189,865, Notice of Allowance mailed May 10, 2023”, 11 pgs. [cited by applicant]
“U.S. Appl. No. 17/189,865, Response filed Apr. 21, 2023 to Final Office Action mailed Dec. 23, 2022”, 9 pgs. [cited by applicant]
“U.S. Appl. No. 17/189,865, Response filed Sep. 19, 2022 to Non Final Office Action mailed Mar. 31, 2022”, 8 pgs. [cited by applicant]
“European Application Serial No. 21763998.8, Response to Communication pursuant to Rules 161(2) and 162 EPC filed Apr. 5, 2023”, 8 pgs. [cited by applicant]
“European Application Serial No. 21763998.8, Extended European Search Report mailed Feb. 23, 2024”, 6 pgs. [cited by applicant]
“European Application Serial No. 21763998.8, Response filed Sep. 10, 2024 to Extended European Search Report mailed Feb. 23, 2024”, 10 pgs. [cited by applicant]
“Canadian Application Serial No. 3,173,977, Office Action mailed Mar. 3, 2026”, 4 pgs. [cited by applicant]
“Canadian Application Serial No. 3,173,977, Response filed Jun. 26, 2026 to Office Action mailed Mar. 3, 2026”, 16 pgs. [cited by applicant]