IP Library › Granted Patent US 12,468,760
Granted Patent B2
US 12,468,760 · App. 17/452,626 · Granted Nov 11, 2025

Customizable framework to extract moments of interest

Inventors: Ali Aminian (San Jose, CA); William Lawrence Marino (Hockessin, DE); Kshitiz Garg (Santa Clara, CA); Aseem Agarwala (Seattle, WA)
Assignee: Adobe Inc.
G06F16/739G06F3/0482G06V20/47G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,468,760
App. No.
17/452,626
Granted
Nov 11, 2025
Kind
B2
Abstract

Embodiments of the present invention provide systems, methods, and computer storage media for extracting moments of interest (e.g., video frames, video segments) from a video. In an example embodiment, independent and/or orthogonal machine learning models are used to extract different types of features considering different modalities, and each frame in the video is assigned an importance score for each model. The importance scores for each model are combined into an aggregated importance score for each frame in the video. Depending on the embodiment, the aggregated importance scores are used to visualize the score per frame, identify moments of interest, automatically crop down the video into a highlight reel, browse or visualize the moments of interest within the video, and/or search across multiple videos.

Claims (42)

1 . A computerized method comprising:

receiving, via one or more inputs into a user interface, a representation of selected modalities of a video;

triggering extraction of one or more moments of interest in the video using an identified machine learning model corresponding to each of the selected modalities, the identified machine learning model being identified from among a plurality of independent machine learning models designated to extract different features from videos;

receiving a representation of the one or more moments of interest in the video; and

causing the user interface to execute an operation associated with the one or more moments of interest.

2 . The computerized method of claim 1 , wherein the representation of the one or more moments of interest in the video corresponds to a summary video that includes only the one or more moments of interest, wherein the operation associated with the one or more moments of interest comprises a download or an upload of the summary video.

3 . The computerized method of claim 1 , wherein the operation associated with the one or more moments of interest comprises updating a video timeline, that provides a browsing functionality, with a visual representation of the one or more moments of interest.

4 . The computerized method of claim 1 , wherein the operation associated with the one or more moments of interest comprises a search for videos that have one or more identified moments of interest.

5 . The computerized method of claim 1 , further comprising receiving, via the one or more inputs into the user interface, a representation of selected classes, wherein the extraction of the one or more moments of interest in the video comprises setting corresponding class weights that prioritize the selected classes over other supported classes.

6 . The computerized method of claim 1 , further comprising receiving, via the one or more inputs into the user interface, a representation of a freeform text query, wherein the extraction of the one or more moments of interest comprises:

encoding the freeform text query into a textual embedding;

encoding each frame of the video into a visual embedding; and

generating a set of importance scores based on cosine similarity between the textual embedding and the visual embedding for each frame.

7 . The computerized method of claim 1 , wherein the extraction of the one or more moments of interest comprises identifying video segments with frames that have corresponding aggregated importance scores above a threshold, the aggregated importance score determined from output scores corresponding to different modalities of the video generated using multiple machine learning models according to the selected modalities.

8 . The computerized method of claim 1 , wherein the extraction of the one or more moments of interest comprises using dynamic programming to accumulate video segments of the video up to a designated duration.

9 . The computerized method of claim 1 , wherein the extraction of the one or more moments of interest comprises generating a signal of importance scores for each of the selected modalities and smoothing the signal by convolving the signal with a Gaussian kernel.

10 . One or more computer storage media storing computer-useable instructions that, when used by one or more computing devices, cause the one or more computing devices to perform operations comprising:

receiving, via one or more inputs into a user interface, a representation of selected modalities of a video;

identifying a machine learning model for each of the selected modalities from a plurality of independent machine learning models designated to extract different features from videos;

triggering extraction of one or more moments of interest in the video using the identified machine learning model corresponding to each of the selected modalities;

receiving a representation of the one or more moments of interest in the video; and

causing the user interface to execute an operation associated with the one or more moments of interest.

11 . The media of claim 10 , wherein the representation of the one or more moments of interest in the video corresponds to a summary video that includes only the one or more moments of interest, wherein the operation associated with the one or more moments of interest comprises a download or an upload of the summary video.

12 . The media of claim 10 , wherein the operation associated with the one or more moments of interest comprises updating a video timeline, that provides a browsing functionality, with a visual representation of the one or more moments of interest.

13 . The media of claim 10 , wherein the operation associated with the one or more moments of interest comprises a search for videos that have one or more identified moments of interest.

14 . The media of claim 10 , further comprising receiving, via the one or more inputs into the user interface, a representation of selected classes, wherein the extraction of the one or more moments of interest in the video comprises setting corresponding class weights that prioritize the selected classes over other supported classes.

15 . The media of claim 10 , further comprising receiving, via the one or more inputs into the user interface, a representation of a freeform text query, wherein the extraction of the one or more moments of interest comprises:

encoding the freeform text query into a textual embedding;

encoding each frame of the video into a visual embedding; and

generating a set of importance scores based on cosine similarity between the textual embedding and the visual embedding for each frame.

16 . A computer system comprising one or more hardware processors configured to cause the system to perform operations comprising:

receiving, via one or more inputs into a user interface, a representation of selected modalities of a video, wherein the modalities are selected from a plurality of modalities that each correspond to a different independent machine learning model;

triggering extraction of one or more moments of interest in the video using an identified machine learning model corresponding to each of the selected modalities;

receiving a representation of the one or more moments of interest in the video; and

causing the user interface to execute an operation associated with the one or more moments of interest.

17 . The system of claim 16 , wherein the representation of the one or more moments of interest in the video corresponds to a summary video that includes only the one or more moments of interest, wherein the operation associated with the one or more moments of interest comprises a download or an upload of the summary video.

18 . The system of claim 16 , wherein the operation associated with the one or more moments of interest comprises updating a video timeline, that provides a browsing functionality, with a visual representation of the one or more moments of interest.

19 . The system of claim 16 , wherein the operation associated with the one or more moments of interest comprises a search for videos that have one or more identified moments of interest.

20 . The system of claim 16 , further comprising receiving, via the one or more inputs into the user interface, a representation of a freeform text query, wherein the extraction of the one or more moments of interest comprises:

encoding the freeform text query into a textual embedding;

encoding each frame of the video into a visual embedding; and

generating a set of importance scores based on cosine similarity between the textual embedding and the visual embedding for each frame.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2021
From: AMINIAN, ALI; MARINO, WILLIAM LAWRENCE; GARG, KSHITIZ; AGARWALA, ASEEM
To: ADOBE INC.
Reel/Frame 057947/0194 →
Continuity (1)
Related Publication 20230140369A1 · May 4, 2023
References Cited (150)
US 6104398A · Cox, Jr. et al. · 2000 [cited by applicant]
US 6400378B1 · Snook · 2002 [cited by applicant]
US 7480442B2 · Girgensohn et al. · 2009 [cited by applicant]
US 7796857B2 · Hiroi et al. · 2010 [cited by applicant]
US 7805678B1 · Niles et al. · 2010 [cited by applicant]
US 8290345B2 · Numoto · 2012 [cited by applicant]
US 8306402B2 · Ishihara · 2012 [cited by applicant]
US 8620893B2 · Howard et al. · 2013 [cited by applicant]
US 8789083B1 · Landers et al. · 2014 [cited by applicant]
US 8874584B1 · Chen et al. · 2014 [cited by applicant]
US 9110562B1 · Eldawy · 2015 [cited by applicant]
US 9583140B1 · Rady · 2017 [cited by applicant]
US 10750245B1 · Zeiler et al. · 2020 [cited by applicant]
US 11080007B1 · Melzer · 2021 [cited by applicant]
US 11120490B1 · Pham et al. · 2021 [cited by applicant]
US 11568900B1 · Achddou et al. · 2023 [cited by applicant]
US 11683290B1 · Stern et al. · 2023 [cited by applicant]
US 11880918B2 · Loui et al. · 2024 [cited by applicant]
US 20020061136A1 · Shibata et al. · 2002 [cited by applicant]
US 20020186234A1 · Van De Streek et al. · 2002 [cited by applicant]
US 20030093790A1 · Logan et al. · 2003 [cited by applicant]
US 20030177503A1 · Sull et al. · 2003 [cited by applicant]
US 20030234805A1 · Toyama et al. · 2003 [cited by applicant]
US 20040125124A1 · Kim et al. · 2004 [cited by applicant]
US 20040135815A1 · Browne et al. · 2004 [cited by applicant]
US 20050203927A1 · Sull et al. · 2005 [cited by applicant]
US 20050220348A1 · Chiu · 2005 [cited by examiner]
US 20060224983A1 · Albrecht et al. · 2006 [cited by applicant]
US 20070022159A1 · Zhu et al. · 2007 [cited by applicant]
US 20070025614A1 · Qian · 2007 [cited by applicant]
US 20070044010A1 · Sull et al. · 2007 [cited by applicant]
US 20070106693A1 · Houh et al. · 2007 [cited by applicant]
US 20070201558A1 · Xu et al. · 2007 [cited by applicant]
US 20080215552A1 · Safoutin · 2008 [cited by applicant]
US 20090080853A1 · Chen et al. · 2009 [cited by applicant]
US 20090172745A1 · Horozov et al. · 2009 [cited by applicant]
US 20090231271A1 · Heubel et al. · 2009 [cited by applicant]
US 20100070483A1 · Delgo et al. · 2010 [cited by applicant]
US 20100088726A1 · Curtis et al. · 2010 [cited by applicant]
US 20100107117A1 · Pearce et al. · 2010 [cited by applicant]
US 20100111417A1 · Ward et al. · 2010 [cited by applicant]
US 20100281372A1 · Lyons et al. · 2010 [cited by applicant]
US 20110307084A1 · Gehring et al. · 2011 [cited by applicant]
US 20120197763A1 · Moreira · 2012 [cited by applicant]
US 20130027412A1 · Roddy · 2013 [cited by applicant]
US 20130097507A1 · Prewett · 2013 [cited by applicant]
US 20130236162A1 · Kim et al. · 2013 [cited by applicant]
US 20130294642A1 · Wang et al. · 2013 [cited by applicant]
US 20140105571A1 · Chang et al. · 2014 [cited by applicant]
US 20140173484A1 · Hicks · 2014 [cited by applicant]
US 20140270708A1 · Girgensohn et al. · 2014 [cited by applicant]
US 20140340204A1 · O'Shea et al. · 2014 [cited by applicant]
US 20140358807A1 · Chinnappan et al. · 2014 [cited by applicant]
US 20150005646A1 · Balakrishnan et al. · 2015 [cited by applicant]
US 20150052465A1 · Altin et al. · 2015 [cited by applicant]
US 20150261389A1 · Abate · 2015 [cited by applicant]
US 20150356355A1 · Oguchi et al. · 2015 [cited by applicant]
US 20150370806A1 · White et al. · 2015 [cited by applicant]
US 20160071524A1 · Tammi et al. · 2016 [cited by applicant]
US 20170133054A1 · Song et al. · 2017 [cited by applicant]
US 20170169128A1 · Batchu Krishnaiahsetty et al. · 2017 [cited by applicant]
US 20170228138A1 · Paluka et al. · 2017 [cited by applicant]
US 20180113579A1 · Johnston et al. · 2018 [cited by applicant]
US 20200066305A1 · Spence et al. · 2020 [cited by applicant]
US 20200090661A1 · Ackerman et al. · 2020 [cited by applicant]
US 20200106965A1 · Malia et al. · 2020 [cited by applicant]
US 20200138321A1 · Niebauer · 2020 [cited by applicant]
US 20200334290A1 · Dontcheva et al. · 2020 [cited by applicant]
US 20210081676A1 · Kim et al. · 2021 [cited by applicant]
US 20210142827A1 · Allen et al. · 2021 [cited by applicant]
US 20210406552A1 · Hong · 2021 [cited by examiner]
US 20220067386A1 · Rotman et al. · 2022 [cited by applicant]
US 20220182577A1 · Kumar · 2022 [cited by examiner]
US 20220198194A1 · Zhang · 2022 [cited by examiner]
US 20220207282A1 · Dong · 2022 [cited by examiner]
CN 111601160A · 2020 [cited by applicant]
First action interview—office action dated Feb. 3, 2022 in U.S. Appl. No. 17/017,362, 3 pages. [cited by applicant]
Final Office Action dated Feb. 28, 2022 in U.S. Appl. No. 17/017,370, 32 pages. [cited by applicant]
Notice of Allowance dated Mar. 23, 2022 in U.S. Appl. No. 17/017,344, 7 pages. [cited by applicant]
Notice of Allowance dated Apr. 6, 2022 in U.S. Appl. No. 17/330,667, 8 pages. [cited by applicant]
Corrected Notice of Allowability dated Apr. 25, 2022 in U.S. Appl. No. 17/017,344, 2 pages. [cited by applicant]
Restriction Requirement dated Apr. 26, 2022 in U.S. Appl. No. 17/330,689, 6 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/017,353, mailed on Jan. 4, 2024, 22 pages. [cited by applicant]
First Preinterview Communication dated Aug. 18, 2022 in U.S. Appl. No. 17/017,353, 31 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/805,075, mailed on Feb. 14, 2024, 5 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/805,076, mailed on Jan. 22, 2024, 3 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/969,536, mailed on Jan. 17, 2024, 2 pages. [cited by applicant]
Shibata, M., “Temporal segmentation method for video sequence”, Proceedings of SPIE, Applications in Optical Science and Engineering, Visual Communications and Image Processing, vol. 1818, pp. 1-13 (1992). [cited by applicant]
Swanberg, D., et al., “Knowledge-guided parsing in video databases”, Proceedings of SPIE, IS&T/SPIE's Symposium on Electronic Imaging: Science and Technology, Storage and retrieval for Image and Video Databases, vol. 19… [cited by applicant]
Mcdarris, J., “Adobe Photoshop Tutorial: EVERY Tool in the Toolbar Explained and Demonstrated”, You Tube, Retrieved from Internet URL: https://www.youtube.com/watch?v=2cQT1ZgvgGI, accessed on Jun. 21, 2023, p. 3. [cited by applicant]
Final Office Action dated Jun. 14, 2023 in U.S. Appl. No. 17/017,353, 34 pages. [cited by applicant]
Final Office Action dated Jun. 20, 2023 in U.S. Appl. No. 17/330,689, 13 pages. [cited by applicant]
Non Final Office Action dated Jun. 14, 2023 in U.S. Appl. No. 17/805,080, 9 pages. [cited by applicant]
Non Final Office Action dated May 19, 2023 in U.S. Appl. No. 17/805,076, 9 pages. [cited by applicant]
Non Final Office Action dated Jun. 30, 2023 in U.S. Appl. No. 17/330,702, 7 pages. [cited by applicant]
Non Final Office Action dated Jun. 12, 2023 in U.S. Appl. No. 17/330,718, 7 pages. [cited by applicant]
Non Final Office Action dated May 18, 2023 in U.S. Appl. No. 17/805,075, 9 pages. [cited by applicant]
Final Office Action dated Jul. 19, 2023 in U.S. Appl. No. 17/969,536, 18 pages. [cited by applicant]
Notice of Allowance dated Aug. 16, 2023 in U.S. Appl. No. 17/805,076, 5 pages. [cited by applicant]
“Apple, Final Cut Pro 7 User Guide”, 2020, Retrieved from Internet URL: https://prohelp.apple.com/finalcutpro_helpr01/English/en/finalcutpro/usermanual/index.html#chapter=7%26section=1), pp. 11 (Year: 2010). [cited by applicant]
Notice of Allowance dated Feb. 8, 2023 in U.S. Appl. No. 17/017,366, 5 pages. [cited by applicant]
Non-Final Office Action dated Jan. 24, 2023 in U.S. Appl. No. 17/330,689, 13 pages. [cited by applicant]
Notice of Allowance dated Feb. 2, 2023 in U.S. Appl. No. 17/017,370, 6 pages. [cited by applicant]
Non-Final Office Action dated Feb. 16, 2023 in U.S. Appl. No. 17/969,536, 16 pages. [cited by applicant]
Final Office Action dated Mar. 9, 2023 in U.S. Appl. No. 17/330,702, 8 pages. [cited by applicant]
Final Office Action dated Mar. 9, 2023 in U.S. Appl. No. 17/330,718, 8 pages. [cited by applicant]
Notice of Allowance dated Apr. 18, 2023 in U.S. Appl. No. 17/330,677, 5 pages. [cited by applicant]
Notice of Allowance dated Jul. 13, 2022 in U.S. Appl. No. 17/017,366, 7 pages. [cited by applicant]
Non-Final Office Action dated Jul. 29, 2022 in U.S. Appl. No. 17/017,370, 37 pages. [cited by applicant]
Final Office Action dated Aug. 15, 2022 in U.S. Appl. No. 17/017,362, 7 pages. [cited by applicant]
Preinterview first office action dated Aug. 18, 2022 in U.S. Appl. No. 17/017,353, 5 pages. [cited by applicant]
Non-Final Office Action dated Sep. 29, 2022 in U.S. Appl. No. 17/330,677, 12 pages. [cited by applicant]
Final Office Action dated May 9, 2022 in U.S. Appl. No. 17/017,366, 14 pages. [cited by applicant]
Non-Final Office Action dated Jun. 2, 2022 in U.S. Appl. No. 17/330,689, 9 pages. [cited by applicant]
Non-Final Office Action dated Oct. 24, 2022 in U.S. Appl. No. 17/330,702, 9 pages. [cited by applicant]
Non-Final Office Action dated Oct. 24, 2022 in U.S. Appl. No. 17/330,718, 10 pages. [cited by applicant]
Non-Final Office Action dated Oct. 26, 2022 in U.S. Appl. No. 17/330,702, 8 pages. [cited by applicant]
Final Office Action dated Oct. 27, 2022 in U.S. Appl. No. 17/330,689, 8 pages. [cited by applicant]
Notice of Allowance dated Nov. 21, 2022 in U.S. Appl. No. 17/017,362, 5 pages. [cited by applicant]
First Action Interview Office Action dated Nov. 22, 2022 in U.S. Appl. No. 17/017,353, 4 pages. [cited by applicant]
Final Office Action dated Jan. 4, 2023 in U.S. Appl. No. 17/330,677, 12 pages. [cited by applicant]
Non-Final Office Action dated Sep. 1, 2023 in U.S. Appl. No. 17/017,353, 32 pages. [cited by applicant]
Non-Final Office Action dated Sep. 21, 2023 in U.S. Appl. No. 17/805,075, 5 pages. [cited by applicant]
Notice of Allowance dated Oct. 4, 2023 in U.S. Appl. No. 17/330,718, 5 pages. [cited by applicant]
Notice of Allowance dated Oct. 18, 2023 in U.S. Appl. No. 17/969,536, 7 pages. [cited by applicant]
Notice of Allowance dated Nov. 1, 2023 in U.S. Appl. No. 17/805,080, 5 pages. [cited by applicant]
Final Office Action dated Nov. 15, 2023 in U.S. Appl. No. 17/330,702, 7 pages. [cited by applicant]
Notice of Allowance dated Sep. 11, 2023 in U.S. Appl. No. 17/330,689, 5 pages. [cited by applicant]
Girgensohn, Andreas, et al. “Locating information in video by browsing and searching.” Interactive Video: Algorithms and Technologies. Berlin, Heidelberg: Springer Berlin Heidelberg, 2006. 207-224. (Year: 2006). [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/805,907, mailed on Mar. 13, 2024, 8 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/017,353, mailed on Mar. 6, 2024, 10 pages. [cited by applicant]
Notice of Allowance received for U.S. Appl. No. 17/330,702, mailed on Mar. 4, 2024, 5 pages. [cited by applicant]
Shipman, Frank, Andreas Girgensohn, and Lynn Wilcox. “Generation of interactive multi-level video summaries.” Proceedings of the eleventh ACM international conference on Multimedia. 2003. (Year: 2003). [cited by applicant]
Alcázar, J.L., et al., “Active Speakers in Context,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), pp. 12465-12474 (2020). [cited by applicant]
Chinfat, “E15—The Tool Bar/The Selection Tool—Adobe Premiere Pro CC 2018,” accessed at https://www.youtube.com/watch?v=IJRhzOqrMzA, p. 1 (May 8, 2018). [cited by applicant]
“Hierarchical Segmentation of video timeline and matching searh data,” accessed https://www.google.com/search?q-hierarchical+segmentation+of+video+timeline+and+matching+search+data&hl=en&blw=1142&bih=547&tbm=pts&s . . .… [cited by applicant]
“Transient oscillation video matching,” accessed at https://www.google.com/search?q=transient+oscillation+video +matching&hl=en&biw=1225&bih=675&tbm=pts&sxsrf=AOaemvJjtgRC6aa4ltvNUFxJa6 . . . , accessed on Dec. 2, 2021,… [cited by applicant]
Co-pending U.S. Appl. No. 16/879,362, filed May 20, 2020. [cited by applicant]
Preinterview First Office Action dated Sep. 27, 2021 in U.S. Appl. No. 17/017,366, 5 pages. [cited by applicant]
Preinterview first Office Action dated Sep. 30, 2021 in U.S. Appl. No. 17/017,362, 4 pages. [cited by applicant]
Restriction Requirement dated Oct. 18, 2021 in U.S. Appl. No. 17/017,344, 5 pages. [cited by applicant]
Non-Final Office Action dated Nov. 9, 2021 in U.S. Appl. No. 17/017,344, 9 pages. [cited by applicant]
First Action Interview—Office Action dated Dec. 3, 2021 in U.S. Appl. No. 17/017,366, 4 pages. [cited by applicant]
Non-Final Office Action dated Dec. 7, 2021 in U.S. Appl. No. 17/017,370, 26 pages. [cited by applicant]
Office Action received for GB Patent Application No. 2206709.4, mailed on Jul. 9, 2024, 3 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/805,907, mailed on Aug. 1, 2024, 8 pages. [cited by applicant]
Final Office Action received for U.S. Appl. No. 17/805,907, mailed on Jan. 10, 2025, 9 pages. [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/805,907, mailed on Oct. 28, 2024, 8 pages. [cited by applicant]
Non-Final Office Action received for U.S. Appl. No. 17/805,907, mailed on May 19, 2025, 9 pages. [cited by applicant]
Office Action received for Chinese Patent Application No. 202210515554.X, mailed on Aug. 27, 2025, 29 pages (15 pages of English Translation and 14 pages of Original Office Action). [cited by applicant]