IP Library › Granted Patent US 12,563,278
Granted Patent B2
US 12,563,278 · App. 18/747,757 · Granted Feb 24, 2026

Systems and methods for video summarization

Inventors: Dawit Mureja Argaw (Daejeon, KR); Seunghyun Yoon (San Jose, CA); Fabian David Caba Heilbron (San Jose, CA); Hanieh Deilamsalehy (Seattle, WA); Trung Huu Bui (San Jose, CA); Zhaowen Wang (San Jose, CA); Franck Dernoncourt (Seattle, WA)
Assignee: ADOBE INC.
H04N21/8549
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,563,278
App. No.
18/747,757
Granted
Feb 24, 2026
Kind
B2
Abstract

A method, apparatus, non-transitory computer readable medium, and system for video summarization include obtaining a video, generating a sequence of contextualized visual representations corresponding to portions of the video, generating a sequence of summary visual representations corresponding to a subset of the portions of the video based on the sequence of contextualized visual representations, and generating a summary video including the subset of the portions of the video based on the sequence of summary visual representations.

Claims (55)

1 . A method for video processing, comprising:

obtaining a video;

generating, using a video encoder of a video summarization model, a sequence of contextualized visual representations corresponding to portions of the video, wherein the sequence of contextualized visual representations comprises a sequence of embeddings of the corresponding portions of the video in a vector space;

generating, using a summary decoder of the video summarization model, a sequence of summary visual representations corresponding to a subset of the portions of the video based on the sequence of contextualized visual representations using an autoregressive prediction scheme; and

generating a summary video including the subset of the portions of the video based on the sequence of summary visual representations.

2 . The method of claim 1 , further comprising:

generating a sequence of preliminary visual representations based on the video, wherein the sequence of contextualized visual representations is generated based on the sequence of preliminary visual representations.

3 . The method of claim 2 , wherein generating the summary video comprises:

matching each of the sequence of summary visual representations to a corresponding portion of the subset of the portions of the video.

4 . The method of claim 1 , further comprising:

obtaining a transcript of the video;

generating a sequence of text features based on the transcript; and

generating a sequence of text-conditioned visual features based on the sequence of text features and the sequence of contextualized visual representations, wherein the sequence of summary visual representations is generated based on the sequence of text-conditioned visual features.

5 . The method of claim 1 , wherein:

the video summarization model is trained using training data including a training video and a text summary of the training video.

6 . A method for training a machine learning model, comprising:

obtaining a training set comprising a training video and a target summary video;

generating, using a video encoder of a video summarization model, a sequence of contextualized visual representations corresponding to portions of the training video, wherein the sequence of contextualized visual representations comprises a sequence of embeddings of the corresponding portions of the video in a vector space;

generating a sequence of target visual representations of the target summary video;

generating, using a summary decoder of the video summarization model, a sequence of summary visual representations corresponding to a subset of the portions of the training video based on the sequence of contextualized visual representations using an autoregressive prediction scheme; and

updating parameters of the video summarization model based on the sequence of summary visual representations and the sequence of target visual representations.

7 . The method of claim 6 , wherein obtaining the training set comprises:

obtaining a text summary of the training video; and

generating the target summary video based on the text summary.

8 . The method of claim 7 , wherein obtaining the training set comprises:

obtaining a transcript of the training video; and

generating, using a language model, the text summary based on the transcript.

9 . The method of claim 7 , wherein generating the target summary video comprises:

matching portions of the text summary to portions of the training video; and

including the matched portions of the training video in the target summary video.

10 . The method of claim 9 , further comprising:

generating a sequence of preliminary visual representations based on the training video, wherein the portions of the text summary are matched to the portions of the training video based on the sequence of preliminary visual representations.

11 . The method of claim 6 , further comprising:

computing a feature reconstruction loss based on a comparison of the sequence of summary visual representations and the sequence of target visual representations, wherein the parameters of the summary decoder are updated based on the feature reconstruction loss.

12 . The method of claim 6 , further comprising:

generating a sequence of preliminary visual representations based on the training video, wherein the sequence of contextualized visual representations is generated based on the sequence of preliminary visual representations.

13 . The method of claim 6 , further comprising:

obtaining a transcript of the training video; and

generating a sequence of text-conditioned visual features based on the transcript and the sequence of contextualized visual representations, wherein the sequence of summary visual representations is generated based on the sequence of text-conditioned visual features.

14 . An apparatus for video processing, comprising:

at least one processor;

at least one memory storing instructions executable by the at least one processor; and

a video summarization model comprising parameters stored in the at least one memory, wherein the video summarization model is trained to generate a sequence of contextualized visual representations corresponding to portions of a video and to generate a sequence of summary visual representations corresponding to a subset of the portions of the video based on the sequence of contextualized visual representations using an autoregressive prediction scheme, wherein the sequence of contextualized visual representations comprises a sequence of embeddings of the corresponding portions of the video in a vector space.

15 . The apparatus of claim 14 , further comprising:

a video summary generation component configured to generate a summary video including the subset of the portions of the video based on the sequence of summary visual representations.

16 . The apparatus of claim 14 , wherein:

the video summarization model comprises a video encoder trained to generate the sequence of contextualized visual representations.

17 . The apparatus of claim 14 , wherein:

the video summarization model comprises a summary decoder trained to generate a sequence of summary visual representations corresponding to a subset of the portions of the video based on the sequence of contextualized visual representations.

18 . The apparatus of claim 14 , wherein:

the video summarization model comprises a text encoder trained to generate a sequence of text features based on a transcript.

19 . The apparatus of claim 14 , wherein:

the video summarization model comprises a cross-modal attention component trained to generate a sequence of text-conditioned visual features based on a sequence of text features and a sequence of contextualized visual representations.

20 . The apparatus of claim 14 , further comprising:

a multi-modal encoder trained to generate a sequence of preliminary visual representations based on the video.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 20, 2024
From: ARGAW, DAWIT MUREJA; YOON, SEUNGHYUN; CABA HEILBRON, FABIAN DAVID; DEILAMSALEHY, HANIEH; BUI, TRUNG HUU; WANG, ZHAOWEN; DERNONCOURT, FRANCK
To: ADOBE INC.
Reel/Frame 067779/0331 →
Continuity (1)
Related Publication 20250392797A1 · Dec 25, 2025
References Cited (77)
US 9286938B1 · Tseytlin · 2016 [cited by examiner]
US 10424341B2 · Thornton · 2019 [cited by examiner]
US 10541000B1 · Karakotsios · 2020 [cited by examiner]
US 10825227B2 · Amer · 2020 [cited by examiner]
US 10945040B1 · Bedi · 2021 [cited by examiner]
US 11910073B1 · Sharma · 2024 [cited by examiner]
US 12283291B1 · Sarfati · 2025 [cited by examiner]
US 12301960B1 · Ben-Cohen · 2025 [cited by examiner]
US 20030152363A1 · Jeannin · 2003 [cited by examiner]
US 20150350747A1 · Jackson · 2015 [cited by examiner]
US 20160014482A1 · Chen · 2016 [cited by examiner]
US 20160070963A1 · Chakraborty · 2016 [cited by examiner]
US 20180189570A1 · Paluri · 2018 [cited by examiner]
US 20180268253A1 · Hoffman · 2018 [cited by examiner]
US 20180295428A1 · Bi · 2018 [cited by examiner]
US 20190180109A1 · Sinha · 2019 [cited by examiner]
US 20190325084A1 · Peng · 2019 [cited by examiner]
US 20190377955A1 · Swaminathan · 2019 [cited by examiner]
US 20200186852A1 · Ramamurthy · 2020 [cited by examiner]
US 20200196028A1 · Kuehne, Jr. · 2020 [cited by examiner]
US 20210392414A1 · Javan Roshtkhari · 2021 [cited by examiner]
US 20220067385A1 · Kaushik · 2022 [cited by examiner]
US 20220078530A1 · Zhu · 2022 [cited by examiner]
US 20220130427A1 · Allibhai · 2022 [cited by examiner]
US 20220353101A1 · Hu · 2022 [cited by examiner]
US 20220414338A1 · Cho · 2022 [cited by examiner]
US 20230401389A1 · Shmuel · 2023 [cited by examiner]
US 20240205520A1 · Carbajo · 2024 [cited by examiner]
US 20240395042A1 · Boiarov · 2024 [cited by examiner]
US 20240406521A1 · Ramesh · 2024 [cited by examiner]
US 20250054306A1 · Cohen · 2025 [cited by examiner]
US 20250124689A1 · Williams · 2025 [cited by examiner]
1Alhojely, et al., “Recent Progress on Text Summarization”, In 2020 International Conference on Computational Science and Computational Intelligence (CSCI), pp. 1503-1509, Dec. 2020, 7 pages. [cited by applicant]
2Bain, et al., “WhisperX: Time-Accurate Speech Transcription of Long-Form Audio”, arXiv preprint arXiv:2303.00747v2 [cs.SD] Jul. 11, 2023, 6 pages. [cited by applicant]
3Bubeck, et al., “Sparks of Artificial General Intelligence: Early Experiments with GPT-4”, arXiv preprint arXiv:2303.12712v5 [cs.CL] Apr. 13, 2023, 155 pages. [cited by applicant]
4Devlin, et al., “Bert: Pre-training of Deep Bidirectional Transformers for Language Understanding”, arXiv preprint arXiv:1810.04805v2 [cs.CL] May 24, 2019, 16 pages. [cited by applicant]
5Elhamifar, et al., “See All by Looking at a Few: Sparse Modeling for Finding Representative Objects”, In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 1600-1607, Jun. 2012, available at https://i… [cited by applicant]
6Gholamrezazadeh, et al., “A Comprehensive Survey on Text Summarization Systems”, In 2009 2nd International Conference on Computer Science and its Applications, pp. 1-6, Jun. 2009, 6 pages. [cited by applicant]
7Gygli, et al., “Creating Summaries from User Videos”, In ECCV 2014, Part VII, LNCS 8695, pp. 505-520, Sep. 2014, 16 pages. [cited by applicant]
8He, et al., “Align and Attend: Multimodal Summarization with Dual Contrastive Losses”, arXiv preprint arXiv:2303.07284v3 [cs.CV] Jun. 12, 2023, 15 pages. [cited by applicant]
9Ji, et al., “Video Summarization with Attention-Based Encoder-Decoder Networks”, arXiv preprint arXiv:1708.09545v2 [cs.CV] Apr. 16, 2018, 9 pages. [cited by applicant]
10Jiang , et al., “Joint Video Summarization and Moment Localization by Cross-Task Sample Transfer”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16388-16398, Jun. 2022, 11 p… [cited by applicant]
11Jung, et al., “Discriminative Feature Learning for Unsupervised Video Summarization”, arXiv preprint arXiv:1811.09791v1 [cs.CV] Nov. 24, 2018, 8 pages. [cited by applicant]
12Kanehira, et al., “Viewpoint-aware Video Summarization”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7435-7444, Jun. 2018, 10 pages. [cited by applicant]
13Kendall, “The treatment of ties in ranking problems”, Biometrika, 33(3): 239-251, 1945, available at https://www.jstor.org/stable/2332303. [cited by applicant]
14Lee, et al., “Discovering Important People and Objects for Egocentric Video Summarization”, Preprint, To Appear In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012, 8 pages. [cited by applicant]
15Lewis, et al., “Bart: Denoising Sequence-to-Sequence Pre-training for Natural Language Generation, Translation, and Comprehension”, arXiv preprint arXiv:1910.13461v1 [cs.CL] Oct. 29, 2019, 10 pages. [cited by applicant]
16Lin, et al., “VideoXum: Cross-modal Visual and Textural Summarization of Videos”, arXiv preprint arXiv:2303.12060v3 [cs.CV] Apr. 23, 2024, 13 pages. [cited by applicant]
17Liu et al., “Text Summarization with Pretrained Encoders”, arXiv preprint arXiv:1908.08345v2 [cs.CL] Sep. 5, 2019, 11 pages. [cited by applicant]
18Liu, et al., “RoBERTa: A Robustly Optimized BERT Pretraining Approach”, arXiv preprint arXiv:1907.11692v1 [cs.CL] Jul. 26, 2019, 13 pages. [cited by applicant]
19Loshchilov, et al., “Decoupled Weight Decay Regularization”, arXiv preprint arXiv:1711.05101v3 [cs.LG] Jan. 4, 2019, 19 pages. [cited by applicant]
20Lu, et al., “A Bag-of-Importance Model With Locality-Constrained Coding Based Feature Learning for Video Summarization”, IEEE Transactions on Multimedia, 16(6): 1497-1509, Oct. 2014, available at https://ieeexplore.ie… [cited by applicant]
21Manyika, “An overview of Bard: an early experiment with generative AI”, 2023, available at https://ai.google/static/documents/google-about-bard.pdf. [cited by applicant]
22Miech, et al., “HowTo100M: Learning a Text-Video Embedding by Watching Hundred Million Narrated Video Clips”, arXiv preprint arXiv:1906.03327v2 [cs.CV] Jul. 31, 2019, 14 pages. [cited by applicant]
23Mihalcea, et al., “TextRank: Bringing Order into Texts”, In Proceedings of the 2004 Conference on Empirical Methods in Natural Language Processing, pp. 404-411, Jul. 2004, 8 pages. [cited by applicant]
24Narasimhan, et al., “CLIP-It! Language-Guided Video Summarization”, arXiv preprint arXiv:2107.00650v2 [cs.CV] Dec. 8, 2021, 15 pages. [cited by applicant]
25Narasimhan, et al., “TL;DW? Summarizing Instructional Videos with Task Relevance and Cross-Modal Saliency”, arXiv preprint arXiv:2208.06773v1 [cs.CV] Aug. 14, 2022, 24 pages. [cited by applicant]
26OpenAI, “GPT-4 Technical Report”, arXiv preprint arXiv:2303.08774v6 [cs.CL] Mar. 4, 2024, 100 pages. [cited by applicant]
27Ouyang, et al., “LLM is Like a Box of Chocolates: the Non-determinism of ChatGPT in Code Generation”, arXiv preprint arXiv:2308.02828v1 [cs.SE] Aug. 5, 2023, 12 pages. [cited by applicant]
28Qiu, et al., “MMSum: A Dataset for Multimodal Summarization and Thumbnail Generation of Videos”, arXiv preprint arXiv:2306.04216v2 [cs.CV] Nov. 19, 2023, 28 pages. [cited by applicant]
29Radford, et al., “Learning Transferable Visual Models from Natural Language Supervision”, arXiv preprint arXiv:2103.00020v1 [cs.CV] Feb. 26, 2021, 48 pages. [cited by applicant]
30Radford, et al., “Robust Speech Recognition via Large-Scale Weak Supervision”, arXiv preprint arXiv:2212.04356v1 [eess.AS] Dec. 6, 2022, 28 pages. [cited by applicant]
31Reimers, et al., “Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks”, arXiv preprint arXiv:1908.10084v1 [cs.CL] Aug. 27, 2019, 11 pages. [cited by applicant]
32Sharghi, et al., “Query-Focused Extractive Video Summarization”, arXiv preprint arXiv:1607.05177v1 [cs.CV] Jul. 18, 2016, 18 pages. [cited by applicant]
33Sharghi, et al., “Query-Focused Video Summarization: Dataset, Evaluation, and A Memory Network Based Approach”, arXiv preprint arXiv:1707.04960v1 [cs.CV] Jul. 16, 2017, 10 pages. [cited by applicant]
34Song, et al., “TVSum: Summarizing Web Videos Using Titles”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5179-5187, 2015, 9 pages. [cited by applicant]
35Touvron, et al., “Training data-efficient image transformers & distillation through attention”, arXiv preprint arXiv:2012.12877v2 [cs.CV] Jan. 15, 2021, 22 pages. [cited by applicant]
36Touvron, et al., “Llama 2: Open Foundation and Fine-Tuned Chat Models”, arXiv preprint arXiv:2307.09288v2 [cs.CL] Jul. 19, 2023, 77 pages. [cited by applicant]
37Vaswani, et al., “Attention Is All You Need”, arXiv:1706.03762v7 [cs.CL] Aug. 2, 2023, 15 pages. [cited by applicant]
38Wang, et al., “Element-aware Summarization with Large Language Models: Expert-aligned Evaluation and Chain-of-Thought Method”, arXiv preprint arXiv:2305.13412v1 [cs.CL] May 22, 2023, 24 pages. [cited by applicant]
39Zhang, et al., “PEGASUS: Pre-training with Extracted Gap-sentences for Abstractive Summarization”, arXiv preprint arXiv:1912.08777v3 [cs.CL] Jul. 10, 2020, 55 pages. [cited by applicant]
40Zhang, et al., “Benchmarking Large Language Models for News Summarization”, arXiv preprint arXiv:2301.13848v1 [cs.CL] Jan. 31, 2023, 14 pages. [cited by applicant]
41Zhao, et al., “Quasi Real-Time Summarization for Consumer Videos”, In CVPR 2014, 8 pages. [cited by applicant]
42Zhao, et al., “HSA-RNN: Hierarchical Structure-Adaptive RNN for Video Summarization”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7405-7414, 2018, 10 pages. [cited by applicant]
43Zhao, et al., “Reconstructive Sequence-Graph Network for Video Summarization”, arXiv preprint arXiv:2105.04066v1 [cs.CV] May 10, 2021, 10 pages. [cited by applicant]
44Zhong, et al., “Can LLM Replace Stack Overflow? A Study on Robustness and Reliability of Large Language Model Code Generation”, arXiv preprint arXiv:2308.10335v5 [cs.CL] Jan. 27, 2024, 9 pages. [cited by applicant]
45Zhu, et al., “DSNet: A Flexible Detect-to-Summarize Network for Video Summarization”, in IEEE Transactions on Image Processing, vol. 30, pp. 948-962, 2021, available at https://ieeexplore.ieee.org/document/9275314. [cited by applicant]