IP Library Granted Patent US 12,254,036
Granted Patent B2
US 12,254,036 · App. 16/398,828 · Granted Mar 18, 2025

System and method for summarizing a multimedia content item

Inventor: Inderjeet Mani (Sunnyvale, CA)
Assignee: YAHOO ASSETS LLC
G06F16/345G10L25/48G10L15/18G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,254,036
App. No.
16/398,828
Granted
Mar 18, 2025
Kind
B2
Abstract

A multimedia content item is summarized based on its audio track and a desired compression budget. The audio track is extracted and processed by an automatic speech recognizer to obtain a time-aligned text transcript. The text-transcript is partitioned into a plurality of segment sequences. An informativeness score based on a salience score and a diversity score is computed for each of the segments. A coherence score is also computed for the segments in the plurality of sequences. A subsequence of one of the segment sequences that optimizes for informativeness and coherence is selected for generating a new content item summarizing the multimedia content item.

Claims (47)

1. A method comprising:

identifying, by a computing device, a multimedia content item comprising digital content and an audio track;

analyzing, via the computing device, the multimedia content item, and based on said analysis, identifying a compression budget, said compression budget comprising information indicating a runtime value of the multimedia content item and a maximum percentage of a total of the runtime value of the multimedia content item;

determining, via the computing device, that the maximum percentage within the compression budget is at least a threshold fraction less than the runtime value;

partitioning, via the computing device, upon said determination, said multimedia content item based on the runtime value and the maximum percentage of the total of the runtime value indicated in said compression budget;

identifying, via the computing device, based on said partitioning, segments of digital content and audio information, the segments corresponding to a portion of the multimedia content item;

generating, via the computing device, a summary content item based on said partitioning, said summary content item comprising said identified segments of the digital content and audio information;

obtaining a time-aligned text transcript of the audio track;

constructing a graph representing the segments, wherein each segment is a node in the graph;

computing an informativeness score for each segment based on a corresponding salience score and diversity score; and

selecting segments for the summary content item based on the informativeness scores.

2. The method of claim 1 , wherein said compression budget indicates a value for a runtime of the summary content item to be generated, said runtime of the summary content item to be generated being a target length of the summary content item.

3. The method of claim 2 , further comprising: identifying a subset of said segments of digital content and audio information, said subset identification based on the runtime of the summary content item indicated by the compression budget.

4. The method of claim 1 , wherein said analysis further comprises: analyzing said multimedia content item by applying a linear regression algorithm, and determining, based on said application, whether said multimedia content item can be summarized, wherein said generation is based on said determination.

5. The method of claim 4 , wherein said determination is further based on a compression rate of said multimedia content item.

6. The method of claim 1 , wherein said analysis further comprises: identifying a length of the digital content, a source of the digital content and information indicating coarse-grained features of the digital content and the audio track.

7. The method of claim 1 , wherein said each segment of the digital content includes information indicating time intervals within the multimedia content item, wherein the time intervals are a sequence of non-contiguous, non-overlapping sub-intervals of equal lengths.

8. A non-transitory computer-readable storage medium tangibly encoded with computer-executable instructions, that when executed by a computing device, perform a method comprising:

identifying, by the computing device, a multimedia content item comprising digital content and an audio track;

analyzing, via the computing device, the multimedia content item, and based on said analysis, identifying a compression budget, said compression budget comprising information indicating a runtime value of the multimedia content item and a maximum percentage of a total of the runtime value of the multimedia content item;

determining, via the computing device, that the maximum percentage within the compression budget is at least a threshold fraction less than the runtime value;

partitioning, via the computing device, upon said determination, said multimedia content item based on the runtime value and the maximum percentage of the total of the runtime value indicated in said compression budget;

identifying, via the computing device, based on said partitioning, segments of digital content and audio information, the segments corresponding to a portion of the multimedia content item;

generating, via the computing device, a summary content item based on said partitioning, said summary content item comprising said identified segments of the digital content and audio information;

obtaining a time-aligned text transcript of the audio track;

constructing a graph representing the segments, wherein each segment is a node in the graph;

computing an informativeness score for each segment based on a corresponding salience score and diversity score; and

selecting segments for the summary content item based on the informativeness scores.

9. The non-transitory computer-readable storage medium of claim 8 , wherein said compression budget indicates a value for a runtime of the summary content item to be generated, said runtime of the summary content item to be generated being a target length of the summary content item.

10. The non-transitory computer-readable storage medium of claim 9 , further comprising: identifying a subset of said segments of digital content and audio information, said subset identification based on the runtime of the summary content item indicated by the compression budget.

11. The non-transitory computer-readable storage medium of claim 8 , wherein said analysis further comprises: analyzing said multimedia content item by applying a linear regression algorithm, and determining, based on said application, whether said multimedia content item can be summarized, wherein said generation is based on said determination.

12. The non-transitory computer-readable storage medium of claim 11 , wherein said determination is further based on a compression rate of said multimedia content item.

13. The non-transitory computer-readable storage medium of claim 8 , wherein said analysis further comprises: identifying a length of the digital content, a source of the digital content and information indicating coarse-grained features of the digital content and the audio track.

14. The non-transitory computer-readable storage medium of claim 8 , wherein said each segment of the digital content includes information indicating time intervals within the multimedia content item, wherein the time intervals are a sequence of non-contiguous, non-overlapping sub-intervals of equal lengths.

15. A computing device comprising:

a processor; and

a non-transitory computer-readable storage medium for tangibly storing thereon program logic for execution by the processor, the program logic comprising:

logic executed by the processor for identifying a multimedia content item comprising digital content and an audio track;

logic executed by the processor for analyzing the multimedia content item, and based on said analysis, identifying a compression budget, said compression budget comprising information indicating a runtime value of the multimedia content item and a maximum percentage of a total of the runtime value of the multimedia content item;

logic executed by the processor for determining that the maximum percentage within the compression budget is at least a threshold fraction less than the runtime value;

logic executed by the processor for partitioning, upon said determination, said multimedia content item based on the runtime value and the maximum percentage of the total of the runtime value indicated in said compression budget;

logic executed by the processor for identifying, based on said partitioning, segments of digital content and audio information, the segments corresponding to a portion of the multimedia content item;

logic executed by the processor for generating a summary content item based on said partitioning, said summary content item comprising said identified segments of the digital content and audio information;

logic executed by the processor for obtaining a time-aligned text transcript of the audio track;

logic executed by the processor for constructing a graph representing the segments, wherein each segment is a node in the graph;

logic executed by the processor for computing an informativeness score for each segment based on a corresponding salience score and diversity score; and

logic executed by the processor for selecting segments for the summary content item based on the informativeness scores.

Assignments (6)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2019
From: MANI, INDERJEET
To: YAHOO! INC.
Reel/Frame 049165/0439 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2019
From: YAHOO! INC.
To: YAHOO HOLDINGS, INC.
Reel/Frame 049165/0631 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 14, 2019
From: YAHOO HOLDINGS, INC.
To: OATH INC.
Reel/Frame 049166/0035 →
Continuity (2)
Continuation 14224511 · Mar 25, 2014
Related Publication 20190258660A1 · Aug 22, 2019
References Cited (16)
US 20030055634A1 · Hidaka · 2003 [cited by examiner]
US 20040085339A1 · Divakaran · 2004 [cited by examiner]
US 20050276570A1 · Reed · 2005 [cited by examiner]
US 20130132374A1 · Olstad · 2013 [cited by examiner]
US 20150370808A1 · Olstad · 2015 [cited by examiner]
WO 2009042340 · 2009 [cited by applicant]
WO 2013006497A2 · 2013 [cited by applicant]
WO 2013066497A1 · 2013 [cited by applicant]
Alemany, “Representing discourse for automatic text summarization via shallow NLP techniques,” (2005). [cited by applicant]
Brin et al., “The Anatomy of a Large-Scale Hypertextual Web Search Engin,” Stanford, CA; 20 pages. [cited by applicant]
Dubey et al., “Diversity in Ranking via Resistive Graph Centers,” (2011). [cited by applicant]
Mani et al., “Using Cohesion and Coherence Models for Text Summarization.” (1998). [cited by applicant]
Mei, Q., et al., “DivRank: the Interplay of Prestige and Diversity in Information Networks;” University of Michigan; Jul. 25-28; 10 pages (2010). [cited by applicant]
Mihalcea, Rad; “Language Independent Extractive Summarization;” University of North Texas; 4 pages. [cited by applicant]
Lin, Chin-Yew; “Rough: A Package for Automatic Evaluation of Summaries;” University of Southern California; 10 pages (2004). [cited by applicant]
Zweig et al., “Continuous Speech Recognition with a TF-IDF Acoustic Model,” (2010). [cited by applicant]