IP Library › Granted Patent US 12,626,729
Granted Patent B2
US 12,626,729 · App. 18/326,720 · Granted May 12, 2026

System and method for video/audio comprehension and automated clipping

Inventors: Andrew Hyde (Burbank, CA); Geoffrey Booth (Burbank, CA); Ognjen Boras (Burbank, CA); Jonathan Flanders (Burbank, CA); Danny Donnell (Burbank, CA)
Assignee: DISNEY ENTERPRISES, INC.
G11B27/34G06F40/205G06F40/295H04N21/8549
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,729
App. No.
18/326,720
Filed
May 31, 2023
Granted
May 12, 2026
Kind
B2
Examiner
TRAN, LOI H
Art Unit
2484
USPC
386/241
Abstract

Systems and Methods for Video/Audio Comprehension and Automated Clipping includes providing at least one media clip (MC) within an event for display or listening on a user device including receiving audio or video media data indicative of the event, transcribing the media data into timestamped text, identifying entities within the text, creating text segments having a begin timestamp and end timestamp and having a minimum number of entity mentions in the text segments, clipping from the media data the at least one media clip having a begin timestamp and end timestamp corresponding to the begin timestamp and end timestamp of a corresponding one of the text segments, and providing the at least one media clip to the user device for viewing or listening by a user. Feedback may also be provided to adjust the logic that identifies MCs. MC Alerts may also be sent to users autonomously or based on user-set parameters.

Claims (73)

1 . An automated computer-based method for providing at least one media clip (MC) from an event for display or listening on a user device, comprising:

receiving media data indicative of the event;

transcribing the audio portion of the media data into text with timestamps;

identifying entities within the text, the entities being named in the content of the text;

performing phonetic correction and co-reference resolution of the entities using predetermined phonetic rules and predetermined co-reference rules, respectively;

segmenting the text into a plurality of text segments based on predetermined text segment creation rules, each of the text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp;

clipping from the media data the at least one media clip having a clip begin timestamp and clip end timestamp that corresponds to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and

providing the at least one media clip for viewing or listening on the user device,

wherein the identifying, performing, segmenting, and clipping are performed contiguously in an automated manner without human intervention after receiving the media data using the predetermined phonetic rules, the predetermined co-reference rules, and the predetermined text segment creation rules.

2 . The method of claim 1 , further comprising determining an entity classification of the at least one entity comprising an amount of time that the at least one entity is mentioned during a given segment or during the entire event, and wherein the media clip includes the entity classification.

3 . The method of claim 1 , wherein the segmenting further comprises creating text clusters from the text based on cluster creation rules, each cluster having at least one entity and having a cluster begin time and a cluster end time.

4 . The method of claim 3 , wherein the cluster creation rules comprises at least one of: maximum entity gap length, minimum mention count, minimum cluster length, and cluster adjustment time, and cluster exclusion rules.

5 . The method of claim 3 , further comprising receiving feedback from a user or an editor on the quality of the at least one media clip and adjusting the cluster creation rules or segment creation rules to improve the quality of media clip.

6 . The method of claim 5 , where in the adjusting is performed using a machine learning model which is trained using prior adjustments.

7 . The method of claim 1 , wherein the segment creation rules comprises at least one of maximum segment length and segment exclusion rules.

8 . The method of claim 1 , wherein the phonetic rules comprises a minimum possible phonetic partial match.

9 . The method of claim 1 , wherein the co-reference rules comprises a co-reference offset maximum.

10 . The method of claim 1 , further comprising aggregating a plurality of the at least one media clip from a plurality of different shows or events.

11 . The method of claim 10 , wherein the plurality of different shows or events corresponds to shows or events selected by the user.

12 . The method of claim 1 , wherein the co-reference resolution comprises associating the entities in the text with corresponding pronouns, relationship words, nicknames, and abbreviations.

13 . The method of claim 1 , wherein the user device comprises a graphic user interface (GUI), which when selected, causes the media clip to play on a device display.

14 . The method of claim 1 , further comprising sending an MC alert message to the user device when a MC is available for viewing or predetermined MC alert criteria are satisfied.

15 . The method of claim 14 , wherein the predetermined MC alert criteria comprise at least one of: MC matching user attributes, MC matching user MC likes, MC matching user Alert settings.

16 . The method of claim 1 , further comprising receiving a settings command from a user and receiving settings inputs from a user.

17 . The method of claim 1 , further comprising receiving user attributes data from a user.

18 . The method of claim 1 , further comprising determining a title for the text segment and providing the title with the media clip for display by the user device.

19 . The method of claim 1 , wherein the media clip is less than 5 min long.

20 . The method of claim 1 , wherein the event comprises a sports show or sporting event.

21 . The method of claim 1 , wherein the media data comprises an audio-only data file.

22 . An automated computer-based method for providing at least one media clip (MC) from an event for display on a user device, comprising:

receiving media data indicative of the event, the media data having video timestamps;

transcribing the audio portion of the media data into timestamped text;

identifying one or more entities within the text, the entities being named in the content of the text;

performing phonetic correction and co-reference resolution of the entities;

segmenting the text into a plurality of text segments, each of the text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp;

determining an entity classification of the at least one entity comprising an amount of time the at least one entity is mentioned during a given segment or during the entire event;

clipping from the media data the at least one media clip having a media clip begin timestamp and media clip end timestamp that corresponds to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and

providing the at least one media clip with the entity classification to the user device, the user device being configured to show the at least one media clip,

wherein the identifying, performing, segmenting, determining, and clipping are performed contiguously in an automated manner without human intervention after receiving the media data, using predetermined rules.

23 . The method of claim 22 wherein the phonetic correction, co-reference resolution, and the segmenting are performed using the predetermined rules.

24 . An automated computer-based method for providing at least one media clip (MC) from a sports event for display on a user device, comprising:

receiving media data indicative of the event, the media data having an audio channel and a video channel, the video channel having video timestamps;

transcribing the audio channel portion of the media data into text with timestamps;

identifying entities being named in the content of the text and associating the entities in the content of the text with corresponding pronouns, relationship words, nicknames, and abbreviations;

creating a plurality of text segments from the text, each of the text segments having the at least one of the entities and having a segment begin timestamp and a segment end timestamp;

extracting from the media data the at least one media clip having a media clip begin timestamp and a media clip end timestamp that corresponds to the segment begin timestamp and segment end timestamp of a corresponding one of the text segments; and

providing the at least one media clip to the user device for viewing by a user,

wherein the identifying, creating, and extracting are performed contiguously in an automated manner without human intervention after receiving the media data, using predetermined rules.

25 . An automated computer-based method for providing at least one media clip (MC) from an event for display or listening on a user device comprising:

receiving media data indicative of the event;

transcribing the media data into timestamped text;

identifying entities within the text, the entities being named in the content of the text;

creating text segments having at least one of the entities and having a segment begin timestamp and a segment end timestamp and having a minimum number of entity mentions in the text segments within a maximum entity gap length time;

clipping from the media data the at least one media clip having a clip begin timestamp and a clip end timestamp corresponding to the segment begin timestamp and the segment end timestamp of a corresponding one of the text segments; and

providing the at least one media clip to the user device for viewing or listening by a user,

wherein the identifying, creating, and clipping are performed contiguously in an automated manner without human intervention after receiving the media data, using predetermined rules.

26 . The method of claim 25 , further comprising determining an entity classification of the at least one entity comprising an amount of time that the at least one entity is mentioned during a given segment or during the entire event, and wherein the media clip includes the entity classification.

27 . The method of claim 25 , wherein the creating text segments further comprises creating text clusters from the text based on cluster creation rules, each cluster having at least one entity and having a cluster begin time and a cluster end time.

28 . The method of claim 27 , wherein the cluster creation rules comprises at least one of: maximum entity gap length, minimum mention count, minimum cluster length, and cluster adjustment time, and cluster exclusion rules.

29 . The method of claim 28 , further comprising receiving feedback from a user or an editor on the quality of the at least one media clip and adjusting the cluster creation rules or segment creation rules to improve the quality of media clips.

30 . The method of claim 29 , where in the adjusting is performed using a machine learning model which is trained using prior adjustments.

31 . The method of claim 27 , wherein the segment creation rules comprises at least one of maximum segment length and segment exclusion rules.

32 . The method of claim 25 , further comprising performing phonetic correction and co-reference resolution of the entities using predetermined phonetic rules and predetermined co-reference rules, respectively; wherein the phonetic rules comprises a minimum possible phonetic partial match.

33 . The method of claim 32 , wherein the co-reference rules comprises a co-reference offset maximum.

34 . The method of claim 25 , wherein the event comprises a sports show or sporting event.

35 . An automated computer-based method for identifying and classifying entities in text and providing entity classification tagged text, comprising:

receiving text data, which is a transcription with timestamps of an audio portion of media data;

identifying at least one entity within the text data, the at least one entity being named in the content of the text;

performing phonetic correction and co-reference resolution of the entities using predetermined phonetic rules and predetermined co-reference rules, respectively;

segmenting the text into a plurality of text segments based on predetermined text segment creation rules, each of the text segments having at least one of the entities and having a segment;

determining an entity classification of the at least one entity comprising an amount that the at least one entity is mentioned during a given segment or during the entire event, and wherein the text segment includes the entity classification; and

providing the text segments with timestamps and with entity classification to a user device for viewing on the user device,

wherein the identifying, performing, segmenting, and determining are performed contiguously in an automated manner without human intervention after receiving the text data, using the predetermined rules.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 5, 2026
From: DONNELL, DANNY; HYDE, ANDREW; BORAS, OGNJEN; BOOTH, GEOFFREY; FLANDERS, JONATHAN
To: DISNEY ENTERPRISES, INC.
Reel/Frame 074860/0678 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2023
From: DONNELL, DANNY; HYDE, ANDREW; BORAS, OGNJEN; BOOTH, GEOFFREY; FLANDERS, JONATHAN
To: DISNEY ENTERPRISES, INC.
Reel/Frame 064198/0670 →
Continuity (1)
Related Publication 20240404563A1 · Dec 5, 2024
References Cited (17)
US 10289963B2 · Chiticariu · 2019 [cited by applicant]
US 11055334B2 · Hu · 2021 [cited by examiner]
US 11354894B2 · Farre Guiu et al. · 2022 [cited by applicant]
US 11501075B1 · Pandey · 2022 [cited by examiner]
US 20050267871A1 · Marchisio · 2005 [cited by examiner]
US 20150088888A1 · Brennan · 2015 [cited by applicant]
US 20150242387A1 · Rachevsky · 2015 [cited by applicant]
US 20160042473A1 · Danielli · 2016 [cited by examiner]
US 20160365121A1 · DeCaprio · 2016 [cited by applicant]
US 20180191852A1 · Brunn · 2018 [cited by examiner]
US 20190065911A1 · Lee · 2019 [cited by applicant]
US 20190294999A1 · Guttmann · 2019 [cited by examiner]
US 20200066271A1 · Li · 2020 [cited by examiner]
US 20200410053A1 · Zhang · 2020 [cited by examiner]
US 20240380945A1 · Schweinsberg · 2024 [cited by examiner]
English Translation of WIPO Publication WO/2023/246395, PCT CN2023/095265 filed May 19, 2023 (Year: 2023). [cited by examiner]
English Translation of Chinese Publication CN116127003 May 16, 2023 (Year: 2023). [cited by examiner]