IP Library Granted Patent US 11,729,478
Granted Patent B2
US 11,729,478 · App. 16/218,351 · Granted Aug 15, 2023

System and method for algorithmic editing of video content

Inventors: Robert Andrew Hitching (Los Altos Hills, CA); Ashley John Wing (San Francisco, CA); Phillip John Wing (Drummoyne, AU)
Assignee: Playable Pty Ltd
H04N21/8549G06F16/71G06F16/75G06F16/784G06F16/7834H04N21/44204
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,729,478
App. No.
16/218,351
Granted
Aug 15, 2023
Kind
B2
Abstract

A computer implemented method for algorithmically editing digital video content is disclosed. A video file containing source video is processed to extract metadata. Label taxonomies are applied to extracted metadata. The labelled metadata is processed to identify higher-level labels. Identified higher-level labels are stored as additional metadata associated with the video file. A clip generating algorithm applies the stored metadata for selectively editing the source video to generate a plurality of different candidate video clips. Responsive to determining a clip presentation trigger on a viewer device, a clip selection algorithm is implemented that applies engagement data and metadata for the candidate video clips to select one of the stored candidate video clips. The engagement data is representative of one or more engagement metrics recorded for at least one of the stored candidate video clips. The selected video clip is presented to one or more viewers via corresponding viewer devices.

Claims (47)

1. A computer implemented method for algorithmically editing digital video content, the method comprising the steps of:

receiving a video file containing source video from a video source;

processing the video file to extract metadata representative of at least one of:

video content within the source video; and

audio content within the source video;

applying labels to the extracted metadata and storing the labelled metadata in a metadata store in association with the video file;

processing the labelled metadata to identify higher-level labels based at least in part on one or more semantic clustering algorithms and one or more statistical proximity algorithms programmed to statistically calculate proximity between labels based on the frequency and proportion of time that two or more labels are detected alone and simultaneously in the video file for the labelled metadata;

storing any identified higher-level labels in the metadata store as additional metadata associated with the video file;

implementing a clip generating algorithm that applies the stored metadata to selectively edit the source video to thereby generate a plurality of different candidate video clips therefrom;

storing the plurality of candidate video clips in a video clip data store;

storing metadata for the plurality of candidate video clips in the metadata store;

determining that a viewer independent of the video source has triggered a clip presentation trigger via a viewer display device;

responsive to determining the clip presentation trigger, implementing a clip selection algorithm that retrieves the metadata for the plurality of candidate clips as well as any available aggregated engagement data for the plurality of candidate video clips and applies the retrieved metadata and available aggregated engagement data to select one of the plurality of stored candidate video clips; and

presenting the selected video clip to the viewer via their associated viewer display device; and

wherein each time one of the stored candidate video dips is presented to a viewer one or more engagement metrics are recorded based on interactions by the viewer during or following that presentation and wherein the one or more engagement metrics recorded for presentations of the same candidate video clip to different viewers is recorded as aggregated engagement data for storing in association with the corresponding candidate video clip and wherein aggregated engagement data associated with the same or similar video file(s) is evaluated and selectively applied by the clip generating algorithm for creating additional candidate video clips for storing in the video clip data store; and wherein the candidate video clips are generated and stored ready for selection and presentation prior to the viewer display device triggering the clip presentation trigger.

2. The method in accordance with claim 1 , further comprising processing the labelled metadata to infer applicable video genres and subgenres that are associated with individual labels or groupings of labels in the labelled metadata and wherein the inferred video genres are stored as additional metadata.

3. The method in accordance with claim 1 , further comprising determining a digital medium by which the selected video clip is to be presented and wherein the digital medium is additionally applied by the clip generating algorithm for determining how to edit the source video for generating the plurality of candidate video clips.

4. The method in accordance with claim 1 , further comprising determining individual viewer profile data for the viewer(s) and wherein the individual viewer profile data is additionally applied by the clip selection algorithm for selecting the candidate video clip to present to the one or more viewers.

5. The method in accordance with claim 1 , wherein the one or more engagement metrics are selected from the following:

(a) a dwell time recorded for viewers that have previously been presented the corresponding candidate video clip; and

(b) a number of times a predefined event has been executed during or following a prior presentation of the corresponding candidate video clip.

6. The method in accordance with claim 1 , wherein the metrics are recorded in association with a context for the prior presentation and wherein the method further comprises determining a current context for the presentation such that clip selection algorithm applies more weight to metrics associated with the current context than metrics that are not associated with the current context.

7. The method in accordance with claim 1 , further comprising determining audience profile data for a collective group of viewers that are to be presented the selected video clip and wherein the audience profile data is applied by at least one of the clip generating and clip selection algorithms.

8. The method in accordance with claim 1 , wherein the step of processing the video file to determine metadata representative of video content comprises analyzing pixels within each frame of the video to measure brightness levels, detect edges, quantify movement from the previous frame, and identify objects, faces, facial expressions, scene changes, text overlays, and/or events.

9. The method in accordance with claim 1 , wherein the step of processing the video file to determine metadata representative of audio content comprises analyzing audio within the video file to determine predefined audio characteristics.

10. The method in accordance with claim 1 , wherein the step of processing the video file to determine metadata representative of audio content comprises performing a speech recognition process to determine speech units within the source video.

11. The method in accordance with claim 1 , further comprising extracting metadata embedded in the video file and storing the embedded metadata in the data store in association with the video file.

12. The method in accordance with claim 1 , wherein the metrics are recorded and applied in real, or near real time, by the clip selection algorithm.

13. A computer system for algorithmically editing digital video content, the system comprising:

a server system implementing:

a video analysis module configured to:

receive a video file containing source video from a video source;

process the video file to extract metadata representative of at least one of:

video content within the source video; and

audio content within the source video;

apply labels to the extracted metadata and storing the labelled metadata in a metadata store in association with the video file;

process the labelled metadata to identify higher-level labels based at least in part on one or more semantic clustering algorithms and one or more statistical proximity algorithms programmed to statistically calculate proximity between labels based on the frequency and proportion of time that two or more labels are detected alone and simultaneously in the video file for the labelled metadata;

store any identified higher-level labels in the metadata store as additional metadata associated with the video file; and

a clip selection module configured to:

implement a clip generating algorithm that applies the stored metadata to selectively edit the source video to thereby generate a plurality of different candidate video clips therefrom;

store the plurality of candidate video clips in a video clip data store;

store metadata for the plurality of candidate video clips in the metadata store;

determine that a viewer independent of the video source has triggered a clip presentation trigger via a viewer display device;

responsive to determining the clip presentation trigger, implementing a clip selection algorithm that retrieves the metadata for the plurality of candidate clips as well as any available aggregated engagement data for the plurality of candidate video clips and applies the retrieved metadata and available aggregated engagement data to select one of the plurality of stored candidate video clips; and

present the selected video clip to the viewer via their associated viewer display device;

wherein each time one of the stored candidate video dips is presented to a viewer one or more engagement metrics are recorded based on interactions by the viewer during or following that presentation and wherein the one or more engagement metrics recorded for presentations of the same candidate video clip to different viewers is recorded as aggregated engagement data for storing in association with the corresponding candidate video clip and wherein aggregated engagement data associated with the same or similar video file(s) is evaluated and selectively applied by the clip generating algorithm for creating additional candidate video clips for storing in the video clip data store.

14. A non-transitory computer readable medium storing at least one instruction, which when executed by a computing system, is operable to carry out the method in accordance with claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2019
From: WING, ASHLEY JOHN; WING, PHILLIP JOHN; HITCHING, ROBERT ANDREW
To: PLAYABLE PTY LTD
Reel/Frame 048448/0032 →
Priority Claims (1)
AU 2017905005 · Dec 13, 2017 · national
Continuity (1)
Related Publication 20190182565A1 · Jun 13, 2019
Cited By (2)
US 12,229,207 US 12,634,563