IP Library Granted Patent US 10,474,903
Granted Patent B2
US 10,474,903 · App. 15/880,077 · Granted Nov 12, 2019

Video segmentation using predictive models trained to provide aesthetic scores

Inventors: Sagar Tandon (Ghaziabadv, IN); Abhishek Shah (Pitam Pura, IN)
Assignee: Adobe Inc.
G06K9/00765G06K9/00255G06K9/00718G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,474,903
App. No.
15/880,077
Granted
Nov 12, 2019
Kind
B2
Abstract

Systems and methods for segmenting video. A segmentation application executing on a computing device receives a video including video frames. The segmentation application calculates, using a predictive model trained to evaluate quality of video frames, a first aesthetic score for a first video frame and a second aesthetic score for a second video frame. The segmentation application determines that the first aesthetic score and the second aesthetic score differ by a quality threshold and that a number of frames between the first video frame and the second video frame exceeds a duration threshold. The segmentation application creates a video segment by merging a subset of video frames ranging from the first video frame to an segment-end frame preceding the second video frame.

Claims (77)

1. A computer-implemented method for segmenting video, the method performed by a computing system and comprising:

receiving a video comprising a plurality of video frames;

calculating, with a predictive model trained to evaluate quality of an input video frame, a first aesthetic score for a first video frame from the plurality of video frames and a second aesthetic score for a second video frame from the plurality of video frames;

determining that (i) the first aesthetic score and the second aesthetic score differ by a quality threshold and (ii) a number of frames between the first video frame and the second video frame exceeds a duration threshold; and

creating a video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames ranging from the first video frame to an segment-end frame preceding the second video frame.

2. The computer-implemented method of claim 1 , further comprising:

calculating, for each of the plurality of video frames, a facial identity of a face depicted in the respective video frame;

determining that a third video frame of the plurality of video frames and a fourth video frame of the plurality of video frames have an identical facial identity;

merging the third video frame and the fourth video frame into a face- detected video segment comprising a facial identity;

determining that the face-detected video segment and the video segment have (i) a number of frames less than a second duration threshold, (ii) the same an identical facial identity, or (iii) aesthetic scores within an aesthetic score threshold; and

responsive to determining that the video segment and the face-detected video segment overlap, creating a combined segment by merging the face-detected video segment and the video segment.

3. The computer-implemented method of claim 2 , further comprising:

calculating, for each of the plurality of video frames, a histogram comprising a distribution of color frequency in the respective video frame;

determining that a fifth video frame of the plurality of video frames and a sixth video frame of the plurality of video frames have histogram scores that differ by more than a histogram threshold;

creating a scene-detected video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames ranging from the fifth video frame to an second segment-end frame preceding the sixth video frame; and

responsive to determining that the scene-detected video segment and the combined segment overlap, merging the scene-detected video segment with the video segment.

4. The computer-implemented method of claim 1 , further comprising:

receiving a training image and a training label indicating an aesthetic score determined by a human; and

training the predictive model by providing the training image and the aesthetic score to the predictive model.

5. computer-implemented method of claim 4 , wherein the first aesthetic score and the second aesthetic score each comprises a component that is a measure of (i) color harmony, (ii) a balance of elements in a frame, (iii) whether content is interesting, (iv) depth of field, (v) whether light in a scene is interesting, (vi) which object is an emphasis of a scene, (vii) whether there is repetition, (viii) a rule of thirds, (ix) vivid colors, or (x) symmetry.

6. The computer-implemented method of claim 1 , further comprising providing the video segment to a user interface.

7. The computer-implemented method of claim 1 , further comprising:

determining an average aesthetic score for the video segment by averaging the aesthetic scores for all of the video frames in the video segment;

identifying, the video segment as a summary segment by determining that the video segment has an average aesthetic score greater than a sixth threshold; and

providing the summary segment to a user interface.

8. A system comprising:

a non-transitory computer-readable medium storing computer-executable program instructions for segmenting video; and

a processing device communicatively coupled to the non-transitory computer-readable medium for executing the computer-executable program instructions, wherein executing the computer-executable program instructions configures the processing device to perform operations comprising:

receiving a video comprising a plurality of video frames;

calculating, with a predictive model trained to evaluate quality of an input video frame, a first aesthetic score for a first video frame from the plurality of video frames and a second aesthetic score for a second video frame from the plurality of video frames, wherein the predictive model is trained to predict aesthetic scores for video frames by receiving a plurality of images comprising training la bels indicating aesthetic scores determined by a human;

determining that the first aesthetic score and the second aesthetic score differ by a quality threshold; and

creating a video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames including the first video frame and an segment-end frame preceding the second video frame.

9. The system of claim 8 , wherein creating the video segment is performed responsive to determining that a number of frames between the first video frame and the second video frame exceeds a duration threshold.

10. The system of claim 8 , wherein the program instructions further configure the processing device to perform operations comprising:

calculating, for each of the plurality of video frames, a facial identity of a face depicted in the respective video frame;

determining that a third video frame of the plurality of video frames and a fourth video frame of the plurality of video frames have an identical facial identity;

merging the third video frame and the fourth video frame into a face-detected video segment comprising a facial identity;

determining that the face-detected video segment and the video segment have (i) a number of frames less than a second duration threshold, (ii) the same an identical facial identity, or (iii) aesthetic scores within an aesthetic score threshold; and

responsive to determining that the video segment and the face-detected video segment overlap, creating a combined segment by merging the face-detected video segment and the video segment.

11. The system of claim 10 , wherein the program instructions further configure the processing device to perform operations comprising:

determining a first average aesthetic score for the face-detected video segment and a second average aesthetic score for the video segment; and

responsive to determining that the first average aesthetic score and the second aesthetic score are within an average aesthetic score threshold, combining the face-detected video segment with the video segment.

12. The system of claim 10 , wherein the program instructions further configure the processing device to perform operations comprising:

responsive to determining that the video segment and the face-detected video segment overlap, creating a combined segment by merging the face-detected video segment and the video segment;

calculating, for each of the plurality of video frames, a histogram comprising a distribution of color frequency in the respective video frame;

determining that a fifth video frame of the plurality of video frames and a sixth video frame of the plurality of video frames have histogram scores that differ by more than a histogram threshold;

creating a scene-detected video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames ranging from the fifth video frame to an second segment-end frame preceding the sixth video frame; and

responsive to determining that the scene-detected video segment and the combined segment overlap, merging the scene-detected video segment with the video segment.

13. The system of claim 11 , wherein the first aesthetic score and the second aesthetic score each comprises a component that is a measure of (i) color harmony, (ii) a balance of elements in a frame, (iii) whether content is interesting, (iv) depth of field, (v) whether light in a scene is interesting, (vi) which object is an emphasis of a scene, (vii) whether there is repetition, (viii) a rule of thirds, (ix) vivid colors, or (x) symmetry.

14. The system of claim 10 , wherein the program instructions further configure the processing device to perform operations comprising:

determining an average aesthetic score for the video segment by averaging the aesthetic scores for all of the video frames in the video segment;

identifying, the video segment as a summary segment by determining that the video segment has an average aesthetic score greater than a sixth threshold; and

providing the summary segment to a user interface.

15. A non-transitory computer-readable storage medium storing computer-executable program instructions for segmenting video, wherein when executed by a processing device, the computer-executable program instructions cause the processing device to perform operations comprising:

a step for receiving a video comprising a plurality of video frames;

a step for calculating, with a predictive model trained to evaluate quality of an input video frame, a first aesthetic score for a first video frame from the plurality of video frames and a second aesthetic score for a second video frame from the plurality of video frames;

a step for determining that (i) the first aesthetic score and the second aesthetic score differ by a quality threshold and (ii) a number of frames between the first video frame and the second video frame exceeds a duration threshold; and

a step for creating a video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames ranging from the first video frame to an segment -end frame preceding the second video frame.

16. The non-transitory computer-readable storage medium of claim 15 , wherein the computer-executable program instructions further cause the processing device to perform operations comprising:

a step for calculating, for each of the plurality of video frames, a facial identity of a face depicted in the respective video frame;

a step for determining that a third video frame of the plurality of video frames and a fourth video frame of the plurality of video frames have an identical facial identity;

a step for merging the third video frame and the fourth video frame into a face-detected video segment comprising a facial identity;

a step for determining that the face-detected video segment and the video segment have (i) a number of frames less than a second duration threshold, (ii) the same an identical facial identity, or (iii) aesthetic scores within an aesthetic score threshold; and

a step for responsive to determining that the video segment and the face-detected video segment overlap, creating a combined segment by merging the face-detected video segment and the video segment.

17. The non-transitory computer-readable storage medium of claim 16 , wherein the computer-executable program instructions further cause the processing device to perform operations comprising:

a step for calculating, for each of the plurality of video frames, a histogram comprising a distribution of color frequency in the respective video frame;

a step for determining that a fifth video frame of the plurality of video frames and a sixth video frame of the plurality of video frames have histogram scores that differ by more than a histogram threshold;

a step for creating a scene-detected video segment by merging a subset of video frames from the plurality of video frames, the subset of video frames ranging from the fifth video frame to an second segment-end frame preceding the sixth video frame; and

a step for responsive to determining that the scene-detected video segment and the combined segment overlap, merging the scene-detected video segment with the video segment.

18. The non-transitory computer-readable storage medium of claim 15 , wherein the computer-executable program instructions further cause the processing device to perform operations comprising:

a step for receiving a training image and a training label indicating an aesthetic score determined by a human; and

a step for training the predictive model by providing the training image and the aesthetic score to the predictive model.

19. The non-transitory computer-readable storage medium of claim 18 , wherein the first aesthetic score and the second aesthetic score each comprises a component that is a measure of (i) color harmony, (ii) a balance of elements in a frame, (iii) whether content is interesting, (iv) depth of field, (v) whether light in a scene is interesting, (vi) which object is an emphasis of a scene, (vii) whether there is repetition, (viii) a rule of thirds, (ix) vivid colors, or (x) symmetry.

20. The computer-readable storage medium of claim 15 , wherein the instructions further cause the processing device to perform operations comprising:

a step for determining an average aesthetic score for the video segment by averaging the aesthetic scores for all of the video frames in the video segment;

a step for identifying, the segment as a summary segment by determining that the video segment has an average aesthetic score greater than a sixth threshold; and

a step for providing the summary segment to a user interface.

Assignments (2)
CHANGE OF NAME Recorded Mar 6, 2019
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 048525/0042 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 25, 2018
From: TANDON, SAGAR; SHAH, ABHISHEK
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 044731/0484 →
Continuity (1)
Related Publication 20190228231A1 · Jul 25, 2019
Cited By (2)
US 12,192,619 US 12,633,122