IP Library Granted Patent US 10,777,228
Granted Patent B1
US 10,777,228 · App. 15/933,008 · Granted Sep 15, 2020

Systems and methods for creating video edits

Inventor: Samuel Wilson (San Mateo, CA)
Assignee: GoPro, Inc.
G11B27/02G06N3/08G11B27/031
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,777,228
App. No.
15/933,008
Granted
Sep 15, 2020
Kind
B1
Abstract

Feature information characterize features of video clips may be obtained. A given video clip may be selected as a segment of a video edit. Other video clips may be iteratively selected as other segments of the video edit based on the feature information of the video clips and recommended feature information of the segments. Recommended feature information of a particular segment may be obtained by processing feature information of a previously selected video clip through a trained recurrent neural network. Video edit information defining the video edit may be generated. The video edit may include the selected video clips as the segments of the video edit.

Claims (51)

1. A system that creates video edits, the system comprising:

a storage medium storing video information defining video content, the video content including video clips, the video clips including a first video clip, a second video clip, and a third video clip; and

one or more physical processors configured by machine-readable instructions to:

obtain feature information of the video clips, the feature information characterizing features of the video clips, the features of the video clips including visuals captured within the video clips, the feature information including first feature information of the first video clip, second feature information of the second video clip, and third feature information of the third video clip, the feature information of the video clips including features values that characterize the visuals captured within the video clips;

select the first video clip as a first segment of a video edit of the video content;

iteratively select at least some of the video clips as other segments of the video edit based on the feature information and recommended feature information, the recommended feature information for next segment of the video edit determined based on processing, through a recurrent neural network, the feature information of a video clip previously selected for inclusion in the video edit as an adjacent segment of the next segment, the recommended feature information characterizing recommended features of a video clip to be selected for inclusion in the video edit as the next segment, the recommended features including recommended visual for the next segment, the recurrent neural network using the feature information from multiple ones of prior video clip selections to output the recommended feature information, wherein iterative selection of the at least some of the video clips as the other segments of the video edit includes:

processing the first feature information through the recurrent neural network, the recurrent neural network outputting first recommended feature information characterizing first recommended visual for a second segment of the video edit based on the first feature information characterizing first visual captured within the first video clip, the second segment adjacent to the first segment in the video edit;

selecting the second video clip as the second segment of the video edit based on a match between the first recommended feature information characterizing the first recommended visual for the second segment and the second feature information characterizing second visual captured within the second video clip;

processing the second feature information through the recurrent neural network, the recurrent neural network outputting second recommended feature information for a third segment of the video edit based on the first feature information and the second feature information, the third segment adjacent to the second segment in the video edit; and

selecting the third video clip as the third segment of the video edit based on a match between the second recommended feature information and the third feature information; and

generate video edit information, the video edit information defining the video edit of the video content, the video edit having a progress length, the video edit including the selected video clips as the segments of the video edit, the selected video clips including the first video clip as the first segment, the second video clip as the second segment, and the third video clip as the third segment of the video edit.

2. The system of claim 1 , wherein the first segment precedes the second segment in the progress length of the video edit.

3. The system of claim 1 , wherein the second segment precedes the first segment in the progress length of the video edit.

4. The system of claim 1 , wherein the match between the first recommended feature information and the second feature information includes the second feature information including same feature values as the first recommended feature information.

5. The system of claim 1 , wherein the match between the first recommended feature information and the second feature information includes feature values of the second feature information being closer to feature values of the first recommended feature information than feature values of the third feature information.

6. The system of claim 1 , wherein the one or more physical processors are further configured by the machine-readable instructions to select another video clip as the first segment of the video edit based on the first recommended feature information for the second segment of the video edit not matching any of the feature information of the video clips.

7. The system of claim 1 , wherein the iterative selection of the at least some of the video clips as the other segments of the video edit ends based on the recommended feature information outputted by the recurrent neural network based on processing of the feature information of the video clip previously selected for inclusion in the video edit recommending ending the video edit with the video clip previously selected for inclusion in the video edit.

8. The system of claim 1 , wherein the feature information of the video clips further includes feature values that characterizes effects used in the video clips, motion of one or more image capture devices that captured the video clips, and audio of the video clips.

9. The system of claim 8 , wherein the one or more physical processors are further configured by the machine-readable instructions to edit the second video clip based on the first recommended feature information to chance one or more features of the second video clip and generate a modified second video clip that is closer to the first recommended feature information than the second video clip, the edited second video clip being characterized by edited second feature information different from the second feature information, wherein the features values of the edited second feature information are closer to the feature values of the first recommended feature information than the feature values of the second feature information.

10. A method for creating video edits, the method performed by a computer system including one or more processors and storage medium storing video information defining video content, the video content including video clips, the video clips including a first video clip, a second video clip, and a third video clip, the method comprising:

obtaining, by the computing system, feature information of the video clips, the feature information characterizing features of the video clips, the features of the video clips including visuals captured within the video clips, the feature information including first feature information of the first video clip, second feature information of the second video clip, and third feature information of the third video clip, the feature information of the video clips including features values that characterize the visuals captured within the video clips;

selecting, by the computing system, the first video clip as a first segment of a video edit of the video content;

iteratively selecting, by the computing system, at least some of the video clips as other segments of the video edit based on the feature information and recommended feature information, the recommended feature information for next segment of the video edit determined based on processing, through a recurrent neural network, the feature information of a video clip previously selected for inclusion in the video edit as an adjacent segment of the next segment, the recommended feature information characterizing recommended features of a video clip to be selected for inclusion in the video edit as the next segment, the recommended features including recommended visual for the next segment, the recurrent neural network using the feature information from multiple ones of prior video clip selections to output the recommended feature information, wherein iterative selection of the at least some of the video clips as the other segments of the video edit includes:

processing the first feature information through the recurrent neural network, the recurrent neural network outputting first recommended feature information characterizing first recommended visual for a second segment of the video edit based on the first feature information characterizing first visual captured within the first video clip, the second segment adjacent to the first segment in the video edit;

selecting the second video clip as the second segment of the video edit based on a match between the first recommended feature information characterizing the first recommended visual for the second segment and the second feature information characterizing second visual captured within the second video clip;

processing the second feature information through the recurrent neural network, the recurrent neural network outputting second recommended feature information for a third segment of the video edit based on the first feature information and the second feature information, the third segment adjacent to the second segment in the video edit; and

selecting the third video clip as the third segment of the video edit based on a match between the second recommended feature information and the third feature information; and

generating, by the computing system, video edit information, the video edit information defining the video edit of the video content, the video edit having a progress length, the video edit including the selected video clips as the segments of the video edit, the selected video clips including the first video clip as the first segment, the second video clip as the second segment, and the third video clip as the third segment of the video edit.

11. The method of claim 10 , wherein the first segment precedes the second segment in the progress length of the video edit.

12. The method of claim 10 , wherein the second segment precedes the first segment in the progress length of the video edit.

13. The method of claim 10 , wherein the match between the first recommended feature information and the second feature information includes the second feature information including same feature values as the first recommended feature information.

14. The method of claim 10 , wherein the match between the first recommended feature information and the second feature information includes feature values of the second feature information being closer to feature values of the first recommended feature information than feature values of the third feature information.

15. The method of claim 10 , further comprising selecting, by the computing system, another video clip as the first segment of the video edit based on the first recommended feature information for the second segment of the video edit not matching any of the feature information of the video clips.

16. The method of claim 10 , wherein the iterative selection of the at least some of the video clips as the other segments of the video edit ends based on the recommended feature information outputted by the recurrent neural network based on processing of the feature information of the video clip previously selected for inclusion in the video edit recommending ending the video edit with the video clip previously selected for inclusion in the video edit.

17. The method of claim 10 , wherein the feature information of the video clips further includes feature values that characterizes effects used in the video clips, motion of one or more image capture devices that captured the video clips, and audio of the video clips.

18. The method of claim 17 , further comprising editing, by the computing system, the second video clip based on the first recommended feature information to change one or more features of the second video clip and generate a modified second video clip that is closer to the first recommended feature information than the second video clip, the edited second video clip being characterized by edited second feature information different from the second feature information, wherein the features values of the edited second feature information are closer to the feature values of the first recommended feature information than the feature values of the second feature information.

19. A system that creates video edits, the system comprising:

a storage medium storing video information defining video content, the video content including video clips, the video clips including a first video clip, a second video clip, and a third video clip; and

one or more physical processors configured by machine-readable instructions to:

obtain feature information of the video clips, the feature information characterizing features of the video clips, the features of the video clips including visuals captured within the video clips, the feature information including first feature information of the first video clip, second feature information of the second video clip, and third feature information of the third video clip, the feature information of the video clips including features values that characterize the visuals captured within the video clips;

select the first video clip as a first segment of a video edit of the video content;

iteratively select at least some of the video clips as other segments of the video edit based on the feature information and recommended feature information, the recommended feature information for next segment of the video edit determined based on processing, through a recurrent neural network, the feature information of a video clip previously selected for inclusion in the video edit as an adjacent segment of the next segment, the recommended feature information characterizing recommended features of a video clip to be selected for inclusion in the video edit as the next segment, the recommended features including recommended visual for the next segment, the recurrent neural network using the feature information from multiple ones of prior video clip selections to output the recommended feature information, wherein iterative selection of the at least some of the video clips as the other segments of the video edit includes:

processing the first feature information through the recurrent neural network, the recurrent neural network outputting first recommended feature information characterizing first recommended visual for a second segment of the video edit based on the first feature information characterizing first visual captured within the first video clip, the second segment adjacent to the first segment in the video edit;

selecting the second video clip as the second segment of the video edit based on a match between the first recommended feature information characterizing the first recommended visual for the second segment and the second feature information characterizing second visual captured within the second video clip;

processing the second feature information through the recurrent neural network, the recurrent neural network outputting second recommended feature information for a third segment of the video edit based on the first feature information and the second feature information, the third segment adjacent to the second segment in the video edit; and

selecting the third video clip as the third segment of the video edit based on a match between the second recommended feature information and the third feature information;

wherein:

the iterative selection of the at least some of the video clips as the other segments of the video edit ends based on the recommended feature information outputted by the recurrent neural network based on processing of the feature information of the video clip previously selected for inclusion in the video edit recommending ending the video edit with the video clip previously selected for inclusion in the video edit; and

the match between the first recommended feature information and the second feature information includes the second feature information including same feature values as the first recommended feature information or feature values of the second feature information being closer to feature values of the first recommended feature information than feature values of the third feature information; and

generate video edit information, the video edit information defining the video edit of the video content, the video edit having a progress length, the video edit including the selected video clips as the segments of the video edit, the selected video clips including the first video clip as the first segment, the second video clip as the second segment, and the third video clip as the third segment of the video edit.

20. The system of claim 19 , wherein the one or more physical processors are further configured by the machine-readable instructions to edit the second video clip based on the first recommended feature information to change one or more features of the second video clip and generate a modified second video clip that is closer to the first recommended feature information than the second video clip, the edited second video clip being characterized by edited second feature information different from the second feature information, wherein the features values of the edited second feature information are closer to the feature values of the first recommended feature information than the feature values of the second feature information.

Assignments (3)
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
SECURITY INTEREST Recorded Sep 5, 2018
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 047016/0417 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 13, 2018
From: WILSON, SAMUEL
To: GOPRO, INC.
Reel/Frame 045538/0588 →
Cited By (3)
US 12,579,596 US 12,592,263 US 12,658,211