IP Library › Granted Patent US 10,740,620
Granted Patent B2
US 10,740,620 · App. 15/782,789 · Granted Aug 11, 2020

Generating a video segment of an action from a video

Inventors: Sudheendra Vijayanarasimhan (Pasadena, CA); Alexis Bienvenu (Mountain View, CA); David Ross (San Jose, CA); Timothy Novikoff (Mountain View, CA); Arvind Balasubramanian (Fremont, CA)
Assignee: Google LLC
G06K9/00751G06F3/0484G06K9/00677G06K9/00718G06N3/04G06K2009/00738
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,740,620
App. No.
15/782,789
Granted
Aug 11, 2020
Kind
B2
Abstract

A computer-implemented method includes receiving a video that includes multiple frames. The method further includes identifying a start time and an end time of each action in the video based on application of one or more of an audio classifier, an RGB classifier, and a motion classifier. The method further includes identifying video segments from the video that include frames between the start time and the end time for each action in the video. The method further includes generating a confidence score for each of the video segments based on a probability that a corresponding action corresponds to one or more of a set of predetermined actions. The method further includes selecting a subset of the video segments based on the confidence score for each of the video segments.

Claims (60)

1. A system comprising: one or more processors; and

a memory with instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising: receiving a video that includes multiple frames;

identifying a start time and an end time of each action in the video based on application of one or more of an audio classifier, an RGB classifier, or a motion classifier;

identifying video segments from the video that include frames between the start time and the end time for each action in the video; generating a confidence score for each of the video segments based on a probability that a corresponding action corresponds to one or more of a set of predetermined actions; and selecting a subset of the video segments based on the confidence score for each of the video segments.

2. The system of claim 1 , wherein the operations further comprise:

generating a video clip from the subset of the video segments, wherein the video clip is displayed in association with the video; and

generating graphical data to display a user interface that includes an option to add an automated effect to the video clip.

3. The system of claim 1 , wherein the set of predetermined actions is associated with a live video and the operations further comprise:

generating a video clip that includes the subset of the video segments, wherein the video clip is a summary of the actions that occurred during the live video.

4. The system of claim 1 , wherein the operations further comprise:

identifying a type of video, wherein the set of predetermined actions correspond to the type of video;

generating a video clip from the subset of the video segments that includes actions that correspond to the type of video; and

adding music to the video clip that corresponds to the type of video.

5. The system of claim 1 , wherein the operations further comprise:

generating graphical data to display a user interface that identifies a time within the video that corresponds to the video segment and a type of action that is performed within the video segment.

6. The system of claim 1 , wherein the operations further comprise:

identifying a first person and a second person in one or more of the video segments, wherein the subset of the video segments is further based on the one or more of the video segments that include the first person;

generating a first video clip that includes the subset of the video segments; and

providing the first video clip to a user as a personalized video clip that includes the first person and an option to generate a second video clip that includes the second person.

7. The system of claim 1 , wherein the operations further comprise:

identifying a type of action for each action in the video;

determining that there are more than one actions in the video for a particular type of action; and

in response to determining that there are more than one actions in the video, selecting a particular video segment of the video segments that correspond to the more than one action, wherein the particular video segment is selected based on the confidence score for the video segment being greater than confidence scores for other video segments that correspond to the more than one actions.

8. The system of claim 1 , wherein the operations further comprise:

generating a machine learning model based on a set of videos where users identified corresponding start times and end times for actions within each video in the set of videos; and

generating one or more of the RGB classifier, the audio classifier, or-and the motion classifier based on the machine learning model.

9. The system of claim 1 , wherein generating the confidence score for each of the video segments is based on applying a mixture of experts model of machine learning.

10. A computer-implemented method comprising:

receiving a video that includes multiple frames;

identifying a start time and an end time of an action in the video based on application of an audio classifier, an RGB classifier, or a motion classifier;

identifying a video segment from the video that includes frames between the start time and the end time for the action in the video;

generating a confidence score for the video segment based on a probability that a corresponding action corresponds to one or more of a set of predetermined actions; and

generating a video clip of the video that includes the video segment.

11. The method of claim 10 , further comprising:

generating graphical data to display a user interface that includes the video clip in association with the video.

12. The method of claim 10 , further comprising:

identifying a type of action and, upon receiving content from a user, an identity of the user in the video segment; and

generating a video clip that includes the video segment with an identification of the type of action and the identity of the user.

13. The method of claim 10 , wherein the set of predetermined actions is associated with a live video and the video clip is a summary of the action that occurred during the live video.

14. The method of claim 10 , further comprising:

receiving a search request from a user for videos that include a particular action;

determining that the particular action matches the action in the video; and

providing the user with the video clip.

15. The method of claim 10 , wherein generating the confidence score for each of the video segments is based on applying a mixture of experts model of machine learning.

16. A non-transitory computer readable medium with instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising: receiving a video that includes multiple frames;

identifying a start time and an end time of each action in the video based on application of one or more of an audio classifier, an RGB classifier, or a motion classifier;

identifying video segments from the video that include frames between the start time and the end time for each action in the video;

generating a confidence score for each of the video segments based on a probability that a corresponding action corresponds to one or more of a set of predetermined actions; and

selecting a subset of the video segments based on the confidence score for each of the video segments.

17. The computer-readable medium of claim 16 , wherein the operations further comprise:

generating a video clip from the subset of the video segments, wherein the video clip is displayed in association with the video; and

generating graphical data to display a user interface that includes an option to add an automated effect to the video clip.

18. The computer-readable medium of claim 16 , wherein the set of predetermined actions is associated with a live video and the operations further comprise:

generating a video clip that includes the subset of the video segments, wherein the video clip is a summary of the actions that occurred during the live video.

19. The computer-readable medium of claim 16 , wherein the operations further comprise:

identifying a type of video, wherein the set of predetermined actions correspond to the type of video;

generating a video clip from the subset of the video segments that includes actions that correspond to the type of video; and

adding music to the video clip that corresponds to the type of video.

20. The computer-readable medium of claim 16 , wherein the operations further comprise:

generating graphical data to display a user interface that identifies a time within the video that corresponds to the video segment and a type of action that is performed within the video segment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2017
From: VIJAYANARASIMHAN, SUDHEENDRA; BIENVENU, ALEXIS; ROSS, DAVID; NOVIKOFF, TIMOTHY; BALASUBRAMANIAN, ARVIND
To: GOOGLE LLC
Reel/Frame 043855/0813 →
Continuity (1)
Related Publication 20190114487A1 · Apr 18, 2019
Cited By (3)
US 12,210,560 US 12,288,342 US 12,555,375