IP Library › Granted Patent US 11,393,209
Granted Patent B2
US 11,393,209 · App. 16/985,342 · Granted Jul 19, 2022

Generating a video segment of an action from a video

Inventors: Sudheendra Vijayanarasimhan (Pasadena, CA); Alexis Bienvenu (Mountain View, CA); David Ross (San Jose, CA); Timothy Novikoff (Mountain View, CA); Arvind Balasubramanian (Fremont, CA)
Assignee: Google LLC
G06V20/47G06F3/0484G06N3/04G06V20/30G06V20/41G06V20/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,393,209
App. No.
16/985,342
Granted
Jul 19, 2022
Kind
B2
Abstract

A computer-implemented method includes receiving a video that includes multiple frames. The method further includes identifying a start time and an end time of each action in the video based on application of one or more of an audio classifier, an RGB classifier, and a motion classifier. The method further includes identifying video segments from the video that include frames between the start time and the end time for each action in the video. The method further includes generating a confidence score for each of the video segments based on a probability that a corresponding action corresponds to one or more of a set of predetermined actions. The method further includes selecting a subset of the video segments based on the confidence score for each of the video segments.

Claims (52)

1. A system comprising:

one or more processors; and

a memory coupled to the one or more processors, with instructions stored thereon that, when executed by the one or more processors, cause the one or more processors to perform operations comprising:

receiving a video that includes multiple frames;

identifying video segments from the video that each include a corresponding action, wherein each video segment of the video segments include two or more contiguous frames of the multiple frames of the video;

generating a confidence score for each of the video segments based on a probability that the corresponding action in the video segment corresponds to a particular action in a set of predetermined actions; and

selecting a subset of the video segments based on the confidence score for each of the video segments.

2. The system of claim 1 , wherein the operations further comprise:

generating a user profile for a user that includes a list of types of videos that the user previously indicated a preference for; and

determining a type for the video, wherein the confidence score is based on the probability that the corresponding action in the video segment corresponds to the type of video in the set of predetermined actions.

3. The system of claim 1 , wherein the operations further comprise:

identifying one or more people in the video segments;

determining that a user has a relationship with the one or more people in a first video segment; and

assigning a higher segment score to the first video segment than one or more other video segments that do not depict a person that is related to the user.

4. The system of claim 1 , wherein the video is a live video and the set of predetermined actions includes at least one of laughing or crying.

5. The system of claim 1 , wherein the operations further include identifying the action in the video based on using an audio classifier that analyzes an audio portion of the video to identify a start time and an end time of the action and to identify a type of action that is associated with audio.

6. The system of claim 1 , wherein the operations further include identifying the action in the video based on using an RGB classifier that employs a feature vector and clustering techniques to identify features that form a pattern associated with the action and a type of action.

7. The system of claim 1 , wherein the operations further include identifying the action in the video based on using a motion classifier that identifies an optical flow in the video that is associated with particular kinds of movement.

8. The system of claim 1 , wherein selecting the subset of the video segments based on the confidence score includes selecting the subset of the video segments that exceed a confidence score threshold and the operations further comprise:

generating a video clip that includes the subset of the video segments.

9. The system of claim 1 , wherein the operations further comprise:

generating a segment score to each of the video segments based on personalization information for a user, wherein selecting the subset of the video segments is further based on a combination of the confidence score and the segment score for each of the video segments exceeding a threshold animation score; and

generating a video clip that includes the subset of the video segments.

10. A method comprising:

receiving a video that includes multiple frames;

identifying video segments from the video that each include a corresponding action, wherein each video segment of the video segments include two or more contiguous frames of the multiple frames of the video;

generating a confidence score for each of the video segments based on a probability that the corresponding action in the video segment corresponds to a particular action in a set of predetermined actions; and

selecting a subset of the video segments based on the confidence score for each of the video segments.

11. The method of claim 10 , further comprising:

generating a user profile for a user that includes a list of types of videos that the user previously indicated a preference for; and

determining a type for the video, wherein the confidence score is based on the probability that the corresponding action in the video segment corresponds to the type of video in the set of predetermined actions.

12. The method of claim 10 , further comprising:

identifying one or more people in the video segments;

determining that a user has a relationship with the one or more people in a first video segment; and

assigning a higher segment score to the first video segment than one or more other video segments that do not depict a person that is related to the user.

13. The method of claim 10 , further comprising identifying the action in the video based on using an audio classifier that analyzes an audio portion of the video to identify a start time and an end time of the action and to identify a type of action that is associated with audio.

14. The method of claim 10 , further comprising identifying the action in the video based on using an RGB classifier that employs a feature vector and clustering techniques to identify features that form a pattern associated with the action and a type of action.

15. The method of claim 10 , further comprising identifying the action in the video based on using a motion classifier that identifies an optical flow in the video that is associated with particular kinds of movement.

16. A non-transitory computer readable medium with instructions that, when executed by one or more computers, cause the one or more computers to perform operations, the operations comprising:

receiving a video that includes multiple frames;

identifying video segments from the video that each include a corresponding action, wherein each video segment of the video segments include two or more contiguous frames of the multiple frames of the video;

generating a confidence score for each of the video segments based on a probability that the corresponding action in the video segment corresponds to a particular action in a set of predetermined actions; and

selecting a subset of the video segments based on the confidence score for each of the video segments.

17. The non-transitory computer readable medium of claim 16 , wherein the operations further comprise:

generating a user profile for a user that includes a list of types of videos that the user previously indicated a preference for; and

determining a type for the video, wherein the confidence score is based on the probability that the corresponding action in the video segment corresponds to the type of video in the set of predetermined actions.

18. The non-transitory computer readable medium of claim 16 , wherein the operations further comprise:

identifying one or more people in the video segments;

determining that a user has a relationship with the one or more people in a first video segment; and

assigning a higher segment score to the first video segment than one or more other video segments that do not depict a person that is related to the user.

19. The non-transitory computer readable medium of claim 16 , wherein the operations further comprise identifying the action in the video based on using an audio classifier that analyzes an audio portion of the video to identify a start time and an end time of the action and to identify a type of action that is associated with audio.

20. The non-transitory computer readable medium of claim 16 , wherein the operations further include identifying the action in the video based on using an RGB classifier that employs a feature vector and clustering techniques to identify features that form a pattern associated with the action and a type of action.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2020
From: VIJAYANARASIMHAN, SUDHEENDRA; BIENVENU, ALEXIS; ROSS, DAVID; NOVIKOFF, TIMOTHY; BALASUBRAMANIAN, ARVIND
To: GOOGLE LLC
Reel/Frame 053404/0158 →
Continuity (2)
Continuation 15782789 · Oct 12, 2017
Related Publication 20200364464A1 · Nov 19, 2020
Cited By (1)
US 12,555,375