IP Library Granted Patent US 10,339,975
Granted Patent B2
US 10,339,975 · App. 15/626,931 · Granted Jul 2, 2019

Voice-based video tagging

Inventors: Nick Hodulik (San Francisco, CA); Jonathan Taylor (San Francisco, CA); Brian Leong (San Jose, CA); Bettina Briz (Palo Alto, CA); Lisa Boghosian (San Francisco, CA); Mike Knott (San Mateo, CA)
Assignee: GoPro, Inc.
G11B27/10G06K9/00718G06K9/00751G10L15/063G10L15/22G10L25/57G11B27/031G11B27/13G11B27/28G11B27/3081G11B27/34H04N5/77H04N5/772H04N5/91H04N9/8205G06K2009/00738G10L25/54G10L2015/0631G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,339,975
App. No.
15/626,931
Granted
Jul 2, 2019
Kind
B2
Abstract

Video and corresponding metadata is accessed. Events of interest within the video are identified based on the corresponding metadata, and best scenes are identified based on the identified events of interest. A video summary can be generated including one or more of the identified best scenes. The video summary can be generated using a video summary template with slots corresponding to video clips selected from among sets of candidate video clips. Best scenes can also be identified by receiving an indication of an event of interest within video from a user during the capture of the video. Metadata patterns representing activities identified within video clips can be identified within other videos, which can subsequently be associated with the identified activities.

Claims (42)

1. A method for identifying an event of interest in a video, the method performed by a camera including one or more processors, the method comprising:

accessing, by the camera, a captured speech pattern, the captured speech pattern captured from a user at a moment during capture of the video;

matching, by the camera, the captured speech pattern to a given stored speech pattern of multiple stored speech patterns, the multiple stored speech patterns corresponding to a command for identifying the event of interest within the video, individual ones of the multiple stored speech patterns stored based on a number of times the individual ones of the multiple stored speech patterns are captured by the camera from a user while the camera is operating in a training mode, wherein the individual ones of the multiple stored speech patterns correspond to an identification of the event of interest as occurring before, during, or after the moment; and

in response to matching the captured speech pattern to the given stored speech pattern, storing, by the camera, event of interest information associated with the video, the event of interest information identifying an event moment during the capture of the video at which the event of interest occurs, the event moment being determined to occur before, during, or after the moment based on the matching of the captured speech pattern to the given stored speech pattern.

2. The method of claim 1 , further comprising:

identifying a portion of the video as a video clip associated with the event of interest based on the event of interest information, the video clip comprising a first amount of the video occurring before the event moment and a second amount of the video occurring after the event moment, the first amount and the second

storing clip information indicating the association of the video clip with the event of interest and the portion of the video included in the video clip.

3. The method of claim 2 , further comprising:

receiving a request to generate a video summary; and

generating the video summary in response to the request, the video summary comprising the video clip associated with the event of interest.

4. The method of claim 1 , wherein the event of interest corresponds to an activity type, wherein the given stored speech pattern identifies the activity type, and wherein storing the event of interest information comprises storing an indication of the activity type in metadata associated with the video.

5. The method of claim 1 , wherein the given stored speech pattern is specific to the user.

6. The method of claim 1 , wherein the given stored speech pattern includes a speech pattern of a spoken word.

7. The method of claim 1 , wherein the given stored speech pattern includes a speech pattern of a spoken phrase.

8. A system for identifying an event of interest in a video, the system comprising:

one or more processors configured by instructions to:

access a captured speech pattern, the captured speech pattern captured from a user at a moment during capture of the video;

match the captured speech pattern to a given stored speech pattern of multiple stored speech patterns, the multiple stored speech patterns corresponding to a command for identifying the event of interest within the video, individual ones of the multiple stored speech patterns stored based on a number of times the individual ones of the multiple stored speech patterns are captured by the camera from a user while the camera is operating in a training mode, wherein the individual ones of the multiple stored speech patterns correspond to an identification of the event of interest as occurring before, during, or after the moment; and

in response a match of the captured speech pattern to the given stored speech pattern, store event of interest information associated with the video, the event of interest information identifying an event moment during the capture of the video at which the event of interest occurs, the event moment being determined to occur before, during, or after the moment based on the matching of the captured speech pattern to the given stored speech pattern.

9. The system of claim 8 , wherein the one or more processors are further configured to:

identify a portion of the video as a video clip associated with the event of interest based on the event of interest information, the video clip comprising a first amount of the video occurring before the event moment and a second amount of the video occurring after the event moment, the first amount and the second amount being determined based on the matching of the captured speech pattern to the given stored speech pattern; and

store clip information indicating the association of the video clip with the event of interest and the portion of the video included in the video clip.

10. The system of claim 9 , wherein the one or more processors are further configured to:

receive a request to generate a video summary; and

generate the video summary in response to the request, the video summary comprising the video clip associated with the event of interest.

11. The system of claim 8 , wherein the event of interest corresponds to an activity type, wherein the given stored speech pattern identifies the activity type, and wherein the event of interest information is stored such that an indication of the activity type is stored in metadata associated with the video.

12. The system of claim 8 , wherein the given stored speech pattern is specific to the user.

13. The system of claim 8 , wherein the given stored speech pattern includes a speech pattern of a spoken word.

14. The system of claim 8 , wherein the given stored speech pattern includes a speech pattern of a spoken phrase.

15. A non-transitory computer-readable storage medium storing instructions for identifying an event of interest in a video, the instructions, when executed, causing one or more processors to:

access a captured speech pattern, the captured speech pattern captured from a user at a moment during capture of the video;

match the captured speech pattern to a given stored speech pattern of multiple stored speech patterns, the multiple stored speech patterns corresponding to a command for identifying the event of interest within the video, individual ones of the multiple stored speech patterns stored based on a number of times the individual ones of the multiple stored speech patterns are captured by the camera from a user while the camera is operating in a training mode, wherein the individual ones of the multiple stored speech patterns correspond to an identification of the event of interest as occurring before, during, or after the moment; and

in response a match of the captured speech pattern to the given stored speech pattern, store event of interest information associated with the video, the event of interest information identifying an event moment during the capture of the video at which the event of interest occurs, the event moment being determined to occur before, during, or after the moment based on the matching of the captured speech pattern to the given stored speech pattern.

16. The computer-readable storage medium of claim 15 , wherein the instructions, when executed, further cause the one or more processors to:

identify a portion of the video as a video clip associated with the event of interest based on the event of interest information, the video clip comprising a first amount of the video occurring before the event moment and a second amount of the video occurring after the event moment, the first amount and the second amount being determined based on the matching of the captured speech pattern to the given stored speech pattern; and

store clip information indicating the association of the video clip with the event of interest and the portion of the video included in the video clip.

17. The computer-readable storage medium of claim 16 , wherein instructions, when executed, further cause the one or more processors to:

receive a request to generate a video summary; and

generate the video summary in response to the request, the video summary comprising the video clip associated with the event of interest.

18. The computer-readable storage medium of claim 15 , wherein the event of interest corresponds to an activity type, wherein the given stored speech pattern identifies the activity type, and wherein the event of interest information is stored such that an indication of the activity type is stored in metadata associated with the video.

19. The computer-readable storage medium of claim 15 , wherein the given stored speech pattern is specific to the user.

20. The computer-readable storage medium of claim 15 , wherein the given stored speech pattern includes a speech pattern of a spoken word or a spoken phrase.

Assignments (6)
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: FARALLON CAPITAL MANAGEMENT, L.L.C., AS AGENT
Reel/Frame 072340/0676 →
SECURITY INTEREST Recorded Aug 4, 2025
From: GOPRO, INC.
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 072358/0001 →
RELEASE OF PATENT SECURITY INTEREST Recorded Jan 25, 2021
From: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
To: GOPRO, INC.
Reel/Frame 055106/0434 →
SECURITY INTEREST Recorded Oct 19, 2020
From: GOPRO, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 054113/0594 →
SECURITY INTEREST Recorded Nov 20, 2017
From: GOPRO, INC.; GOPRO CARE, INC.; GOPRO CARE SERVICES, INC.; WOODMAN LABS CAYMAN, INC.
To: JPMORGAN CHASE BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 044180/0430 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 19, 2017
From: HODULIK, NICK; TAYLOR, JONATHAN; LEONG, BRIAN; BRIZ, BETTINA; BOGHOSIAN, LISA; KNOTT, MIKE
To: GOPRO, INC.
Reel/Frame 042750/0405 →
Continuity (5)
Continuation 14530245 · Oct 31, 2014
Continuation 14513151 · Oct 13, 2014
Provisional Application 62039849 · Aug 20, 2014
Provisional Application 62028254 · Jul 23, 2014
Related Publication 20170287523A1 · Oct 5, 2017
Cited By (1)
US 12,243,307