IP Library Granted Patent US 10,482,168
Granted Patent B2
US 10,482,168 · App. 15/472,033 · Granted Nov 19, 2019

Method and apparatus for annotating video content with metadata generated using speech recognition technology

Inventors: Kevin M. McKoen (San Diego, CA); Michael A. Grossman (San Diego, CA)
Assignee: Google Technology Holdings LLC
G06F17/241G06F16/70G06F16/78G06F16/7844G10L15/26G10L15/265
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,482,168
App. No.
15/472,033
Granted
Nov 19, 2019
Kind
B2
Abstract

A method and apparatus is provided for annotating video content with metadata generated using speech recognition technology. The method begins by rendering video content on a display device. A segment of speech is received from a user such that the speech segment annotates a portion of the video content currently being rendered. The speech segment is converted to a text-segment and the text-segment is associated with the rendered portion of the video content. The text segment is stored in a selectively retrievable manner so that it is associated with the rendered portion of the video content.

Claims (68)

1. A computer-implemented method comprising:

while rendering video content on a display device:

receiving, at a processor in communication with the display device, audio data from a user, the audio data annotating a particular portion of the video content;

converting, by the processor, the audio data into text; and

generating, by the processor, annotation metadata associating the text with the particular portion of the video content;

after rendering the video content on the display device, receiving, at the processor, a search query comprising one or more terms present in the text; and

providing, by the processor, the particular portion of the video content for output based on the search query and the annotation metadata associating the text with the particular portion of the video content.

2. The method of claim 1 , further comprising, prior to receiving the audio data from the user while rendering the video content on the display device:

superimposing, by the processor, a user interface over the video content rendered on the display device, the user interface comprising an annotation selectable control to enter an annotation mode;

in response to receiving a selection indication indicating selection of the annotation selectable control, entering, by the processor, the annotation mode.

3. The method of claim 2 , wherein the user interface further comprises an analyze selectable control to enter a mode to analyze previous annotations and an edit selectable control to enter a mode to edit previous annotations.

4. The method of claim 1 , further comprising:

while the particular portion of the video content is being output, and while an annotation mode is activated, providing, by the processor, for output, an interface that includes a request for information related to the particular portion of the video content;

receiving, at the processor, the information related to the particular portion of the video content; and

including, by the processor, the information related to the particular portion of the video content in the text.

5. The method of claim 1 , further comprising:

while the particular portion of the video content is being output, and while an annotation mode is activated, providing, by the processor, for output, a selectable control to activate a voice recognition mode; and

in response to receiving data indicating a selection of the selectable control to enter the annotation mode, entering, by the processor, the voice recognition mode.

6. The method of claim 1 , further comprising, after the annotation mode is deactivated, providing, by the processor, for output, the particular portion of the video content overlaid with the text.

7. The method of claim 6 , wherein providing, for output, the particular portion of the video content overlaid with the text comprises:

providing, for output, a selectable control to edit the text; and

the method further comprises:

receiving data indicating a selection of the selectable control to edit the text; and

providing, for output, a user interface to edit the text.

8. A system comprising:

one or more computers and one or more storage devices storing instructions that are operable, when executed by the one or more computers, to cause the one or more computers to perform operations comprising:

while rendering video content on a display device:

receiving audio data from a user, the audio data annotating a particular portion of the video content;

converting the audio data into text; and

generating annotation metadata associating the text with the particular portion of the video content;

after rendering the video content on the display device, receiving a search query comprising one or more terms present in the text; and

providing the particular portion of the video content for output based on the search query and the annotation metadata associating the text with the particular portion of the video content.

9. The system of claim 8 , wherein the operations further comprise, prior to receiving the audio data from the user while rendering the video content on the display device:

superimposing a user interface over the video content rendered on the display device, the user interface comprising an annotation selectable control to enter an annotation mode; and

in response to receiving a selection indication indicating selection of the annotation selectable control, entering the annotation mode.

10. The system of claim 9 , wherein the user interface further comprises an analyze selectable control to enter a mode to analyze previous annotations and an edit selectable control to enter a mode to edit previous annotations.

11. The system of claim 8 , wherein the operations further comprise:

while the particular portion of the video content is being output, and while an annotation mode is activated, providing, for output, an interface that includes a request for information related to the particular portion of the video content;

receiving the information related to the particular portion of the video content; and

including the information related to the particular portion of the video content in the text.

12. The system of claim 8 , wherein the operations further comprise:

while the particular portion of the video content is being output, and while an annotation mode is activated, providing, for output, a selectable control to activate a voice recognition mode; and

in response to receiving data indicating a selection of the selectable control to enter the annotation mode, entering the voice recognition mode.

13. The system of claim 8 , wherein the operations further comprise, after the annotation mode is deactivated, providing, for output, the particular portion of the video content overlaid with the text.

14. The system of claim 13 , wherein providing, for output, the particular portion of the video content overlaid with the text further comprises:

providing, for output, a selectable control to edit the text; and

wherein the operations further comprise:

receiving data indicating a selection of the selectable control to edit the text; and

providing, for output, a user interface to edit the text.

15. A non-transitory computer-readable medium storing software comprising instructions executable by one or more computers which, upon such execution, cause the one or more computers to perform operations comprising:

while rendering video content on a display device:

receiving audio data from a user, the audio data annotating a particular portion of the video content;

converting the audio data into text; and

generating annotation metadata associating the text with the particular portion of the video content;

after rendering the video content on the display device, receiving a search query comprising one or more terms present in the text; and

providing the particular portion of the video content for output based on the search query and the annotation metadata associating the text with the particular portion of the video content.

16. The medium of claim 15 , wherein the operations further comprise, prior to receiving the audio data from the user while rendering the video content on the display device:

superimposing a user interface over the video content rendered on the display device, the user interface comprising an annotation selectable control to enter an annotation mode; and

in response to receiving a selection indication indicating selection of the annotation selectable control, entering the annotation mode.

17. The medium of claim 16 , wherein the user interface further comprises an analyze selectable control to enter a mode to analyze previous annotations and an edit selectable control to enter a mode to edit previous annotations.

18. The medium of claim 15 , wherein the operations further comprise:

while the particular portion of the video content is being output, and while the annotation mode is activated, providing, for output, an interface that includes a request for information related to the particular portion of the video content;

receiving the information related to the particular portion of the video content; and

including the information related to the particular portion of the video content in the text.

19. The medium of claim 15 , wherein the operations further comprise:

while the particular portion of the video content is being output, and while the annotation mode is activated, providing, for output, a selectable control to activate a voice recognition mode; and

in response to receiving data indicating a selection of the selectable control to enter the annotation mode, entering the voice recognition mode.

20. The system of claim 15 , wherein the operations further comprise, after the annotation mode is deactivated, providing, for output, the particular portion of the video content overlaid with the text.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2017
From: MCKOEN, KEVIN M.; GROSSMAN, MICHAEL A.
To: GENERAL INSTRUMENT CORPORATION
Reel/Frame 041784/0271 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2017
From: GENERAL INSTRUMENT CORPORATION
To: GENERAL INSTRUMENT HOLDINGS, INC.
Reel/Frame 042110/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2017
From: GENERAL INSTRUMENT HOLDINGS, INC.
To: MOTOROLA MOBILITY LLC
Reel/Frame 042110/0060 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2017
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 042110/0126 →
Continuity (4)
Continuation 14336063 · Jul 21, 2014
Continuation 13654327 · Oct 17, 2012
Continuation 11747584 · May 11, 2007
Related Publication 20170199856A1 · Jul 13, 2017
Cited By (1)
US 12,462,806