IP Library Granted Patent US 8,316,302
Granted Patent B2
US 8,316,302 · App. 11/747,584 · Granted Nov 20, 2012

Method and apparatus for annotating video content with metadata generated using speech recognition technology

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,316,302
App. No.
11/747,584
Granted
Nov 20, 2012
Kind
B2
Abstract

A method and apparatus is provided for annotating video content with metadata generated using speech recognition technology. The method begins by rendering video content on a display device. A segment of speech is received from a user such that the speech segment annotates a portion of the video content currently being rendered. The speech segment is converted to a text-segment and the text-segment is associated with the rendered portion of the video content. The text segment is stored in a selectively retrievable manner so that it is associated with the rendered portion of the video content.

Claims (32)

1. At least one non-transitory computer-readable medium encoded with instructions which, when executed by a processor, performs a method including:

rendering video content on a display device;

receiving a segment of speech from a user such that the speech segment annotates a portion of the video content currently being rendered;

converting the speech segment to a text-segment;

associating the text-segment with the portion of the video content; and

storing in a selectively retrievable manner the text-segment so that it is associated with the rendered portion of the video content.

2. The computer-readable medium of claim 1 further comprising receiving a signal from the user selecting an operational state before receiving the segment of speech.

3. The computer-readable medium of claim 2 wherein the operational state is selected from the group consisting of an annotate state, a narrate state, a commentary state, an analyze state and a review/edit state.

4. The computer-readable medium of claim 2 wherein the user-selectable operational states are presented as a GUI on the display device.

5. The computer-readable medium of claim 4 wherein the GUI is superimposed over the video content being rendered.

6. The computer-readable medium of claim 1 wherein the video content is rendered by a set top box.

7. The computer-readable medium of claim 1 wherein the video content is rendered by a DVR.

8. The computer-readable medium of claim 6 wherein the set top box receives the video content from a video camera.

9. The computer-readable medium of claim 7 wherein the DVR receives the video content from a video camera.

10. The computer-readable medium of claim 1 further comprising presenting the user with a plurality of different user-selectable operational states defining a mode in which the speech request is to be received.

11. An apparatus for rendering a video program, comprising:

a computer-readable storage medium; and

a processor responsive to the computer-readable storage medium and to a software program, the software program, when loaded into the processor, operative to:

render the video content on a display device;

receive a segment of speech from a user such that the speech segment annotates a portion of the video content currently being rendered;

convert the speech segment to a text-segment;

associate the text-segment with the portion of the video content; and

store in a selectively retrievable manner the text-segment so that it is associated with the rendered portion of the video content.

12. The apparatus of claim 11 further comprising:

a receiver/tuner for receiving programming content over a broadband communications system; and

a decoder for decoding programming content provided by the receiver/tuner.

13. The apparatus of claim 11 wherein the processor is further configured to receive a signal from the user selecting an operational state before receiving the segment of speech.

14. The apparatus of claim 13 wherein the operational state is selected from the group consisting of an annotate state, a narrate state, a commentary state, an analyze state and a review/edit state.

15. The apparatus of claim 11 wherein the processor is further configured to receive the video content from a video camera.

16. The apparatus of claim 11 wherein the processor is further configured to present the user with a plurality of different user-selectable operational states defining a mode in which the speech request is to be received.

17. The apparatus of claim 16 wherein the user-selectable operational states are presented as a GUI on the display device.

18. The apparatus of claim 17 wherein the GUI is superimposed over the video content being rendered.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2014
From: MOTOROLA MOBILITY LLC
To: GOOGLE TECHNOLOGY HOLDINGS LLC
Reel/Frame 034244/0014 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2013
From: GENERAL INSTRUMENT CORPORATION
To: GENERAL INSTRUMENT HOLDINGS, INC.
Reel/Frame 030764/0575 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 9, 2013
From: GENERAL INSTRUMENT HOLDINGS, INC.
To: MOTOROLA MOBILITY LLC
Reel/Frame 030866/0113 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2007
From: MCKOEN, KEVIN M.; GROSSMAN, MICHAEL A.
To: GENERAL INSTRUMENT CORPORATION
Reel/Frame 019282/0756 →