IP Library Granted Patent US 8,707,381
Granted Patent B2
US 8,707,381 · App. 12/886,769 · Granted Apr 22, 2014

Caption and/or metadata synchronization for replay of previously or simultaneously recorded live programs

Inventors: Richard T. Polumbus (Englewood, CO); Michael W. Homyack (Highlands Ranch, CO)
Assignee: Caption Colorado L.L.C.
H04N21/440236G10L15/22G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,707,381
App. No.
12/886,769
Granted
Apr 22, 2014
Kind
B2
Abstract

A synchronization process between captioning data and/or corresponding metatags and the associated media file parses the media file, correlates the caption information and/or metatags with segments of the media file, and provides a capability for textual search and selection of particular segments. A time-synchronized version of the captions is created that is synchronized to the moment that the speech is uttered in the recorded media. The caption data is leveraged to enable search engines to index not merely the title of a video, but the entirety of what was said during the video as well as any associated metatags relating to contents of the video. Further, because the entire media file is indexed, a search can request a particular scene or occurrence within the event recorded by the media file, and the exact moment within the media relevant to the search can be accessed and played for the requester.

Claims (51)

1. A method in a computer system for caption synchronization of media programs, comprising:

receiving, utilizing at least one processing unit, at least a portion of a media stream and at least a portion of a caption data stream, the media stream and the caption data stream corresponding to an event, the media stream comprising at least an audio component;

tuning, utilizing the least one processing unit, speech recognition functionality of an audio mining engine with the at least the portion of the caption data stream;

converting, utilizing the at least one processing unit, at least a portion of the audio component to text by executing the speech recognition functionality of the tuned audio mining engine on the audio component;

producing, utilizing the at least one processing unit, an audio mining transcript including timing information from the audio component and the text generated by the speech recognition functionality of the tuned audio mining engine;

aligning, utilizing the at least one processing unit, the audio mining transcript with the at least the portion of the caption data stream; and

generating a time synchronized caption data stream, utilizing the at least one processing unit, by applying the timing information from the audio mining transcript to the at least the portion of the caption stream based on the aligning.

2. The method of claim 1 , further comprising:

providing at least one search engine with access to the time synchronized caption data stream.

3. The method of claim 1 , further comprising:

the first output and the first input are combined as a first bi-directional port; and the second output comprises a second port different than the first port.

4. The method of claim 1 , further comprising:

searching the time synchronized caption data stream in response to a text search query; and

providing access to a point in the at least the portion of the media stream that corresponds to a point in the time synchronized caption data stream at least partially matching the text search query.

5. The method of claim 4 , wherein the caption data stream includes at least one metadata annotation and searching the time synchronized caption data stream comprises at least one of searching text of the time synchronized caption data stream and searching the at least one metadata annotation.

6. The method of claim 5 , further comprising:

indexing at least one of the text of the time synchronized caption data stream and the at least one metadata annotation.

7. The method of claim 4 , further comprising:

transmitting the point in the at least the portion of the media stream to a display device.

8. The method of claim 1 , wherein said aligning, utilizing the at least one processing unit, the audio mining transcript with the at least the portion of the caption data stream comprises:

dividing the at least the portion of the caption data stream into a plurality of discrete caption segments and the audio mining transcript into a plurality of overlapping transcript segments;

comparing at least one of the plurality of discrete caption segments to at least one of the plurality of overlapping transcript segments; and

associating the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments if the comparison indicates there is a strong correlation between the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments.

9. The method of claim 8 , wherein said comparing at least one of the plurality of discrete caption segments to at least one of the plurality of overlapping transcript segments comprises:

developing a language model of quasi-phonemes of the at least one of the plurality of discrete caption segments;

building an acoustic model of the at least one of the plurality of overlapping transcript segments; and

correlating the language model of quasi-phonemes of the at least one of the plurality of discrete caption segments with the acoustic model of the at least one of the plurality of overlapping transcript segments.

10. The method of claim 9 , wherein the comparison indicates there is the strong correlation between the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments when there is a strong correlation between the language model of quasi-phonemes of the at least one of the plurality of discrete caption segments and the acoustic model of the at least one of the plurality of overlapping transcript segments.

11. A system for caption synchronization of media programs, comprising:

at least one processing unit coupled to at least one storage media,

at least one communication component, coupled to the at least one processing unit, operable to receive at least a portion of a media stream and at least a portion of a caption data stream, the media stream and the caption data stream corresponding to an event, the media stream comprising at least an audio component, and

an audio mining engine, executable by the at least one processing unit, operable to produce an audio mining transcript including timing information wherein the at least one processing unit

tunes speech recognition functionality of the audio mining engine utilizing the at least the portion of the caption data stream,

converts at least a portion of the audio component to text by executing the speech recognition functionality of the tuned audio mining engine on the audio component,

aligns the audio mining transcript with the at least the portion of the caption data stream, and

generates a time synchronized caption data stream by applying the timing information to the at least the portion of the caption stream based on the aligning wherein the timing information is from the audio component and the text generated by the speech recognition functionality of the tuned audio mining engine.

12. The system of claim 11 , wherein the at least one communication component transmits at least one of the at least the portion of the media stream and the time synchronized caption data stream substantially in real time with receiving the at least the portion of the media stream.

13. The system of claim 11 , wherein the at least one processing unit provides at least one search engine with access to the time synchronized caption data stream via the at least one communication component.

14. The system of claim 11 , wherein the processing unit is operable to search the time synchronized caption data stream in response to a text search query received by the at least one communication component and provide access to a point in the at least the portion of the media stream that corresponds to a point in the time synchronized caption data stream at least partially matching the text search query.

15. The system of claim 14 , wherein the caption data stream includes at least one metadata annotation and the at least one processing unit searches the time synchronized caption data stream by searching at least one of text of the time synchronized caption data stream and the at least one metadata annotation.

16. The system of claim 15 , wherein the at least one processing unit indexes at least one of the text of the time synchronized caption data stream and the at least one metadata annotation.

17. The system of claim 14 , wherein the at least one processing unit transmits the point in the at least the portion of the media stream to a display device.

18. The system of claim 11 , wherein the at least one processing unit aligns the audio mining transcript with the at least a portion of the caption data stream by:

dividing the at least the portion of the caption data stream into a plurality of discrete caption segments and the audio mining transcript into a plurality of overlapping transcript segments;

comparing at least one of the plurality of discrete caption segments to at least one of the plurality of overlapping transcript segments; and

associating the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments if the comparison indicates there is a strong correlation between the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments.

19. The system of claim 18 , wherein said comparing at least one of the plurality of discrete caption segments to at least one of the plurality of overlapping transcript segments comprises:

developing a language model of quasi-phonemes of the at least one of the plurality of discrete caption segments;

building an acoustic model of the at least one of the plurality of overlapping transcript segments; and

correlating the language model of quasi-phonemes of the at least one of the plurality of discrete caption segments with the acoustic model of the at least one of the plurality of overlapping transcript segments.

20. The system of claim 19 , wherein the comparison indicates there is the strong correlation between the at least one of the plurality of discrete caption segments to the at least one of the plurality of overlapping transcript segments when there is a strong correlation between the language model of quasi-phonemes of the at least one of the plurality of discrete caption segments and the acoustic model of the at least one of the plurality of overlapping transcript segments.

Assignments (5)
SECURITY INTEREST Recorded Sep 14, 2021
From: VITAC CORPORATION
To: SILICON VALLEY BANK
Reel/Frame 057477/0671 →
RELEASE OF SECURITY INTEREST Recorded May 11, 2021
From: ABACUS FINANCE GROUP, LLC
To: VITAC CORPORATION
Reel/Frame 056202/0563 →
SECURITY INTEREST Recorded Feb 10, 2017
From: VITAC CORPORATION
To: ABACUS FINANCE GROUP, LLC
Reel/Frame 041228/0786 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 31, 2017
From: CAPTION COLORADO, L.L.C.
To: VITAC CORPORATION
Reel/Frame 041134/0404 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2010
From: POLUMBUS, RICHARD T.; HOMYACK, MICHAEL W.
To: CAPTION COLORADO L.L.C.
Reel/Frame 025081/0392 →
Continuity (2)
Provisional Application 61244823 · Sep 22, 2009
Related Publication 20110069230A1 · Mar 24, 2011