IP Library Granted Patent US 9,332,319
Granted Patent B2
US 9,332,319 · App. 12/890,744 · Granted May 3, 2016

Amalgamating multimedia transcripts for closed captioning from a plurality of text to speech conversions

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,332,319
App. No.
12/890,744
Granted
May 3, 2016
Kind
B2
Abstract

Methods and systems for converting speech to text are disclosed. One method includes analyzing multimedia content to determine the presence of closed captioning data. The method includes, upon detecting closed captioning data, indexing the closed captioning data as associated with the multimedia content. The method also includes, upon failure to detect closed captioning data in the multimedia content, extracting audio data from multimedia content, the audio data including speech data, performing a plurality of speech to text conversions on the speech data to create a plurality of transcripts of the speech data, selecting text from one or more of the plurality of transcripts to form an amalgamated transcript, and indexing the amalgamated transcript as associated with the multimedia content.

Claims (52)

1. A method of converting speech to text, the method comprising:

analyzing multimedia content using one or more computing devices to determine the presence of dosed captioning data;

upon detecting closed captioning data, causing at least one of the one or more computing devices to begin:

i) indexing the closed captioning data as associated with the multimedia content;

upon failure to detect closed captioning data in the multimedia content causing at least one of the one or more computing devices to begin:

i) extracting audio data from multimedia content, the audio data including speech data;

ii) performing a plurality of different speech to text conversions on the speech data to create a plurality of transcripts of the speech data, the plurality of different speech to text conversions include speech to text conversion processes from different software vendors, wherein at least one of the plurality of transcripts is different from a remainder of the plurality of transcripts, wherein at least one of the speech to text conversions uses a context-sensitive speech to text dictionary selected according to the subject matter of the multimedia content;

iii) selecting text from among the plurality of transcripts to form an amalgamated transcript; and

iv) indexing the amalgamated transcript as associated with the multimedia content, wherein indexing the amalgamated transcript includes storing metadata associating text in the amalgamated transcript to timestamps associated with the multimedia content.

2. The method of claim 1 , further comprising:

analyzing the multimedia content to detect subtitle information; and

upon detecting subtitle information, indexing the subtitle information as associated with the multimedia content.

3. The method of claim 1 , further comprising associating the indexed closed captioning data or amalgamated transcript with metadata describing the multimedia content.

4. The method of claim 1 , wherein indexing the amalgamated transcript comprises associating indexed portions of the multimedia content with a query engine.

5. The method of claim 1 , wherein each of the plurality of speech to text conversions is performed by a corresponding speech to text conversion program.

6. The method of claim 5 , further comprising training the speech to text program.

7. The method of claim 6 , wherein training the speech to text conversion program includes training the speech to text conversion program using speech patterns specific to a presenter whose speech data is included in the multimedia content.

8. The method of claim 6 , wherein training the speech to text conversion program includes training the speech to text conversion program using a speech to text dictionary in a training mode of the speech to text conversion program.

9. The method of claim 8 , wherein the speech to text dictionary is a context-sensitive dictionary selected according to the subject matter of the multimedia content.

10. The method of claim 1 , further comprising selecting a subject-specific speech to text dictionary related to the multimedia content.

11. The method of claim 1 , further comprising receiving a script alongside the multimedia content and including at least a portion of the script in a transcript indexed to the multimedia content.

12. The method of claim 1 wherein the one or more computing devices comprise a single computing device.

13. A system for converting speech to text, the system comprising:

one or more computing systems each including a programmable circuit and a memory, the one or more computing systems executing program instructions, which, when executed, cause the one or more computing systems to:

i) analyze multimedia content to determine the presence of closed captioning data;

ii) upon detecting closed captioning data: index the closed captioning data as associated with the multimedia content;

iii) upon failure to detect closed captioning data in the multimedia content: extract audio data from multimedia content, the audio data including speech data; perform a plurality of different speech to text conversions on the speech data to create a plurality of transcripts of the speech data, the plurality of different speech to text conversions include speech to text conversion processes from different software vendors, wherein at least one of the plurality of transcripts is different from a remainder of the plurality of transcripts, and wherein at least one of the speech to text conversions uses a context-sensitive speech to text dictionary selected according to the subject matter of the multimedia content; select text from among the plurality of transcripts to form an amalgamated transcript; and index the amalgamated transcript as associated with the multimedia content, wherein indexing the amalgamated transcript includes storing metadata associating text in the amalgamated transcript to timestamps associated with the multimedia content.

14. The system of claim 13 , wherein the closed captioning data or amalgamated transcript is associated with metadata describing the multimedia content.

15. A system for converting speech to text, the system comprising:

an analysis module operating on one or more computing systems and configured to analyze multimedia content to determine the presence of closed captioning data;

an audio extraction module operating on one or more computing systems, the audio extraction module configured to extract audio data from multimedia content, the audio data including speech data;

a plurality of speech to text conversion programs operating on one or more computing systems, each of the plurality of speech to text conversion programs operating on the speech data to create a plurality of transcripts of the speech data, the plurality of speech to text conversion programs include programs from different software vendors, and at least one of the plurality of speech to text conversion programs uses a context-sensitive speech to text dictionary selected according to the subject matter of the multimedia content, wherein at least one of the plurality of transcripts is different from a remainder of the plurality of transcripts;

a transcript selection module configured to select text from one or more of the plurality of transcripts to form an amalgamated transcript; and

an indexing module operating on the one or more computing systems and configured to,

upon detecting closed captioning data:

index the closed captioning data as associated with the multimedia content; and

upon failure to detect closed captioning data in the speech data:

index the amalgamated transcript as associated with the multimedia content, wherein indexing the amalgamated transcript includes storing metadata associating text in the amalgamated transcript to timestamps associated with the multimedia content.

16. The system of claim 15 , wherein the analysis module is further configured to analyze the multimedia content to detect subtitle information, and the indexing module is further configured to, upon detecting subtitle information, index the subtitle information as associated with the multimedia content.

17. The system of claim 15 , further comprising a training module operating on the one or more computing systems and configured to train the one or more speech to text programs using speech patterns specific to a presenter whose speech data is included in the multimedia content.

18. The system of claim 15 , wherein the training module selects a subject-specific speech to text dictionary related to the multimedia content.

19. The system of claim 15 , wherein the closed captioning data or amalgamated transcript is associated with metadata describing the multimedia content.

20. The system of claim 15 , wherein the indexing module is configured to store metadata associating text in the amalgamated transcript to timestamps associated with the multimedia content.

21. A method of converting speech to text, the method comprising:

using one or more processors to train one or more speech to text programs using a context-sensitive speech to text dictionary selected according to the subject matter of the multimedia content;

analyzing the extracted speech data using at least one of the one or more processors to determine the presence of closed captioning data;

upon detecting closed captioning data: using at least one of the one or more processors to begin indexing the closed captioning data as associated with the multimedia content;

upon failure to detect closed captioning data in the speech data, causing at least one of the one or more processors to begin:

i) extracting audio data from multimedia content, the audio data including speech data; performing a plurality of different speech to text conversions on the speech data using the one or more speech to text programs to create a plurality of transcripts of the speech data, wherein each of the plurality of transcripts is different from a remainder of the plurality of transcripts and the speech to text programs each associated with different software vendors;

ii) selecting text from one or more of the plurality of transcripts to form an amalgamated transcript; and

iii) indexing the amalgamated transcript as associated with the multimedia content by storing metadata associating text in the amalgamated transcript to timestamps associated with the multimedia content.

22. The method of claim 21 wherein the one or more processors comprise a single processor.

Assignments (12)
AMENDED AND RESTATED PATENT SECURITY AGREEMENT Recorded Jun 27, 2025
From: UNISYS CORPORATION; UNISYS HOLDING CORPORATION; UNISYS NPL, INC.; UNISYS AP INVESTMENT COMPANY I
To: COMPUTERSHARE TRUST COMPANY, N.A., AS COLLATERAL TRUSTEE
Reel/Frame 071759/0527 →
RELEASE OF SECURITY INTEREST Recorded Oct 28, 2020
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: UNISYS CORPORATION
Reel/Frame 054231/0496 →
RELEASE OF SECURITY INTEREST Recorded Nov 9, 2017
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: UNISYS CORPORATION
Reel/Frame 044416/0114 →
RELEASE OF SECURITY INTEREST Recorded Nov 9, 2017
From: WELLS FARGO BANK, NATIONAL ASSOCIATION (SUCCESSOR TO GENERAL ELECTRIC CAPITAL CORPORATION)
To: UNISYS CORPORATION
Reel/Frame 044416/0358 →
SECURITY INTEREST Recorded Oct 6, 2017
From: UNISYS CORPORATION
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 044144/0081 →
PATENT SECURITY AGREEMENT Recorded Apr 27, 2017
From: UNISYS CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS COLLATERAL TRUSTEE
Reel/Frame 042354/0001 →
SECURITY INTEREST Recorded Oct 6, 2016
From: UNISYS CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS AGENT
Reel/Frame 039960/0057 →
RELEASE OF SECURITY INTEREST Recorded Mar 26, 2013
From: DEUTSCHE BANK TRUST COMPANY AMERICAS, AS COLLATERAL TRUSTEE
To: UNISYS CORPORATION
Reel/Frame 030082/0545 →
RELEASE OF SECURITY INTEREST Recorded Mar 15, 2013
From: DEUTSCHE BANK TRUST COMPANY
To: UNISYS CORPORATION
Reel/Frame 030004/0619 →
SECURITY AGREEMENT Recorded Jun 27, 2011
From: UNISYS CORPORATION
To: GENERAL ELECTRIC CAPITAL CORPORATION, AS AGENT
Reel/Frame 026509/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 23, 2010
From: TSAI, JOHNEY; MILLER, MATTHEW; STRONG, DAVID
To: UNISYS CORPORATION
Reel/Frame 025397/0675 →
SECURITY AGREEMENT Recorded Nov 2, 2010
From: UNISYS CORPORATION
To: DEUTSCHE BANK NATIONAL TRUST COMPANY
Reel/Frame 025227/0391 →