IP Library Granted Patent US 8,572,488
Granted Patent B2
US 8,572,488 · App. 12/748,695 · Granted Oct 29, 2013

Spot dialog editor

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,572,488
App. No.
12/748,695
Granted
Oct 29, 2013
Kind
B2
Abstract

Automated methods are used to augment an original script of a time-based media program that contains speech with timing metadata and an indication of a confidence level with which script portions are matched with the speech in the program audio track. An editor uses a spot dialog editor tool to produce a word-accurate spot dialog master by reviewing and editing script portions displayed next to the corresponding portions of the time-based media. The spot dialog editor tool selects for review and editing the parts of the script having a low confidence match with the speech in the program.

Claims (42)

1. In a computer-based system, a method of conforming a script or transcript comprising words and phrases corresponding to a time-based media program that includes recorded speech, the method comprising:

receiving the script or transcript;

receiving timing information comprising, for each of a plurality of words or phrases within the script or transcript, a temporal location within the time-based media program where that word or phrase has been matched with a corresponding spoken word or phrase recognized in the recorded speech, and a confidence level for a match of the word or phrase from the selected portion of the script or transcript to the word or phrase recognized in the recorded speech;

displaying a selected portion of the script or transcript;

using the timing information, displaying, along with the word or phrases of the displayed selected portion of the script or transcript to which the confidence level corresponds, a graphical representation of the confidence levels for the matches of the words or phrases from the displayed selected portion of the script or transcript to the corresponding words or phrases recognized in the recorded speech from the temporal location within the time-based media program according to the received timing information, wherein at least one of the words or phrases within the selected portion has a confidence level equal to or lower than a predetermined threshold confidence level indicating that said at least one of the words or phrases in the selected portion do not match the words or phrases recognized in the recorded speech;

retrieving a portion of the time-based media program corresponding to the temporal location associated with the selected portion of the script or transcript;

playing back the retrieved portion of the time-based media program, while the selected portion of the script or transcript and graphical representation of the corresponding confidence levels are displayed, to enable a user to compare the retrieved portion of the time-based media program and the selected portion of the script or transcript; and

in response to input received from the user, making corrections as indicated by the user to the text of the selected portion of the script or transcript so as to have the corrected text match the words or phrases in the recorded speech, such corrections being made while text from the selected portion of the script or transcript and corresponding graphical representations of the confidence levels are displayed and the corresponding portion of the time-based media is played back.

2. The method of claim 1 , wherein the time-based media includes a video component synchronized with the recorded speech.

3. The method of claim 1 , wherein the time-based media is an audio-only program.

4. The method of claim 1 , wherein the script or transcript includes metadata associated with the time-based media program in addition to the timing information.

5. The method of claim 1 , the method further comprising outputting a spot dialog master corresponding to the time-based media program, wherein the spot dialog master includes:

the edited script or transcript;

for each word or phrase in the edited script or transcript, a name of a character speaking that word or phrase in the time-based media program; and

the timing information.

6. The method of claim 5 , wherein the spot dialog master is represented as an XML document.

7. The method of claim 5 , further comprising generating from the spot dialog master a selected character dialog master that includes only words or phrases of the edited script or transcript spoken by the selected character, and timing information corresponding to the words of phrases spoken by the selected character.

8. The method of claim 1 wherein the graphical representation of the confidence level comprises an indication of an amount of editing to be done to words or phrases in the script or transcript.

9. The method of claim 1 wherein the graphical representation of the confidence level comprises a formatting applied to displayed text.

10. The method of claim 1 wherein the graphical representation of the confidence level includes an indication of high, medium and low confidence levels.

11. The method of claim 1 wherein the graphical representation of the confidence level includes a different formatting applied to text for each of high, medium and low confidence levels.

12. In a computer-based system, a method of conforming a transcript comprising words and phrases of dialog to a dialog audio track of a time-based media program, the method comprising:

receiving an augmented version of the transcript and, for each word or phrase of a plurality of the words and phrases within the transcript, timing information comprising a temporal location within the time-based media program where the word or phrase is spoken in the dialog audio track, and a confidence level indicating a quality of a match between each of the words or phrases from the transcript and words or phrases recognized in their corresponding identified temporal location within the time-based media;

receiving the time-based media program; and

providing an interactive graphical interface for a user, the graphical interface including a transcript display portion for displaying text from a portion of the transcript and a media display portion, simultaneously displayed with the transcript display portion, for displaying a portion of the time-based media program spanning the identified temporal location corresponding to the displayed text from the portion of the transcript according to the timing information;

displaying text from the portion of the transcript in the transcript display portion, wherein the text from the transcript is displayed with a visual attribute corresponding to the confidence levels for matches of the text from the portion of the transcript according to the timing information;

in response to a request from the user, playing the portion of the time-based media in the media display portion while the corresponding text and visual attribute corresponding to the confidence levels are displayed in the transcript display portion; and

enabling the user to make corrections to the displayed text from the transcript in the transcript display portion so as to have the text match the words or phrases spoken in the dialog audio track while the text from the portion of the transcript and the corresponding visual attribute corresponding to the corresponding confidence level are displayed and the corresponding portion of the time-based media is played.

13. The method of claim 12 , wherein the time-based media includes a video component.

14. A computer program product, comprising:

a computer-readable medium; and

computer program instructions encoded on the computer-readable medium, wherein the computer program instructions, when processed by a computer, instruct the computer to perform a method for editing time-based media that includes recorded speech, the method comprising:

receiving the script or transcript;

receiving timing information comprising, for each of a plurality of words or phrases within the script or transcript, a temporal location within the time-based media program where that word or phrase has been matched with a corresponding spoken word or phrase recognized in the recorded speech, and a confidence level for a match of the word or phrase from the selected portion of the transcript to the spoken word or phrase recognized in the recorded speech from the associated temporal location within the time-based media program;

displaying a selected portion of the script or transcript;

using the timing information, displaying, along with the word or phrases of the displayed selected portion of the script or transcript to which the confidence level corresponds, a graphical representation of the confidence levels for the matches of the words or phrases from the displayed selected portion of the script or transcript to the corresponding words or phrases recognized in the recorded speech from the temporal location within the time-based media program according to the received timing information, wherein at least one of the words or phrases within the selected portion has a confidence level equal to or lower than a predetermined threshold confidence level, indicating that said at least one of the words or phrases in the selected portion do not match the words or phrases recognized in the recorded speech;

retrieving a portion of the time-based media program corresponding to the temporal location associated with the selected portion of the script or transcript;

playing back the retrieved portion of the time-based media program, while the selected portion of the script or transcript and graphical representation of the corresponding confidence levels are displayed, to enable a user to compare the retrieved portion of the time-based media program and the selected portion of the script or transcript; and

in response to input received from the user, making corrections as indicated by the user to the text of the selected portion of the script or transcript, so as to have the corrected text match the words or phrases in the recorded speech, such corrections being made while text from the selected portion of the script or transcript and corresponding graphical representations of the confidence levels are displayed and the corresponding portion of the time-based media is played back.

15. The computer program product of claim 14 , wherein the time-based media includes a video component synchronized with the recorded speech.

16. The computer program product of claim 14 , wherein the time-based media is an audio-only program.

17. The computer program product of claim 14 , wherein the script or transcript includes metadata associated with the time-based media program in addition to the timing information.

Assignments (8)
RELEASE OF SECURITY INTEREST IN PATENTS (REEL/FRAME 054900/0716) Recorded Nov 8, 2023
From: JPMORGAN CHASE BANK, N.A.
To: AVID TECHNOLOGY, INC.
Reel/Frame 065523/0146 →
PATENT SECURITY AGREEMENT Recorded Nov 8, 2023
From: AVID TECHNOLOGY, INC.
To: SIXTH STREET LENDING PARTNERS, AS ADMINISTRATIVE AGENT
Reel/Frame 065523/0194 →
SECURITY INTEREST Recorded Jan 5, 2021
From: AVID TECHNOLOGY, INC.
To: JPMORGAN CHASE BANK, N.A., AS ADMINISTRATIVE AGENT
Reel/Frame 054900/0716 →
RELEASE OF SECURITY INTEREST Recorded Jan 5, 2021
From: CERBERUS BUSINESS FINANCE, LLC
To: AVID TECHNOLOGY, INC.
Reel/Frame 055731/0019 →
RELEASE OF SECURITY INTEREST IN UNITED STATES PATENTS Recorded Mar 1, 2016
From: KEYBANK NATIONAL ASSOCIATION
To: AVID TECHNOLOGY, INC.
Reel/Frame 037970/0201 →
ASSIGNMENT FOR SECURITY -- PATENTS Recorded Feb 26, 2016
From: AVID TECHNOLOGY, INC.
To: CERBERUS BUSINESS FINANCE, LLC, AS COLLATERAL AGENT
Reel/Frame 037939/0958 →
PATENT SECURITY AGREEMENT Recorded Jun 23, 2015
From: AVID TECHNOLOGY, INC.
To: KEYBANK NATIONAL ASSOCIATION, AS THE ADMINISTRATIVE AGENT
Reel/Frame 036008/0824 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2010
From: PHILLIPS, MICHAEL E.; LEA, GLENN
To: AVID TECHNOLOGY, INC.
Reel/Frame 024154/0001 →