IP Library Granted Patent US 9,020,817
Granted Patent B2
US 9,020,817 · App. 13/744,585 · Granted Apr 28, 2015

Using speech to text for detecting commercials and aligning edited episodes with transcripts

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,020,817
App. No.
13/744,585
Granted
Apr 28, 2015
Kind
B2
Abstract

Methods and apparatus, including computer program products, for using speech to text for detecting commercials and aligning edited episodes with transcripts. A method includes, receiving an original video or audio having a transcript, receiving an edited video or audio of the original video or audio, applying a speech-to-text process to the received original video or audio having a transcript, applying a speech-to-text process to the received edited video or audio, and applying an alignment to determine locations of the edits.

Claims (30)

1. A method comprising:

in a computer system having a processor and a memory, receiving an original video or audio having a transcript;

receiving an edited video or audio of the original video or audio;

applying a speech-to-text process to the received original video or audio having a transcript;

applying a speech-to-text process to the received edited video or audio; and

applying an alignment to determine locations of the edits.

2. The method of claim 1 further comprising:

aligning an output of the speech-to-text process to the received original video or audio having a transcript and the output of the speech-to-text process to the received edited video or audio;

examining alignments for large sections that exist only in the original and not in the edited version; and

for each pair of segments of the remaining segments aside from the commercials in the original, and the corresponding segments in the edited version, executing a closed-caption alignment process.

3. The method of claim 1 wherein the output of the speech-to-text process to the received original video or audio having a transcript and the output of the speech-to-text process to the received edited video or audio include timestamps.

4. The method of claim 1 wherein the transcript is derived from closed-captioning.

5. The method of claim 1 wherein the transcript is derived from metadata associated with the video or audio.

6. The method of claim 5 wherein the metadata comprises director's commentary.

7. The method of claim 5 wherein the metadata comprises scene descriptions.

8. The method of claim 1 wherein the original video or audio is selected from the group consisting of broadcast media, cable media and satellite media.

9. The method of claim 2 wherein aligning comprises performing a match on words with dynamic programming.

10. The method of claim 2 wherein aligning comprises performing a match on words minimizing an error metric.

11. The method of claim 10 wherein the error metric comprises a number of substitutions, insertions and deletions.

12. The method of claim 2 wherein examining comprises considering each position as a putative commercial boundary and looking at an accuracy within a window before/after the putative commercial boundary.

13. The method of claim 12 wherein an accuracy is derived from a combination of one or more of the number of correct, substitutions, insertions, and deletions.

14. The method of claim 2 wherein examining comprises using multiple window sizes.

15. The method of claim 2 wherein examining comprises using finite state machine logic.

16. The method of claim 2 wherein examining comprises using statistical classification or other machine learning.

17. The method of claim 2 wherein examining comprises imposing minimum or maximum length restrictions.

18. The method of claim 2 wherein examining comprises using confidence scores.

19. The method of claim 2 wherein examining comprises looking at time stamps to a left and a right.

20. The method of claim 2 further comprising performing closed-captioning alignment using the detected commercials to cut the original and edited media into segments to be aligned.

21. The method of claim 1 wherein applying a speech-to-text process to the received original video or audio having the transcript is replaced with audio fingerprinting the received original video or audio having the transcript.

22. The method of claim 1 wherein applying a speech-to-text process to the received edited video or audio is replaced with audio fingerprinting the received edited video or audio.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Jan 2, 2025
From: SIXTH STREET SPECIALTY LENDING INC.
To: PIANO SOFTWARE B.V.
Reel/Frame 069722/0531 →
SECURITY INTEREST Recorded Oct 3, 2022
From: PIANO SOFTWARE B.V.
To: SIXTH STREET SPECIALTY LENDING, INC.
Reel/Frame 061290/0590 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 30, 2015
From: RAMP HOLDINGS INC.
To: CXENSE ASA
Reel/Frame 037018/0816 →