IP Library Granted Patent US 10,091,543
Granted Patent B2
US 10,091,543 · App. 15/489,095 · Granted Oct 2, 2018

Monitoring audio-visual content with captions

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,091,543
App. No.
15/489,095
Granted
Oct 2, 2018
Kind
B2
Abstract

To monitor audio-visual content which includes captions, caption fingerprints are derived from a length of each word in the caption, without regard to the identity of the character or characters forming the word. Audio-visual content is searched to identify a caption event having a matching fingerprint and missing captions; caption timing errors and caption discrepancies are measured.

Claims (37)

1. A method of monitoring audio-visual content which includes a succession of video images and a plurality of caption events, each caption event being associated with and intended to be co-timed with a respective string of successive images, the method comprising the steps in at least one processor of:

processing a caption event to derive a caption event fingerprint;

searching audio-visual content to identify a caption event matching a defined caption event fingerprint;

analysing any matching caption event; and

measuring any caption event error;

in which the caption event comprises a plurality of words, each formed from one or more characters, wherein the caption event fingerprint is derived from a length of each word in the caption event, without regard to the identity of the character or characters forming the word.

2. The method of claim 1 where the measured caption event error is selected from the group consisting of a missing caption event; a caption event timing error and a caption discrepancy.

3. The method of claim 1 , further comprising the step of determining the timing relative to the succession of video images of the identified caption event.

4. The method of claim 1 in which the caption event comprises a caption image, wherein the length of each word in the caption event is determined by:

analysing the caption image to identify caption image regions corresponding respectively with words in the caption; and

determining a horizontal dimension of each such caption image region.

5. The method of claim 4 , where the caption image is analysed to identify caption image regions corresponding respectively with lines of words in the caption and the length of a word is represented as a proportion of the length of a line.

6. The method of claim 5 , where the length of a word is represented as a proportion of the length of a line containing the word.

7. The method of claim 5 , where a measurement window of audio-visual content is defined containing a plurality of caption events and the length of a word is represented as a proportion of the representative line length derived from the measurement window.

8. The method of claim 7 , where the representative line length from the measurement window is selected from the group consisting of: the average line length; the length of a representative line; the length of the longest line; the length of line with the greatest number of words, or the length of the temporally closest line.

9. The method of claim 1 where the text of a caption event cannot be derived from the caption event fingerprint.

10. A system for monitoring audio-visual content which includes a succession of video images and a plurality of caption events, each caption event comprising a plurality of words, each word formed from one or more characters, each caption event being associated with and intended to be co-timed with a respective string of successive images, the system comprising:

at least first and second fingerprint generators operating in a content delivery chain at respective locations upstream and downstream of defined content manipulation process or processes, each fingerprint generator serving to process a caption event to derive a caption event fingerprint from a length of each word in the caption event, without regard to the identity of the character or characters forming the word; and

a fingerprint processor serving to compare caption event fingerprints from the respective first and second fingerprint generators to identify matching caption events; and to measure any caption event error selected from the group consisting of a missing caption event; a caption event timing error and a caption discrepancy.

11. The system of claim 10 , each fingerprint generator serving to record the timing of that caption event and the fingerprint processor serving to determine the timing relative to the succession of video images of each matched caption event.

12. The system of claim 10 in which the caption event comprises a caption image, wherein the length of each word in the caption event is determined by:

analysing the caption image to identify caption image regions corresponding respectively with words in the caption; and

determining a horizontal dimension of each such caption image region.

13. The system of claim 12 , where the caption image is analysed to identify caption image regions corresponding respectively with lines of words in the caption and the length of a word is represented as a proportion of the length of a line.

14. The system of claim 13 , where the length of a word is represented as a proportion of the length of a line containing the word.

15. The system of claim 13 , where a measurement window of audio-visual content is defined containing a plurality of caption events and the length of a word is represented as a proportion of the representative line length derived from the measurement window.

16. The system of claim 15 , where the representative line length is the average line length in the measurement window or the length of a representative line in the measurement window, for example the longest line, the line with the greatest number of words, or the temporally closest line.

17. A non-transient computer readable medium containing instructions causing a processor to implement a method of monitoring audio-visual content which includes a succession of video images and a plurality of caption events, each caption event being associated with and intended to be co-timed with a respective string of successive images, the method comprising the steps in at least one processor of:

processing a caption event to derive a caption event fingerprint;

searching audio-visual content to identify a caption event matching a defined caption event fingerprint;

analysing any matching caption event; and

measuring any caption event error, where the measured caption event error is selected from the group consisting of a missing caption event; a caption event timing error and a caption discrepancy;

in which the caption event comprises a plurality of words, each formed from one or more characters, wherein the caption event fingerprint is derived from a length of each word in the caption event, without regard to the identity of the character or characters forming the word.

18. The computer readable medium of claim 17 , in which the caption event comprises a plurality of words, each formed from one or more characters, wherein the caption event fingerprint is derived from a length of each word in the caption event, without regard to the identity of the character or characters forming the word.

19. The computer readable medium of claim 18 in which the caption event comprises a caption image, wherein the length of each word in the caption event is determined by:

analysing the caption image to identify caption image regions corresponding respectively with words in the caption; and

determining a horizontal dimension of each such caption image region.

Assignments (6)
ASSIGNMENT OF INTELLECTUAL PROPERTY SECURITY AGREEMENTS Recorded Dec 12, 2025
From: MS PRIVATE CREDIT ADMINISTRATIVE SERVICES LLC
To: MGG INVESTMENT GROUP LP
Reel/Frame 073959/0584 →
TERMINATION AND RELEASE OF PATENT SECURITY AGREEMENT Recorded Mar 21, 2024
From: MGG INVESTMENT GROUP LP
To: GRASS VALLEY USA, LLC; GRASS VALLEY CANADA; GRASS VALLEY LIMITED
Reel/Frame 066867/0336 →
SECURITY INTEREST Recorded Mar 20, 2024
From: GRASS VALLEY CANADA; GRASS VALLEY LIMITED
To: MS PRIVATE CREDIT ADMINISTRATIVE SERVICES LLC
Reel/Frame 066850/0869 →
GRANT OF SECURITY INTEREST - PATENTS Recorded Jul 2, 2020
From: GRASS VALLEY USA, LLC; GRASS VALLEY CANADA; GRASS VALLEY LIMITED
To: MGG INVESTMENT GROUP LP, AS COLLATERAL AGENT
Reel/Frame 053122/0666 →
CHANGE OF NAME Recorded Mar 9, 2020
From: SNELL ADVANCED MEDIA LIMITED
To: GRASS VALLEY LIMITED
Reel/Frame 052127/0795 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2017
From: DIGGINS, JONATHAN
To: SNELL ADVANCED MEDIA LIMITED
Reel/Frame 042990/0313 →