IP Library Patent Application 14623354
Patent Application
App. No. 14/623,354

METHOD AND APPARATUS FOR MANAGING AUDIO VISUAL, AUDIO OR VISUAL CONTENT

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
14/623,354
Abstract

To manage audio visual content, a stream of fingerprints is derived in a fingerprint generator and received at a fingerprint processor that is physically separate from the fingerprint generator. Metadata is generated by processing the fingerprints to detect the sustained occurrence of low values of an audio fingerprint to generate metadata indicating silence; comparing the pattern of differences between temporally succeeding values of a fingerprint with expected patterns of film cadence to generate metadata indicating a film cadence; and comparing differences between temporally succeeding values of a fingerprint with a threshold to generate metadata indicating a still image or freeze frame.

Claims (46)

1 . A method of managing audio visual, audio or visual content, comprising the steps of:

receiving a stream of fingerprints, derived in a fingerprint generator by an irreversible data reduction process from respective temporal regions within a particular audio visual, audio or visual content stream, at a fingerprint processor that is physically separate from the fingerprint generator via a communication network; and

processing said fingerprints in the fingerprint processor to generate metadata which is not directly encoded in the fingerprints, with one or more processes selected from the group consisting of:

detecting the sustained occurrence of low values of an audio fingerprint to generate metadata indicating silence;

comparing the pattern of differences between temporally succeeding values of a fingerprint with expected patterns of film cadence to generate metadata indicating a film cadence; and

comparing differences between temporally succeeding values of a fingerprint with a threshold to generate metadata indicating a still image or freeze frame.

2 . The method according to claim 1 , wherein said communication network comprises the Internet.

3 . The method according to claim 1 , wherein an audio fingerprint stream has a data rate of less than about 500 byte/s per audio channel.

4 . The method according to claim 1 , wherein an audio fingerprint stream has a data rate of less than about 250 byte/s per audio channel.

5 . The method according to claim 1 , wherein a video fingerprint stream has a data rate of less than about 500 byte per field.

6 . The method according to claim 1 , wherein a video fingerprint stream has a data rate of less than about 200 byte per field.

7 . The method according to claim 1 , wherein said content comprises a video stream of video frames and wherein a fingerprint is generated for substantially every frame in the video stream.

8 . A method of managing audio visual, audio or visual content, comprising the steps of:

receiving a stream of fingerprints, derived in a fingerprint generator by an irreversible data reduction process from respective temporal regions within a particular audio visual, audio or visual content stream, at a fingerprint processor that is physically separate from the fingerprint generator via a communication network; and

processing said fingerprints in the fingerprint processor to generate metadata which is not directly encoded in the fingerprints; wherein said processing includes

windowing the stream of fingerprints with a time window,

deriving frequencies of occurrence of particular fingerprint values or ranges of fingerprint values within each time window,

determining statistical moments or entropy values of said frequencies of occurrence,

comparing said statistical moments or entropy values with expected values for particular types of content, and

generating metadata representing the type of the audio visual, audio or visual content.

9 . The method according to claim 8 , wherein said statistical moment comprises one or more of the mean; variance; skew or kurtosis of said frequencies of occurrence.

10 . The method according to claim 8 , wherein said communication network comprises the Internet.

11 . The method according to claim 8 , wherein a video fingerprint stream has a data rate of less than about 500 byte per field.

12 . The method according to claim 8 , wherein a video fingerprint stream has a data rate of less than about 200 byte per field.

13 . The method according to claim 8 , wherein said content comprises a video stream of video frames and wherein a fingerprint is generated for substantially every frame in the video stream.

14 . An apparatus for use in managing audio visual, audio or visual content, the apparatus comprising:

a fingerprint processor configured to receive via a communication network a stream of fingerprints derived in a fingerprint generator that is physically separate from the fingerprint processor by an irreversible data reduction process from respective temporal regions within a particular audio visual, audio or visual content stream, at a fingerprint processor generator; the fingerprint processor including

a window unit configure to receive said stream of fingerprints and apply a time window,

a frequency of occurrence histogram unit configured to derive the frequencies of occurrence of particular fingerprint values in each time window,

a statistical moment unit configured to derive statistical moments of said frequencies of occurrence, and

a classifier configured to generate from said statistical moments metadata representing the type of the audio visual, audio or visual content.

15 . The apparatus according to claim 14 , further comprising an entropy unit configured to derive entropy values for histograms of frequencies of occurrence and wherein said classifier is configured to generate said metadata representing the type of the audio visual, audio or visual content additionally from said entropy values.

16 . A non-transitory computer program product adapted to cause programmable apparatus to implement a method of managing audio visual, audio or visual content, comprising the steps of:

receiving a stream of fingerprints, derived in a fingerprint generator by an irreversible data reduction process from respective temporal regions within a particular audio visual, audio or visual content stream, at a fingerprint processor that is physically separate from the fingerprint generator via a communication network; and

processing said fingerprints in the fingerprint processor to generate metadata which is not directly encoded in the fingerprints, with one or more processes selected from the group consisting of,

detecting the sustained occurrence of low values of an audio fingerprint to generate metadata indicating silence,

comparing the pattern of differences between temporally succeeding values of a fingerprint with expected patterns of film cadence to generate metadata indicating a film cadence, and

comparing differences between temporally succeeding values of a fingerprint with a threshold to generate metadata indicating a still image or freeze frame.

17 . A non-transitory computer program product adapted to cause programmable apparatus to implement a method of managing audio visual, audio or visual content, comprising the steps of:

receiving a stream of fingerprints, derived in a fingerprint generator by an irreversible data reduction process from respective temporal regions within a particular audio visual, audio or visual content stream, at a fingerprint processor that is physically separate from the fingerprint generator via a communication network; and

processing said fingerprints in the fingerprint processor to generate metadata which is not directly encoded in the fingerprints; wherein said processing includes

windowing the stream of fingerprints with a time window,

deriving frequencies of occurrence of particular fingerprint values or ranges of fingerprint values within each time window,

determining statistical moments or entropy values of said frequencies of occurrence,

comparing said statistical moments or entropy values with expected values for particular types of content, and

generating metadata representing the type of the audio visual, audio or visual content.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 14, 2015
From: DIGGINS, JONATHAN
To: SNELL LIMITED
Reel/Frame 035408/0416 →