IP Library Granted Patent US 11,025,985
Granted Patent B2
US 11,025,985 · App. 16/421,391 · Granted Jun 1, 2021

Audio processing for detecting occurrences of crowd noise in sporting event television programming

Inventors: Mihailo Stojancic (San Jose, CA); Warren Packard (Palo Alto, CA)
Assignee: STATS LLC
H04N21/4394G10L21/0232G10L25/18G10L25/51G11B27/031H04N21/433
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,025,985
App. No.
16/421,391
Granted
Jun 1, 2021
Kind
B2
Abstract

Metadata for highlights of audiovisual content depicting a sporting event or other event are extracted from audiovisual content. The highlights may be segments of the content, such as a broadcast of a sporting event, that are of particular interest. Audio data for the audiovisual content is stored, and portions of the audio data indicating crowd excitement (noise) is automatically identified by analyzing an audio signal in the joint time and frequency domains. Multiple indicators are derived and subsequently processed to detect, validate, and render occurrences of crowd noise. Metadata are automatically generated, including time of occurrence, level of noise (excitement), and duration of cheering. Metadata may be stored, comprising at least a time index indicating a time, within the audiovisual content, at which each of the portions occurs. Periods of intense crowd noise may be used to identify highlights and/or to indicate crowd excitement during viewing of a highlight.

Claims (58)

1. A method for extracting metadata from depiction of an event, the method comprising:

at a data store, storing audio data depicting at least part of the event;

at a processor, automatically pre-processing the audio data to generate a spectrogram, in a spectral domain, for at least part of the audio data;

at the processor, automatically identifying one or more portions of the audio data that indicate crowd excitement at the event; and

at the data store, storing metadata comprising at least a time index indicating a time, within the depiction of the event, at which each of the one or more portions occurs;

wherein automatically identifying the one or more portions comprises:

identifying spectral magnitude peaks in each position of a sliding two-dimensional time-frequency analysis window of the spectrogram;

for each position of the sliding two-dimensional time-frequency analysis window, generating a spectral indicator representing an average spectral peak magnitude; and

using the spectral indicators to form a vector of spectral indicators with associated time portions.

2. The method of claim 1 , further comprising:

identifying runs of pairs of spectral indicators and associated analysis window time positions with contiguous time spacing below a threshold;

capturing the identified runs in a set of R vectors; and

forming a vector E with R vectors as its elements.

3. The method of claim 2 , further comprising extracting a run length for each of the R vectors by counting elements of each R vector.

4. The method of claim 2 , further comprising processing elements of the R vectors to obtain a maximum magnitude indicator for each R vector.

5. The method of claim 4 , further comprising extracting the time index for each of the R vectors.

6. The method of claim 5 , further comprising generating a preliminary event vector by replacing each of the R vectors in the vector E with a parameter triplet representing the maximum magnitude indicator, the time index, and a run length.

7. The method of claim 6 , further comprising processing the preliminary event vector to generate crowd noise event information comprising the time index.

8. A non-transitory computer-readable medium for extracting metadata from depiction of an event, comprising instructions stored thereon, that when executed by a processor, perform steps comprising:

causing a data store to store audio data depicting at least part of the event;

automatically pre-processing the audio data to generate a spectrogram, in a spectral domain, for at least part of the audio data prior to automatic identification of one or more portions of the audio data that indicate crowd excitement at the event;

automatically identifying one or more portions of the audio data that indicate crowd excitement at the event; and

causing the data store to store metadata comprising at least a time index indicating a time, within the depiction of the event, at which each of the one or more portions occurs;

wherein automatically identifying the one or more portions comprises:

identifying spectral magnitude peaks in each position of a sliding two-dimensional time-frequency analysis window of the spectrogram;

for each position of the sliding two-dimensional time-frequency analysis window, generating a spectral indicator representing an average spectral peak magnitude; and

using the spectral indicators to form a vector of spectral indicators with associated time portions.

9. The non-transitory computer-readable medium of claim 8 , further comprising instructions stored thereon, that when executed by the processor, further perform the steps comprising:

identifying runs of pairs of spectral indicators and associated analysis window time positions with contiguous time spacing below a threshold;

capturing the identified runs in a set of R vectors; and

forming a vector E with R vectors as its elements.

10. The non-transitory computer-readable medium of claim 9 , further comprising instructions stored thereon, that when executed by the processor, further perform the steps comprising:

extracting a run length for each of the R vectors by counting elements of each R vector;

process elements of the R vectors to obtain a maximum magnitude indicator for each R vector;

extract the time index for each of the R vectors;

generate a preliminary event vector by replacing each of the R vectors in the vector E with a parameter triplet representing the maximum magnitude indicator, the time index, and a run length; and

process the preliminary event vector to generate crowd noise event information comprising the time index.

11. A system for extracting metadata from depiction of an event, the system comprising:

a data store configured to store audio data depicting at least part of the event; and

a processor configured to:

automatically pre-process the audio data to generate a spectrogram, in a spectral domain, for at least part of the audio data; and

automatically identify one or more portions of the audio data that indicate crowd excitement at the event;

wherein:

the data store is further configured to store metadata comprising at least a time index indicating a time, within the depiction of the event, at which each of the one or more portions occurs; and

automatically identifying the one or more portions comprises:

identifying spectral magnitude peaks in each position of a sliding two-dimensional time-frequency analysis window of the spectrogram;

for each position of the sliding two-dimensional time-frequency analysis window, generating a spectral indicator representing an average spectral peak magnitude; and

using the spectral indicators to form a vector of spectral indicators with associated time portions.

12. The system of claim 11 , wherein the processor is further configured to:

identify runs of pairs of spectral indicators and associated analysis window time positions with contiguous time spacing below a threshold;

capture the identified runs in a set of R vectors; and

form a vector E with R vectors as its elements.

13. The system of claim 12 , wherein the processor is further configured to:

extract a run length for each of the R vectors by counting elements of each R vector;

process elements of the R vectors to obtain a maximum magnitude indicator for each R vector;

extract the time index for each of the R vectors;

generate a preliminary event vector by replacing each of the R vectors in the vector E with a parameter triplet representing the maximum magnitude indicator, the time index, and the run length; and

process the preliminary event vector to generate crowd noise event information comprising the time index.

Assignments (7)
RELEASE OF SECURITY INTEREST Recorded Apr 23, 2026
From: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
To: STATS LLC
Reel/Frame 074456/0116 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY'S NAME PREVIOUSLY RECORDED AT REEL: 055709 FRAME: 0376. ASSIGNOR(S) HEREBY CONFIRMS THE SECOND LIEN PATENT SECURITY AGREEMENT. Recorded May 14, 2021
From: STATS LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 056319/0270 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CONVEYING PARTY NAME PREVIOUSLY RECORDED AT REEL: 055709 FRAME: 0367. ASSIGNOR(S) HEREBY CONFIRMS THE FIRST LIEN PATENT SECURITY AGREEMENT. Recorded Mar 30, 2021
From: STATS LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 056021/0568 →
SECOND LIEN PATENT SECURITY AGREEMENT Recorded Mar 24, 2021
From: STATS INTERMEDIATE HOLDINGS, LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 055709/0376 →
FIRST LIEN PATENT SECURITY AGREEMENT Recorded Mar 24, 2021
From: STATS INTERMEDIATE HOLDINGS, LLC
To: MORGAN STANLEY SENIOR FUNDING, INC., AS COLLATERAL AGENT
Reel/Frame 055709/0367 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 4, 2021
From: THUUZ, INC.
To: STATS LLC
Reel/Frame 055490/0394 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 23, 2019
From: STOJANCIC, MIHAILO; PACKARD, WARREN
To: THUUZ, INC.
Reel/Frame 049273/0800 →
Continuity (4)
Provisional Application 62680955 · Jun 5, 2018
Provisional Application 62712041 · Jul 30, 2018
Provisional Application 62746454 · Oct 16, 2018
Related Publication 20190373310A1 · Dec 5, 2019
Cited By (1)
US 12,537,909