Methods and systems for detecting content within media streams
Systems and methods are provided for detecting a content type of content within a media stream. A computing device may receive a media stream and define a set of media segments that each represent a portion of the media stream. The computing device may identify a first media segment that includes a first boundary and a second media segment that includes a second boundary. The computing device may predict whether the subset of media segments that are positioned between the first media segment and the second media segment include content of a particular content type. The computing device may then transmit an indication of the prediction.
1 . A system comprising:
one or more processors; and
a machine-readable storage medium storing instructions that when executed by the one or more processors, cause the one or more processors to perform operations including:
defining a set of audio segments from an audio stream of media, each audio segment representing a portion of the audio stream;
executing one or more neural networks, wherein the one or more neural networks generate, for each audio segment of the set of audio segments, a first probability that an audio segment corresponds to a boundary and a second probability that the audio segment corresponds to a content type;
identifying a first audio segment associated with a first boundary and a second audio segment associated with a second boundary, wherein the first boundary and the second boundary define a media segment corresponding to a portion of the media between the first boundary and the second boundary;
determining that the media segment corresponds to the content type based on the second probability of at least one audio segment between the first audio segment and the second audio segment and based on a time interval of the media segment being less than or equal to a threshold time interval, wherein the threshold time interval is defined based on the content type; and
facilitating a transmission including an indication of the content type of the media segment.
2 . The system of claim 1 , wherein at least one neural network of the one or more neural networks is a recurrent neural network.
3 . The system of claim 1 , further comprising:
translating the set of audio segments into a frequency-domain representation of the set of audio segments, wherein the one or more neural networks are executed using the frequency-domain representation of the set of audio segments.
4 . The system of claim 1 , wherein defining the set of audio segments includes:
generating a spectrogram from a portion of the audio stream; and
defining audio segments using a tensor derived from the spectrogram.
5 . The system of claim 1 , wherein determining that the media segment corresponds to the content type is further based on an average probability derived from the second probability of the first audio segment and the second probability of the second audio segment.
6 . The system of claim 1 , further comprising:
determining an identification of the media segment.
7 . A method comprising:
defining a set of audio segments from an audio stream of media, each audio segment representing a portion of the audio stream;
executing one or more neural networks, wherein the one or more neural networks generate, for each audio segment of the set of audio segments, a first probability that an audio segment corresponds to a boundary and a second probability that the audio segment corresponds to a content type;
identifying a first audio segment associated with a first boundary and a second audio segment associated with a second boundary, wherein the first boundary and the second boundary define a media segment corresponding to a portion of the media between the first boundary and the second boundary;
determining that the media segment corresponds to the content type based on the second probability of at least one audio segment between the first audio segment and the second audio segment and based on a time interval of the media segment being less than or equal to a threshold time interval, wherein the threshold time interval is defined based on the content type; and
facilitating a transmission including an indication of the content type of the media segment.
8 . The method of claim 7 , wherein at least one neural network of the one or more neural networks is a recurrent neural network.
9 . The method of claim 7 , further comprising:
translating the set of audio segments into a frequency-domain representation of the set of audio segments, wherein the one or more neural networks are executed using the frequency-domain representation of the set of audio segments.
10 . The method of claim 7 , wherein defining the set of audio segments includes:
generating a spectrogram from a portion of the audio stream; and
defining audio segments using a tensor derived from the spectrogram.
11 . The method of claim 7 , wherein determining that the media segment corresponds to the content type is further based on an average probability derived from the second probability of the first audio segment and the second probability of the second audio segment.
12 . The method of claim 7 , further comprising:
determining an identification of the media segment.
13 . A non-transitory computer-readable medium storing instructions that when executed by one or more processors, cause the one or more processors to perform operations including:
defining a set of audio segments from an audio stream of media, each audio segment representing a portion of the audio stream;
executing one or more neural networks, wherein the one or more neural networks generate, for each audio segment of the set of audio segments, a first probability that an audio segment corresponds to a boundary and a second probability that the audio segment corresponds to a content type;
identifying a first audio segment associated with a first boundary and a second audio segment associated with a second boundary, wherein the first boundary and the second boundary define a media segment corresponding to a portion of the media between the first boundary and the second boundary;
determining that the media segment corresponds to the content type based on the second probability of at least one audio segment between the first audio segment and the second audio segment and based on a time interval of the media segment being less than or equal to a threshold time interval, wherein the threshold time interval is defined based on the content type; and
facilitating a transmission including an indication of the content type of the media segment.
14 . The non-transitory computer-readable medium of claim 13 , wherein at least one neural network of the one or more neural networks is a recurrent neural network.
15 . The non-transitory computer-readable medium of claim 13 , further comprising:
translating the set of audio segments into a frequency-domain representation of the set of audio segments, wherein the one or more neural networks are executed using the frequency-domain representation of the set of audio segments.
16 . The non-transitory computer-readable medium of claim 13 , wherein defining the set of audio segments includes:
generating a spectrogram from a portion of the audio stream; and
defining audio segments using a tensor derived from the spectrogram.
17 . The non-transitory computer-readable medium of claim 13 , wherein determining that the media segment corresponds to the content type is further based on an average probability derived from the second probability of the first audio segment and the second probability of the second audio segment.