Automatically moderating content of media programs using multi-tiered machine learning solutions
As a media program is aired to listeners, a control system monitors audio data transmitted to the listeners and interactions received from the listeners to determine whether the media program has violated or may violate one or more rules. The audio data is processed to identify words expressed therein and features of the audio data. Additionally, features of users (e.g., a creator or any listeners or guests) may be calculated based on any information or data available regarding such users. An embedding is formed with data representing the words, the audio features and the user features, and provided to a model trained to determine whether a media program is at risk of violating any rules. One or more actions are selected and executed or recommended based on a score generated by the model representing a level of risk that a rule has been, is being or will be violated.
1 . A first computer system, comprising:
at least one computer processor; and
at least one data store including one or more sets of instructions stored thereon that, when executed by the at least one computer processor, cause the first computer system to perform actions comprising:
receiving media content from at least one of:
a second computer system associated with a creator of a media program; or
a third computer system associated with a media source;
causing the media content to be transmitted to at least a fourth computer system associated with a listener in accordance with a first media program;
processing the media content to identify a first portion of the media content corresponding to words and a second portion of the media content not corresponding to the words;
recognizing at least one of the words within the first portion of the media content using at least a first machine learning model;
generating an audio feature based at least in part on the second portion of the media content using at least the first machine learning model;
determining information regarding the first media program, wherein the information regarding the first media program comprises at least one of:
a text-based description of the first media program;
a rating of the first media program;
an acoustic feature of the first media program;
an identifier of at least one of the plurality of listeners to the first media program; or
an image associated with the first media program;
generating an embedding based at least in part on:
the at least one of the words;
the audio feature; and
the information regarding the first media program;
providing the embedding to at least a second machine learning model as an input;
receiving an output from the second machine learning model in response to the input;
calculating a score representative of a risk that the first media program has violated at least one rule based at least in part on the output; and
terminating, based at least in part on the score and prior to completion of playback of the first media program, further transmission, by at least the first computing device, of the media content from the second computer system associated with the creator.
2 . The first computer system of claim 1 , wherein the at least one rule is a restriction on a selected type or category of content.
3 . The first computer system of claim 1 , further comprising:
creating a task for at least one human actor to evaluate at least one of the creator or the first media program based at least in part on the score.
4 . A computer-implemented method comprising:
under control of at least a first computer system configured with executable instructions,
receiving, by at least the first computer system from a second computer system, first data representing first media content of a first media program, wherein the second computer system is associated with a creator of the first media program;
transmitting, by at least the first computer system, second data to at least a third computer system, wherein the second data comprises at least some of the first data, and wherein the third computer system is associated with one of a plurality of listeners to the first media program;
receiving, by at least the first computer system from the third computer system, information representing at least one interaction with the third computer system by the one of the plurality of listeners during a playing of at least the first media content by the third computer system;
identifying, by at least the first computer system, a first portion of the second data representing a set of words spoken by at least one participant during the first media program;
determining, by at least the first computer system, a first feature based at least in part on a second portion of the second data;
determining, by at least the first computer system, a second feature based at least in part on the information representing the at least one interaction;
providing, by at least the first computer system, information regarding at least the set of words, the first feature and the second feature to a first machine learning model as a first input;
identifying, by at least the first computer system, a first output generated by the first machine learning model in response to the first input;
determining, by at least the first computer system, a first level of risk that at least one rule associated with the first media program has been violated based at least in part on the first output; and
terminating, based at least in part on the first level of risk and prior to completion of playback of the first media program, further transmission, by at least the first computer system, of the first data representing the first media content of the first media program to the third computer system.
5 . The computer-implemented method of claim 4 , wherein determining the first level of risk associated with the first media program based at least in part on the first output comprises:
calculating, by at least the first computer system, a first score representing a likelihood that the at least one rule associated with the first media program has been violated, wherein the indication of the first level of risk comprises the first score.
6 . The computer-implemented method of claim 5 , further comprising:
determining, by at least the first computer system, that the first score exceeds a predetermined threshold.
7 . The computer-implemented method of claim 5 , further comprising:
determining, by at least the first computer system that the first score does not exceed a predetermined threshold; and
in response to determining that the score does not exceed the predetermined threshold,
providing the warning message to the second computer system associated with the creator of the media program.
8 . The computer-implemented method of claim 5 , further comprising:
prior to transmitting the second data to at least the third computer system,
determining, by at least the first computer system, information regarding the first media program, wherein the information regarding the first media program comprises:
a text-based description of the first media program;
a rating of the first media program;
an acoustic feature of the first media program;
an identifier of at least one of the plurality of listeners to the first media program; or
an image associated with the first media program;
determining, by at least the first computer system, information regarding the creator, wherein the information regarding the creator comprises:
a history of violations of at least one rule associated with the creator; or
feedback received from the at least one of the plurality of listeners to one of the first media program or a second media program associated with the creator,
providing, by at least the first computer system, information regarding the first media program and the information regarding the creator to the first machine learning model as a second input;
identifying, by at least the first computer system, a second output generated by the first machine learning model in response to the second input;
determining, by at least the first computer system, a second level of risk that the at least one rule associated with the first media program has been violated based at least in part on the second output; and
calculating, by at least the first computer system, a second score representing a likelihood that the at least one rule associated with the first media program will be violated, wherein the indication of the second level of risk comprises the second score, and
wherein the first input further comprises the second score.
9 . The computer-implemented method of claim 4 , further comprising:
providing, by at least the first computer system, at least one of the first portion of the second data or the second portion of the second data to at least a second machine learning model as a second input;
identifying, by at least the first computer system, a second output generated by the second machine learning model in response to the second input; and
identifying, by at least the first computer system, at least one of the set of words or the first feature based at least in part on the second output.
10 . The computer-implemented method of claim 4 , wherein the first machine learning model is one of:
an artificial neural network; or
a bidirectional encoder representation from transformers.
11 . The computer-implemented method of claim 4 , further comprising:
determining, by at least the first computer system, information regarding at least one of the third computer system or the one of the plurality of listeners, wherein the information identifies at least one of:
a second media program to which the one of the plurality of listeners is a listener; and
an interaction with the third computer system by the one of the plurality of listeners during the second media program; and
calculating, by at least the first computer system, a score representative of a level of confidence in the at least one interaction,
wherein the first input comprises the score.
12 . The computer-implemented method of claim 4 , wherein the at least one interaction is one of:
an action for playing, pausing, stopping, advancing or rewinding media content;
a chat message received from one of the plurality of listeners;
a voice sample received from one of the plurality of listeners; or
an expression of at least one emotion of one of the plurality of listeners.
13 . The computer-implemented method of claim 4 , wherein the at least one rule is a content-based restriction.
14 . The computer-implemented method of claim 4 , wherein the first input is an embedding having a plurality of sets of values comprising:
a first set of values corresponding to at least some of the set of words;
a second set of values corresponding to the first feature; and
a third set of values corresponding to the second feature.
15 . The computer-implemented method of claim 4 , further comprising:
prior to providing at least the set of words, the first feature and the second feature to the first machine learning model as the first input,
training, by at least the first computer system, the first machine learning model based at least in part on a plurality of embeddings, wherein each of the plurality of embeddings represents a media program, and
wherein each of the embeddings comprises:
a first set of values corresponding to at least some of a set of words spoken or sung during one of a plurality of media programs;
a second set of values corresponding to an audio feature derived based at least in part on media content of the one of the plurality of media programs;
a third set of values corresponding to a user feature derived based at least in part on information regarding a listener to the one of the plurality of media programs; and
a fourth set of values indicating whether at least one rule was violated during the one of the plurality of media programs.
16 . The computer-implemented method of claim 4 , further comprising:
receiving, by at least the first computer system from a fourth computer system, fourth data representing at least one of:
an advertisement;
a media entity;
a news program;
a sports program; or
a weather report,
wherein the second data comprises at least some of the first data and at least some of the fourth data.
17 . The computer-implemented method of claim 4 , wherein the second computer system is at least one component of one of an automobile, a desktop computer, a laptop computer, a media player, a smartphone, a smart speaker, a tablet computer, a television, or a wristwatch.
18 . A method comprising:
under control of one or more computing systems configured with executable instructions,
determining, by at least a first computer system, information regarding a first media program, wherein the information regarding the first media program comprises:
a text-based description of the first media program;
a rating of the first media program;
an acoustic feature of the first media program;
an identifier of at least one of the plurality of listeners to the first media program; or
an image associated with the first media program;
determining, by at least the first computer system, information regarding a creator associated with the first media program, wherein the information regarding the creator comprises:
a history of violations of at least one rule associated with the creator; or
feedback received from the at least one of the plurality of listeners to one of the first media program or a second media program associated with the creator,
providing, by at least the first computer system, information regarding the first media program and the information regarding the creator to a first machine learning model as a first input;
identifying, by at least the first computer system, a first output generated by the first machine learning model in response to the first input;
calculating, by at least the first computer system, a first score representing a likelihood that at least one rule associated with the first media program will be violated based at least in part on the first output;
transmitting, by at least the first computer system, first data to at least a second computer system associated with one of a plurality of listeners to the first media program, wherein the first data comprises media content in accordance with the first media program;
identifying, by at least the first computer system, a first portion of the first data representing a set of words spoken by at least one participant during the first media program;
determining, by at least the first computer system, a first feature based at least in part on a second portion of the first data;
providing, by at least the first computer system, information regarding at least the set of words, the first feature and the first score to the first machine learning model as a second input;
identifying, by at least the first computer system, a second output generated by the first machine learning model in response to the second input;
calculating, by at least the first computer system, a second score representing a likelihood that at least one rule associated with the first media program has been violated based at least in part on the second output; and
terminating, based at least in part on the second score and prior to completion of playback of the first media program, further distribution by at least the first computing device of the first media program associated with the creator.
19 . The method of claim 18 , further comprising:
prior to transmitting the first data to at least the second computer system,
providing, by at least the first computer system, at least the first portion of the first data to at least a second machine learning model as a third input;
identifying, by at least the first computer system, a third output generated by the second machine learning model in response to the third input; and
identifying, by at least the first computer system, at least one of the set of words or the first feature based at least in part on the third output.
20 . The method of claim 18 , wherein the second input is an embedding having a plurality of sets of values comprising:
a first set of values corresponding to at least some of the set of words;
a second set of values corresponding to the first feature; and
a third set of values corresponding to the first score.