IP Library › Granted Patent US 12,647,492
Granted Patent B1
US 12,647,492 · App. 17/490,934 · Granted Jun 2, 2026

Automatically moderating content of media programs using multi-tiered machine learning solutions

Inventors: Juan Martin Borgnino (Marina Del Rey, CA); Sanjeev Kumar (Redmond, WA); Shenshen Liang (Menlo Park, CA); Ayman Mahfouz (Culver City, CA); Robert Eicher Simmering (Santa Ana, CA); Harshal Dilip Wanjari (Issaquah, WA); Muhammad Yahia (Anaheim, CA)
Assignee: Amazon Technologies, Inc.
H04L67/535G10L15/16H04H60/65H04L65/1083
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,647,492
App. No.
17/490,934
Filed
Sep 30, 2021
Granted
Jun 2, 2026
Kind
B1
Art Unit
2449
USPC
709/219
Abstract

As a media program is aired to listeners, a control system monitors audio data transmitted to the listeners and interactions received from the listeners to determine whether the media program has violated or may violate one or more rules. The audio data is processed to identify words expressed therein and features of the audio data. Additionally, features of users (e.g., a creator or any listeners or guests) may be calculated based on any information or data available regarding such users. An embedding is formed with data representing the words, the audio features and the user features, and provided to a model trained to determine whether a media program is at risk of violating any rules. One or more actions are selected and executed or recommended based on a score generated by the model representing a level of risk that a rule has been, is being or will be violated.

Claims (133)

1 . A first computer system, comprising:

at least one computer processor; and

at least one data store including one or more sets of instructions stored thereon that, when executed by the at least one computer processor, cause the first computer system to perform actions comprising:

receiving media content from at least one of:

a second computer system associated with a creator of a media program; or

a third computer system associated with a media source;

causing the media content to be transmitted to at least a fourth computer system associated with a listener in accordance with a first media program;

processing the media content to identify a first portion of the media content corresponding to words and a second portion of the media content not corresponding to the words;

recognizing at least one of the words within the first portion of the media content using at least a first machine learning model;

generating an audio feature based at least in part on the second portion of the media content using at least the first machine learning model;

determining information regarding the first media program, wherein the information regarding the first media program comprises at least one of:

 a text-based description of the first media program;

a rating of the first media program;

an acoustic feature of the first media program;

an identifier of at least one of the plurality of listeners to the first media program; or

an image associated with the first media program;

generating an embedding based at least in part on:

 the at least one of the words;

the audio feature; and

the information regarding the first media program;

providing the embedding to at least a second machine learning model as an input;

receiving an output from the second machine learning model in response to the input;

calculating a score representative of a risk that the first media program has violated at least one rule based at least in part on the output; and

terminating, based at least in part on the score and prior to completion of playback of the first media program, further transmission, by at least the first computing device, of the media content from the second computer system associated with the creator.

2 . The first computer system of claim 1 , wherein the at least one rule is a restriction on a selected type or category of content.

3 . The first computer system of claim 1 , further comprising:

creating a task for at least one human actor to evaluate at least one of the creator or the first media program based at least in part on the score.

4 . A computer-implemented method comprising:

under control of at least a first computer system configured with executable instructions,

receiving, by at least the first computer system from a second computer system, first data representing first media content of a first media program, wherein the second computer system is associated with a creator of the first media program;

transmitting, by at least the first computer system, second data to at least a third computer system, wherein the second data comprises at least some of the first data, and wherein the third computer system is associated with one of a plurality of listeners to the first media program;

receiving, by at least the first computer system from the third computer system, information representing at least one interaction with the third computer system by the one of the plurality of listeners during a playing of at least the first media content by the third computer system;

identifying, by at least the first computer system, a first portion of the second data representing a set of words spoken by at least one participant during the first media program;

determining, by at least the first computer system, a first feature based at least in part on a second portion of the second data;

determining, by at least the first computer system, a second feature based at least in part on the information representing the at least one interaction;

providing, by at least the first computer system, information regarding at least the set of words, the first feature and the second feature to a first machine learning model as a first input;

identifying, by at least the first computer system, a first output generated by the first machine learning model in response to the first input;

determining, by at least the first computer system, a first level of risk that at least one rule associated with the first media program has been violated based at least in part on the first output; and

terminating, based at least in part on the first level of risk and prior to completion of playback of the first media program, further transmission, by at least the first computer system, of the first data representing the first media content of the first media program to the third computer system.

5 . The computer-implemented method of claim 4 , wherein determining the first level of risk associated with the first media program based at least in part on the first output comprises:

calculating, by at least the first computer system, a first score representing a likelihood that the at least one rule associated with the first media program has been violated, wherein the indication of the first level of risk comprises the first score.

6 . The computer-implemented method of claim 5 , further comprising:

determining, by at least the first computer system, that the first score exceeds a predetermined threshold.

7 . The computer-implemented method of claim 5 , further comprising:

determining, by at least the first computer system that the first score does not exceed a predetermined threshold; and

in response to determining that the score does not exceed the predetermined threshold,

providing the warning message to the second computer system associated with the creator of the media program.

8 . The computer-implemented method of claim 5 , further comprising:

prior to transmitting the second data to at least the third computer system,

determining, by at least the first computer system, information regarding the first media program, wherein the information regarding the first media program comprises:

a text-based description of the first media program;

a rating of the first media program;

an acoustic feature of the first media program;

an identifier of at least one of the plurality of listeners to the first media program; or

an image associated with the first media program;

determining, by at least the first computer system, information regarding the creator, wherein the information regarding the creator comprises:

a history of violations of at least one rule associated with the creator; or

feedback received from the at least one of the plurality of listeners to one of the first media program or a second media program associated with the creator,

providing, by at least the first computer system, information regarding the first media program and the information regarding the creator to the first machine learning model as a second input;

identifying, by at least the first computer system, a second output generated by the first machine learning model in response to the second input;

determining, by at least the first computer system, a second level of risk that the at least one rule associated with the first media program has been violated based at least in part on the second output; and

calculating, by at least the first computer system, a second score representing a likelihood that the at least one rule associated with the first media program will be violated, wherein the indication of the second level of risk comprises the second score, and

wherein the first input further comprises the second score.

9 . The computer-implemented method of claim 4 , further comprising:

providing, by at least the first computer system, at least one of the first portion of the second data or the second portion of the second data to at least a second machine learning model as a second input;

identifying, by at least the first computer system, a second output generated by the second machine learning model in response to the second input; and

identifying, by at least the first computer system, at least one of the set of words or the first feature based at least in part on the second output.

10 . The computer-implemented method of claim 4 , wherein the first machine learning model is one of:

an artificial neural network; or

a bidirectional encoder representation from transformers.

11 . The computer-implemented method of claim 4 , further comprising:

determining, by at least the first computer system, information regarding at least one of the third computer system or the one of the plurality of listeners, wherein the information identifies at least one of:

a second media program to which the one of the plurality of listeners is a listener; and

an interaction with the third computer system by the one of the plurality of listeners during the second media program; and

calculating, by at least the first computer system, a score representative of a level of confidence in the at least one interaction,

wherein the first input comprises the score.

12 . The computer-implemented method of claim 4 , wherein the at least one interaction is one of:

an action for playing, pausing, stopping, advancing or rewinding media content;

a chat message received from one of the plurality of listeners;

a voice sample received from one of the plurality of listeners; or

an expression of at least one emotion of one of the plurality of listeners.

13 . The computer-implemented method of claim 4 , wherein the at least one rule is a content-based restriction.

14 . The computer-implemented method of claim 4 , wherein the first input is an embedding having a plurality of sets of values comprising:

a first set of values corresponding to at least some of the set of words;

a second set of values corresponding to the first feature; and

a third set of values corresponding to the second feature.

15 . The computer-implemented method of claim 4 , further comprising:

prior to providing at least the set of words, the first feature and the second feature to the first machine learning model as the first input,

training, by at least the first computer system, the first machine learning model based at least in part on a plurality of embeddings, wherein each of the plurality of embeddings represents a media program, and

wherein each of the embeddings comprises:

a first set of values corresponding to at least some of a set of words spoken or sung during one of a plurality of media programs;

a second set of values corresponding to an audio feature derived based at least in part on media content of the one of the plurality of media programs;

a third set of values corresponding to a user feature derived based at least in part on information regarding a listener to the one of the plurality of media programs; and

a fourth set of values indicating whether at least one rule was violated during the one of the plurality of media programs.

16 . The computer-implemented method of claim 4 , further comprising:

receiving, by at least the first computer system from a fourth computer system, fourth data representing at least one of:

an advertisement;

a media entity;

a news program;

a sports program; or

a weather report,

wherein the second data comprises at least some of the first data and at least some of the fourth data.

17 . The computer-implemented method of claim 4 , wherein the second computer system is at least one component of one of an automobile, a desktop computer, a laptop computer, a media player, a smartphone, a smart speaker, a tablet computer, a television, or a wristwatch.

18 . A method comprising:

under control of one or more computing systems configured with executable instructions,

determining, by at least a first computer system, information regarding a first media program, wherein the information regarding the first media program comprises:

a text-based description of the first media program;

a rating of the first media program;

an acoustic feature of the first media program;

an identifier of at least one of the plurality of listeners to the first media program; or

an image associated with the first media program;

determining, by at least the first computer system, information regarding a creator associated with the first media program, wherein the information regarding the creator comprises:

a history of violations of at least one rule associated with the creator; or

feedback received from the at least one of the plurality of listeners to one of the first media program or a second media program associated with the creator,

providing, by at least the first computer system, information regarding the first media program and the information regarding the creator to a first machine learning model as a first input;

identifying, by at least the first computer system, a first output generated by the first machine learning model in response to the first input;

calculating, by at least the first computer system, a first score representing a likelihood that at least one rule associated with the first media program will be violated based at least in part on the first output;

transmitting, by at least the first computer system, first data to at least a second computer system associated with one of a plurality of listeners to the first media program, wherein the first data comprises media content in accordance with the first media program;

identifying, by at least the first computer system, a first portion of the first data representing a set of words spoken by at least one participant during the first media program;

determining, by at least the first computer system, a first feature based at least in part on a second portion of the first data;

providing, by at least the first computer system, information regarding at least the set of words, the first feature and the first score to the first machine learning model as a second input;

identifying, by at least the first computer system, a second output generated by the first machine learning model in response to the second input;

calculating, by at least the first computer system, a second score representing a likelihood that at least one rule associated with the first media program has been violated based at least in part on the second output; and

terminating, based at least in part on the second score and prior to completion of playback of the first media program, further distribution by at least the first computing device of the first media program associated with the creator.

19 . The method of claim 18 , further comprising:

prior to transmitting the first data to at least the second computer system,

providing, by at least the first computer system, at least the first portion of the first data to at least a second machine learning model as a third input;

identifying, by at least the first computer system, a third output generated by the second machine learning model in response to the third input; and

identifying, by at least the first computer system, at least one of the set of words or the first feature based at least in part on the third output.

20 . The method of claim 18 , wherein the second input is an embedding having a plurality of sets of values comprising:

a first set of values corresponding to at least some of the set of words;

a second set of values corresponding to the first feature; and

a third set of values corresponding to the first score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2021
From: BORGNINO, JUAN MARTIN; KUMAR, SANJEEV; LIANG, SHENSHEN; MAHFOUZ, AYMAN; SIMMERING, ROBERT EICHER; WANJARI, HARSHAL DILIP; YAHIA, MUHAMMAD
To: AMAZON TECHNOLOGIES, INC.
Reel/Frame 057660/0950 →
References Cited (150)
US 8023800B2 · Concotelli · 2011 [cited by applicant]
US 8112720B2 · Curtis · 2012 [cited by applicant]
US 8560683B2 · Funk et al. · 2013 [cited by applicant]
US 8572243B2 · Funk et al. · 2013 [cited by applicant]
US 8768782B1 · Myslinski · 2014 [cited by applicant]
US 8850301B1 · Rose · 2014 [cited by applicant]
US 9003032B2 · Funk et al. · 2015 [cited by applicant]
US 9369740B1 · Funk et al. · 2016 [cited by applicant]
US 9613636B2 · Gibbon et al. · 2017 [cited by applicant]
US 9706253B1 · Funk et al. · 2017 [cited by applicant]
US 9729596B2 · Sanghavi et al. · 2017 [cited by applicant]
US 9781491B2 · Wilson · 2017 [cited by applicant]
US 9872069B1 · Funk et al. · 2018 [cited by applicant]
US 10083169B1 · Ghosh · 2018 [cited by examiner]
US 10091547B2 · Sheppard et al. · 2018 [cited by applicant]
US 10110952B1 · Gupta et al. · 2018 [cited by applicant]
US 10135887B1 · Esser et al. · 2018 [cited by applicant]
US 10140364B1 · Diamondstein · 2018 [cited by applicant]
US 10178422B1 · Panchaksharaiah et al. · 2019 [cited by applicant]
US 10178442B2 · Shkedi · 2019 [cited by applicant]
US 10313726B2 · Woods et al. · 2019 [cited by applicant]
US 10356476B2 · Dharmaji · 2019 [cited by applicant]
US 10432335B2 · Bretherton · 2019 [cited by applicant]
US 10489395B2 · Akkur et al. · 2019 [cited by applicant]
US 10685050B2 · Krishna et al. · 2020 [cited by applicant]
US 10698906B2 · Hargreaves et al. · 2020 [cited by applicant]
US 10719837B2 · Kolowich et al. · 2020 [cited by applicant]
US 10769678B2 · Li · 2020 [cited by applicant]
US 10846330B2 · Shilo · 2020 [cited by applicant]
US 10893329B1 · Trim et al. · 2021 [cited by applicant]
US 10985853B2 · Bretherton · 2021 [cited by applicant]
US 10986064B2 · Siegel et al. · 2021 [cited by applicant]
US 10997240B1 · Aschner et al. · 2021 [cited by applicant]
US 11431660B1 · Leeds et al. · 2022 [cited by applicant]
US 11451863B1 · Benjamin et al. · 2022 [cited by applicant]
US 11463772B1 · Wanjari et al. · 2022 [cited by applicant]
US 11521179B1 · Shetty · 2022 [cited by applicant]
US 11580982B1 · Karnawat et al. · 2023 [cited by applicant]
US 11586344B1 · Balagurunathan et al. · 2023 [cited by applicant]
US 20020042920A1 · Thomas et al. · 2002 [cited by applicant]
US 20020056087A1 · Berezowski et al. · 2002 [cited by applicant]
US 20060268667A1 · Jellison et al. · 2006 [cited by applicant]
US 20070124756A1 · Covell et al. · 2007 [cited by applicant]
US 20070271518A1 · Tischer et al. · 2007 [cited by applicant]
US 20070271580A1 · Tischer et al. · 2007 [cited by applicant]
US 20080086742A1 · Aldrey et al. · 2008 [cited by applicant]
US 20090044217A1 · Lutterbach et al. · 2009 [cited by applicant]
US 20090076917A1 · Jablokov et al. · 2009 [cited by applicant]
US 20090100098A1 · Feher et al. · 2009 [cited by applicant]
US 20090254934A1 · Grammens · 2009 [cited by applicant]
US 20100088187A1 · Courtney et al. · 2010 [cited by applicant]
US 20100280641A1 · Harkness et al. · 2010 [cited by applicant]
US 20110063406A1 · Albert et al. · 2011 [cited by applicant]
US 20110067044A1 · Albo · 2011 [cited by applicant]
US 20120040604A1 · Amidon et al. · 2012 [cited by applicant]
US 20120191774A1 · Bhaskaran et al. · 2012 [cited by applicant]
US 20120304206A1 · Roberts et al. · 2012 [cited by applicant]
US 20120311444A1 · Chaudhri · 2012 [cited by applicant]
US 20120311618A1 · Blaxland · 2012 [cited by applicant]
US 20120331168A1 · Chen · 2012 [cited by applicant]
US 20130074109A1 · Skelton et al. · 2013 [cited by applicant]
US 20130247081A1 · Vinson et al. · 2013 [cited by applicant]
US 20130253934A1 · Parekh et al. · 2013 [cited by applicant]
US 20140019225A1 · Guminy et al. · 2014 [cited by applicant]
US 20140040494A1 · Deinhard et al. · 2014 [cited by applicant]
US 20140068432A1 · Kucharz et al. · 2014 [cited by applicant]
US 20140073236A1 · Iyer · 2014 [cited by applicant]
US 20140108531A1 · Klau · 2014 [cited by applicant]
US 20140123191A1 · Hahn et al. · 2014 [cited by applicant]
US 20140228010A1 · Barbulescu et al. · 2014 [cited by applicant]
US 20140325557A1 · Evans et al. · 2014 [cited by applicant]
US 20140372179A1 · Ju et al. · 2014 [cited by applicant]
US 20150095014A1 · Marimuthu · 2015 [cited by applicant]
US 20150163184A1 · Kanter et al. · 2015 [cited by applicant]
US 20150242068A1 · Losey · 2015 [cited by examiner]
US 20150248798A1 · Howe et al. · 2015 [cited by applicant]
US 20150289021A1 · Miles · 2015 [cited by applicant]
US 20150319472A1 · Kotecha et al. · 2015 [cited by applicant]
US 20150326922A1 · Givon et al. · 2015 [cited by applicant]
US 20160027196A1 · Schiffer et al. · 2016 [cited by applicant]
US 20160093289A1 · Pollet · 2016 [cited by applicant]
US 20160188728A1 · Gill et al. · 2016 [cited by applicant]
US 20160217488A1 · Ward et al. · 2016 [cited by applicant]
US 20160266781A1 · Dandu et al. · 2016 [cited by applicant]
US 20160293036A1 · Niemi et al. · 2016 [cited by applicant]
US 20160330529A1 · Byers · 2016 [cited by applicant]
US 20170127136A1 · Roberts et al. · 2017 [cited by applicant]
US 20170164357A1 · Fan et al. · 2017 [cited by applicant]
US 20170193531A1 · Fatourechi et al. · 2017 [cited by applicant]
US 20170213248A1 · Jing et al. · 2017 [cited by applicant]
US 20170289617A1 · Song et al. · 2017 [cited by applicant]
US 20170329466A1 · Krenkler et al. · 2017 [cited by applicant]
US 20170366854A1 · Puntambekar et al. · 2017 [cited by applicant]
US 20180025078A1 · Quennesson · 2018 [cited by applicant]
US 20180035142A1 · Rao et al. · 2018 [cited by applicant]
US 20180205797A1 · Faulkner · 2018 [cited by applicant]
US 20180227632A1 · Rubin et al. · 2018 [cited by applicant]
US 20180255114A1 · Dharmaji · 2018 [cited by applicant]
US 20180293221A1 · Finkelstein et al. · 2018 [cited by applicant]
US 20180322411A1 · Wang et al. · 2018 [cited by applicant]
US 20180367229A1 · Gibson et al. · 2018 [cited by applicant]
US 20190065610A1 · Singh · 2019 [cited by applicant]
US 20190132636A1 · Gupta et al. · 2019 [cited by applicant]
US 20190156196A1 · Zoldi et al. · 2019 [cited by applicant]
US 20190171762A1 · Luke et al. · 2019 [cited by applicant]
US 20190273570A1 · Bretherton · 2019 [cited by applicant]
US 20190327103A1 · Niekrasz · 2019 [cited by applicant]
US 20190385600A1 · Kim · 2019 [cited by applicant]
US 20200021888A1 · de Mello Brandao · 2020 [cited by examiner]
US 20200106885A1 · Koster et al. · 2020 [cited by applicant]
US 20200160458A1 · Bodin et al. · 2020 [cited by applicant]
US 20200226418A1 · Dorai-Raj · 2020 [cited by examiner]
US 20200279553A1 · McDuff et al. · 2020 [cited by applicant]
US 20200364727A1 · Scott-Green et al. · 2020 [cited by applicant]
US 20210090224A1 · Zhou et al. · 2021 [cited by applicant]
US 20210104245A1 · Alas et al. · 2021 [cited by applicant]
US 20210105149A1 · Roedel et al. · 2021 [cited by applicant]
US 20210125054A1 · Banik et al. · 2021 [cited by applicant]
US 20210160588A1 · Joseph et al. · 2021 [cited by applicant]
US 20210210102A1 · Huh et al. · 2021 [cited by applicant]
US 20210217413A1 · Tushinskiy et al. · 2021 [cited by applicant]
US 20210232577A1 · Ogawa et al. · 2021 [cited by applicant]
US 20210256086A1 · Askarian et al. · 2021 [cited by applicant]
US 20210281925A1 · Shaikh et al. · 2021 [cited by applicant]
US 20210366462A1 · Yang et al. · 2021 [cited by applicant]
US 20210407520A1 · Neckermann et al. · 2021 [cited by applicant]
US 20220038783A1 · Wee · 2022 [cited by applicant]
US 20220038790A1 · Duan et al. · 2022 [cited by applicant]
US 20220159377A1 · Wilberding et al. · 2022 [cited by applicant]
US 20220223286A1 · Lach et al. · 2022 [cited by applicant]
US 20220230632A1 · Maitra et al. · 2022 [cited by applicant]
US 20220254348A1 · Tay et al. · 2022 [cited by applicant]
US 20220286748A1 · Dyer et al. · 2022 [cited by applicant]
US 20220369034A1 · Kumar et al. · 2022 [cited by applicant]
US 20230036192A1 · Alakoye · 2023 [cited by applicant]
US 20230085683A1 · Turner · 2023 [cited by applicant]
US 20230217195A1 · Poltorak · 2023 [cited by applicant]
AU 2013204532B2 · 2014 [cited by applicant]
CA 2977959A1 · 2016 [cited by applicant]
CN 104813305A · 2015 [cited by applicant]
KR 20170079496A · 2017 [cited by applicant]
WO 2019089028A1 · 2019 [cited by applicant]
Github, “Spotify iOS SDK,” GitHub.com, GitHub Inc. and GitHub B.V., Feb. 17, 2021, available at URL: https://github.com/spotify/ios-sdk#how-do-app-remote-calls-work, 10 pages. [cited by applicant]
Stack Overflow, “Audio mixing of Spotify tracks in IOS app,” stackoverflow.com, Stack Overflow Network, Jul. 2012, available at URL: https://stackoverflow.com/questions/11396348/audio-mixing-of-spotify-tracks-in-ios-app… [cited by applicant]
Tengeh, R. K., & Udoakpan, N. (2021). Over-the-Top Television Services and Changes in Consumer Viewing Patterns in South Africa. Management Dynamics in the Knowledge Economy. 9(2), 257-277. DOI 10.2478/mdke-2021-0018 IS… [cited by applicant]
Arora, S. et al., “A Practical Algorithm for Topic Modeling with Provable Guarantees,” Proceedings in the 30th International Conference on Machine Learning, JMLR: W&CP vol. 28, published 2013 (Year: 2013), 9 pages. [cited by applicant]
Hoegen, Rens, et al. “An End-to-End Conversational Style Matching Agent.” Proceedings of the 19th ACM International Conference on Intelligent Virtual Agents. 2019, pp. 1-8. (Year: 2019). [cited by applicant]
B. Subin, “Spotify for Android Tests New Floating Mini Player UI / Beebom,” URL: https://beebom.com/spotify-tests-new-mini-player-android/, retrieved on Aug. 26, 2023, 3 pages. [cited by applicant]
Matt Ellis, “Desktop vs. mobile app design: how to optimize your user experience—99 designs,” URL: https://99designs.com/blog/web-digital/desktop-vs-mobile-app-design/, retrieved Aug. 26, 2023, 12 pages. [cited by applicant]
Salesforce, “Introducing a profile page as sleek as a Tableau Public viz,” https://www.tableau.com/, Tableau Software, LLC, a Salesforce Company, Jul. 21, 2021. Accessed Aug. 31, 2023. URL: https://www.tableau.com/blog/… [cited by applicant]