Simulating crowd noise for live events through emotional analysis of distributed inputs
Methods and systems are provided for generating crowd noise related to a media event being presented using a cloud service is provided. The method includes receiving audio data captured from a viewer of the media event. The method includes processing the audio data to identify utterances of the viewer. In one embodiment, features of the utterances are classified to build a reaction model for identifying reaction states of the viewer. The method includes producing a soundscape for the crowd noise, the soundscape blends together audio of generic crowd noise related to the media event and audio corresponding to one or more of said reaction states of the viewer. In one embodiment, the soundscape is output to a speaker associated with presentation of the media event to the viewer.
1 . A method comprising:
receiving audio data captured from a viewer of a media event;
processing the audio data to identify utterances of the viewer, wherein features of the utterances are classified to build a reaction model for identifying reaction states of the viewer, the reaction states including one or more emotion types, wherein the one or more emotion types are: (I) associated with the utterances and (II) scored by the reaction model, the scores corresponding to an intensity associated with corresponding utterances;
producing a soundscape for crowd noise, wherein the soundscape blends together first audio of generic crowd noise related to the media event and second audio that simulates noises based at least in part on one or more of said reaction states of the viewer, wherein the second audio is generated from one or more sounds of a plurality of sounds, the one or more sounds corresponding to a meaning of the utterances of the viewer and selected from a database based at least in part on the score; and
transmitting the soundscape for output at a speaker.
2 . The method of claim 1 , further comprising:
the media event is being presented to a plurality of additional viewers;
receiving audio data captured from said plurality of additional viewers;
processing the audio data of the additional viewers to identify utterances of the additional viewers, wherein the reaction model is used to identify reaction states of the additional viewers; and
augmenting the produced soundscape for the crowd noise to blend additional audio corresponding to said reaction states of said additional viewers.
3 . The method of claim 2 , wherein the soundscape, as augmented, is output to respective speakers associated with presentation of the media event to the additional viewers.
4 . The method of claim 1 , wherein the second audio is customizable based on received preferences of the viewer.
5 . The method of claim 1 , further comprising:
processing additional reaction states of other viewers of the media event;
identifying audio corresponding to the additional reaction states; and
augmenting the produced soundscape to additionally include blending of said audio corresponding to the additional reaction states of the other viewers.
6 . The method of claim 5 , wherein the soundscape is presented to said viewer and one or more of said other viewers as output to speakers when viewing said media event.
7 . The method of claim 1 , wherein the media event is a live event or an event being viewed as a group by the viewer and other viewers.
8 . The method of claim 1 , wherein the second audio is generated based on one or more semantic equivalents of the utterances without replicating the utterances of the viewer.
9 . The method of claim 1 , wherein the second audio is accessed from a database of pre-recorded audio files, wherein said pre-recorded audio files are tagged with an emotional score and used for selecting the audio corresponding to one or more of said reaction states of the viewer.
10 . The method of claim 1 , wherein the reaction model implements a machine learning engine that is configured to identify the features of the utterances to classify attributes of the viewer, wherein the attributes of the viewer are used to identify the reaction states of the viewer.
11 . A system comprising:
one or more processors; and
one or more memories storing instructions that, upon execution by the one or more processors, configure the system to:
receive audio data captured from a plurality of viewers of a media event;
process the audio data to identify utterances of the plurality of viewers, wherein features of the utterances are classified to build a reaction model for identifying reaction states of the plurality of viewers, wherein the reaction states include one or more emotion types, wherein the emotion types are: (I) associated with the utterances and (II) scored by the reaction model, the scores corresponding to an intensity associated with corresponding utterances; and
produce a soundscape for crowd noise, wherein the soundscape blends together first audio of generic crowd noise related to the media event and second audio that simulates noises that correspond to one or more of said reaction states of the plurality of viewers, wherein the second audio corresponds to a semantic equivalent of the utterances of a viewer of the plurality of viewers, the second audio generated from one or more sounds of a plurality of sounds selected from a database based at least in part on the score.
12 . The system of claim 11 further comprising transmitting the soundscape for output at a speaker.
13 . The system of claim 11 , wherein the soundscape is customizable based on received preferences of the plurality of viewers.
14 . The system of claim 11 , wherein the media event is a live event or a recorded event being viewed by the plurality of viewers as a group or separately in different geographical locations.