Systems and methods for adjusting dubbed speech based on context of a scene
Systems and methods are disclosed herein for detecting dubbed speech in a media asset and receiving metadata corresponding to the media asset. The systems and methods may determine a plurality of scenes in the media asset based on the metadata, retrieve a portion of the dubbed speech corresponding to the first scene, and process the retrieved portion of the dubbed speech corresponding to the first scene to identify a speech characteristic of a character featured in the first scene. Further, the systems and methods may determine whether the speech characteristic of the character featured in the first scene matches the context of the first scene, and if the match fails, perform a function to adjust the portion of the dubbed speech so that the speech characteristic of the character featured in the first scene matches the context of the first scene.
1 . A method comprising:
analyzing metadata of a media asset to determine a genre of the media asset;
detecting a first text in the media asset and a second text in the media asset;
receiving a selection of a first emotion associated with the first text;
receiving a selection of a second emotion associated with the second text;
identifying a first portion of the media asset associated with the first text and a second portion of the media asset associated with the second text;
modifying a first waveform of a first audio for the first portion of the media asset based on: (a) the genre of the media asset, (b) the first text, and (c) the first emotion; and
modifying a second waveform of a second audio for the second portion of the media asset based on: (a) the genre of the media asset, (b) the second text, and (c) the second emotion.
2 . The method of claim 1 , wherein the receiving input from the user device selecting the first emotion comprises:
generating for display on a user device a first plurality of emotions associated with the first text;
receiving a user interface selection of the first emotion from the first plurality of emotions; and wherein the receiving input from the user device selecting the second emotion comprises:
generating for display on a user device a second plurality of emotions associated with the second text; and
receiving a user selection of the first emotion from the second plurality of emotions.
3 . The method of claim 1 , wherein the receiving input from the user device selecting the first emotion with the first text and the receiving input from the user device selecting the second emotion with the second text comprises:
receiving a first input from an input interface, wherein the first input indicates the first selection of the first emotion; and
receiving a second input from an input interface, wherein the second input indicates the second selection of the second emotion.
4 . The method of claim 1 , wherein the modifying of the first waveform of the first audio for the first portion of the media asset and the second waveform of the second audio for the second portion of the media asset comprises:
retrieving a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset;
generating the first audio for the first portion of the media asset based on the first text and the second audio for the second portion of the media asset based on the second text; and
modifying the first waveform of the first audio by inserting the first set of vocal characteristics into the first audio; and
modifying the second waveform of the second audio by inserting the second set of vocal characteristics into the second audio.
5 . The method of claim 4 , wherein the first vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof, and wherein the second vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof.
6 . The method of claim 1 , wherein the first emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof, and wherein the second emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.
7 . The method of claim 1 , wherein the modifying of the first waveform of the first audio for the first portion of the media asset based on the first text and the first emotion comprises:
retrieving a first personality metadata corresponding to the first portion of the media asset;
identifying a first speech characteristic from the first personality metadata; and
adjusting the audio for the first portion of the media asset from the identified first speech characteristic.
8 . The method of claim 1 , wherein the modifying of the second waveform of the second audio for the second portion of the media asset based on the second text and the second emotion comprise:
retrieving a second personality metadata corresponding to the second portion of the media asset;
identifying a second speech characteristic from the second personality metadata; and
adjusting the second audio for the second portion of the media asset from the identified second speech characteristic.
9 . A method comprising:
analyzing metadata of a media asset to determine a genre of the media asset;
detecting a first text in the media asset and a second text in the media asset;
receiving a selection of a first emotion associated with the first text;
receiving a selection of a second emotion associated with the second text;
identifying a first portion of the media asset associated with the first text and a second portion of the media asset associated with the second text;
modifying a first audio for the first portion of the media asset based on: (a) the genre of the media asset, (b) the first text, and (c) the first emotion, wherein the modifying the first audio for the first portion of the media asset comprises:
retrieving a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset; and
in response to identifying that a first vocal characteristic from the first set of vocal characteristics does not match a second vocal characteristic from a second set of vocal characteristics, adjusting the first vocal characteristic from the first set of vocal characteristics to match the second vocal characteristic from the second set of vocal characteristics; and
modifying a second audio for the second portion of the media asset based on: (a) the genre of the media asset, (b) the second text, and (c) the second emotion.
10 . A system comprising:
input/output circuitry configured to:
receive a selection of a first emotion associated with the first text; and
receive a selection of a second emotion associated with the second text; and
control circuitry configured to:
analyze metadata of a media asset to determine a genre of the media asset;
detect a first text in the media asset and a second text in the media asset;
identify a first portion of the media asset associated with the first text and a second portion of the media asset associated with the second text;
modify a first waveform of a first audio for the first portion of the media asset based on: (a) the genre of the media asset, (b) the first text, and (c) the first emotion; and
modify a second waveform of a second audio for the second portion of the media asset based on: (a) the genre of the media asset, (b) the second text, and (c) the second emotion.
11 . The system of claim 10 , wherein the input/output circuitry configured to receive input from the user device selecting the first emotion is configured to:
generate for display on a user device a first plurality of emotions associated with the first text;
receive a user interface selection of the first emotion from the first plurality of emotions;
and wherein the receiving input from the user device selecting the second emotion comprises:
generate for display on a user device a second plurality of emotions associated with the second text; and
receive a user selection of the first emotion from the second plurality of emotions.
12 . The system of claim 10 , wherein the control circuitry configured to modify the first waveform of the first audio for the first portion of the media asset is configured to:
retrieve a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset; and
in response to identifying that a first vocal characteristic from the first set of vocal characteristics does not match a second vocal characteristic from a second set of vocal characteristics, adjust the first vocal characteristic from the first set of vocal characteristics to match the second vocal characteristic from the second set of vocal characteristics.
13 . The system of claim 10 , wherein the input/output circuitry configured to receive input from the user device selecting the first emotion with the first text and to receive input from the user device selecting the second emotion with the second text is configured to:
receive a first input from an input interface, wherein the first input indicates the first selection of the first emotion; and
receive a second input from an input interface, wherein the second input indicates the second selection of the second emotion.
14 . The system of claim 10 , wherein the control circuitry configured to modify the first waveform of the first audio for the first portion of the media asset and the second waveform of the second audio for the second portion of the media asset is configured to:
retrieve a first set of vocal characteristics corresponding to the first portion of the media asset and a second set of vocal characteristics corresponding to the second portion of the media asset;
generate the first audio for the first portion of the media asset based on the first text and the second audio for the second portion of the media asset based on the second text; and
modify the first audio by inserting the first set of vocal characteristics into the first audio; and
modify the second audio by inserting the second set of vocal characteristics into the second audio.
15 . The system of claim 14 , wherein the first vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof, and wherein the second vocal characteristic is one of a pitch, pause, rate, and rhythm or combinations thereof.
16 . The system of claim 10 , wherein the first emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof, and wherein the second emotion is one of an angry, calm, gentle, loving, alarmed, scared, comic, confused, excited, doubtful, urgent, and accusing or combinations thereof.
17 . The system of claim 10 , wherein the control circuitry configured to modify the first waveform of the first audio for the first portion of the media asset based on the first text and the first emotion is configured to:
retrieve a first personality metadata corresponding to the first portion of the media asset;
identify a first speech characteristic from the first personality metadata; and
adjust the audio for the first portion of the media asset from the identified first speech characteristic.
18 . The system of claim 10 , wherein the control circuitry configured to modify the second waveform of the second audio for the second portion of the media asset based on the second text and the second emotion is configured to:
retrieve a second personality metadata corresponding to the second portion of the media asset;
identify a second speech characteristic from the second personality metadata; and
adjust the second audio for the second portion of the media asset from the identified second speech characteristic.