Information processing device, information processing method, and program
A device and method are provided that adjust voices of a plurality of sound sources included in distribution content from an information processing device and make it easier to hear sound of each sound source by a reception terminal that receives and reproduces distribution content. A first output voice adjustment unit executes adjustment processing on an output voice of each of the plurality of sound sources, and content including synthesized voice data obtained by synthesizing the output voices corresponding to the sound sources adjusted by the first output voice adjustment unit is output. The first output voice adjustment unit executes output voice adjustment processing for matching a maximum value of a volume level corresponding to a frequency of the output voice of each sound source to a target level. Moreover, a second output voice adjustment unit executes the output voice adjustment processing in accordance with a content type or scene.
1 . An information processing device comprising:
circuitry configured to
input an output voice of each of a plurality of sound sources and execute adjustment processing on the output voice of each of a first sound source, a second sound source, and a third sound source of the plurality of sound sources, wherein
the output voice associated with the first sound source corresponds to a distribution user voice input via a microphone,
the output voice associated with the second sound source corresponds to operation of an application, and
the output voice associated with the third sound source corresponds to a viewing user comment input via a network;
synthesis synthesize the output voice corresponding to each of the plurality of sound sources having been adjusted to generate synthesized voice data; and
output content including the synthesized voice data, wherein
the adjustment processing includes:
analyzing a volume level corresponding to a frequency regarding the output voice of each of the first, second, and third sound sources of the plurality of sound sources, and
setting a maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level for the output of the content that includes the synthesized voice data,
wherein output volume corresponding to the synthesized voice data is limited to no greater than the target level for each of the first, second, and third sound sources of the plurality of sound sources.
2 . The information processing device according to claim 1 , wherein
the output voice adjustment processing is executed to match the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a single target level common to the first, second, and third sound sources of the plurality of sound sources.
3 . The information processing device according to claim 1 , wherein
the output voice adjustment processing is executed for matching the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level specific for each of the first, second, and third sound sources.
4 . The information processing device according to claim 1 , wherein the output voice adjustment processing is executed to reduce a difference in the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources.
5 . The information processing device according to claim 1 , wherein the circuitry is configured to:
execute the output voice adjustment processing on the output voice of each of the first, second, and third sound sources in accordance with a type of the content or a scene of the content output via the circuitry.
6 . The information processing device according to claim 5 , wherein the circuitry is configured to
select one of the first, second, or third sound sources that is an execution target of the output voice adjustment processing in accordance with the type of the content output via the circuitry, and
execute the output voice adjustment processing only on the output voice of the selected one sound source from among the first, second, and third sound sources.
7 . The information processing device according to claim 5 , wherein the circuitry is configured to
in a case where the type of the content output via the circuitry is game content execute output voice adjustment processing to emphasize the distribution user voice.
8 . The information processing device according to claim 5 , wherein the circuitry is configured to
in a case where the type of the content output via the circuitry is music content, execute output voice adjustment processing to emphasize music reproduced sound of the music content.
9 . The information processing device according to claim 5 , wherein the circuitry is configured to
select one of the first, second, or third sound sources that is an execution target of the output voice adjustment processing in accordance with the scene of the content output via the circuitry, and
execute the output voice adjustment processing only on the output voice of the selected one sound source from among the first, second, and third sound sources.
10 . The information processing device according to claim 5 , wherein the circuitry is configured to
determine the scene of the content output via the circuitry,
analyze attribute information of the application being executed by the information processing device,
display information of a display,
execute the voice adjustment processing in accordance with the determined scene determined by the circuitry.
11 . The information processing device according to claim 1 , wherein the output voice associated with the third sound source corresponding to the viewing user comment is based on a text input via the network.
12 . The information processing device according to claim 1 , wherein the application is one of a gaming application or a musical content hosting application.
13 . An information processing method comprising:
inputting an output voice of each of a plurality of sound sources and execute adjustment processing on the output voice of each of a first sound source, a second sound source, and a third sound source of the plurality of sound sources, wherein
the output voice associated with the first sound source corresponds to a distribution user voice input via a microphone,
the output voice associated with the second sound source corresponds to operation of an application, and
the output voice associated with the third sound source corresponds to a viewing user comment input via a network;
synthesizing the output voice corresponding to each of the plurality of sound sources having been adjusted to generate synthesized voice data; and
outputting content including the synthesized voice data, wherein
said adjustment processing includes:
analyzing a volume level corresponding to a frequency regarding the output voice of each of the first, second, and third sound sources of the plurality of sound sources, and
setting a maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level for the output of the content that includes the synthesized voice data,
wherein output volume corresponding to the synthesized voice data is limited to no greater than the target level for each of the first, second, and third sound sources of the plurality of sound sources.
14 . The information processing method according to claim 13 , wherein
said executing the output voice adjustment processing matches the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a single target level common to the first, second, and third sound sources of the plurality of sound sources.
15 . The information processing method according to claim 13 , wherein
said executing the output voice adjustment matches the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level specific for each of the first, second, and third sound sources.
16 . The information processing method according to claim 13 , wherein
said executing the output voice adjustment reduces a difference in the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources.
17 . The information processing method according to claim 13 , wherein
said executing the output voice adjustment in accordance with a type of the content or a scene of the content output for said outputting.
18 . The information processing method according to claim 13 , wherein the output voice associated with the third sound source corresponding to the viewing user comment is based on a text input via the network.
19 . The information processing method according to claim 13 , wherein the application is one of a gaming application or a musical content hosting application.