IP Library › Granted Patent US 12,744,050
Granted Patent B1
US 12,744,050 · App. 18/684,196 · Granted Sep 22, 2026

Information processing device, information processing method, and program

Inventor: Kouichiro Takashima (Tokyo, JP)
Assignee: SONY GROUP CORPORATION
G10L21/0364G10L13/033G10L21/034
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,744,050
App. No.
18/684,196
Granted
Sep 22, 2026
Kind
B1
Abstract

A device and method are provided that adjust voices of a plurality of sound sources included in distribution content from an information processing device and make it easier to hear sound of each sound source by a reception terminal that receives and reproduces distribution content. A first output voice adjustment unit executes adjustment processing on an output voice of each of the plurality of sound sources, and content including synthesized voice data obtained by synthesizing the output voices corresponding to the sound sources adjusted by the first output voice adjustment unit is output. The first output voice adjustment unit executes output voice adjustment processing for matching a maximum value of a volume level corresponding to a frequency of the output voice of each sound source to a target level. Moreover, a second output voice adjustment unit executes the output voice adjustment processing in accordance with a content type or scene.

Claims (57)

1 . An information processing device comprising:

circuitry configured to

input an output voice of each of a plurality of sound sources and execute adjustment processing on the output voice of each of a first sound source, a second sound source, and a third sound source of the plurality of sound sources, wherein

the output voice associated with the first sound source corresponds to a distribution user voice input via a microphone,

the output voice associated with the second sound source corresponds to operation of an application, and

the output voice associated with the third sound source corresponds to a viewing user comment input via a network;

synthesis synthesize the output voice corresponding to each of the plurality of sound sources having been adjusted to generate synthesized voice data; and

output content including the synthesized voice data, wherein

the adjustment processing includes:

analyzing a volume level corresponding to a frequency regarding the output voice of each of the first, second, and third sound sources of the plurality of sound sources, and

setting a maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level for the output of the content that includes the synthesized voice data,

wherein output volume corresponding to the synthesized voice data is limited to no greater than the target level for each of the first, second, and third sound sources of the plurality of sound sources.

2 . The information processing device according to claim 1 , wherein

the output voice adjustment processing is executed to match the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a single target level common to the first, second, and third sound sources of the plurality of sound sources.

3 . The information processing device according to claim 1 , wherein

the output voice adjustment processing is executed for matching the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level specific for each of the first, second, and third sound sources.

4 . The information processing device according to claim 1 , wherein the output voice adjustment processing is executed to reduce a difference in the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources.

5 . The information processing device according to claim 1 , wherein the circuitry is configured to:

execute the output voice adjustment processing on the output voice of each of the first, second, and third sound sources in accordance with a type of the content or a scene of the content output via the circuitry.

6 . The information processing device according to claim 5 , wherein the circuitry is configured to

select one of the first, second, or third sound sources that is an execution target of the output voice adjustment processing in accordance with the type of the content output via the circuitry, and

execute the output voice adjustment processing only on the output voice of the selected one sound source from among the first, second, and third sound sources.

7 . The information processing device according to claim 5 , wherein the circuitry is configured to

in a case where the type of the content output via the circuitry is game content execute output voice adjustment processing to emphasize the distribution user voice.

8 . The information processing device according to claim 5 , wherein the circuitry is configured to

in a case where the type of the content output via the circuitry is music content, execute output voice adjustment processing to emphasize music reproduced sound of the music content.

9 . The information processing device according to claim 5 , wherein the circuitry is configured to

select one of the first, second, or third sound sources that is an execution target of the output voice adjustment processing in accordance with the scene of the content output via the circuitry, and

execute the output voice adjustment processing only on the output voice of the selected one sound source from among the first, second, and third sound sources.

10 . The information processing device according to claim 5 , wherein the circuitry is configured to

determine the scene of the content output via the circuitry,

analyze attribute information of the application being executed by the information processing device,

display information of a display,

execute the voice adjustment processing in accordance with the determined scene determined by the circuitry.

11 . The information processing device according to claim 1 , wherein the output voice associated with the third sound source corresponding to the viewing user comment is based on a text input via the network.

12 . The information processing device according to claim 1 , wherein the application is one of a gaming application or a musical content hosting application.

13 . An information processing method comprising:

inputting an output voice of each of a plurality of sound sources and execute adjustment processing on the output voice of each of a first sound source, a second sound source, and a third sound source of the plurality of sound sources, wherein

the output voice associated with the first sound source corresponds to a distribution user voice input via a microphone,

the output voice associated with the second sound source corresponds to operation of an application, and

the output voice associated with the third sound source corresponds to a viewing user comment input via a network;

synthesizing the output voice corresponding to each of the plurality of sound sources having been adjusted to generate synthesized voice data; and

outputting content including the synthesized voice data, wherein

said adjustment processing includes:

analyzing a volume level corresponding to a frequency regarding the output voice of each of the first, second, and third sound sources of the plurality of sound sources, and

setting a maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level for the output of the content that includes the synthesized voice data,

wherein output volume corresponding to the synthesized voice data is limited to no greater than the target level for each of the first, second, and third sound sources of the plurality of sound sources.

14 . The information processing method according to claim 13 , wherein

said executing the output voice adjustment processing matches the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a single target level common to the first, second, and third sound sources of the plurality of sound sources.

15 . The information processing method according to claim 13 , wherein

said executing the output voice adjustment matches the maximum value of the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources to a target level specific for each of the first, second, and third sound sources.

16 . The information processing method according to claim 13 , wherein

said executing the output voice adjustment reduces a difference in the volume level corresponding to the frequency of the output voice of each of the first, second, and third sound sources.

17 . The information processing method according to claim 13 , wherein

said executing the output voice adjustment in accordance with a type of the content or a scene of the content output for said outputting.

18 . The information processing method according to claim 13 , wherein the output voice associated with the third sound source corresponding to the viewing user comment is based on a text input via the network.

19 . The information processing method according to claim 13 , wherein the application is one of a gaming application or a musical content hosting application.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: TAKASHIMA, KOUICHIRO
To: SONY GROUP CORPORATION
Reel/Frame 067613/0417 →
Priority Claims (1)
JP 2021-136869 · Aug 25, 2021 · national
References Cited (30)
US 10390097B1 · Lee · 2019 [cited by examiner]
US 10466959B1 · Yang · 2019 [cited by applicant]
US 10917061B2 · Wu · 2021 [cited by examiner]
US 20020111204A1 · Lupo · 2002 [cited by examiner]
US 20030203746A1 · Iwase · 2003 [cited by examiner]
US 20050117761A1 · Sato · 2005 [cited by examiner]
US 20080028094A1 · Kang · 2008 [cited by examiner]
US 20100111313A1 · Namba · 2010 [cited by examiner]
US 20100280638A1 · Matsuda · 2010 [cited by applicant]
US 20130108056A1 · Noro · 2013 [cited by examiner]
US 20140133681A1 · Mizuta · 2014 [cited by examiner]
US 20140328487A1 · Hiroe · 2014 [cited by examiner]
US 20160098912A1 · Mori · 2016 [cited by examiner]
US 20160371051A1 · Rowe · 2016 [cited by applicant]
US 20170151501A1 · Kitamura · 2017 [cited by examiner]
US 20190082276A1 · Crow · 2019 [cited by examiner]
US 20190118096A1 · Jeon · 2019 [cited by examiner]
US 20200225844A1 · Lerner · 2020 [cited by applicant]
US 20220193549A1 · Wakeland · 2022 [cited by examiner]
US 20220286801A1 · Azizian · 2022 [cited by examiner]
JP S4826057B1 · 1973 [cited by applicant]
JP 2002258842A · 2002 [cited by applicant]
JP 2003243952A · 2003 [cited by applicant]
JP 2008228184A · 2008 [cited by applicant]
JP 2012054863A · 2012 [cited by applicant]
JP 2019180073A · 2019 [cited by applicant]
WO 2018096954A1 · 2018 [cited by applicant]
WO WO2020062922A1 · 2020 [cited by examiner]
International Search Report and Written Opinion mailed on Jun. 28, 2022, received for PCT Application PCT/JP2022/013429, filed on Mar. 23, 2022, 13 pages including English Translation. [cited by applicant]
Enrique Perez-Gonzalez et al: “Automatic Gain and Fader Control for Live Mixing”, IEEE Workshop on Applications of Signal Processing to Audio and Acoustics, Oct. 18, 2009 (Oct. 18, 2009), pp. 1-4, XP055555402. [cited by applicant]