IP Library Granted Patent US 9,711,161
Granted Patent B2
US 9,711,161 · App. 14/455,070 · Granted Jul 18, 2017

Voice processing apparatus, voice processing method, and program

Inventors: Yuhki Mitsufuji (Tokyo, JP); Toru Chinen (Kanagawa, JP)
Assignee: SONY CORPORATION
G10L21/003G10L17/04G10L25/60G10L2021/0135
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,711,161
App. No.
14/455,070
Granted
Jul 18, 2017
Kind
B2
Abstract

A voice processing apparatus includes a voice quality determining unit configured to determine a target speaker determining method used for a voice quality conversion in accordance with a determining method control value for instructing the target speaker determining method of determining a target speaker whose voice quality is targeted to the voice quality conversion, and determine the target speaker in accordance with the target speaker determining method.

Claims (32)

1. A voice processing apparatus for voice quality conversion, comprising:

one or more processors configured to:

determine a target speaker determining method used for the voice quality conversion based on a determining method control value to instruct the target speaker determining method that determines a target speaker whose voice quality is targeted to the voice quality conversion;

determine the target speaker based on the target speaker determining method for a voice of a reference speaker whose voice quality is to be converted;

determine, as the target speaker determining method, a method that randomly samples a voice quality parameter distribution that a voice quality parameter of the reference speaker belongs to, wherein the voice quality parameter distribution is a distribution of a voice quality parameter calculated based on voices of a plurality of speakers in a voice quality space of the voice quality parameter to represent the voice quality;

determine, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to a sampling point obtained as a result of the sampling based on the determining method control value; and

convert the voice quality of the voice of the reference speaker to a voice of the target speaker based on the determined method.

2. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to generate the voice of the voice quality of the target speaker from the voice of the reference speaker.

3. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine the target speaker based on the voice quality parameter distribution.

4. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine, as the target speaker determining method used for the voice quality conversion, a method that determines, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter distributed in the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, based on the determining method control value.

5. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine, as the target speaker determining method used for the voice quality conversion, a method that determines, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to a point where a point that corresponds to the voice quality parameter of the reference speaker in the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, is shifted in a point symmetry direction with respect to a determined point based on the determining method control value.

6. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine, as the target speaker determining method used for the voice quality conversion, a method that determines, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter distributed in the voice quality parameter distribution different from the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, based on the determining method control value.

7. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine, as the target speaker determining method used for the voice quality conversion, one of:

a method that utilizes the voice quality parameter distribution which is the distribution of the voice quality parameter calculated based on the voices of the plurality of speakers in the voice quality space of the voice quality parameter to represent the voice quality, or

a method that utilizes the voice quality parameter distribution different from the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, based on the determining method control value.

8. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to determine, as the target speaker determining method used for the voice quality conversion, one of:

the method that randomly samples the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to and determine, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to the sampling point obtained as the result of the sampling,

a method that determines, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to a point where a point that corresponds to the voice quality parameter of the reference speaker in the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, is shifted in a point symmetry direction with respect to a determined point, or

a method that determines, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter distributed in the voice quality parameter distribution different from the voice quality parameter distribution that the voice quality parameter of the reference speaker belongs to, based on the determining method control value.

9. The voice processing apparatus according to claim 1 , wherein the one or more processors are further configured to execute the voice quality conversion based on the determined target speaker.

10. A voice processing method for voice quality conversion, the method comprising:

determining, by one or more processors, a target speaker determining method used for the voice quality conversion based on a determining method control value for instructing the target speaker determining method of determining a target speaker whose voice quality is targeted to the voice quality conversion;

determining, by the one or more processors, the target speaker based on the target speaker determining method for a voice of a reference speaker whose voice quality is to be converted;

determining, as the target speaker determining method, a method of randomly sampling a voice quality parameter distribution that a voice quality parameter of the reference speaker belongs to, wherein the voice quality parameter distribution is a distribution of a voice quality parameter calculated based on voices of a plurality of speakers in a voice quality space of the voice quality parameter to represent the voice quality;

determining, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to a sampling point obtained as a result of the sampling based on the determining method control value; and

converting, by the one or more processors, the voice quality of the voice of the reference speaker to a voice of the target speaker based on the determined method.

11. A non-transitory computer-readable medium having stored thereon computer-readable instructions, which when executed by a computer, cause the computer to execute operations, the operations comprising:

determining a target speaker determining method used for voice quality conversion based on a determining method control value for instructing the target speaker determining method of determining a target speaker whose voice quality is targeted to the voice quality conversion;

determining the target speaker based on the target speaker determining method for a voice of a reference speaker whose voice quality is to be converted;

determining, as the target speaker determining method, a method of randomly sampling a voice quality parameter distribution that a voice quality parameter of the reference speaker belongs to, wherein the voice quality parameter distribution is a distribution of a voice quality parameter calculated based on voices of a plurality of speakers in a voice quality space of the voice quality parameter to represent the voice quality;

determining, as the voice quality of the target speaker, the voice quality represented by the voice quality parameter that corresponds to a sampling point obtained as a result of the sampling based on the determining method control value; and

converting the voice quality of the voice of the reference speaker to a voice of the target speaker based on the determined method.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2014
From: MITSUFUJI, YUHKI; CHINEN, TORU
To: SONY CORPORATION
Reel/Frame 033494/0903 →
Priority Claims (1)
JP 2013-170504 · Aug 20, 2013 · national
Continuity (1)
Related Publication 20150058015A1 · Feb 26, 2015