IP Library Granted Patent US 10,157,608
Granted Patent B2
US 10,157,608 · App. 15/433,690 · Granted Dec 18, 2018

Device for predicting voice conversion model, method of predicting voice conversion model, and computer program product

Inventors: Yamato Ohtani (Kanagawa, JP); Yu Nasu (Tokyo, JP); Masatsune Tamura (Kangawa, JP); Masahiro Morita (Kanagawa, JP)
Assignee: KABUSHIKI KAISHA TOSHIBA
G10L13/0335G10L13/033G10L13/047G10L13/08G10L21/003
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,157,608
App. No.
15/433,690
Granted
Dec 18, 2018
Kind
B2
Abstract

According to an embodiment, a voice processing device includes an interface system, a determining processor, and a predicting processor. The interface system configured to receive neutral voice data representing audio in a neutral voice of a user. The determining processor configured to determine a predictive parameter based at least in part on the neutral voice data. The predicting processor configured to predict a voice conversion model for converting the neutral voice of the speaker to a target voice using at least the predictive parameter.

Claims (30)

1. A device for predicting a voice conversion model, the device comprising:

an interface system configured to receive neutral voice data representing audio in a neutral voice of a user;

a determining processor, implemented in computer hardware, configured to determine a predictive parameter based at least in part on the neutral voice data; and

a predicting processor, implemented in computer hardware, configured to predict a voice conversion model for converting the neutral voice of the speaker to a target voice tone using at least the predictive parameter, wherein

a plurality of neutral voice predictive models are respectively associated with voice conversion predictive models each of which is optimized for converting the corresponding neutral voice predictive model to a voice model of the target voice,

the neutral voice data comprises acoustic feature quantity data representing a feature of the voice obtained by analyzing the audio in the neutral voice of the user and language attribute date representing an attribute of a language obtained by analyzing the audio in the neutral voice of the user, and

the determining processor is configured to:

calculate a likelihood of a linear sum of a vector based at least in part on the neutral voice predictive models with respect to the acoustic feature quantity data and the language attribute data,

determine, as a weight, a coefficient of the linear sum comprising the highest calculated likelihood, and

determine the predictive parameter generated by adding, to a model parameter of each voice conversion predictive model, the weight determined with respect to the corresponding neutral voice predictive model.

2. A method of predicting a voice conversion model, the method comprising:

receiving, by an interface system, neutral voice data representing audio in a calm voice tone of a user;

determining, by a determining processor implemented in computer hardware, a predictive parameter based at least in part on the neutral voice data; and

predicting, by a predicting processor implemented in computer hardware, a voice conversion model for converting the neutral voice of the speaker to a target voice using at least the predictive parameter, wherein

a plurality of neutral voice predictive models are respectively associated with voice conversion predictive models each of which is optimized for converting the corresponding neutral voice predictive model to a voice model of the target voice,

the neutral voice data comprises acoustic feature quantity data representing a feature of the voice obtained by analyzing the audio in the neutral voice of the user and language attribute date representing an attribute of a language obtained by analyzing the audio in the neutral voice of the user, and

the determining includes:

calculating a likelihood of a linear sum of a vector based at least in part on the neutral voice predictive models with respect to the acoustic feature quantity data and the language attribute data,

determining, as a weight, a coefficient of the linear sum comprising the highest calculated likelihood, and

determining the predictive parameter generated by adding, to a model parameter of each voice conversion predictive model, the weight determined with respect to the corresponding neutral voice predictive model.

3. A computer program product comprising a non-transitory computer-readable medium containing a computer program that causes a computer to function as:

an interface system configured to receive neutral voice data representing audio in a neutral voice of a user;

a determining processor configured to determine a predictive parameter at least in part on the neutral voice data; and

a predicting processor configured to predict a voice conversion model for converting the neutral voice of the speaker to a target voice, wherein

a plurality of neutral voice predictive models are respectively associated with voice conversion predictive models each of which is optimized for converting the corresponding neutral voice predictive model to a voice model of the target voice,

the neutral voice data comprises acoustic feature quantity data representing a feature of the voice obtained by analyzing the audio in the neutral voice of the user and language attribute date representing an attribute of a language obtained by analyzing the audio in the neutral voice of the user, and

the determining processor is configured to:

calculate a likelihood of a linear sum of a vector based at least in part on the neutral voice predictive models with respect to the acoustic feature quantity data and the language attribute data,

determine, as a weight, a coefficient of the linear sum comprising the highest calculated likelihood, and

determine the predictive parameter generated by adding, to a model parameter of each voice conversion predictive model, the weight determined with respect to the corresponding neutral voice predictive model.

Assignments (4)
CORRECTIVE ASSIGNMENT TO CORRECT THE RECEIVING PARTY'S ADDRESS PREVIOUSLY RECORDED ON REEL 048547 FRAME 0187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT OF ASSIGNORS INTEREST. Recorded May 6, 2020
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 052595/0307 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ADD SECOND RECEIVING PARTY PREVIOUSLY RECORDED AT REEL: 48547 FRAME: 187. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Aug 13, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: KABUSHIKI KAISHA TOSHIBA; TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 050041/0054 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2019
From: KABUSHIKI KAISHA TOSHIBA
To: TOSHIBA DIGITAL SOLUTIONS CORPORATION
Reel/Frame 048547/0187 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 16, 2017
From: OHTANI, YAMATO; NASU, YU; TAMURA, MASATSUNE; MORITA, MASAHIRO
To: KABUSHIKI KAISHA TOSHIBA
Reel/Frame 041276/0987 →
Continuity (2)
Continuation PCTJP2014074581 · Sep 17, 2014
Related Publication 20170162187A1 · Jun 8, 2017