IP Library Granted Patent US 8,751,239
Granted Patent B2
US 8,751,239 · App. 11/867,196 · Granted Jun 10, 2014

Method, apparatus and computer program product for providing text independent voice conversion

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,751,239
App. No.
11/867,196
Granted
Jun 10, 2014
Kind
B2
Abstract

An apparatus for providing text independent voice conversion may include a first voice conversion model and a second voice conversion model. The first voice conversion model may be trained with respect to conversion of training source speech to synthetic speech corresponding to the training source speech. The second voice conversion model may be trained with respect to conversion to training target speech from synthetic speech corresponding to the training target speech. An output of the first voice conversion model may be communicated to the second voice conversion model to process source speech input into the first voice conversion model into target speech corresponding to the source speech as the output of the second voice conversion model.

Claims (51)

1. A method comprising:

training, at a user terminal, a first voice conversion model with respect to a training source speech of a first speaker and a second voice conversion model with respect to a training target speech of a second speaker;

wherein training the first voice conversion model further comprises determining a first conversion function for transforming any source speech into corresponding synthetic speech, the first conversion function receiving the training source speech of the first speaker and a training source synthetic speech of the first speaker as inputs, and

wherein training the second voice conversion model further comprises determining a second conversion function for transforming synthetic speech into corresponding target speech, the second conversion function receiving the training target speech of the second speaker and a training target synthetic speech of the second speaker as inputs, and wherein said training target synthetic speech is produced from said training target speech;

processing, at the user terminal, source speech of the first speaker using the first voice conversion model to convert the source speech to synthetic speech; and

processing, at the user terminal, an output of the first voice conversion model at the second voice conversion model to produce target speech corresponding to the source speech.

2. The method of claim 1 , wherein utterances of the training source speech are not parallel to utterances of the training target speech.

3. The method of claim 2 , wherein training the first voice conversion model further comprises training the first voice conversion model to convert the training source speech to synthetic speech corresponding to the training source speech in which the synthetic speech is generated by a text-to-speech device having parallel text corresponding to the training source speech.

4. The method of claim 2 , wherein training the second voice conversion model comprises training the second voice conversion model for conversion to the training target speech from synthetic speech corresponding to the training target speech in which the synthetic speech is generated by a text-to-speech device having parallel text corresponding to the training target speech.

5. The method of claim 1 , wherein processing the source speech at the first voice conversion model comprises converting the source speech to intermediate synthetic speech based on the first voice conversion model.

6. The method of claim 5 , wherein processing the output of the first voice conversion model comprises converting the intermediate synthetic speech to the target speech based on the second voice conversion model.

7. The method of claim 1 , further comprising concatenating the first and second voice conversion models.

8. A computer program product comprising at least one computer-readable storage medium having computer-readable program code portions stored therein, the computer-readable program code portions comprising:

a first executable portion for training, at a user terminal, a first voice conversion model with respect to a training source speech of a first speaker and a second voice conversion model with respect to a training target speech of a second speaker;

wherein training the first voice conversion model further comprises determining a first conversion function for transforming any source speech into corresponding synthetic speech, the first conversion function receiving the training source speech of the first speaker and a training source synthetic speech of the first speaker as inputs, and

wherein training the second voice conversion model further comprises determining a second conversion function for transforming synthetic speech into corresponding target speech, the second conversion function receiving the training target speech of the second speaker and a training target synthetic speech of the second speaker as inputs, and wherein said training target synthetic speech is produced from said training target speech;

a second executable portion for processing, at the user terminal, source speech using the first voice conversion model to convert the source speech to synthetic speech; and

a third executable portion for processing, at the user terminal, an output of the first voice conversion model at the second voice conversion model to produce target speech corresponding to the source speech.

9. The computer program product of claim 8 , wherein utterances of the training source speech are not parallel to utterances of the training target speech.

10. The computer program product of claim 9 , wherein the first executable portion further comprises instructions for training the first voice conversion model to convert the training source speech to synthetic speech corresponding to the training source speech in which the synthetic speech is generated by a text-to-speech device having parallel text corresponding to the training source speech.

11. The computer program product of claim 9 , wherein the third executable portion includes instructions for training the second voice conversion model for conversion to the training target speech from synthetic speech corresponding to the training target speech in which the synthetic speech is generated by a text-to-speech device having parallel text corresponding to the training target speech.

12. The computer program product of claim 8 , wherein the first executable portion includes instructions for converting the source speech to intermediate synthetic speech based on the first voice conversion model.

13. The computer program product of claim 12 , further comprising a fourth executable portion comprising instructions for converting the intermediate synthetic speech to the target speech based on the second voice conversion model.

14. The computer program product of claim 8 , wherein the second executable portion includes instructions for concatenating the first and second voice conversion models.

15. An apparatus comprising a processor and memory storing computer program code, the memory and computer program code configured to, with the processor cause the apparatus at least to:

train, at a user terminal, a first voice conversion model with respect to a training source speech of a first speaker and a second voice conversion model with respect to a training target speech of a second speaker;

wherein training the first voice conversion model further comprises determining a first conversion function for transforming any source speech into corresponding synthetic speech, the first conversion function receiving the training source speech of the first speaker and a training source synthetic speech of the first speaker as inputs, and

wherein training the second voice conversion model further comprises determining a second conversion function for transforming synthetic speech into corresponding target speech, the second conversion function receiving the training target speech of the second speaker and a training target synthetic speech of the second speaker as inputs, and wherein said training target synthetic speech is produced from said training target speech;

process, at the user terminal, source speech of the first speaker using the first voice conversion model to convert the source speech to synthetic speech; and

process, at the user terminal, an output of the first voice conversion model at the second voice conversion model to produce target speech corresponding to the source speech.

16. The apparatus of claim 15 , wherein utterances of the training source speech are not parallel to utterances of the training target speech.

17. The apparatus of claim 16 , further comprising a text-to-speech device in communication with the first and second voice conversion models and wherein the memory and the computer program code are further configured to, with the processor, cause the apparatus to train the first voice conversion model to convert the training source speech to synthetic speech corresponding to the training source speech in which the synthetic speech is generated by the text-to-speech device having parallel text corresponding to the training source speech.

18. The apparatus of claim 16 , further comprising a text-to-speech device in communication with the first and second voice conversion models and wherein the memory and the computer program code are further configured to, with the processor, cause the apparatus to train the second voice conversion model for conversion to the training target speech from synthetic speech corresponding to the training target speech in which the synthetic speech is generated by the text-to-speech device having parallel text corresponding to the training target speech.

19. The apparatus of claim 15 , wherein the memory and the computer program code are further configured to, with the processor, cause the apparatus to convert the source speech to intermediate synthetic speech based on the first voice conversion model.

20. The apparatus of claim 19 , wherein the memory and the computer program code are further configured to, with the processor, cause the apparatus to convert the intermediate synthetic speech to the target speech based on the second voice conversion model.

21. The apparatus of claim 15 , wherein the first and second voice conversion models are concatenated.

22. An apparatus comprising:

means for training, at a user terminal, a first voice conversion model with respect to a training source speech of a first speaker and a second voice conversion model with respect to a training target speech of a second speaker;

wherein training the first voice conversion model further comprises determining a first conversion function for transforming any source speech into corresponding synthetic speech, the first conversion function receiving the training source speech of the first speaker and a training source synthetic speech of the first speaker as inputs, and

wherein training the second voice conversion model further comprises determining a second conversion function for transforming synthetic speech into corresponding target speech, the second conversion function receiving the training target speech of the second speaker and a training target synthetic speech of the second speaker as inputs; and wherein said training targets synthetic speech is produced from said training target speech;

means for processing, at the user terminal, source speech of the first speaker using the first voice conversion model to convert the source speech to synthetic speech; and

means for processing, at a user terminal, an output of the first voice conversion model at the second voice conversion model to produce target speech corresponding to the source speech.

23. The apparatus of claim 22 , wherein utterances of the training source speech are not parallel to utterances of the training target speech.

24. A method comprising;

generating, at a user terminal, synthetic speech corresponding to training source speech based on parallel text corresponding to the training source speech;

training a first voice conversion model with respect to converting source speech to first synthetic speech based on the training source speech and the synthetic speech corresponding to the training source speech, the first voice conversion model being trained at the user terminal;

wherein training the first voice conversion model further comprises determining a first conversion function for transforming any source speech into corresponding synthetic speech, the first conversion function receiving the training source speech of the first speaker and a training source synthetic speech of the first speaker as inputs, and

generating, at the user terminal, synthetic speech corresponding to the training target speech based on parallel text corresponding to the training target speech; and

training a second voice conversion model with respect to converting second synthetic speech to target speech based on the training target speech and the synthetic speech corresponding to the training target speech, the second voice conversion model being trained at the user terminal,

wherein training the second voice conversion model further comprises determining a second conversion function for transforming synthetic speech into corresponding target speech, the second conversion function receiving the training target speech of the second speaker and training target synthetic speech of the second speaker as inputs and wherein said training target synthetic speech is produced from said training target speech.

25. The method of claim 24 , further comprising concatenating the first and second voice conversion models to enable the production of the target speech corresponding to input source speech.

Assignments (9)
RELEASE OF SECURITY INTEREST Recorded Apr 13, 2021
From: CPPIB CREDIT INVESTMENTS INC.
To: CONVERSANT WIRELESS LICENSING S.A R.L.
Reel/Frame 055910/0698 →
AMENDED AND RESTATED U.S. PATENT SECURITY AGREEMENT (FOR NON-U.S. GRANTORS) Recorded Aug 22, 2018
From: CONVERSANT WIRELESS LICENSING S.A R.L.
To: CPPIB CREDIT INVESTMENTS, INC.
Reel/Frame 046897/0001 →
CHANGE OF NAME Recorded Oct 20, 2017
From: CORE WIRELESS LICENSING S.A.R.L.
To: CONVERSANT WIRELESS LICENSING S.A R.L.
Reel/Frame 044250/0398 →
UCC FINANCING STATEMENT AMENDMENT - DELETION OF SECURED PARTY Recorded Aug 30, 2016
From: NOKIA CORPORATION
To: MICROSOFT CORPORATION
Reel/Frame 039872/0112 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2012
From: 2011 INTELLECTUAL PROPERTY ASSET TRUST
To: CORE WIRELESS LICENSING S.A.R.L
Reel/Frame 027485/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2011
From: NOKIA CORPORATION
To: NOKIA 2011 PATENT TRUST
Reel/Frame 027120/0608 →
CHANGE OF NAME Recorded Oct 26, 2011
From: NOKIA 2011 PATENT TRUST
To: 2011 INTELLECTUAL PROPERTY ASSET TRUST
Reel/Frame 027121/0353 →
SHORT FORM PATENT SECURITY AGREEMENT Recorded Sep 13, 2011
From: CORE WIRELESS LICENSING S.A.R.L.
To: NOKIA CORPORATION; MICROSOFT CORPORATION
Reel/Frame 026894/0665 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 4, 2007
From: TIAN, JILEI; POPA, VICTOR; NURMINEN, JANI K.
To: NOKIA CORPORATION
Reel/Frame 019921/0355 →