Cross-lingual voice conversion system and method
A cross-lingual voice conversion system and method comprises a voice feature extractor configured to receive a first voice audio segment in a first language and a second voice audio segment in a second language, and extract, respectively, audio features comprising first-voice, speaker-dependent acoustic features and second-voice, speaker-independent linguistic features. One or more generators are configured to receive extracted features, and produce therefrom a third voice candidate keeping the first-voice, speaker-dependent acoustic features and the second-voice, speaker-independent linguistic features, wherein the third voice candidate speaks the second language. One or more discriminators are configured to compare the third voice candidate with the ground truth data, and provide results of the comparison back to the generator for refining the third voice candidate.
1. A method comprising:
receiving a first voice audio segment and a second voice audio segment;
extracting a speaker-dependent acoustic feature from the first voice audio segment;
extracting a speaker-independent linguistic feature from the second voice audio;
generating a plurality of voice candidates with a machine learning system, wherein each of the plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and
storing the voice candidates in a database connected to the machine learning system.
2. The method of claim 1 further comprising:
generating a third voice audio segment based on a voice candidate from the plurality of voice candidates,
wherein the first voice audio segment is in a first language and the second voice audio segment is in a second language, and
wherein the third voice audio segment is in the second language translated based on the first language.
3. The method of claim 1 ,
wherein the speaker-dependent acoustic feature is based on a segmental feature, and
wherein the speaker-independent linguistic feature is based on a supra-segmental feature.
4. The method of claim 2 , further comprising:
comparing the voice candidate of the plurality of voice candidates with ground truth data comprising the speaker-dependent acoustic feature and the speaker-independent linguistic feature; and
refining the voice candidate based on a comparison between the third voice audio segment and the ground truth data.
5. The method of claim 1 ,
wherein the speaker-dependent acoustic feature comprises a short-term segmental feature corresponding to vocal tract characteristics, and
wherein the speaker-independent linguistic feature comprises a supra-segmental feature corresponding to acoustic properties over more than one segment.
6. The method of claim 1 , further comprising:
selecting a voice candidate of the plurality of voice candidates for voice conversion.
7. The method of claim 1 , wherein the database comprises a plurality of trained voice candidates.
8. The method of claim 1 , further comprising:
generating a plurality of dubbed version audio files based on each of the plurality of voice candidates, respectively, wherein each of the plurality of dubbed version audio files comprises the level of the speaker-dependent acoustic feature and the level of the speaker-independent linguistic feature of its respective voice candidate.
9. A method of training a machine learning system comprising:
learning a forward mapping function and an inverse mapping function, wherein the forward mapping function comprises:
receiving, by a voice feature extractor, a first voice audio segment;
extracting, by the voice feature extractor, a speaker-dependent acoustic feature;
sending the speaker-dependent acoustic feature to a first speaker generator;
receiving, by the first speaker generator, a speaker-independent linguistic feature from an inverse mapping function;
generating, by the first speaker generator, a first plurality of voice candidates based on the speaker-dependent acoustic feature and the speaker-independent linguistic feature, wherein each of the first plurality of the voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and
determining, by a first discriminator, whether there is a discrepancy between a level of the speaker-dependent acoustic feature associated with the first plurality of voice and the speaker-dependent acoustic feature, and
wherein the inverse mapping function comprises:
receiving, by the feature extractor, a second voice audio segment;
extracting, by the feature extractor, the speaker-independent linguistic feature;
sending the speaker-independent linguistic feature to a second voice candidate generator;
receiving, by the second voice candidate generator, the speaker-dependent acoustic feature from the forward mapping function;
generating, by the second voice candidate generator, a second plurality of voice candidates using the speaker-independent linguistic feature and the speaker-dependent acoustic feature, wherein each of the second plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and
storing the second plurality of voice candidates in a database connected to the machine learning system.
10. The method of claim 9 , wherein each of the second plurality of voice candidates includes a second language translated based on a first language, and wherein the method further comprises determining, by a second discriminator, whether there is a discrepancy between a level of the speak-independent acoustic feature associated with the second plurality of voice candidates and the speaker-independent linguistic feature.
11. The method of claim 9 , further comprising:
determining, by the first discriminator, the level of the speaker-dependent acoustic feature associated with the first plurality of voice candidates and the speaker-dependent acoustic feature are not consistent;
providing first inconsistency information to the first voice candidate generator for refining the first plurality of voice candidates;
sending the first plurality of voice candidates to a third speaker generator;
generating a converted speaker-dependent acoustic feature; and
sending the converted speaker-dependent acoustic feature to the first voice candidate generator; and
determining, by the second discriminator, a level of the speaker-independent feature associated with the second plurality of voice candidates and the speaker-independent linguistic feature are not consistent;
providing second inconsistency information to the second voice candidate generator for refining the second plurality of voice candidates;
sending the second plurality of voice candidates to a fourth speaker generator;
generating a converted speaker-independent linguistic feature; and
sending the converted speaker-independent linguistic feature to the second voice candidate generator.
12. The method of claim 9 , further comprising employing identity mapping loss for preserving identity-related features of each of the first and second voice audio segments.
13. A system comprising:
a machine learning system stored in memory of a server computer system and being implemented by at least one processor, the machine learning system being configured to execute instructions to perform operations comprising:
receiving a first voice audio segment and a second voice audio segment;
extracting a speaker-dependent acoustic feature from the first voice audio segment;
extracting a speaker-independent linguistic feature from the second voice audio segment;
generating a plurality of voice candidates, wherein each of the plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and
storing the voice candidates in a database connected to the machine learning system.
14. The system of claim 13 , wherein the operations further comprise:
generating a third voice audio segment based on a voice candidate from the plurality of voice candidates,
wherein the first voice audio segment is in a first language and the second voice audio segment is in a second language, and
wherein the third voice audio segment is in the second language translated based on the first language.
15. The system of claim 13 ,
wherein the speaker-dependent acoustic feature is based on a segmental feature, and
wherein the speaker-independent linguistic feature is based on a supra-segmental feature.
16. The system of claim 14 , the operations further comprising:
comparing the voice candidate of the plurality of voice candidates with ground truth data comprising the speaker-dependent acoustic feature and the speaker-independent linguistic feature; and
refining the voice candidate based on a comparison between the third voice audio segment and the ground truth data.
17. The system of claim 13 ,
wherein the speaker-dependent acoustic feature comprises a short-term segmental feature corresponding to vocal tract characteristics, and
wherein the speaker-independent linguistic feature comprises a supra-segmental feature corresponding to acoustic properties over more than one segment.