IP Library Granted Patent US 12,354,616
Granted Patent B2
US 12,354,616 · App. 18/374,583 · Granted Jul 8, 2025

Cross-lingual voice conversion system and method

Inventor: Cevat Yerli (Dubai, AE)
Assignee: TMRW Group IP
G10L21/003G06F40/58G10L13/00G10L15/02G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,616
App. No.
18/374,583
Granted
Jul 8, 2025
Kind
B2
Abstract

A cross-lingual voice conversion system and method comprises a voice feature extractor configured to receive a first voice audio segment in a first language and a second voice audio segment in a second language, and extract, respectively, audio features comprising first-voice, speaker-dependent acoustic features and second-voice, speaker-independent linguistic features. One or more generators are configured to receive extracted features, and produce therefrom a third voice candidate keeping the first-voice, speaker-dependent acoustic features and the second-voice, speaker-independent linguistic features, wherein the third voice candidate speaks the second language. One or more discriminators are configured to compare the third voice candidate with the ground truth data, and provide results of the comparison back to the generator for refining the third voice candidate.

Claims (72)

1. A method comprising:

receiving a first voice audio segment and a second voice audio segment;

extracting a speaker-dependent acoustic feature from the first voice audio segment;

extracting a speaker-independent linguistic feature from the second voice audio;

generating a plurality of voice candidates with a machine learning system, wherein each of the plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and

storing the voice candidates in a database connected to the machine learning system.

2. The method of claim 1 further comprising:

generating a third voice audio segment based on a voice candidate from the plurality of voice candidates,

wherein the first voice audio segment is in a first language and the second voice audio segment is in a second language, and

wherein the third voice audio segment is in the second language translated based on the first language.

3. The method of claim 1 ,

wherein the speaker-dependent acoustic feature is based on a segmental feature, and

wherein the speaker-independent linguistic feature is based on a supra-segmental feature.

4. The method of claim 2 , further comprising:

comparing the voice candidate of the plurality of voice candidates with ground truth data comprising the speaker-dependent acoustic feature and the speaker-independent linguistic feature; and

refining the voice candidate based on a comparison between the third voice audio segment and the ground truth data.

5. The method of claim 1 ,

wherein the speaker-dependent acoustic feature comprises a short-term segmental feature corresponding to vocal tract characteristics, and

wherein the speaker-independent linguistic feature comprises a supra-segmental feature corresponding to acoustic properties over more than one segment.

6. The method of claim 1 , further comprising:

selecting a voice candidate of the plurality of voice candidates for voice conversion.

7. The method of claim 1 , wherein the database comprises a plurality of trained voice candidates.

8. The method of claim 1 , further comprising:

generating a plurality of dubbed version audio files based on each of the plurality of voice candidates, respectively, wherein each of the plurality of dubbed version audio files comprises the level of the speaker-dependent acoustic feature and the level of the speaker-independent linguistic feature of its respective voice candidate.

9. A method of training a machine learning system comprising:

learning a forward mapping function and an inverse mapping function, wherein the forward mapping function comprises:

receiving, by a voice feature extractor, a first voice audio segment;

extracting, by the voice feature extractor, a speaker-dependent acoustic feature;

sending the speaker-dependent acoustic feature to a first speaker generator;

receiving, by the first speaker generator, a speaker-independent linguistic feature from an inverse mapping function;

generating, by the first speaker generator, a first plurality of voice candidates based on the speaker-dependent acoustic feature and the speaker-independent linguistic feature, wherein each of the first plurality of the voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and

determining, by a first discriminator, whether there is a discrepancy between a level of the speaker-dependent acoustic feature associated with the first plurality of voice and the speaker-dependent acoustic feature, and

wherein the inverse mapping function comprises:

receiving, by the feature extractor, a second voice audio segment;

extracting, by the feature extractor, the speaker-independent linguistic feature;

sending the speaker-independent linguistic feature to a second voice candidate generator;

receiving, by the second voice candidate generator, the speaker-dependent acoustic feature from the forward mapping function;

generating, by the second voice candidate generator, a second plurality of voice candidates using the speaker-independent linguistic feature and the speaker-dependent acoustic feature, wherein each of the second plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and

storing the second plurality of voice candidates in a database connected to the machine learning system.

10. The method of claim 9 , wherein each of the second plurality of voice candidates includes a second language translated based on a first language, and wherein the method further comprises determining, by a second discriminator, whether there is a discrepancy between a level of the speak-independent acoustic feature associated with the second plurality of voice candidates and the speaker-independent linguistic feature.

11. The method of claim 9 , further comprising:

determining, by the first discriminator, the level of the speaker-dependent acoustic feature associated with the first plurality of voice candidates and the speaker-dependent acoustic feature are not consistent;

providing first inconsistency information to the first voice candidate generator for refining the first plurality of voice candidates;

sending the first plurality of voice candidates to a third speaker generator;

generating a converted speaker-dependent acoustic feature; and

sending the converted speaker-dependent acoustic feature to the first voice candidate generator; and

determining, by the second discriminator, a level of the speaker-independent feature associated with the second plurality of voice candidates and the speaker-independent linguistic feature are not consistent;

providing second inconsistency information to the second voice candidate generator for refining the second plurality of voice candidates;

sending the second plurality of voice candidates to a fourth speaker generator;

generating a converted speaker-independent linguistic feature; and

sending the converted speaker-independent linguistic feature to the second voice candidate generator.

12. The method of claim 9 , further comprising employing identity mapping loss for preserving identity-related features of each of the first and second voice audio segments.

13. A system comprising:

a machine learning system stored in memory of a server computer system and being implemented by at least one processor, the machine learning system being configured to execute instructions to perform operations comprising:

receiving a first voice audio segment and a second voice audio segment;

extracting a speaker-dependent acoustic feature from the first voice audio segment;

extracting a speaker-independent linguistic feature from the second voice audio segment;

generating a plurality of voice candidates, wherein each of the plurality of voice candidates comprises a different level of the speaker-dependent acoustic feature of the first voice audio segment and a different level of the speaker-independent linguistic feature of the second voice audio segment; and

storing the voice candidates in a database connected to the machine learning system.

14. The system of claim 13 , wherein the operations further comprise:

generating a third voice audio segment based on a voice candidate from the plurality of voice candidates,

wherein the first voice audio segment is in a first language and the second voice audio segment is in a second language, and

wherein the third voice audio segment is in the second language translated based on the first language.

15. The system of claim 13 ,

wherein the speaker-dependent acoustic feature is based on a segmental feature, and

wherein the speaker-independent linguistic feature is based on a supra-segmental feature.

16. The system of claim 14 , the operations further comprising:

comparing the voice candidate of the plurality of voice candidates with ground truth data comprising the speaker-dependent acoustic feature and the speaker-independent linguistic feature; and

refining the voice candidate based on a comparison between the third voice audio segment and the ground truth data.

17. The system of claim 13 ,

wherein the speaker-dependent acoustic feature comprises a short-term segmental feature corresponding to vocal tract characteristics, and

wherein the speaker-independent linguistic feature comprises a supra-segmental feature corresponding to acoustic properties over more than one segment.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 22, 2025
From: YERLI, CEVAT
To: TMRW FOUNDATION IP S.ÀR.L.
Reel/Frame 070914/0481 →
CHANGE OF NAME Recorded Apr 18, 2025
From: TMRW FOUNDATION IP S.À R.L.
To: TMRW GROUP IP
Reel/Frame 070891/0005 →
Continuity (3)
Continuation 17138642 · Dec 30, 2020
Provisional Application 62955227 · Dec 30, 2019
Related Publication 20240028843A1 · Jan 25, 2024
References Cited (46)
US 8930183B2 · Chun · 2015 [cited by examiner]
US 10930263B1 · Mahyar · 2021 [cited by examiner]
US 11797782B2 · Yerli · 2023 [cited by examiner]
US 20020009043A1 · Ko · 2002 [cited by examiner]
US 20120069974A1 · Zhu et al. · 2012 [cited by applicant]
US 20120173241A1 · Li et al. · 2012 [cited by applicant]
US 20120253794A1 · Chun et al. · 2012 [cited by applicant]
US 20150127349A1 · Agiomyrgiannakis · 2015 [cited by examiner]
US 20170040016A1 · Cui · 2017 [cited by examiner]
US 20180342256A1 · Huffman · 2018 [cited by examiner]
US 20180342257A1 · Huffman · 2018 [cited by examiner]
US 20190354592A1 · Musham · 2019 [cited by examiner]
US 20200058289A1 · Gabryjelski · 2020 [cited by examiner]
US 20200365166A1 · Zhang · 2020 [cited by examiner]
US 20200380952A1 · Zhang · 2020 [cited by examiner]
US 20200410976A1 · Zhou · 2020 [cited by examiner]
US 20220358905A1 · Gupta · 2022 [cited by examiner]
US 20220405492A1 · Rathnam · 2022 [cited by examiner]
CN 109147758A · 2019 [cited by applicant]
CN 109671442A · 2019 [cited by applicant]
CN 110060691A · 2019 [cited by applicant]
CN 110246488A · 2019 [cited by applicant]
CN 110459232A · 2019 [cited by applicant]
CN 110600046A · 2019 [cited by applicant]
JP 2009186820A · 2009 [cited by applicant]
JP 2019101391A · 2019 [cited by applicant]
JP 2019109306A · 2019 [cited by applicant]
KR 20190094315A · 2019 [cited by applicant]
KR 20190114938A · 2019 [cited by applicant]
WO 2019182346A1 · 2019 [cited by applicant]
Sündermann, David, and Hermann Ney. “An automatic segmentation and mapping approach for voice conversion parameter training.” Proc. of the AST'03, Maribor, Slovenia (2003). (Year: 2003). [cited by examiner]
Sundermann, David, Hermann Ney, and H. Hoge. “VTLN-based cross-language voice conversion.” 2003 IEEE workshop on automatic speech recognition and understanding (IEEE Cat. No. 03EX721). IEEE, 2003. (Year: 2003). [cited by examiner]
Extended European Search Report mailed on Jul. 27, 2021, EP Patent Application No. 20217111, 9 pages. [cited by applicant]
Saito, Y., et al., “Statistical Parametric Speech Synthesis Incorporating Generative Adversarial Networks”, IEEE/ACM Transactions on Audio, Speech and Language Processing, vol. 26, No. 1, Jan. 2018, pp. 84-96. [cited by applicant]
Kiran Reddy, M., et al. “DNN-Based Cross Lingual Voice Conversion Using Bottleneck Features,” Neural Processing Letters, Kluwer Academic Publishers, Norwell, MA vol. 51, No. 2, Nov. 7, 2019, p. 2029-2042. [cited by applicant]
Office Action mailed on Jul. 27, 2021, Indian Patent Application No. 202014056471, filed on Dec. 24, 2020, 6 pages. [cited by applicant]
Office Action mailed on Apr. 26, 2022, Japanese Patent Application No. JP2020215179 (Japanese version), 2 pages. [cited by applicant]
Lorenzo-Trueba, J., et al., “Can We Steal Your Voice Identity from the Internet?” Initial Investigation of Cloning Obama's Voice Using GAN Wavenet and low-quality found data, Odyssey 2018 the Speaker and Language Recogn… [cited by applicant]
Rallabandi, S.S. et al., An Approach to Cross-Lingual Voice Conversion, Jul. 19, 2019, 7 pages. [cited by applicant]
Fang, F., et al., “High-Quality Nonparallel Voice Conversion Based on Cycle-Consistent Adversarial Network,” 2018 IEEE Intl., Conf. on Acoustics, Speech, Signal Processing (ICASSP) Calgary AB, Canada, Apr. 15-20, 2018, … [cited by applicant]
Hsu, C., et al., “Voice Conversion from Unaligned Corpora Using Variational Autoencoding Wasserstein Generative Adversarial Networks,” Interspeech 2017, Stockholm Sweden, Aug. 20-24, 2017, p. 3364-3368. [cited by applicant]
Kaneko, T., et al., “Parallel-Data-Tree Voice Conversion Using Cycle Consistent Adversarial Networks,” ArXiv, 5 pg. Dec. 20, 2017. [cited by applicant]
Sisman, B., et al., “On the Study of Generative Adversarial Networks for Cross-Lingual Voice Conversion,” 2019 IEEE Automatic Speech Recognition and Understanding Workshop (ASRU) Sentosa, Singapore, Dec. 14-18, 2019, 8 … [cited by applicant]
Tian, X., Voice Conversion with Parallel Non-Parallel Data and Synthetic Speech Detection Doctoral Thesis, Nanyang Tech University, Jurong West, Singapore, 2018, p. 1-116. [cited by applicant]
Zhu, J. et al., “Unpaired Image-to-Image Translation Using Cycle-Consisten Adversarial Networks,” 2017 IEEE International Conference on Computer Vision (ICCV), Venice, Italy, Oct. 22-29, 2017, pp. 1-18. [cited by applicant]
Rallabandi, S.S. et al., “An Approach to Cross-Lingual Voice Conversion,” Jul. 19, 2019, 7 page. [cited by applicant]