IP Library › Granted Patent US 12,586,590
Granted Patent B2
US 12,586,590 · App. 17/826,987 · Granted Mar 24, 2026

Techniques for improved zero-shot voice conversion with a conditional disentangled sequential variational auto-encoder

Inventors: Chunlei Zhang (Bellevue, WA); Jiachen Lian (Palo Alto, CA); Dong Yu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
G10L19/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,590
App. No.
17/826,987
Granted
Mar 24, 2026
Kind
B2
Abstract

A method, system, apparatus, and computer-readable medium for voice conversion using a conditional disentangled sequential variational auto-encoder (C-DSVAE) is provided. The method, performed by at least one processor, includes receiving input speech segments, encoding the input speech segments via a shared encoder to generate a speaker embedding and a content embedding, and encoding a posterior distribution of the speaker embedding via a speaker encoder and encoding a posterior distribution of the content embedding via a content encoder to obtain encoded results. The method further includes enabling a content bias, reshaping the content embedding using the content bias, and generating a reconstructed speech output based on the encoded results and the reshaped content embedding.

Claims (40)

1 . A method for voice conversion using a conditional disentangled sequential variational auto-encoder (C-DSVAE), performed by at least one processor and comprising:

receiving input speech segments;

encoding the input speech segments via a shared encoder to generate a speaker embedding and a content embedding;

encoding a posterior distribution of the speaker embedding via a speaker encoder and encoding a posterior distribution of the content embedding via a content encoder to obtain encoded results;

enabling a content bias, and reshaping the content embedding by preserving phonetic information using the content bias sampled from the posterior distribution; and

generating a reconstructed speech output based on the encoded results and the reshaped content embedding,

wherein a total loss of the voice conversion is based on a reconstruction loss between the input speech segments and the reconstructed speech output, a prior distribution and the posterior distribution of the speaker embedding, and a Kullback-Leibler (KL)-Divergence between a conditional prior and the posterior distribution of the content embedding,

wherein the content bias is forced alignment, and the content bias is concatenated to prior modeling during a training process, and

wherein during the training process, the content embedding and a target embedding are concatenated when inferencing to obtain a voice conversion speech output, and the concatenated results of the content embedding and the target embedding are provided to a decoder to generate the reconstructed speech in a form of a spectrogram.

2 . The method of claim 1 , wherein the content bias is pseudo labels.

3 . The method of claim 1 , wherein the method is performed on voice cloning toolkit (VCTK) datasets.

4 . The method of claim 1 , wherein segments are randomly selected from the input speech segments for training the C-DSVAE.

5 . The method of claim 1 , further comprising converting the reconstructed speech output into a waveform.

6 . An apparatus for voice conversion using a conditional disentangled sequential variational auto-encoder (C-DSVAE), the apparatus comprising:

at least one memory configured to store program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code including:

receiving code configured to cause the at least one processor to receive input speech segments;

first encoding code configured to cause the at least one processor to encode the input speech segments via a shared encoder to generate a speaker embedding and a content embedding;

second encoding code configured to cause the at least one processor to encode a posterior distribution of the speaker embedding via a speaker encoder and encode a posterior distribution of the content embedding via a content encoder to obtain encoded results;

enabling code configured to cause the at least one processor to enable a content bias, and reshape the content embedding by preserving phonetic information using the content bias sampled from the posterior distribution; and

generating code configured to cause the at least one processor to generate a reconstructed speech output based on the encoded results and the reshaped content embedding,

wherein a total loss of the voice conversion is based on a reconstruction loss between the input speech segments and the reconstructed speech output, a prior distribution and the posterior distribution of the speaker embedding, and a Kullback-Leibler (KL)-Divergence between a conditional prior and the posterior distribution of the content embedding, and

wherein the content bias is forced alignment, and the content bias is concatenated to prior modeling during a training process, and

wherein during the training process, the content embedding and a target embedding are concatenated when inferencing to obtain a voice conversion speech output, and the concatenated results of the content embedding and the target embedding are provided to a decoder to generate the reconstructed speech in a form of a spectrogram.

7 . The apparatus of claim 6 , wherein the content bias is pseudo labels.

8 . The apparatus of claim 6 , wherein the method is performed on voice cloning toolkit (VCTK) datasets.

9 . The apparatus of claim 6 , wherein segments are randomly selected from the input speech segments for training the C-DSVAE.

10 . The apparatus of claim 6 , further comprising converting code configured to cause the at least one processor to convert the reconstructed speech output into a waveform.

11 . A non-transitory computer-readable medium storing instructions, the instructions comprising: one or more instructions that, when executed by at least one processor of an apparatus for voice conversion using a conditional disentangled sequential variational auto-encoder (C-DSVAE) storing instructions that, cause the at least one processor to:

receive input speech segments;

encode the input speech segments via a shared encoder to generate a speaker embedding and a content embedding;

encode a posterior distribution of the speaker embedding via a speaker encoder and encode a posterior distribution of the content embedding via a content encoder to obtain encoded results;

enable a content bias, and reshape the content embedding by preserving phonetic information using the content bias sampled from the posterior distribution; and

generate a reconstructed speech output based on the encoded results and the reshaped content embedding to generate the reconstructed speech output,

wherein a total loss of the voice conversion is based on a reconstruction loss between the input speech segments and the reconstructed speech output, a prior distribution and the posterior distribution of the speaker embedding, and a Kullback-Leibler (KL)-Divergence between a conditional prior and the posterior distribution of the content embedding,

wherein the content bias is forced alignment, and the content bias is concatenated to prior modeling during a training process, and

wherein during the training process, the content embedding and a target embedding are concatenated when inferencing to obtain a voice conversion speech output, and the concatenated results of the content embedding and the target embedding are provided to a decoder to generate the reconstructed speech in a form of a spectrogram.

12 . The non-transitory computer-readable medium of claim 11 , wherein the content bias is pseudo labels.

13 . The non-transitory computer-readable medium of claim 11 , wherein the method is performed on voice cloning toolkit (VCTK) datasets.

14 . The non-transitory computer-readable medium of claim 11 , wherein segments are randomly selected from the input speech segments for training the C-DSVAE.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2022
From: LIAN, JIACHEN; ZHANG, CHUNLEI; YU, DONG
To: TENCENT AMERICA LLC
Reel/Frame 060133/0341 →
Continuity (1)
Related Publication 20230386479A1 · Nov 30, 2023
References Cited (45)
US 5598507A · Kimber · 1997 [cited by examiner]
US 10347241B1 · Meng · 2019 [cited by examiner]
US 11580965B1 · Sunkara · 2023 [cited by examiner]
US 12272371B1 · Giri · 2025 [cited by examiner]
US 20140058731A1 · Tyagi · 2014 [cited by examiner]
US 20150127337A1 · Heigold · 2015 [cited by examiner]
US 20150279351A1 · Nguyen · 2015 [cited by examiner]
US 20160099007A1 · Alvarez · 2016 [cited by examiner]
US 20160232440A1 · Gregor · 2016 [cited by examiner]
US 20160267924A1 · Terao · 2016 [cited by examiner]
US 20180012602A1 · Komissarchik · 2018 [cited by examiner]
US 20180012603A1 · Komissarchik · 2018 [cited by examiner]
US 20190122101A1 · Lei · 2019 [cited by examiner]
US 20200082916A1 · Polykovskiy · 2020 [cited by examiner]
US 20200090069A1 · Mandt · 2020 [cited by examiner]
US 20200372339A1 · Che · 2020 [cited by examiner]
US 20200380952A1 · Zhang et al. · 2020 [cited by applicant]
US 20200401916A1 · Rolfe · 2020 [cited by examiner]
US 20210034969A1 · Wayne · 2021 [cited by examiner]
US 20210043216A1 · Wang · 2021 [cited by examiner]
US 20210049460A1 · Ahn · 2021 [cited by examiner]
US 20210065066A1 · Xue · 2021 [cited by examiner]
US 20210142120A1 · Min et al. · 2021 [cited by applicant]
US 20210174784A1 · Min · 2021 [cited by examiner]
US 20210183373A1 · Moritz · 2021 [cited by examiner]
US 20210272571A1 · Balasubramaniam · 2021 [cited by examiner]
US 20210280202A1 · Wang et al. · 2021 [cited by applicant]
US 20210350786A1 · Chen · 2021 [cited by examiner]
US 20210358577A1 · Zhang · 2021 [cited by examiner]
US 20210397895A1 · Sun · 2021 [cited by examiner]
US 20210406764A1 · Sha · 2021 [cited by examiner]
US 20220084371A1 · Semichev · 2022 [cited by examiner]
US 20220121934A1 · Duan · 2022 [cited by examiner]
US 20220171989A1 · Min · 2022 [cited by examiner]
US 20220189456A1 · Pang · 2022 [cited by examiner]
US 20220223066A1 · Chen · 2022 [cited by examiner]
US 20220246132A1 · Zhang · 2022 [cited by examiner]
US 20230084333A1 · Clinchant · 2023 [cited by examiner]
US 20230206898A1 · Stanton · 2023 [cited by examiner]
US 20230252972A1 · Harazi · 2023 [cited by examiner]
US 20230326445A1 · Adam · 2023 [cited by examiner]
US 20240062743A1 · Elias · 2024 [cited by examiner]
International Search Report dated Feb. 17, 2023 in International Application No. PCT/US22/43314. [cited by applicant]
Written Opinion dated Feb. 17, 2023 in International Application No. PCT/US22/43314. [cited by applicant]
Bai et al., “Contrastively Disentangled Sequential Variational Autoencoder”, NeurIPS, 2021, <https://arxiv.org/pdf/2110.12091.pdf>, pp. 1-25 (25 pages total). [cited by applicant]