IP Library › Granted Patent US 12,658,175
Granted Patent B2
US 12,658,175 · App. 18/494,640 · Granted Jun 16, 2026

Controllable diffusion-based speech generative model

Inventors: Kyungguen Byun (San Diego, CA); Sunkuk Moon (San Diego, CA); Erik Visser (San Diego, CA)
Assignee: QUALCOMM Incorporated
G10L13/10G10L13/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,658,175
App. No.
18/494,640
Granted
Jun 16, 2026
Kind
B2
Abstract

Systems and techniques described herein relate to a diffusion-based model for generating converted speech from a source speech based on target speech. For example, a device may extract first prosody data from input data and may generate a content embedding based on the input data. The device may extract second prosody data from target speech, generate a speaker embedding from the target speech, and generate a prosody embedding from the second prosody data. The device may generate, based on the first prosody data and the prosody embedding, converted prosody data. The device may then generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding.

Claims (70)

1 . An apparatus to generate output speech from input data, comprising:

one or more memories configured to store the input data; and

one or more processors coupled to the one or more memories and configured to:

extract first prosody data from the input data;

generate a content embedding based on the input data;

extract second prosody data from target speech;

generate a speaker embedding from the target speech;

generate a prosody embedding from the second prosody data;

generate, based on the first prosody data and the prosody embedding, converted prosody data;

generate a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and

generate converted speech based on converted spectrogram.

2 . The apparatus of claim 1 , wherein the input data comprises one or more of speech data or text data.

3 . The apparatus of claim 2 , wherein the input data comprises one of speech data and text data.

4 . The apparatus of claim 1 , wherein the first prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.

5 . The apparatus of claim 1 , wherein the second prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.

6 . The apparatus of claim 1 , wherein the one or more processors are configured to:

generate the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding via a decoder comprising a diffusion decoder or a non-diffusion decoder.

7 . The apparatus of claim 1 , wherein the one or more processors are configured to:

generate, based on the converted prosody data, a predicted global speaking rate and via a rate control engine, a speaking rate for the converted spectrogram; and

generating, via a vocoder, the converted speech based on the input data.

8 . The apparatus of claim 7 , wherein the vocoder comprises a neural vocoder.

9 . The apparatus of claim 7 , wherein the rate control engine is configured to manipulate a speaking rate depending upon a predicted speed.

10 . The apparatus of claim 7 , wherein the one or more processors are configured to:

generate, based on the converted prosody data via the rate control engine, the speaking rate for the converted spectrogram independent of an automatic speech recognition model.

11 . The apparatus of claim 1 , wherein the one or more processors are configured to:

extract the first prosody data from the input data via a first prosody extractor engine;

generate the content embedding based on the input data via a content encoder;

extract the second prosody data from target speech via a second prosody extractor engine;

generate the speaker embedding from the target speech via a speaker encoder;

generate the prosody embedding from the second prosody data via a prosody encoder;

generate, based on the first prosody data and the prosody embedding, converted prosody data via a prosody conversion engine; and

generate the converted spectrogram based on the converted prosody data, the speaker embedding and the content embedding via a decoder.

12 . The apparatus of claim 11 , wherein the apparatus comprises the decoder, and wherein the decoder is configured to synthesize a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosody data.

13 . The apparatus of claim 11 , the prosody encoder is configured to generate the prosody embedding at one or more of a frame-level or a sentence-level.

14 . The apparatus of claim 13 , wherein the apparatus comprises the prosody encoder, and wherein the prosody encoder is configured to generate the prosody embedding at the frame-level to enable frame-level intonation control.

15 . The apparatus of claim 1 , wherein the input data comprises speech data, the apparatus further comprising one or more microphones configured to capture the speech data.

16 . The apparatus of claim 1 , further comprising one or more speakers configured to output speech data comprising the converted prosody data.

17 . A method of generating output speech from input, the method comprising:

extracting first prosody data from input data;

generating a content embedding based on the input data;

extracting second prosody data from target speech;

generating a speaker embedding from the target speech;

generating a prosody embedding from the second prosody data;

generating, based on the first prosody data and the prosody embedding, converted prosody data;

generating a converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding; and

generating converted speech based on converted spectrogram.

18 . The method of claim 17 , wherein the input data comprises one or more of speech data or text data.

19 . The method of claim 18 , wherein the input data comprises one of speech data and text data.

20 . The method of claim 17 , wherein the first prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.

21 . The method of claim 17 , wherein the second prosody data comprises one or more of a fundamental frequency, an energy value and a speed value.

22 . The method of claim 17 , further comprising:

generating the converted spectrogram based on the converted prosody data, the speaker embedding, and the content embedding via a decoder comprising a diffusion decoder or a non-diffusion decoder.

23 . The method of claim 17 , further comprising:

generating, based on the converted prosody data, a predicted global speaking rate and via a rate control engine, a speaking rate for the converted spectrogram; and

generating, via a vocoder, the converted speech based on the input data.

24 . The method of claim 23 , wherein the vocoder comprises a neural vocoder.

25 . The method of claim 23 , further comprising manipulating, using the rate control engine, a speaking rate depending upon a predicted speed.

26 . The method of claim 23 , further comprising:

generating, based on the converted prosody data via the rate control engine, the speaking rate for the converted spectrogram independent of an automatic speech recognition model.

27 . The method of claim 17 , further comprising:

extracting the first prosody data from the input data via a first prosody extractor engine;

generating the content embedding based on the input data via a content encoder;

extracting the second prosody data from target speech via a second prosody extractor engine;

generating the speaker embedding from the target speech via a speaker encoder;

generating the prosody embedding from the second prosody data via a prosody encoder;

generating, based on the first prosody data and the prosody embedding, converted prosody data via a prosody conversion engine; and

generating the converted spectrogram based on the converted prosody data, the speaker embedding and the content embedding via a decoder.

28 . The method of claim 27 , further comprising synthesizing, using a decoder, a speech spectrum conditioned on the content embedding, the speaker embedding, and the converted prosody data.

29 . The method of claim 27 , further comprising generating, using the prosody encoder, the prosody embedding at one or more of a frame-level or a sentence-level.

30 . The method of claim 29 , further comprising generating, using a prosody encoder, the prosody embedding at the frame-level to enable frame-level intonation control.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2023
From: BYUN, KYUNGGUEN; MOON, SUNKUK; VISSER, ERIK
To: QUALCOMM INCORPORATED
Reel/Frame 065737/0744 →
Continuity (2)
Provisional Application 63580660 · Sep 5, 2023
Related Publication 20250078810A1 · Mar 6, 2025
References Cited (7)
US 11605388B1 · Gupta · 2023 [cited by examiner]
US 20200342852A1 · Kim · 2020 [cited by examiner]
US 20250349282A1 · Li · 2025 [cited by examiner]
Byun K., et al., “Highly Controllable Diffusion-Based Any-to-Any Voice Conversion Model with Frame-Level Prosody Feature”, arXiv:2309.03364v1 [cs.SD], Arxiv.Org, Cornell University Library, 201 Olin Library Cornell Univ… [cited by applicant]
International Search Report and Written Opinion—PCT/US2024/044535—ISA/EPO—Oct. 25, 2024. [cited by applicant]
Lee Y., et al., “Robust and Fine-Grained Prosody Control of End-to-End Speech Synthesis”, ICASSP 2019—2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), IEEE, May 12, 2019, pp. 5911-… [cited by applicant]
Lyu Z., et al., “Enriching Style Transfer in Multi-Scale Control Based Personalized End-to-End Speech Synthesis”, 2022 12th International Conference on Information Science and Technology (ICIST), IEEE, Kaifeng, China, O… [cited by applicant]