IP Library › Granted Patent US 12,646,496
Granted Patent B2
US 12,646,496 · App. 18/608,583 · Granted Jun 2, 2026

Methods and systems of text-conditioned audio-visual speech generation with multi-modal latent diffusion models

Inventors: Gaurav Sharma (Newark, CA); Siddharth Srivastava (New Delhi, IN); Neeraj Matiyali (Haldwani, IN)
Assignee: TensorType Inc.
G10L13/02G06T13/40G10L13/08G10L19/02G10L21/0208G10L25/18G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,496
App. No.
18/608,583
Granted
Jun 2, 2026
Kind
B2
Abstract

Methods, systems, and computer programs are presented for audio-visual speech generation with multi-modal latent diffusion models. One method includes encoding raw audio signals and video frames into respective latent spaces using audio and visual autoencoders. A text transcript is processed into phoneme sequences using a text transcript processor. The audio and visual latent spaces are conditioned using the text transcript and a conditioning variable. Joint distributions of the visual and audio latent spaces, text transcript, and conditioning variable are learned using a multi-modal latent diffusion model. The model adds noise to the latent audio-visual representations and predicts the noise through denoising neural networks. An inverted diffusion process is utilized to generate diverse speech content and speaker characteristics, resulting in realistic audio-visual speech. The technology presented provides a novel approach to conditional speech generation with potential applications in speech synthesis, voice conversion, and speech recognition.

Claims (42)

1 . A computer-implemented method, for conditional audio-visual speech generation, comprising:

encoding raw audio signals into an audio latent space using an audio autoencoder;

encoding raw video frames into a visual latent space using a visual autoencoder;

processing text transcripts associated with the raw audio signals into phoneme sequences using a text transcript processor;

conditioning the audio latent space and the visual latent space using the text transcripts and a conditioning variable;

learning a generative model based on joint distributions of the visual latent space, the audio latent space, the text transcripts, and the conditioning variable using a multi-modal latent diffusion model by adding noise to latent audio-visual representations of the audio latent space and the visual latent space and predicting said noise through denoising neural networks;

receiving a text input and a reference video speech of a speaker; and

generating, by the generative model, an audio-visual output based on the text input and the reference video speech of the speaker, wherein an audio signal of the audio-visual output contains speech based on the text input and a video signal of the audio-visual output includes a video of the speaker coordinated with the audio signal.

2 . The computer-implemented method as claimed in claim 1 , wherein the conditioning variable includes parameters specifying emotional expression in the raw audio signals and raw video frames.

3 . The computer-implemented method as claimed in claim 2 , wherein the conditioning variable includes parameters related to gender of the speaker, allowing the generated audio-visual output to exhibit gender-specific characteristics.

4 . The computer-implemented method as claimed in claim 1 , wherein the audio autoencoder employs convolutional neural networks (CNNs) to transform the raw audio signals into mel-spectrogram representations before encoding into the audio latent space.

5 . The computer-implemented method as claimed in claim 1 , wherein the visual autoencoder utilizes 3D convolutional networks to encode the raw video frames into the visual latent space.

6 . The computer-implemented method as claimed in claim 1 , wherein a style encoder encodes a reference audio-visual sample to extract specific style characteristics to be applied to the generated audio-visual output.

7 . The computer-implemented method as claimed in claim 1 , wherein the denoising neural networks utilize cross-modal connections to align audio and visual modalities.

8 . The computer-implemented method as claimed in claim 1 , wherein the inverted diffusion process includes generating diverse speech content by manipulating the audio latent space and the visual latent space representations based on the text transcripts and additional conditioning variables.

9 . The computer-implemented method as claimed in claim 1 , wherein the inverted diffusion process includes adjusting noise levels based on a complexity of desired speech content, ensuring realistic generation in varied scenarios.

10 . The computer-implemented method as claimed in claim 1 , wherein the denoising neural networks employ attention mechanisms to focus on specific phonemes and facial expressions, enhancing an accuracy of noise prediction and speech generation.

11 . The computer-implemented method as claimed in claim 1 , wherein the joint distributions are learned iteratively, refining the generation process based on user feedback and preferences for specific speaker characteristics and speech styles.

12 . A system comprising:

a memory comprising instructions; and

one or more computer processors, wherein the instructions, when executed by the one or more computer processors, cause the system to perform operations comprising:

encoding raw audio signals into an audio latent space using an audio autoencoder;

encoding raw video frames into a visual latent space using a visual autoencoder;

processing text transcripts associated with the raw audio signals into phoneme sequences using a text transcript processor;

conditioning the audio latent space and the visual latent space using the text transcripts and a conditioning variable;

learning a generative model based on joint distributions of the visual latent space, the audio latent space, the text transcripts transcript, and the conditioning variable using a multi-modal latent diffusion model by adding noise to latent audio-visual representations of the audio latent space and the visual latent space and predicting said noise through denoising neural networks;

receiving a text input and a reference video speech of a speaker; and

generating, by the generative model, an audio-visual output based on the text input and the reference video speech of the speaker, wherein an audio signal of the audio-visual output contains speech based on the text input and a video signal of the audio-visual output includes a video of the speaker coordinated with the audio signal.

13 . The system as recited in claim 12 , wherein the conditioning variable includes parameters specifying emotional expression in the raw audio signals and raw video frames.

14 . The system as recited in claim 13 , wherein the conditioning variable includes parameters related to gender of the speaker, allowing the generated audio-visual output to exhibit gender-specific characteristics.

15 . The system as recited in claim 12 , wherein the audio autoencoder employs convolutional neural networks (CNNs) to transform the raw audio signals into mel-spectrogram representations before encoding into the audio latent space.

16 . The system as recited in claim 12 , wherein the visual autoencoder utilizes 3D convolutional networks to encode the raw video frames into the visual latent space.

17 . A non-transitory machine-readable storage medium including instructions that, when executed by a machine, cause the machine to perform operations comprising:

encoding raw audio signals into an audio latent space using an audio autoencoder;

encoding raw video frames into a visual latent space using a visual autoencoder;

processing text transcripts associated with the raw audio signals into phoneme sequences using a text transcript processor;

conditioning the audio latent space and the visual latent space using the text transcripts and a conditioning variable;

learning a generative model based on joint distributions of the visual latent space, the audio latent space, the text transcripts, and the conditioning variable using a multi-modal latent diffusion model by adding noise to latent audio-visual representations of the audio latent space and the visual latent space and predicting said noise through denoising neural networks;

receiving a text input and a reference video speech of a speaker; and

generating, by the generative model, an audio-visual output based on the text input and the reference video speech of the speaker, wherein an audio signal of the audio-visual output contains speech based on the text input and a video signal of the audio-visual output includes a video of the speaker coordinated with the audio signal.

18 . The non-transitory machine-readable storage medium as recited in claim 17 , wherein the conditioning variable includes parameters specifying emotional expression in the raw audio signals and raw video frames.

19 . The non-transitory machine-readable storage medium as recited in claim 18 , wherein the conditioning variable includes parameters related to gender of the speaker, allowing the generated audio-visual output to exhibit gender-specific characteristics.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2024
From: TENSORTOUR INC.
To: TENSORTYPE INC.
Reel/Frame 067953/0375 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 10, 2024
From: SHARMA, GAURAV; SRIVASTAVA, SIDDHARTH; MATIYALI, NEERAJ
To: TENSORTOUR INC.
Reel/Frame 067953/0500 →
Continuity (2)
Provisional Application 63442821 · Feb 2, 2023
Related Publication 20250292763A1 · Sep 18, 2025
References Cited (7)
US 7117155B2 · Cosatto · 2006 [cited by examiner]
US 20250292763A1 · Sharma · 2025 [cited by examiner]
CN 112041924A · 2020 [cited by examiner]
KR 20220090586A · 2022 [cited by examiner]
WO WO2005031654A1 · 2005 [cited by examiner]
Song, Hyoung-Kyu, et al. “Talking face generation with multilingual tts.” Proceedings of the ieee/cvf conference on computer vision and pattern recognition. 2022. (Year: 2022). [cited by examiner]
Shen, Shuai, et al. “Difftalk: Crafting diffusion models for generalized audio-driven portraits animation.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2023. (Year: 2023). [cited by examiner]