IP Library Granted Patent US 12,573,372
Granted Patent B2
US 12,573,372 · App. 18/051,507 · Granted Mar 10, 2026

Text-to-speech system with variable frame rate

Inventors: Steve Pearson (Felton, CA); Jon Grossman (Cupertino, CA)
Assignee: SoundHound AI IP, LLC
G10L13/047G10L13/06
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,573,372
App. No.
18/051,507
Granted
Mar 10, 2026
Kind
B2
Abstract

A neural TTS system is trained to generate key acoustic frames at variable rates while omitting other frames. The frame skipping depends on the acoustic features to be generated for the input text. The TTS system can interpolate frames between the key frames at a target rate for a vocoder to synthesis audio samples.

Claims (30)

1 . A computer-implemented method of speech synthesis, the method comprising:

receiving a sequence of symbols; and

synthesizing from the sequence of symbols, by a speech synthesis model, a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and one or more skipped frames is associated with the average key frame rate input.

2 . The computer-implemented method of claim 1 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.

3 . The computer-implemented method of claim 1 , wherein a frame rate of the plurality of key frames is variable and the at least one interpolation parameter indicates a length of time between the key frames.

4 . The computer-implemented method of claim 1 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.

5 . The computer-implemented method of claim 1 , further comprising:

interpolating, by an interpolation model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and

generating, by a vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.

6 . The computer-implemented method of claim 1 , further comprising:

interpolating, by a vocoder model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and

generating, by the vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.

7 . A computer-implemented method of speech synthesis, the method comprising:

receiving, at a speech synthesis model, a sequence of symbols for speech synthesis;

synthesizing a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and the one or more skipped frames is associated with the average key frame rate input;

interpolating, by an interpolation model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and

generating, by a vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.

8 . The computer-implemented method of claim 7 , wherein the plurality of key frames have a variable frame rate.

9 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.

10 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter indicates a length of time between the plurality of key frames.

11 . The computer-implemented method of claim 7 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.

12 . A computer-implemented method of speech synthesis, the method comprising:

receiving, at a speech synthesis model, a sequence of symbols for speech synthesis;

synthesizing a plurality of key frames based on an average key frame rate input, wherein a key frame comprises at least one interpolation parameter indicating the ratio of the number of the plurality of key frames and the one or more skipped frames is associated with the average key frame rate input;

interpolating, by a vocoder model, one or more interpolated frames based on the plurality of key frames and the at least one interpolation parameter; and

generating, by the vocoder model, speech waveforms of the sequence of symbols based on the plurality of key frames and the one or more interpolated frames.

13 . The computer-implemented method of claim 12 , wherein the plurality of key frames have a variable frame rate.

14 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter indicates one or more skipped frames between the plurality of key frames.

15 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter indicates a length of time between the plurality of key frames.

16 . The computer-implemented method of claim 12 , wherein the at least one interpolation parameter comprises an indicated interpolation mode.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jun 11, 2024
From: ACP POST OAK CREDIT II LLC, AS COLLATERAL AGENT
To: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
Reel/Frame 067698/0845 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 27, 2023
From: SOUNDHOUND AI IP HOLDING, LLC
To: SOUNDHOUND AI IP, LLC
Reel/Frame 064205/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 23, 2023
From: SOUNDHOUND, INC.
To: SOUNDHOUND AI IP HOLDING, LLC
Reel/Frame 064083/0484 →
SECURITY INTEREST Recorded Apr 17, 2023
From: SOUNDHOUND, INC.; SOUNDHOUND AI IP, LLC
To: ACP POST OAK CREDIT II LLC
Reel/Frame 063349/0355 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 2, 2022
From: PEARSON, STEVE; GROSSMAN, JON
To: SOUNDHOUND, INC.
Reel/Frame 061622/0271 →
Continuity (1)
Related Publication 20240144910A1 · May 2, 2024
References Cited (4)
US 8527276B1 · Senior · 2013 [cited by examiner]
US 20230335110A1 · Kenter · 2023 [cited by examiner]
Shen, Jonathan, Ruoming Pang, Ron J. Weiss, Mike Schuster, Navdeep Jaitly, Zongheng Yang, Zhifeng Chen et al. “Natural tts synthesis by conditioning wavenet on mel spectrogram predictions.” In 2018 IEEE international co… [cited by applicant]
Valin, Jean-Marc, and Jan Skoglund. “LPCNet: Improving neural speech synthesis through linear prediction.” In ICASSP 2019-2019 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), pp. 5891-… [cited by applicant]