IP Library › Granted Patent US 11,302,300
Granted Patent B2
US 11,302,300 · App. 16/953,114 · Granted Apr 12, 2022

Method and apparatus for forced duration in neural speech synthesis

Inventors: Nick Rossenbach (Aachen, DE); Mudar Yaghi (McLean, VA)
Assignee: Applications Technology (AppTek), LLC
G10L13/02G06F40/279G06F40/58G10L15/26G10L19/02G10L25/78G10L2025/783
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,302,300
App. No.
16/953,114
Granted
Apr 12, 2022
Kind
B2
Abstract

A system and method enable one to set a target duration of a desired synthesized utterance without removing or adding spoken content. Without changing the spoken text, the voice characteristics may be kept the same or substantially the same. Silence adjustment and interpolation may be used to alter the duration while preserving speech characteristics. Speech may be translated prior to a vocoder step, pursuant to which the translated speech is constrained by the original audio duration, while mimicking the speech characteristics of the original speech.

Claims (36)

1. A system for constraining the duration of audio associated with text in a text-to-speech system, comprising:

ASR system producing a stream of text, and a time duration associated with the stream, as part of an end-to-end system;

a feature model configured to:

receive a text stream; and

produce a spectral feature output in frames associated with the text;

a duration constraint processor configured to:

receive the spectral feature output and the duration associated with the text stream; and

process the spectral feature output and the duration associated with the text stream to:

determine frames representing silence or an absence of text;

determine whether the stream is longer or shorter than the desired duration;

remove silence frames when required to reduce the duration of the spectral feature output;

add silence frames when required to increase the duration of the spectral feature output; and

perform interpolation on the spectral feature output after adjusting silence frames to make the duration of the spectral feature output match the required duration; and

a vocoder configured to:

receive the updated spectral feature output frames; and

produce synthesized audio.

2. The system of claim 1 , further comprising:

a machine translation engine, coupled to the ASR system and the feature model, configured to translate the text stream from the ASR system into a translated text stream,

wherein the feature model is further configured to:

receive the translated text stream, and

produce the spectral feature output corresponding to the translated text stream.

3. A method for constraining the duration of audio associated with text in a text-to-speech system, the method comprising:

producing a stream of text and a time duration associated with portions of the stream from an ASR system that is part of an end-to-end system;

receiving the text stream at a feature model;

generating a stream of spectral feature output in frames associated with the text;

determining frames representing silence or an absence of text;

determining whether the stream of spectral feature output is longer or shorter than the time duration associated with the text;

removing silence frames when required to reduce the duration of the spectral feature output;

adding silence frames when required to increase the duration of the spectral feature output;

performing interpolation on the spectral feature output after adjusting silence frames to make the duration of the spectral feature output match the required duration; and

synthesizing audio from the spectral feature output.

4. The method of claim 3 , wherein:

the silence frames are determined in segments of at least five frames; and

when silence frames are removed from the stream, the silence frames are removed evenly from the center of the segments.

5. The method of claim 4 , wherein, when silence frames are added to the stream, the silence frames are added evenly to the center of the segments.

6. The method of claim 3 , further comprising translating the text stream into another language using machine translation prior to receiving the text stream by the feature model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2021
From: ROSSENBACH, NICK; YAGHI, MUDAR
To: APPLICATIONS TECHNOLOGY (APPTEK), LLC
Reel/Frame 055465/0017 →
Continuity (2)
Provisional Application 62937449 · Nov 19, 2019
Related Publication 20210151028A1 · May 20, 2021