IP Library Granted Patent US 11,335,321
Granted Patent B2
US 11,335,321 · App. 17/005,974 · Granted May 17, 2022

Building a text-to-speech system from a small amount of speech data

Inventors: Ye Jia (Santa Clara, CA); Byungha Chun (Tokyo, JP); Yusuke Oda (Mountain View, CA); Norman Casagrande (Mountain View, CA); Tejas Iyer (Mountain View, CA); Fan Luo (Mountain View, CA); Russell John Wyatt Skerry-Ryan (Mountain View, CA); Jonathan Shen (Mountain View, CA); Yonghui Wu (Fremont, CA); Yu Zhang (Mountain View, CA)
Assignee: Google LLC
G10L13/04G10L13/033G10L13/086G10L15/063G10L2015/0635G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,335,321
App. No.
17/005,974
Granted
May 17, 2022
Kind
B2
Abstract

A method of building a text-to-speech (TTS) system from a small amount of speech data includes receiving a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker where the assortment of speakers does not include the target speaker. The method further includes training a TTS model using the first plurality of recorded speech samples from the assortment of speakers. Here, the trained TTS model is configured to output synthetic speech as an audible representation of a text input. The method also includes re-training the trained TTS model using the second plurality of recorded speech samples from the target speaker combined with the first plurality of recorded speech samples from the assortment of speakers. Here, the re-trained TTS model is configured to output synthetic speech resembling speaking characteristics of the target speaker.

Claims (28)

1. A method comprising:

receiving, at data processing hardware, a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker, the assortment of speakers not including the target speaker;

training, by the data processing hardware, a text-to-speech (TTS) model using the first plurality of recorded speech samples from the assortment of speakers, the trained TTS model configured to output synthetic speech as an audible representation of a text input; and

re-training, by the data processing hardware, the trained TTS model using retraining speech data, the retraining speech data comprising the second plurality of recorded speech samples from the target speaker combined with the first plurality of recorded speech samples from the assortment of speakers, the second plurality of recorded speech samples from the target speaker corresponds to less than fifty percent of the retraining speech data, the re-trained TTS model configured to output synthetic speech resembling speaking characteristics of the target speaker.

2. The method of claim 1 , wherein the TTS model comprises an encoder, a decoder, and an attention mechanism.

3. The method of claim 2 , wherein re-training the trained TTS model using retraining speech data comprises retraining the decoder and the attention mechanism of the trained TTS model, but not retraining the encoder of the trained TTS model.

4. The method of claim 2 , wherein the TTS model comprises an additive attention mechanism.

5. The method of claim 2 , wherein the TTS model comprises a location sensitive attention mechanism.

6. The method of claim 2 , wherein the TTS model comprises a dynamic convolution attention mechanism.

7. The method of claim 1 , wherein the second plurality of recorded speech samples from the target speaker corresponds to ten percent of the retraining speech data.

8. The method of claim 1 , further comprising processing, by the data processing hardware, the first plurality of recorded speech samples of the assortment of speakers to have:

consistent loudness; and

an equal duration of leading silence and trailing silence.

9. A system comprising:

data processing hardware; and

memory hardware in communication with the data processing hardware, the memory hardware storing instructions that when executed on the data processing hardware cause the data processing hardware to perform operations comprising:

receiving a first plurality of recorded speech samples from an assortment of speakers and a second plurality of recorded speech samples from a target speaker, the assortment of speakers not including the target speaker;

training a text-to-speech (TTS) model using the first plurality of recorded speech samples from the assortment of speakers, the second plurality of recorded speech samples from the target speaker corresponds to less than fifty percent of the retraining speech data, the trained TTS model configured to output synthetic speech as an audible representation of a text input; and

re-training the trained TTS model using retraining speech data, the retraining speech data comprising the second plurality of recorded speech samples from the target speaker combined with the first plurality of recorded speech samples from the assortment of speakers, the re-trained TTS model configured to output synthetic speech resembling speaking characteristics of the target speaker.

10. The system of claim 9 , wherein the TTS model comprises an encoder, a decoder, and an attention mechanism.

11. The system of claim 10 , wherein re-training the trained TTS model using retraining speech data comprises retraining the decoder and the attention mechanism of the trained TTS model, but not retraining the encoder of the trained TTS model.

12. The system of claim 10 , wherein the TTS model comprises an additive attention mechanism.

13. The system of claim 10 , wherein the TTS model comprises a location sensitive attention mechanism.

14. The system of claim 10 , wherein the TTS model comprises a dynamic convolution attention mechanism.

15. The system of claim 9 , wherein the second plurality of recorded speech samples from the target speaker corresponds to ten percent of the retraining speech data.

16. The system of claim 9 , further comprising processing, by the data processing hardware, the first plurality of recorded speech samples of the assortment of speakers to have:

consistent loudness; and

an equal duration of leading silence and trailing silence.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 15, 2020
From: JIA, YE; CHUN, BYUNGHA; ZHANG, YU; ODA, YUSUKE; CASAGRANDE, NORMAN; IYER, TEJAS; LUO, FAN; SKERRY-RYAN, RUSSELL JOHN WYATT; SHEN, JONATHAN; WU, YONGHUI
To: GOOGLE LLC
Reel/Frame 053782/0128 →
Continuity (1)
Related Publication 20220068256A1 · Mar 3, 2022
Cited By (1)
US 12,586,567