IP Library › Granted Patent US 11,488,575
Granted Patent B2
US 11,488,575 · App. 17/055,951 · Granted Nov 1, 2022

Synthesis of speech from text in a voice of a target speaker using neural networks

Inventors: Ye Jia (Santa Clara, CA); Zhifeng Chen (Sunnyvale, CA); Yonghui Wu (Fremont, CA); Jonathan Shen (Mountain View, CA); Ruoming Pang (New York, NY); Ron J. Weiss (New York, NY); Ignacio Lopez Moreno (Brooklyn, NY); Fei Ren (Mountain View, CA); Yu Zhang (Mountain View, CA); Quan Wang (Hoboken, NJ); Patrick Nguyen (Mountain View, CA)
Assignee: Google LLC
G10L13/04G10L17/04G10L19/00G06N3/08G10L2013/021
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,488,575
App. No.
17/055,951
Filed
Nov 16, 2020
Granted
Nov 1, 2022
Kind
B2
Art Unit
2677
USPC
704/259
Abstract

Methods, systems, and apparatus, including computer programs encoded on a computer storage medium, for speech synthesis. The methods, systems, and apparatus include actions of obtaining an audio representation of speech of a target speaker, obtaining input text for which speech is to be synthesized in a voice of the target speaker, generating a speaker vector by providing the audio representation to a speaker encoder engine that is trained to distinguish speakers from one another, generating an audio representation of the input text spoken in the voice of the target speaker by providing the input text and the speaker vector to a spectrogram generation engine that is trained using voices of reference speakers to generate audio representations, and providing the audio representation of the input text spoken in the voice of the target speaker for output.

Claims (25)

1. A computer-implemented method comprising:

obtaining an audio representation of speech of a target speaker;

obtaining input text for which speech is to be synthesized in a voice of the target speaker;

generating a speaker embedding vector by providing the audio representation to a speaker verification neural network that is trained to distinguish speakers from one another;

generating an audio representation of the input text spoken in the voice of the target speaker by providing the input text and the speaker embedding vector to a spectrogram generation neural network that is trained using voices of reference speakers to generate audio representations; and

providing the audio representation of the input text spoken in the voice of the target speaker to a vocoder to generate a time domain representation of the input text spoken in the voice of the target speaker; and

providing the time domain representation for playback to a user.

2. The method of claim 1 , wherein the speaker verification neural network is trained to generate speaker embedding vectors of audio representations of speech from the same speaker that are close together in an embedding space while generating speaker embedding vectors of audio representations of speech from different speakers that are distant from each other.

3. The method of claim 1 , wherein the speaker verification neural network is trained separately from the spectrogram generation neural network.

4. The method of claim 1 , wherein the speaker verification neural network is a long short-term memory (LSTM) neural network.

5. A computer-implemented method comprising:

obtaining an audio representation of speech of a target speaker;

obtaining input text for which speech is to be synthesized in a voice of the target speaker;

generating a speaker embedding vector by:

providing a plurality of overlapping sliding windows of the audio representation to a speaker verification neural network to generate a plurality of individual vector embeddings, the speaker verification neural network trained to distinguish speakers from one another; and

generating the speaker embedding vector by computing an average of the individual vector embeddings;

generating an audio representation of the input text spoken in the voice of the target speaker by providing the input text and the speaker embedding vector to a spectrogram generation neural network that is trained using voices of reference speakers to generate audio representations; and

providing the audio representation of the input text spoken in the voice of the target speaker for output.

6. The method of claim 1 , wherein the vocoder comprises a vocoder neural network.

7. The method of claim 1 , wherein the spectrogram generation neural network is a sequence-to-sequence attention neural network that is trained to predict mel spectrograms from a sequence of phoneme or grapheme inputs.

8. The method of claim 7 , wherein the spectrogram generation neural network includes an encoder neural network, an attention layer, and a decoder neural network.

9. The method of claim 8 , wherein the spectrogram generation neural network concatenates the speaker embedding vector with outputs of the encoder neural network that are provided as input to the attention layer.

10. The method of claim 1 , wherein the speaker embedding vector is different from any speaker embedding vectors used during the training of the speaker verification neural network or the spectrogram generation neural network.

11. The method of claim 1 , wherein, during the training of the spectrogram generation neural network, parameters of the speaker verification neural network are fixed.

12. A system comprising one or more computers and one or more storage devices storing instructions that when executed by the one or more computers cause the one or more computers to perform the operations of the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 17, 2020
From: JIA, YE; CHEN, ZHIFENG; WU, YONGHUI; SHEN, JONATHAN; PANG, RUOMING; WEISS, RON J.; MORENO, IGNACIO LOPEZ; REN, FEI; ZHANG, YU; WANG, QUAN; NGUYEN, PATRICK AN PHU
To: GOOGLE LLC
Reel/Frame 054386/0652 →
Continuity (2)
Provisional Application 62672835 · May 17, 2018
Related Publication 20210217404A1 · Jul 15, 2021