IP Library › Granted Patent US 11,276,389
Granted Patent B1
US 11,276,389 · App. 16/700,654 · Granted Mar 15, 2022

Personalizing a DNN-based text-to-speech system using small target speech corpus

Inventor: Sandesh Aryal (Kathmandu, NP)
Assignee: OBEN, INC.
G10L13/02G10L15/02G10L15/16G10L25/18G10L25/90G10L2015/025G10L2015/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,276,389
App. No.
16/700,654
Granted
Mar 15, 2022
Kind
B1
Abstract

A personalized text-to-speech system configured to perform speaker adaption is disclosed. The TTS system includes an acoustic model comprising a base neural network and a differential neural network. The base neural network is configured to generate acoustic parameters corresponding to a base speaker or voice actor, while the differential neural network is configured to generate acoustic parameters corresponding to differences between acoustic parameters of the base speaker and a particular target speaker. The output of the acoustic model is then a weighted linear combination of the output from the base neural network and differential neural network. The base neural network and differential neural network share a first input layer and first plurality of hidden layers. Thereafter, the base neural network further comprises a second plurality of hidden layers and output layer. In parallel, the differential neural network further comprises a third plurality of hidden layers and separate output layer.

Claims (22)

1. An acoustic model for a speaker adaptation text-to-speech system, wherein the acoustic model comprises:

a base neural network comprising:

a) a first input layer;

b) a first plurality of hidden layers;

c) a second plurality of hidden layers; and

d) a first output layer;

a differential neural network comprising:

a) the first input layer;

b) the first plurality of hidden layers;

c) a third plurality of hidden layers; and

d) a second output layer; and

a summing circuit configured to generate a weighted linear combination from the first output layer and second output layer;

wherein the base neural network is configured to generate acoustic parameters corresponding to a base speaker, and the differential neural network is configured to generate acoustic parameters corresponding to differences between acoustic parameters of the base speaker and a target speaker.

2. The acoustic model of claim 1 , wherein the acoustic model comprises a deep neural network.

3. The acoustic model of claim 1 , wherein the first input layer of the acoustic model is to receive a plurality of frame-level linguistic feature vectors, each comprising a numerical representation specifying an identity of a phoneme, a phonetic context, and a position of the phoneme within a syllable or word.

4. The acoustic model of claim 1 , wherein the summing circuit is configured to output a plurality of acoustic feature vector, each acoustic feature vector comprising spectral features, a pitch feature, and band aperiodicity features.

5. The acoustic model of claim 1 , wherein the summing circuit is configured to generate a weighted linear combination from the first output layer and second output layer.

6. The acoustic model of claim 5 , wherein the summing circuit is configured to apply a first weight to the first output layer and second weight to the second output layer, wherein the first and second weights depend on an amount of target speaker training data used to train the differential neural network.

7. A method of generating an acoustic model for a speaker adaptation text-to-speech system, wherein the method comprises:

training a base neural network, wherein the base neural network comprises a first input layer, a first plurality of hidden layers, a second plurality of hidden layers, and a first output layer; wherein the base neural network is configured to generate acoustic parameters corresponding to a base speaker,

after training the base neural network, then training a differential neural network, wherein the differential neural network comprises a third plurality of hidden layers connected to the first plurality of hidden layers, and a second output layer; wherein the differential neural network is configured to generate acoustic parameters corresponding to differences between acoustic parameters of the base speaker and a target speaker; and

generating a summing circuit for producing a weighted linear combination from the first output layer and second output layer.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2020
From: ARYAL, SANDESH
To: OBEN, INC.
Reel/Frame 051956/0936 →
Continuity (1)
Provisional Application 62774065 · Nov 30, 2018
Cited By (1)
US 12,536,987