IP Library Patent Application 13240886
Patent Application
App. No. 13/240,886

OBJECTIVE EVALUATION OF SYNTHESIZED SPEECH ATTRIBUTES

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/240,886
Abstract

A method of evaluating attributes of synthesized speech. The method includes processing a text input into a synthesized speech utterance using a processor of a text-to-speech system, applying a human speech utterance to a speech model to obtain a reference wherein the human speech utterance corresponds to the text input, applying the synthesized speech utterance to at least one of the speech model or an other speech model to obtain a test, and calculating a difference between the test and the reference. The method also can be used in a speech synthesis method.

Claims (36)

1 . A method of evaluating attributes of synthesized speech, comprising the steps of:

(a) processing a text input into a synthesized speech utterance using a processor of a text-to-speech system;

(b) applying a human speech utterance to a speech model to obtain a reference, wherein the human speech utterance corresponds to the text input of step (a);

(c) applying the synthesized speech utterance to at least one of the speech model or an other speech model to obtain a test; and

(d) calculating a difference between the test and the reference.

2 . The method of claim 1 , wherein step (b) includes training the speech model with a human speech utterance having subjective speech data associated therewith, step (c) includes training the other speech model with the synthesized speech from step (a), wherein the reference is a human speech model, and the test is a synthesized speech model.

3 . The method of claim 2 , wherein the human speech model is a human speech Hidden Markov Model (HMM), and the synthesized speech model is a synthesized speech HMM.

4 . The method of claim 1 , wherein the speech model is a codebook that is vector quantized from a corpus of human speech utterances having subjective speech data associated therewith, and in step (b) the human speech utterance is applied to the codebook and in step (c) the synthesized speech utterance is applied to the codebook, wherein the reference is a human speech sequence of clusters of the codebook, and the test is a synthesized speech sequence of clusters of the codebook.

5 . A method of evaluating attributes of synthesized speech, comprising the steps of:

(a) processing a text input into synthesized speech using a processor of a text-to-speech system;

(b) training a first speech model with a human speech utterance having subjective speech data associated therewith and corresponding to the text input of step (a);

(c) training a second speech model with the synthesized speech from step (b); and

(d) calculating a statistical distance between the speech models.

6 . The method of claim 5 , further comprising the steps of:

(e) repeating steps (a) through (d) for a plurality of text inputs and corresponding first and second speech models to generate a plurality of statistical differences; and

(f) producing a correlation between the statistical distances calculated in step (e) and the subjective speech data.

7 . The method of claim 6 , further comprising the step of:

(g) predicting attributes of synthesized speech produced by the text-to-speech system based on the correlation from step (f).

8 . The method of claim 5 , wherein the human speech utterances and the associated subjective speech data are from a corpus of phonemically and lexically transcribed human speech annotated with subjective speech data.

9 . The method of claim 5 , wherein the speech models are at least one of Hidden Markov Models or Gaussian Mixture Models.

10 . A method of evaluating attributes of synthesized speech, comprising the steps of:

(a) processing a text input into a synthesized speech utterance using a processor of a text-to-speech system;

(b) applying a human speech utterance to a codebook that is vector quantized from a corpus of human speech utterances having subjective speech data associated therewith to obtain a human speech sequence of clusters of the codebook, wherein the human speech utterance corresponds to the text input of step (a);

(c) applying the synthesized speech utterance to the speech model to obtain a synthesized speech sequence of clusters of the codebook; and

(d) calculating a statistical distance between the synthesized speech sequence of clusters of the codebook and the human speech sequence of clusters of the codebook.

11 . A method of speech synthesis, comprising the steps of:

(a) receiving a text input in a text-to-speech system;

(b) processing the text input into a synthesized speech utterance using a processor of the system; and

(c) evaluating attributes of the synthesized speech, including:

(c1) applying to a speech model a human speech utterance corresponding to the text input to obtain a reference;

(c2) applying the synthesized speech utterance to at least one of the speech model or an other speech model to obtain a test; and

(c3) calculating a statistical distance between the test and the reference.

12 . The method of claim 11 , wherein step (c) also includes a sub-step (c4) correlating the statistical distance with a subjective speech score.

13 . The method of claim 12 , further comprising the step of:

(d) outputting the synthesized speech to a user via a loudspeaker, if the subjective speech score is greater than a predetermined acceptable level.

14 . The method of claim 11 , wherein the speech model is a codebook that is vector quantized from a corpus of human speech utterances having subjective speech data associated therewith, and the human speech utterance and the synthesized speech utterance are applied to the codebook, wherein the reference is a human speech sequence of clusters of the codebook, and the test is a synthesized speech sequence of clusters of the codebook.

Assignments (3)
RELEASE OF SECURITY INTEREST Recorded Nov 7, 2014
From: WILMINGTON TRUST COMPANY
To: GENERAL MOTORS LLC
Reel/Frame 034183/0436 →
SECURITY AGREEMENT Recorded Jun 22, 2012
From: GENERAL MOTORS LLC
To: WILMINGTON TRUST COMPANY
Reel/Frame 028423/0432 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 23, 2011
From: TALWAR, GAURAV; ZHAO, XUFANG
To: GENERAL MOTORS LLC
Reel/Frame 026954/0755 →