IP Library Patent Application 11704522
Patent Application
App. No. 11/704,522

Line Spectrum pair density modeling for speech applications

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
11/704,522
Abstract

Novel techniques for providing superior performance and sound quality in speech applications, such as speech synthesis, speech coding, and automatic speech recognition, are hereby disclosed. In one illustrative embodiment, a method includes modeling a speech signal with parameters comprising line spectrum pairs. Density parameters are provided based on the density of the line spectrum pairs. A speech application output, such as synthesized speech, is provided based at least in part on the line spectrum pair density parameters. The line spectrum pair density parameters use computing resources efficiently while providing improved performance and sound quality in the speech application output.

Claims (31)

1 . A method comprising:

modeling a speech signal with parameters comprising line spectrum pairs;

providing density parameters based on a measure of density of two or more of the line spectrum pairs; and

providing a speech application output based at least in part on the density parameters.

2 . The method of claim 1 , wherein providing the speech application output based at least in part on the density parameters comprises modeling the speech signal with greater clarity in segments of the speech signal associated with an increased density of the line spectrum pairs.

3 . The method of claim 1 , further providing dynamic parameters based at least in part on changes in the density of two or more of the line spectrum pairs, from one part of the speech signal to another, and providing the speech application output based at least in part on the dynamic parameters.

4 . The method of claim 1 , wherein the speech application output comprises an automatic speech synthesis output.

5 . The method of claim 1 , wherein the speech application output comprises a speech coding output.

6 . The method of claim 1 , wherein the speech application output comprises a speech recognition output.

7 . The method of claim 1 , wherein modeling the speech signal comprises using at least one hidden Markov model trained at least in part with the parameters comprising line spectrum pairs.

8 . The method of claim 7 , further comprising using a maximum likelihood function to determine the parameters.

9 . The method of claim 1 , wherein modeling the speech signal comprises converting the speech signal to a sequence of feature vectors, wherein the line spectrum pairs are comprised in the feature vectors.

10 . The method of claim 1 , further comprising providing dynamical density parameters based on a measure of changes in the density of two or more of the line spectrum pairs over time, and providing the speech application output based also at least in part on the dynamical density parameters.

11 . The method of claim 1 , further comprising selecting a fixed number of line spectrum pair frequencies per frame used for modeling the speech signal based at least in part on an evaluation of computing resources available for the modeling.

12 . The method of claim 1 , further comprising sharpening one or more formant frequencies prior to determining the line spectrum pairs.

13 . The method of claim 1 , wherein modeling the speech signal with parameters comprising line spectrum pairs, comprises transforming observation feature vectors extracted from the speech signal, using a block matrix that provides the observation feature vectors, differences between adjacent observation feature vectors, and rates of change in the difference between the adjacent observation feature vectors.

14 . The method of claim 13 , wherein providing the density parameters comprises modifying the block matrix to compare the observation feature vectors between two adjacent line spectrum pair frequencies, and using the comparison to evaluate a frequency difference between the two adjacent line spectrum pair frequencies.

15 . The method of claim 1 , wherein modeling the speech signal with parameters comprising line spectrum pairs and providing the density parameters based on the measure of density of the two or more of the line spectrum pairs, are performed by a training portion of a system, and providing the speech application output based at least in part on the density parameters is performed by a speech application output portion of a system.

16 . The method of claim 1 , further comprising using at least 24 line spectrum pairs for modeling the speech signal.

17 . The method of claim 1 , wherein the speech signal is modeled with parameters that further comprise one or more of: gain, duration, pitch, or a voiced/unvoiced distinction.

18 . A medium comprising instructions that are readable and executable by a computing system, wherein the instructions configure the computing system to train and implement a speech application system, comprising configuring the computing system to:

extract features from a set of speech signals, wherein the features comprise line spectrum pairs;

evaluate differences between the frequencies of adjacent line spectrum pairs;

use the extracted features, including the differences between the frequencies of adjacent line spectrum pairs, for training one or more hidden Markov models; and

synthesize a speech signal having enhanced signal clarity in one or more portions of a frequency spectrum in which the differences between the frequencies of adjacent line spectrum pairs are indicated to be relatively small.

19 . The medium of claim 18 , further comprising configuring the computing system to assign at least one of: a number of line spectrum pairs, or a frame size for the synthesized speech signal, based in part on computing resources available to the computing system.

20 . A computing system configured to synthesize speech, the system comprising:

means for modeling information content of speech signals using hidden Markov modeling;

means for evaluating line spectrum pairs in a linear predictive coding power spectrum representing the speech signals;

means for evaluating density of the line spectrum pairs; and

means for concentrating the information content of the speech signals in frequency ranges in which the density of the line spectrum pairs is concentrated.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 15, 2015
From: MICROSOFT CORPORATION
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 034766/0509 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 12, 2007
From: SOONG, FRANK KAO-PING; QIAN, YAO
To: MICROSOFT CORPORATION
Reel/Frame 018992/0020 →