IP Library Granted Patent US 7,702,503
Granted Patent B2
US 7,702,503 · App. 12/183,591 · Granted Apr 20, 2010

Voice model for speech processing based on ordered average ranks of spectral features

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,702,503
App. No.
12/183,591
Granted
Apr 20, 2010
Kind
B2
Abstract

Methods and arrangements for generating a voice model in speech processing. Upon accepting at least two input vectors with spectral features, vectors of ranks are created via ranking values of the spectral features of each input vector, ordered vectors are created via arranging the values of each input vector according to rank, and a vector of ordered average values is created via determining the average of corresponding values of the ordered vectors. Thence, a vector of ordered average ranks is created via determining the sum or average of the vectors of ranks, a vector of ordered ranks is created via ranking the values of the ordered average ranks and a spectral feature vector is created via employing the rank order represented by the vector of ordered ranks to reorder the vector of ordered average ranks.

Claims (64)

1. An apparatus comprising a computer processor configured for executing operations on encoded computer program instructions stored in computer memory and arranged for generating a spectral feature vector data output, said program instructions comprising:

an arrangement for accepting at least two input vectors, each input vector including speech or audio or voice spectral features;

an arrangement for creating vectors of ranks via ranking values of the spectral features of each input vector;

an arrangement for creating ordered vectors via arranging the values of each input vector according to rank;

an arrangement for creating a vector of ordered average values via determining the average of corresponding values of the ordered vectors;

an arrangement for creating a vector of ordered average ranks via determining the sum or average of the vectors of ranks;

an arrangement for creating a vector of ordered ranks via ranking the values of the ordered average ranks;

an arrangement for creating a spectral feature vector via employing the rank order represented by the vector of ordered ranks to reorder the vector of ordered average values.

2. The apparatus according to claim 1 , further comprising:

an arrangement for accepting speech or audio or voice input;

said arrangement for accepting at least two input vectors being adapted to develop input vectors associated with the speech or audio or voice input spectral features;

said apparatus further comprising an arrangement for providing probabilities that each input vector belongs to one or more classes;

said arrangements for creating a vector of ordered average values and creating a vector of ordered average ranks being adapted to assign the probabilities as weights to the input vectors; and

said apparatus further comprising an arrangement for developing a voice model based on the spectral feature vector.

3. The apparatus according to claim 2 , wherein said arrangement for providing probabilities is adapted to probabilities that each vector belongs to particular context-dependent sub-phone units.

4. The apparatus according to claim 2 , further comprising:

an arrangement for generating mel frequency log spectra associated with the speech or audio or voice input;

said arrangement for generating mel frequency log spectra being adapted to:

segment the speech or audio or voice input into a plurality of segments;

apply a Fourier transform and generating a Fourier spectrum for each segment;

bin the Fourier spectrum into channels based on mel frequency; and

determine the logarithm of each channel.

5. The apparatus according to claim 4 , wherein:

said arrangement for creating vectors of ranks and creating ordered vectors being adapted to sort the mel bins to provide a set of ranks and values;

said arrangement for creating a spectral feature vector being further adapted to weight the sorted values and ranks with the weights and accumulate the weighted ranks and weights as two sets of n-dimensional sums, and develop a total weight for each context-dependent sub-phone unit; and

said arrangements for creating a vector of ordered average values and creating a vector of ordered average ranks being adapted to divide the sums corresponding to each context-dependent sub-phone unit by corresponding total weights to yield a n-dimensional vector of average values and a vector of average ranks.

6. The apparatus according to claim 4 , wherein the input is speech input.

7. The apparatus according to claim 6 , wherein: said arrangement for providing probabilities is adapted to:

determine the probabilities via providing a transcription of the audio input;

expand words associated with the transcription into phonetic sequences; and

join the phonetic sequences based on a word sequence of the transcription;

determine context-dependent sub-phonetic units for each phone and employ a speech recognition model to align the result with the segmented audio input.

8. The apparatus according to claim 2 , wherein said at least two input vectors correspond to at least two voice models, whereby said apparatus is adapted for creating from at least two voice models an additional voice model having predetermined characteristics.

9. The apparatus according to claim 2 , wherein the voice model is used as a vector quantization codebook for speech coding.

10. The apparatus according to claim 2 , wherein the voice model is employed in voice transformation via regenerating speech from a sequence of model vectors corresponding to the time sequence of classes corresponding to the audio input.

11. The apparatus according to claim 2 , wherein the voice model is used for voice morphing via:

determining at least one time sequence of classes corresponding to the input speech;

determining a time varying weighted average of spectral feature vectors of:

at least two voice models; or

at least one input vector and at least one voice model; and

regenerating speech from the resulting sequence.

12. The apparatus according to claim 2 , wherein the voice model is used for speech synthesis via:

expanding a sequence of words into phonetic sequences;

joining the phonetic sequences based on a word sequence;

expanding the phonetic sequence into one or more classes;

sequencing vectors associated with the voice model according to a class sequence, and

generating speech from the resulting sequence.

13. The apparatus according to claim 1 , wherein said at least two input vectors correspond to at least two spectral models, whereby said apparatus is adapted for creating from at least two spectral models an additional spectral model having predetermined characteristics.

14. The apparatus according to claim 1 , wherein said apparatus is adapted for voice morphing via:

determining a first time sequence of classes corresponding to one instance of input speech;

determining a second time sequence of classes corresponding to another instance of input speech;

determining a time-varying weighted average of the vectors of corresponding classes; and

regenerating speech from the resulting sequence.

15. The apparatus according to claim 1 , wherein:

the set of input vectors is a contiguous set of vectors taken from a time sequence of spectral feature vectors; and

said apparatus further comprises an arrangement for assigning weights for averaging based on the relative position in time of the input spectral feature vectors.

16. A computer program storage device medium readable by a computer processor machine, tangibly embodying an encoded program of instructions executable by the machine to perform method steps for generating a spectral feature vector data output, said method comprising the steps of:

accepting at least two input vectors, each input vector including speech or audio or voice spectral features;

creating vectors of ranks via ranking values of the spectral features of each input vector;

creating ordered vectors via arranging the values of each input vector according to rank;

creating a vector of ordered average values via determining the average of corresponding values of the ordered vectors;

creating a vector of ordered average ranks via determining the sum or average of the vectors of ranks;

creating a vector of ordered ranks via ranking the values of the ordered average ranks;

creating a spectral feature vector via employing the rank order represented by the vector of ordered ranks to reorder the vector of ordered average values.

Assignments (8)
RELEASE (REEL 052935 / FRAME 0584) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0818 →
CORRECTIVE ASSIGNMENT TO CORRECT THE REPLACE THE CONVEYANCE DOCUMENT WITH THE NEW ASSIGNMENT PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Apr 19, 2022
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 059804/0186 →
SECURITY AGREEMENT Recorded Jun 15, 2020
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A.
Reel/Frame 052935/0584 →
RELEASE OF SECURITY INTEREST Recorded Jun 12, 2020
From: BARCLAYS BANK PLC
To: CERENCE OPERATING COMPANY
Reel/Frame 052927/0335 →
SECURITY AGREEMENT Recorded Nov 7, 2019
From: CERENCE OPERATING COMPANY
To: BARCLAYS BANK PLC
Reel/Frame 050953/0133 →
CORRECTIVE ASSIGNMENT TO CORRECT THE ASSIGNEE NAME PREVIOUSLY RECORDED AT REEL: 050836 FRAME: 0191. ASSIGNOR(S) HEREBY CONFIRMS THE INTELLECTUAL PROPERTY AGREEMENT. Recorded Oct 29, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE OPERATING COMPANY
Reel/Frame 050871/0001 →
INTELLECTUAL PROPERTY AGREEMENT Recorded Oct 23, 2019
From: NUANCE COMMUNICATIONS, INC.
To: CERENCE INC.
Reel/Frame 050836/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 2, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022330/0088 →