IP Library Granted Patent US 8,484,023
Granted Patent B2
US 8,484,023 · App. 12/889,845 · Granted Jul 9, 2013

Sparse representation features for speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,484,023
App. No.
12/889,845
Granted
Jul 9, 2013
Kind
B2
Abstract

Techniques are disclosed for generating and using sparse representation features to improve speech recognition performance. In particular, principles of the invention provide sparse representation exemplar-based recognition techniques. For example, a method comprises the following steps. A test vector and a training data set associated with a speech recognition system are obtained. A subset of the training data set is selected. The test vector is mapped with the selected subset of the training data set as a linear combination that is weighted by a sparseness constraint such that a new test feature set is formed wherein the training data set is moved more closely to the test vector subject to the sparseness constraint. An acoustic model is trained on the new test feature set. The acoustic model trained on the new test feature set may be used to decode user speech input to the speech recognition system.

Claims (39)

1. A method, comprising:

obtaining a test vector and a training data set associated with a speech recognition system;

selecting a subset of the training data set;

mapping the test vector with the selected subset of the training data set as a linear combination that is weighted by a sparseness constraint such that a new test feature set is formed wherein the training data set is moved more closely to the test vector subject to the sparseness constraint; and

training, using a processor, an acoustic model on the new test feature set.

2. The method of claim 1 , further comprising using the acoustic model trained on the new test feature set to decode user speech input to the speech recognition system.

3. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set as the k nearest neighbors to the test vector in the training data set.

4. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on a trigram language model.

5. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on a unigram language model.

6. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on only acoustic information.

7. The method of claim 6 , wherein the acoustic information selecting step further comprises using acoustic information with unique phoneme identities.

8. The method of claim 6 , wherein the acoustic information comprises a given number of top scoring Gaussian Mixture Models.

9. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on Gaussian means.

10. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on random sampling.

11. The method of claim 1 , wherein the selecting step further comprises selecting the subset of the training data set based on cosine similarity sampling.

12. The method of claim 1 , wherein the mapping step further comprises solving an equation y=Hβ where y is the test vector, H is the selected subset of the training data set, and β is the sparseness constraint value.

13. The method of claim 12 , wherein β is computed using an approximate Bayesian compressive sensing method.

14. An apparatus, comprising:

a memory; and

a processor operatively coupled to the memory and configured to:

obtain a test vector and a training data set associated with a speech recognition system;

select a subset of the training data set;

map the test vector with the selected subset of the training data set as a linear combination that is weighted by a sparseness constraint such that a new test feature set is formed wherein the training data set is moved more closely to the test vector subject to the sparseness constraint; and

train an acoustic model on the new test feature set.

15. The apparatus of claim 14 , wherein the processor is further configured to use the acoustic model trained on the new test feature set to decode user speech input to the speech recognition system.

16. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set as the k nearest neighbors to the test vector in the training data set.

17. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on a trigram language model.

18. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on a unigram language model.

19. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on only acoustic information.

20. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on Gaussian means.

21. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on random sampling.

22. The apparatus of claim 14 , wherein the selecting step further comprises selecting the subset of the training data set based on cosine similarity sampling.

23. The apparatus of claim 14 , wherein the mapping step further comprises solving an equation y=Hβ where y is the test vector, H is the selected subset of the training data set, and β is the sparseness constraint value.

24. The apparatus of claim 23 , wherein β is computed using an approximate Bayesian compressive sensing method.

25. A non-transitory computer readable storage medium having tangibly embodied thereon computer readable program code which, when executed, causes a processor device to:

obtain a test vector and a training data set associated with a speech recognition system;

select a subset of the training data set;

map the test vector with the selected subset of the training data set as a linear combination that is weighted by a sparseness constraint such that a new test feature set is formed wherein the training data set is moved more closely to the test vector subject to the sparseness constraint; and

train an acoustic model on the new test feature set.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065532/0152 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 1, 2013
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 030323/0965 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 24, 2010
From: KANEVSKY, DIMITRI; NAHAMOO, DAVID; RAMABHADRAN, BHUVANA; SAINATH, TARA N.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 025038/0925 →