IP Library Granted Patent US 8,600,749
Granted Patent B2
US 8,600,749 · App. 12/633,334 · Granted Dec 3, 2013

System and method for training adaptation-specific acoustic models for automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,600,749
App. No.
12/633,334
Granted
Dec 3, 2013
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer-readable storage media for training adaptation-specific acoustic models. A system practicing the method receives speech and generates a full size model and a reduced size model, the reduced size model starting with a single distribution for each speech sound in the received speech. The system finds speech segment boundaries in the speech using the full size model and adapts features of the speech data using the reduced size model based on the speech segment boundaries and an overall centroid for each speech sound. The system then recognizes speech using the adapted features of the speech. The model can be a Hidden Markov Model (HMM). The reduced size model can also be of a reduced complexity, such as having fewer mixture components than a model of full complexity. Adapting features of speech can include moving the features closer to an overall feature distribution center.

Claims (40)

1. A method comprising:

receiving speech data, to yield received speech data;

generating, via a computing device, a full size adaptation-specific acoustic model and a reduced size adaptation-specific acoustic model, the reduced size adaptation-specific acoustic model starting with a single distribution for each speech sound in the received speech data;

finding speech segment boundaries in the received speech data using the full size adaptation-specific acoustic model;

increasing a size of the reduced size adaptation-specific acoustic model using the full size adaptation-specific acoustic model to meet a desired performance level, to yield a modified reduced size adaptation-specific acoustic model;

after the finding of the speech segment boundaries, adapting features of the received speech data using the modified reduced size adaptation-specific acoustic model based on the speech segment boundaries and an overall centroid for each speech sound, to yield adapted features; and

recognizing the received speech data based on the adapted features.

2. The method of claim 1 , wherein the full size adaptation-specific acoustic model and the reduced size adaptation-specific model are hidden markov models.

3. The method of claim 1 , wherein the reduced size adaptation-specific acoustic model is a model of reduced complexity.

4. The method of claim 3 , wherein the reduced size adaptation-specific acoustic model has less mixture components than a model of full complexity.

5. The method of claim 3 , wherein the reduced size adaptation-specific acoustic model has less mixture components per state than a model of full complexity.

6. The method of claim 1 , wherein adapting features of the received speech data further comprises moving the features closer to a center of an overall feature distribution.

7. The method of claim 1 , further comprising recursively training a recognition model based on the adapted features, to yield a recursively trained model.

8. The method of claim 7 , further comprising adapting speech of a new speaker at recognition time using the recursively trained model.

9. A system comprising:

a processor; and

a computer-readable storage device having instructions stored which, when executed by the processor, result in the processor performing operations comprising:

receiving speech data, to yield received speech data;

generating a full size adaptation-specific acoustic model and a reduced size adaptation-specific acoustic model, the reduced size adaptation-specific acoustic model starting with a single distribution for each speech sound in the received speech data;

finding speech segment boundaries in the received speech data using the full size adaptation-specific acoustic model;

increasing a size of the reduced size adaptation-specific acoustic model using the full size adaptation-specific acoustic model to meet a desired performance level, to yield a modified reduced size adaptation-specific acoustic model;

after the finding of the speech segment boundaries, adapting features of the received speech data using the modified reduced size adaptation-specific acoustic model based on the speech segment boundaries and an overall centroid for each speech sound, to yield adapted features; and

recognizing the received speech data based on the adapted features.

10. The system of claim 9 , wherein the full size adaptation-specific acoustic model and the reduced size adaptation-specific model are hidden markov models.

11. The system of claim 9 , wherein the reduced size adaptation-specific acoustic model is a model of reduced complexity.

12. The system of claim 11 , wherein the reduced size adaptation-specific acoustic model has less mixture components than a model of full complexity.

13. The system of claim 11 , wherein the reduced size adaptation-specific acoustic model has less mixture components per state than a model of full complexity.

14. The system of claim 9 , wherein the adapting of the features of the received speech data further comprises moving the features closer to a center of an overall feature distribution.

15. The system of claim 9 , the computer-readable storage device having additional instructions stored which result in the operations further comprising recursively training a recognition model based on the adapted features, to yield a recursively trained model.

16. The system of claim 15 , the computer-readable storage device having additional instructions stored which result in the operations further comprising adapting speech of a new speaker at recognition time using the recursively trained model.

17. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving speech data, to yield received speech data;

generating a full size adaptation-specific acoustic model and a reduced size adaptation-specific acoustic model, the reduced size adaptation-specific acoustic model starting with a single distribution for each speech sound in the received speech data;

finding speech segment boundaries in the received speech data using the full size adaptation-specific acoustic model;

increasing a size of the reduced size adaptation-specific acoustic model using the full size adaptation-specific acoustic model to meet a desired performance level, to yield a modified reduced size adaptation-specific acoustic model;

after the finding of the speech segment boundaries, adapting features of the received speech data using the modified reduced size adaptation-specific acoustic model based on the speech segment boundaries and an overall centroid for each speech sound, to yield adapted features; and

recognizing the received speech data based on the adapted features.

18. The computer-readable storage device of claim 17 , wherein the full size adaptation-specific acoustic model and the reduced size adaptation-specific model are hidden markov models.

19. The computer-readable storage device of claim 18 , wherein the reduced size adaptation-specific acoustic model is a model of reduced complexity.

20. The computer-readable storage device of claim 19 , wherein the reduced size adaptation-specific acoustic model has less mixture components than a model of full complexity.

Assignments (15)
RELEASE OF SECURITY INTEREST Recorded Sep 4, 2025
From: RUNWAY GROWTH FINANCE CORP., AS AGENT
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 072802/0931 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE APPLICATION NUMBER PREVIOUSLY RECORDED AT REEL: 060445 FRAME: 0733. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Feb 1, 2023
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 062919/0063 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 27, 2022
From: INTERACTIONS LLC; INTERACTIONS CORPORATION
To: RUNWAY GROWTH FINANCE CORP.
Reel/Frame 060445/0733 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2017
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 044071/0203 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →