IP Library Granted Patent US 9,431,011
Granted Patent B2
US 9,431,011 · App. 14/488,844 · Granted Aug 30, 2016

System and method for pronunciation modeling

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,431,011
App. No.
14/488,844
Granted
Aug 30, 2016
Kind
B2
Abstract

Systems, computer-implemented methods, and tangible computer-readable media for generating a pronunciation model. The method includes identifying a generic model of speech composed of phonemes, identifying a family of interchangeable phonemic alternatives for a phoneme in the generic model of speech, labeling the family of interchangeable phonemic alternatives as referring to the same phoneme, and generating a pronunciation model which substitutes each family for each respective phoneme. In one aspect, the generic model of speech is a vocal tract length normalized acoustic model. Interchangeable phonemic alternatives can represent a same phoneme for different dialectal classes. An interchangeable phonemic alternative can include a string of phonemes.

Claims (53)

1. A method comprising:

receiving user speech;

identifying a user dialect in the user speech;

selecting a set of phoneme alternatives representing the user dialect from a library of alternative pronunciation models, each alternative generated by the process of:

(i) identifying a generic model of speech composed of phonemes;

(ii) identifying a family of interchangeable phonemic alternatives for a phoneme in the generic model of speech;

(iii) labeling each interchangeable phonemic alternative in the family as referring to the phoneme; and

(iv) generating a pronunciation model which substitutes the family of interchangeable phonemic alternatives for the phoneme; and

recognizing, via a processor, the user speech using the selected set of phoneme alternatives in the pronunciation model.

2. The method of claim 1 , wherein identifying the user dialect further comprises:

recognizing the user speech with a plurality of dialect models;

eliminating dialect models in the plurality of dialect models which do not phonemically match the user speech until a single dialect model remains; and

identifying the remaining single dialect model as the user dialect model.

3. The method of claim 1 , wherein the generic model of speech is a vocal tract length normalized acoustic model.

4. The method of claim 3 , wherein the generic model does not include non-dialectal acoustic variation.

5. The method of claim 1 , wherein interchangeable phonemic alternatives represent a same phoneme for different dialectal classes.

6. The method of claim 1 , wherein an interchangeable phonemic alternative comprises a string of phonemes.

7. The method of claim 1 , further comprising repeating the receiving, the identifying, the selecting, and the recognizing for each speaker in a group of speakers, to yield a different model for each speaker.

8. A system comprising:

a processor; and

a computer-readable storage medium having instructions stored which, when executed by the processor, cause the processor to perform operations comprising:

receiving user speech;

identifying a user dialect in the user speech;

selecting a set of phoneme alternatives representing the user dialect from a library of alternative pronunciation models, each alternative generated by the process of:

(i) identifying a generic model of speech composed of phonemes;

(ii) identifying a family of interchangeable phonemic alternatives for a phoneme in the generic model of speech;

(iii) labeling each interchangeable phonemic alternative in the family as referring to the phoneme; and

(iv) generating a pronunciation model which substitutes the family of interchangeable phonemic alternatives for the phoneme; and

recognizing the user speech using the selected set of phoneme alternatives in the pronunciation model.

9. The system of claim 8 , wherein identifying the user dialect further comprises:

recognizing the user speech with a plurality of dialect models;

eliminating dialect models in the plurality of dialect models which do not phonemically match the user speech until a single dialect model remains; and

identifying the remaining single dialect model as the user dialect model.

10. The system of claim 8 , wherein the generic model of speech is a vocal tract length normalized acoustic model.

11. The system of claim 10 , wherein the generic model does not include non-dialectal acoustic variation.

12. The system of claim 8 , wherein interchangeable phonemic alternatives represent a same phoneme for different dialectal classes.

13. The system of claim 8 , wherein an interchangeable phonemic alternative comprises a string of phonemes.

14. The system of claim 8 , the computer-readable storage medium having additional instructions stored which, when executed by the processor, result in operations comprising repeating the receiving, the identifying, the selecting, and the recognizing for each speaker in a group of speakers, to yield a different pronunciation model for each speaker.

15. A computer-readable storage device having instructions stored which, when executed by a computing device, cause the computing device to perform operations comprising:

receiving user speech;

identifying a user dialect in the user speech;

selecting a set of phoneme alternatives representing the user dialect from a library of alternative pronunciation models, each alternative generated by the process of:

(i) identifying a generic model of speech composed of phonemes;

(ii) identifying a family of interchangeable phonemic alternatives for a phoneme in the generic model of speech;

(iii) labeling each interchangeable phonemic alternative in the family as referring to the phoneme; and

(iv) generating a pronunciation model which substitutes the family of interchangeable phonemic alternatives for the phoneme; and

recognizing the user speech using the selected set of phoneme alternatives in the pronunciation model.

16. The computer-readable storage device of claim 15 , wherein identifying the user dialect further comprises:

recognizing the user speech with a plurality of dialect models;

eliminating dialect models in the plurality of dialect models which do not phonemically match the user speech until a single dialect model remains; and

identifying the remaining single dialect model as the user dialect model.

17. The computer-readable storage device of claim 15 , wherein the generic model of speech is a vocal tract length normalized acoustic model.

18. The computer-readable storage device of claim 17 , wherein the generic model does not include non-dialectal acoustic variation.

Assignments (15)
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 036100/0925 Recorded Jul 1, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060559/0576 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 043039/0808 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060557/0636 →
RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY RECORDED AT REEL/FRAME: 049388/0082 Recorded Jun 30, 2022
From: SILICON VALLEY BANK
To: INTERACTIONS LLC
Reel/Frame 060558/0474 →
TERMINATION AND RELEASE OF SECURITY INTEREST IN INTELLECTUAL PROPERTY Recorded May 23, 2022
From: ORIX GROWTH CAPITAL, LLC
To: INTERACTIONS CORPORATION; INTERACTIONS LLC
Reel/Frame 061749/0825 →
RELEASE OF SECURITY INTEREST Recorded May 18, 2020
From: BEARCUB ACQUISITIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 052693/0866 →
INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 5, 2019
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 049388/0082 →
ASSIGNMENT OF IP SECURITY AGREEMENT Recorded Nov 17, 2017
From: ARES VENTURE FINANCE, L.P.
To: BEARCUB ACQUISITIONS LLC
Reel/Frame 044481/0034 →
AMENDED AND RESTATED INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jun 29, 2017
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 043039/0808 →
CORRECTIVE ASSIGNMENT TO CORRECT THE CHANGE PATENT 7146987 TO 7149687 PREVIOUSLY RECORDED ON REEL 036009 FRAME 0349. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded Nov 17, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 037134/0712 →
FIRST AMENDMENT TO INTELLECTUAL PROPERTY SECURITY AGREEMENT Recorded Jul 13, 2015
From: INTERACTIONS LLC
To: SILICON VALLEY BANK
Reel/Frame 036100/0925 →
SECURITY INTEREST Recorded Jun 23, 2015
From: INTERACTIONS LLC
To: ARES VENTURE FINANCE, L.P.
Reel/Frame 036009/0349 →
SECURITY INTEREST Recorded Dec 19, 2014
From: INTERACTIONS LLC
To: ORIX VENTURES, LLC
Reel/Frame 034677/0768 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2014
From: AT&T ALEX HOLDINGS, LLC
To: INTERACTIONS LLC
Reel/Frame 034642/0640 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 10, 2014
From: AT&T INTELLECTUAL PROPERTY I, L.P.
To: AT&T ALEX HOLDINGS, LLC
Reel/Frame 034462/0764 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 17, 2014
From: LJOLJE, ANDREJ; CONKIE, ALISTAIR D.; SYRDAL, ANN K.
To: AT&T INTELLECTUAL PROPERTY I, L.P.
Reel/Frame 033764/0041 →