IP Library Granted Patent US 10,186,256
Granted Patent B2
US 10,186,256 · App. 15/109,321 · Granted Jan 22, 2019

Method and apparatus for exploiting language skill information in automatic speech recognition

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,186,256
App. No.
15/109,321
Granted
Jan 22, 2019
Kind
B2
Abstract

Typical speech recognition systems usually use speaker-specific speech data to apply speaker adaptation to models and parameters associated with the speech recognition system. Given that speaker-specific speech data may not be available to the speech recognition system, information indicative of language skills is employed in adapting configurations of a speech recognition system. According to at least one example embodiment, a method and corresponding apparatus, for speech recognition comprise maintaining information indicative of language skills of users of the speech recognition system. A configuration of the speech recognition system for a user is determined based at least in part on corresponding information indicative of language skills of the user. Upon receiving speech data from the user, the configuration of the speech recognition system determined is employed in performing speech recognition.

Claims (48)

1. A method of speech recognition comprising:

employing, at a speech recognition system, a first model corresponding to a first language and first degree of proficiency of a user and a second model corresponding to a second language and second degree of proficiency of the user based on language skills included in a user profile of the user, the language skills indicating languages spoken by the user and, for each language spoken by the user, a degree of proficiency, the language skills further determined based on personal information of the user, wherein configuring the speech recognition system loads the first language, first degree of proficiency, second language, and second degree of proficiency from the user profile of the user; and

performing speech recognition on speech data representing speech by the user's speaking in the first language using the first model, and responsive to the user's switching from speaking in the first language to the second language, automatically performing speech recognition using the second model.

2. The method as recited in claim 1 , further comprising acquiring information indicative of the language skills by inferring the information from personal data of the users.

3. The method as recited in claim 2 , wherein inferring the information from personal data of the users includes at least one of:

periodically inferring the information from personal data of the users; and

inferring the information in response to updates made to the personal data of the users.

4. The method as recited in claim 1 , wherein employing the first model and second model of the speech recognition system further includes at least one of:

determining one or more acoustic models for the user;

restricting the languages supported by the speech recognition system to languages indicated in the information indicative of language skills of the user;

allowing the speech recognition system to switch between languages indicated in the information indicative of language skills of the user; and

setting rules for pronunciation generation associated with the speech recognition system.

5. The method as recited in claim 4 , wherein switching between languages is triggered seamlessly based on automatic language identification among languages indicated in the information indicative of language skills of the user.

6. The method as recited in claim 4 , wherein the switch between languages is triggered manually by the user.

7. The method as recited in claim 1 , wherein the information indicative of language skills of the user includes languages spoken by the user.

8. The method as recited in claim 1 , wherein the information indicative of language skills of the user includes an indication of a degree of proficiency in a language spoken by the user.

9. An apparatus for speech recognition comprising:

a processor; and

a memory, with computer code instructions stored thereon,

the processor and the memory, with the computer code instructions, being configured to cause the apparatus to:

employing, at a speech recognition system, a first model corresponding to a first language and first degree of proficiency of a user and a second model corresponding to a second language and second degree of proficiency of the user based on language skills included in a user profile of the user, the language skills indicating languages spoken by the user and, for each language spoken by the user, a degree of proficiency, the language skills further determined based on personal information of the user, wherein configuring the speech recognition system loads the first language, first degree of proficiency, second language, and second degree of proficiency from the user profile of the user; and

performing speech recognition on speech data representing speech by the user's speaking in a first language using a first model of the plurality of models, the first model being configured to recognize speech of the first language, and responsive to the user's switching from speaking in the first language to speaking in a second language, automatically performing speech recognition using a second model of the plurality of models, the second model being configured to recognize speech of the second language.

10. The apparatus as recited in claim 9 , wherein the instructions are further configured to cause the apparatus to acquire the information indicative of language skills by inferring the information from personal data of the users.

11. The apparatus as recited in claim 10 , wherein inferring the information from personal data of the users, the processor and the memory, with the computer code instructions, are further configured to cause the apparatus to:

periodically infer the information from the personal data; or

infer the information in response to updates made to the personal data of the users.

12. The apparatus as recited in claim 9 , wherein employing the first and second model of the speech recognition system, the processor and the memory, with the computer code instructions, are further configured to cause the apparatus to:

determine one or more acoustic models for the user;

restrict languages supported by the speech recognition system to languages indicated in the information indicative of language skills of the user;

switch between languages indicated in the information indicative of language skills of the user; or

set rules for pronunciation generation associated with the speech recognition system.

13. The apparatus as recited in claim 12 , wherein switching between languages, the processor and the memory, with the computer code instructions, are further configured to cause the apparatus to switch between languages upon automatic language identification among languages indicated in the information indicative of language skills of the user.

14. The apparatus as recited in claim 12 , wherein switching between languages, the processor and the memory, with the computer code instructions, are further configured to cause the apparatus to switch between languages upon manual trigger by the user.

15. The apparatus as recited in claim 9 , wherein the information indicative of language skills of the user includes languages spoken by the user or an indication of a degree of proficiency in a language spoken by the user.

16. A non-transitory computer readable medium having computer software instructions for speech recognition stored thereon, the computer software instructions when executed by a processor cause an apparatus to:

employing, at a speech recognition system, a first model corresponding to a first language and first degree of proficiency of a user and a second model corresponding to a second language and second degree of proficiency of the user based on language skills included in a user profile of the user, the language skills indicating languages spoken by the user and, for each language spoken by the user, a degree of proficiency, the language skills further determined based on personal information of the user, wherein configuring the speech recognition system loads the first language, first degree of proficiency, second language, and second degree of proficiency from the user profile of the user; and

performing speech recognition on speech data representing speech by the user's speaking in a first language using a first model of the plurality of models, the first model being configured to recognize speech of the first language, and responsive to the user's switching from speaking in the first language to speaking in a second language, automatically performing speech recognition using a second model of the plurality of models, the second model being configured to recognize speech of the second language;

wherein the automatic switching of languages is informed by a list of languages of the user profile of the user.

17. The non-transitory computer readable medium of claim 16 , further comprising acquiring information indicative of language skills of the users by inferring the information from personal data of the users.

18. The non-transitory computer readable medium of claim 17 , wherein inferring the information from personal data of the users includes at least one of:

periodically inferring the information from personal data of the users; and

inferring the information in response to updates made to the personal data of the users.

19. The non-transitory computer readable medium of claim 16 , wherein employing the first and second model of the speech recognition system further includes at least one of:

determining one or more acoustic models for the user;

restricting the languages supported by the speech recognition system to languages indicated in the information indicative of language skills of the user;

allowing the speech recognition system to switch between languages indicated in the information indicative of language skills of the user; and

setting rules for pronunciation generation associated with the speech recognition system.

20. The non-transitory computer readable medium of claim 19 , wherein switching between languages is triggered seamlessly based on automatic language identification among languages indicated in the information indicative of language skills of the user.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 9, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065530/0871 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065258/0771 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 18, 2016
From: LI, WEIYING; WILLETT, DANIEL
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 039177/0220 →