IP Library Granted Patent US 8,160,876
Granted Patent B2
US 8,160,876 · App. 10/953,712 · Granted Apr 17, 2012

Interactive speech recognition model

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,160,876
App. No.
10/953,712
Granted
Apr 17, 2012
Kind
B2
Abstract

A method and apparatus for updating a speech model on a multi-user speech recognition system with a personal speech model for a single user. A speech recognition system, for instance in a car, can include a generic speech model for comparison with the user speech input. A way of identifying a personal speech model, for instance in a mobile phone, is connected to the system. A mechanism is included for receiving personal speech model components, for instance a BLUETOOTH connection. The generic speech model is updated using the received personal speech model components. Speech recognition can then be performed on user speech using the updated generic speech model.

Claims (63)

1. A method for updating a first speech model in a speech recognition system, comprising:

identifying that a user device of a user is in communication with the speech recognition system via a network connection;

receiving from the user device, the user device comprising a personal speech model trained for the user through previous speech recognition operations, one or more personal speech model components of the personal speech model trained for the user through previous speech recognition operations, the one or more personal speech model components describing personal speech characteristics of the user, wherein the one or more personal speech model components are received from the user device over the network connection in response to identifying that the user device is in communication with the speech recognition system;

updating the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

performing speech recognition on user speech using the first speech model updated with the at least some of the one or more personal speech model components.

2. The method of claim 1 , wherein the user device is a mobile device.

3. The method of claim 2 , wherein the network connection includes a BLUETOOTH® connection between the mobile telephone and the speech recognition system.

4. The method of claim 3 , wherein the speech recognition system is part of a navigation system located in a vehicle.

5. The method of claim 1 , wherein the one or more personal speech model components are received over the network connection in response to the speech recognition system performing a check in response to identifying that the user device is in communication with the speech recognition system, wherein the check evaluates at least one criteria to determine whether an update to the first speech model is to be received.

6. The method of claim 5 , wherein the criteria comprises a time period since a previous update to the first speech model.

7. The method of claim 1 , wherein the personal speech model components are personal language model components.

8. The method of claim 1 , wherein the personal speech model components are personal acoustic model components.

9. At least one non-transitory computer readable medium encoded with instructions that, when executed on at least one computer, performs a method for updating a first speech model in a speech recognition system, comprising:

identifying that a user device of a user is in communication with a speech recognition system via a network connection;

receiving from the user device, the user device comprising a personal speech model trained for the user through previous speech recognition operations, one or more personal speech model components of the personal speech model trained for the user through previous speech recognition operations, the one or more personal speech model components describing personal voice characteristics of the user, wherein the one or more personal speech model components are received from the user device over the network connection in response to identifying that the user device is in communication with the speech recognition system;

updating the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

performing speech recognition on user speech using the first speech model updated with the at least some of the one or more personal speech model components.

10. The at least one non-transitory computer readable medium of claim 9 , wherein the user device is a mobile device.

11. The at least one non-transitory computer readable medium of claim 10 , wherein the network connection includes a BLUETOOTH® connection between the mobile telephone and the speech recognition system.

12. The at least one non-transitory computer readable medium of claim 11 , wherein the speech recognition system is part of a navigation system located in a vehicle.

13. The at least one non-transitory computer readable medium of claim 9 , wherein the one or more personal speech model components are received over the network connection in response to the speech recognition system performing a check in response to identifying that the user device is in communication with the speech recognition system, wherein the check evaluates at least one criteria to determine whether an update to the first speech model is to be received.

14. The at least one non-transitory computer readable medium of claim 13 , wherein the criteria comprises a time period since a previous update to the first speech model.

15. The at least one non-transitory computer readable medium of claim 9 , wherein the personal speech model components are personal language model components.

16. The at least one non-transitory computer readable medium of claim 9 , wherein the personal speech model components are personal acoustic model components.

17. A speech recognition system accessible over a network, comprising:

a first speech model;

at least one processor programmed to implement a speech recognition engine capable of recognizing speech data based, at least in part, on the first speech model; and

a speech model controller to:

identify that a user device of a user is in communication with the speech recognition system via a network connection;

receive from the user device, the user device comprising a personal speech model trained for the user through previous speech recognition operations, one or more personal speech model components of the personal speech model trained for the user through previous speech recognition operations, the one or more personal speech model components describing personal voice characteristics of the user, wherein the one or more personal speech model components are received from the user device over the network connection in response to identifying that the user device is in communication with the speech recognition system;

update the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

transmit user speech to be recognized by the speech recognition engine using the first speech model updated with the at least some of the one or more personal speech model components.

18. The speech recognition system of claim 17 , wherein the user device is a mobile device.

19. The speech recognition system of claim 18 , wherein the network connection includes a BLUETOOTH® connection between the mobile telephone and the speech recognition system.

20. The speech recognition system of claim 19 , wherein the speech recognition system is part of a navigation system located in a vehicle.

21. The speech recognition system of claim 17 , wherein the one or more personal speech model components are received over the network connection in response to the speech recognition system performing a check in response to identifying that the user device is in communication with the speech recognition system, wherein the check evaluates at least one criteria to determine whether an update to the first speech model is to be received.

22. The speech recognition system of claim 21 , wherein the criteria comprises a time period since a previous update to the first speech model.

23. The speech recognition system of claim 17 , wherein the personal speech model components are personal language model components.

24. The speech recognition system of claim 17 , wherein the personal speech model components are personal acoustic model components.

25. A method for updating a first speech model in a speech recognition system, comprising:

identifying a voice application to be used by a user;

identifying a list of words that the voice application is programmed to recognize in a voice input;

requesting, from a personal device of the user, personal speech model components of a personal speech model trained by the user through previous speech recognition operations, the personal speech model components describing personal speech characteristics of the user, wherein only personal speech model components associated with words in the list of words are requested;

receiving one or more personal speech model components from the personal device, wherein each of the one or more personal speech model components received by the speech recognition system is associated with at least one word in the list of words, and wherein the personal speech model comprises at least one other personal speech component that is not received by the speech recognition system;

updating the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

performing speech recognition on user speech using the first speech model updated with the at least some of the one or more personal speech model components.

26. At least one non-transitory computer readable medium encoded with instructions that, when executed on at least one computer, performs a method for updating a first speech model in a speech recognition system, comprising:

identifying a voice application to be used by a user;

identifying a list of words that the voice application is programmed to recognize in a voice input;

requesting, from a personal device of the user, personal speech model components of a personal speech model trained by the user through previous speech recognition operations, the personal speech model components describing personal speech characteristics of the user, wherein only personal speech model components associated with words in the list of words are requested;

receiving one or more personal speech model components from the personal device, wherein each of the one or more personal speech model components received by the speech recognition system is associated with at least one word in the list of words, and wherein the personal speech model comprises at least one other personal speech component that is not received by the speech recognition system;

updating the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

performing speech recognition on user speech using the first speech model updated with the at least some of the one or more personal speech model components.

27. A speech recognition system accessible over a network, comprising:

a first speech model;

at least one processor programmed to implement a speech recognition engine capable of recognizing speech data based, at least in part, on the first speech model; and

a speech model controller to:

identify a voice application to be used by the user;

identify a list of words that the voice application is programmed to recognize in a voice input;

request, from a personal device of a user, personal speech model components of a personal speech model trained by a user through previous speech recognition operations, the one or more personal speech model components describing personal voice characteristics of the user, wherein only personal speech model components associated with words in the list of words are requested;

receive, from the personal device, one or more personal speech model components, wherein each of the one or more personal speech model components received by the speech recognition system is associated with at least one word in the list of words, and wherein the personal speech model comprises at least one other personal speech component that is not received by the speech recognition system;

update the first speech model using at least some of the one or more personal speech model components, by modifying at least one speech model component of the first speech model and/or adding at least one speech model component to the first speech model; and

transmit user speech to be recognized by the speech recognition engine using the first speech model updated with the at least some of the one or more personal speech model components.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →