IP Library Granted Patent US 7,716,051
Granted Patent B2
US 7,716,051 · App. 11/427,029 · Granted May 11, 2010

Distributed voice recognition system and method

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,716,051
App. No.
11/427,029
Granted
May 11, 2010
Kind
B2
Abstract

A distributed voice recognition system ( 500 ) and method employs principles of bottom-up (i.e., raw input) and top-down (i.e., prediction based on past experience) processing to perform client-side and server-side processing by (i) at the client-side, replacing application data by a phonotactic table ( 504 ); (ii) at the server-side, tracking separate confidence scores for matches against an acoustic model and comparison to a grammar; and (iii) at the server-side using a contention resolver ( 514 ) to weight the client-side and server-side results to establish a single output which represents the collaboration between client-side processing and server-side processing.

Claims (42)

1. A voice recognition system, comprising:

an input device configured to receive voice information and produce a first result by comparing the voice information to first predetermined data;

a remote computer configured to process the voice information to produce a second result by comparing the voice information to second predetermined data, wherein the second result comprises a plurality of possible matches; and

a contention resolver configured to receive the first result and the second result and to select one of the plurality of possible matches of the second result as an output result, wherein the output result is selected based, at least in part, on the first result;

wherein the input device is configured to compare the voice information to the first predetermined data so as to produce the first result as at least one sound combination corresponding to less than the entirety of a given word for which the remote computer produces the plurality of possible matches of the second result.

2. The system of claim 1 , wherein the input device comprises a first acoustic front end and a first decoder, and the remote computer comprises a second acoustic front end and a second decoder.

3. The system of claim 2 , wherein the second decoder is configured to produce the plurality possible matches based on confidence scores.

4. The system of claim 2 , wherein the second decoder produces a grammar confidence score for each of the plurality of possible matches; and

wherein the contention resolver is configured to validate, by using the first result, ones of the plurality of possible matches associated with a grammar confidence score less than a rejection threshold.

5. The system of claim 1 , wherein the input device comprises a phonotactic table holding the first predetermined data.

6. The system of claim 1 , wherein the remote computer is configured to calculate the confidence scores for the plurality of possible matches against an acoustic model and comparison to a grammar.

7. The system of claim 1 , wherein the contention resolver is configured to weight the first result and the plurality of possible matches to select the output result.

8. The system of claim 1 , wherein the first result comprises a plurality of sound combinations, each sound combination being associated with a confidence value;

wherein the contention resolver selects the output result based, at least in part, on the confidence values for sound combinations in each of the plurality of possible matches.

9. A voice recognition method comprising:

producing, by an input device, a first result by comparing received voice information to first predetermined data;

producing, by a remote computer, a second result by comparing the voice information to second predetermined data, wherein the second result comprises a plurality of possible matches; and

selecting one of the plurality of possible matches of the second result as an output result, wherein selection of the output result is based, at least in part, on the first result;

wherein producing the first result comprises comparing the received voice information to the first predetermined data so as to produce the first result as at least one sound combination corresponding to less than the entirety of a given word for which the remote computer produces the plurality of possible matches of the second result.

10. The method of claim 9 , wherein producing the first result comprises performing, by the input device, first acoustic front end processing and performing first decoding, and producing the second result comprises performing, by the remote server second acoustic front end processing and performing second decoding.

11. The method of claim 10 , wherein performing second decoding comprises producing the plurality of possible matches based on confidence scores.

12. The method of claim 10 , wherein performing second decoding comprises producing a grammar confidence score for each of the plurality of possible matches; and

wherein selecting one of the plurality of possible matches comprises validating, by using the first result, ones of the plurality of possible matches associated with a grammar confidence score less than a rejection threshold.

13. The method of claim 9 , wherein producing the first result comprises using a phonotactic table holding the first predetermined data.

14. The method of claim 9 , wherein producing the second result comprises calculating confidence scores for the plurality of possible matches against an acoustic model and comparison to a grammar.

15. The method of claim 9 , wherein selecting one of the plurality of possible matches comprises weighting the first result and the plurality of possible matches to select the output result.

16. The method of claim 9 , wherein the first result comprises a plurality of sound combinations, each sound combination being associated with a confidence value;

wherein selecting one of the plurality of possible matches further comprises selecting the output result based, at least in part, on the confidence values for sound combinations in each of the plurality of possible matches.

17. A computer-readable storage medium encoded with a plurality of instructions that, when executed by a computer, perform a method of:

producing a first result by comparing received voice information to first predetermined data;

producing a second result by comparing the voice information to second predetermined data, wherein the second result comprises a plurality of possible matches; and

selecting one of the plurality of possible matches of the second result as an output result, wherein selection of the output result is based, at least in part, on the first result;

wherein producing the first result comprises comparing the received voice information to the first predetermined data so as to produce the first result as at least one sound combination corresponding to less than the entirety of a given word for which the remote computer produces the plurality of possible matches of the second result.

18. The computer-readable storage medium of claim 17 , wherein the producing the first result comprises performing first acoustic front end processing and performing first decoding, and producing the second result comprises performing second acoustic front end processing and performing second decoding.

19. The computer-readable storage medium of claim 18 , wherein performing second decoding comprises producing the plurality of possible matches based on confidence scores.

20. The computer-readable storage medium of claim 18 , wherein performing second decoding comprises producing a grammar confidence score for each of the plurality of possible matches; and

wherein selecting one of the plurality of possible matches comprises validating, by using the first result, ones of the plurality of possible matches associated with a grammar confidence score less than a rejection threshold.

21. The computer-readable storage medium of claim 17 , wherein producing the first result comprises using a phonotactic table holding the first predetermined data.

22. The computer-readable storage medium of claim 17 , wherein producing the second result comprises calculating confidence scores for the plurality of possible matches against an acoustic model and comparison to a grammar.

23. The computer-readable storage medium of claim 17 , wherein selecting one of the plurality of possible matches comprises weighting the first result and the plurality of possible matches to select the output result.

24. The computer-readable storage medium of claim 17 , wherein the first result comprises a plurality of sound combinations, each sound combination being associated with a confidence value;

wherein selecting one of the plurality of possible matches further comprises selecting the output result based, at least in part, on the confidence values for sound combinations in each of the plurality of possible matches.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 13, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065552/0934 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 13, 2009
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 022689/0317 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 13, 2006
From: DOW, BARRY NEIL; LAWRENCE, STEPHEN GRAHAM; PICKERING, JOHN BRIAN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 017929/0948 →