IP Library Granted Patent US 10,580,406
Granted Patent B2
US 10,580,406 · App. 15/807,025 · Granted Mar 3, 2020

Unified N-best ASR results

Inventor: Darrin Kenneth John Fry (Kanata, CA)
Assignee: 2236008 Ontario Inc.
G10L15/22G10L15/1815G10L15/32G10L15/1822G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,580,406
App. No.
15/807,025
Granted
Mar 3, 2020
Kind
B2
Abstract

A system and method receives a spoken utterance and converts the spoken utterance into recognized speech results through automatic speech recognition modules. The system and method renders a composite recognition speech result comprising the recognized speech results joined in a return function. The system and method interprets the recognized speech results joined in a return function from each of the automatic speech recognition modules through multiple conversation modules.

Claims (45)

1. A system of automatic speech recognition, comprising:

an audio capture unit for capturing a spoken utterance;

a processor coupled to the audio capture unit, the processor being configured to:

convert, via one or more recognition modules, the spoken utterance into a plurality of recognized speech results;

obtain, by one or more conversation modules, ratings for the plurality of recognized speech results based on each recognized speech result's fitness for use by the conversation modules;

identify a first recognized speech result having the highest rating by the one or more conversation modules;

identify a preferred conversation module for consuming a recognition speech result for the spoken utterance based on the ratings for the plurality of recognized speech results;

in response to determining that an n-best processing of the first recognized speech results is selected:

generate a composite recognition speech result that includes the first recognized speech result supplemented with alternate values for at least one recognized data field for the first recognized speech result, the alternate values being extracted from other ones of the recognized speech results; and

provide, to the preferred conversation module, the composite recognition speech result.

2. The system of claim 1 where the composite recognition speech result incorporates n-best results into a single structure signifying n-best alternate recognition values for the at least one recognized data field.

3. The system of claim 2 where the n-best results comprise a plurality of hypotheses as to the meaning of the spoken utterance.

4. The system of claim 1 where the processor is further configured to render a composite intent result.

5. The system of claim 4 where the composite intent result incorporates n-best results into a single intent result indicating one or more alternate intent values.

6. The system of claim 4 , wherein the composite intent result is added to the composite recognition speech result as output.

7. The system of claim 1 where the system comprises a vehicle.

8. The system of claim 1 , wherein outputting the composite recognition speech result comprises performing, by the processor, an action based on the composite recognition speech result.

9. A method of processing speech recognition results, the method comprising:

receiving, via an audio capture unit, a spoken utterance;

converting, via one or more recognition modules, the spoken utterance into a plurality of recognized speech results;

obtaining, by one or more conversation modules, ratings for the plurality of recognized speech results based on each recognized speech result's fitness for use by the conversation modules;

identifying a first recognized speech result having the highest rating by the one or more conversation modules;

identifying a preferred conversation module for consuming a recognition speech result for the spoken utterance based on the ratings for the plurality of recognized speech results;

in response to determining that an n-best processing of the first recognized speech results is selected:

generating a composite recognition speech result that includes the first recognized speech result supplemented with alternate values for at least one recognized data field for the first recognized speech result, the alternate values being extracted from other ones of the recognized speech results; and

providing, to the preferred conversation module, the composite recognition speech result.

10. The method of claim 9 where the composite recognition speech result incorporates n-best results into a single structure signifying n-best alternate recognition values for the at least one recognized data field.

11. The method of claim 10 where the n-best results comprise a plurality of hypotheses as to the meaning of the spoken utterance.

12. The method of claim 9 where the processor is further configured to render a composite intent result.

13. The method of claim 12 where the composite intent result incorporates n-best results into a single intent result indicating one or more alternate intent values.

14. The method of claim 12 , wherein the composite intent result is added to the composite recognition speech result as output.

15. The method of claim 9 where the automatic speech recognition system comprises a vehicle.

16. The method of claim 9 , wherein outputting the composite recognition speech result comprises performing, by the processor, an action based on the composite recognition speech result.

17. A non-transitory machine-readable medium encoded with machine-executable instructions which, when executed by a processor, will cause the processor to:

receive, via an audio capture unit, a spoken utterance;

convert, via one or more recognition modules, the spoken utterance into a plurality of recognized speech results;

obtain, by one or more conversation modules, ratings for the plurality of recognized speech results based on each recognized speech result's fitness for use by the conversation modules;

identify a first recognized speech result having the highest rating by the one or more conversation modules;

identify a preferred conversation module for consuming a recognition speech result for the spoken utterance based on the ratings for the plurality of recognized speech results;

in response to determining that an n-best processing of the first recognized speech results is selected:

generate a composite recognition speech result that includes the first recognized speech result supplemented with alternate values for at least one recognized data field for the first recognized speech result, the alternate values being extracted from other ones of the recognized speech results; and

provide, to the preferred conversation module, the composite recognition speech result.

18. The non-transitory machine-readable medium of claim 17 where the composite recognition speech result incorporates n-best results into a single structure signifying n-best alternate recognition values for the at least one recognized data field.

19. The non-transitory machine-readable medium claim 18 where the n-best results comprise a plurality of hypotheses as to the meaning of the spoken utterance.

20. The non-transitory machine-readable medium claim 17 where the non-transitory machine-readable medium is encoded in a vehicle.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 22, 2020
From: 2236008 ONTARIO INC.
To: BLACKBERRY LIMITED
Reel/Frame 053313/0315 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 22, 2018
From: QNX SOFTWARE SYSTEMS LIMITED
To: 2236008 ONTARIO INC.
Reel/Frame 045008/0689 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 15, 2017
From: FRY, DARRIN KENNETH JOHN
To: QNX SOFTWARE SYSTEMS LIMITED
Reel/Frame 044137/0520 →
Continuity (2)
Provisional Application 62547422 · Aug 18, 2017
Related Publication 20190057691A1 · Feb 21, 2019