IP Library Granted Patent US 11,929,073
Granted Patent B2
US 11,929,073 · App. 17/958,663 · Granted Mar 12, 2024

Hybrid arbitration system

Inventor: Min Tang (Yarrow Point, WA)
Assignee: Cerence Operating Company
G10L15/22G06N3/08G10L15/02G10L15/16G10L15/1815G10L15/30G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,929,073
App. No.
17/958,663
Granted
Mar 12, 2024
Kind
B2
Abstract

A method for selecting a speech recognition result on a computing device includes receiving a first speech recognition result determined by the computing device, receiving first features, at least some of the features being determined using the first speech recognition result, determining whether to select the first speech recognition result or to wait for a second speech recognition result determined by a cloud computing service based at least in part on the first speech recognition result and the first features.

Claims (45)

1. A method for selecting a speech recognition result on a computing device, the method comprising:

acquiring first speech data for a first utterance;

soliciting a first speech recognition result from the computing device including sending the first speech data to the computing device;

soliciting a second speech recognition result from a cloud computing service including sending the first speech data to the cloud computing service;

receiving the first speech recognition result determined by the computing device;

receiving a first plurality of features, at least some of the features being determined using the first speech recognition result;

determining, prior to receiving the second speech recognition result from the cloud computing service, whether to select the first speech recognition result or to wait for the second speech recognition result based at least in part on the first speech recognition result and the first plurality of features; and

selecting the first speech recognition result based on the determining.

2. The method of claim 1 wherein the determining includes computing, based at least in part on the first speech recognition result and the first plurality of features, a confidence value for the first speech recognition result and comparing the confidence value to a first predetermined threshold.

3. The method of claim 2 wherein selecting the first speech recognition result based on the determining includes selecting the first speech recognition result if the confidence value exceeds the first predetermined threshold.

4. The method of claim 2 wherein computing the first confidence value includes processing the first speech recognition result and the first plurality of features in a classifier.

5. The method of claim 4 wherein the classifier is implemented as a neural network.

6. The method of claim 1 wherein at least some of the first plurality of features characterize a quality of the first speech recognition result.

7. The method of claim 1 further comprising providing the selected first speech recognition result for subsequent processing by the computing device.

8. The method of claim 1 wherein the first plurality of features includes one or more of a natural language understanding dialogue context, a natural language understanding recognized domain name, a natural language understanding recognized context name, a natural language understanding recognized action name, one or more natural language understanding recognized entity names, and an NBest word sequence.

9. The method of claim 1 further comprising performing an action based at least in part on the selected speech recognition result.

10. The method of claim 1 further comprising:

acquiring second speech data for a second utterance;

soliciting a third speech recognition result from the computing device including sending the second speech data to the computing device;

soliciting a fourth speech recognition result from the cloud computing service including sending the second speech data to the cloud computing service;

receiving the third speech recognition result determined by the computing device;

receiving a second plurality of features, at least some of the features being determined using the third speech recognition result;

determining, prior to receiving the fourth speech recognition result from the cloud computing service, whether to select the third speech recognition result or to wait for the fourth speech recognition result based at least in part on the third speech recognition result and the second plurality of features; and

waiting for the fourth speech recognition result based on the determining.

11. The method of claim 10 further comprising:

receiving the fourth speech recognition result from the cloud computing device;

receiving a third plurality of features, at least some of the features being determined using the fourth speech recognition result;

computing, based at least in part on the third speech recognition result and the second plurality of features, a second confidence value for the third speech recognition result;

computing, based at least in part on the fourth speech recognition result and the third plurality of features, a third confidence value for the fourth speech recognition result; and

selecting the fourth speech recognition result based at least in part on a comparison of the second confidence value to the third confidence value.

12. The method of claim 11 wherein selecting the fourth speech recognition result includes determining that the third confidence value is greater than the second confidence value.

13. The method of claim 11 further comprising providing the selected fourth speech recognition result for subsequent processing by the computing device.

14. A system for selecting a speech recognition result on a computing device, the system comprising:

a first input for acquiring first speech data for a first utterance;

an output for soliciting a first speech recognition result from the computing device including sending the first speech data to the computing device and soliciting a second speech recognition result from a cloud computing service including sending the first speech data to the cloud computing service;

a second input for receiving the first speech recognition result determined by the computing device and receiving a first plurality of features, at least some of the features being determined using the first speech recognition result;

one or more processors for determining, prior to receiving the second speech recognition result from the cloud computing service, whether to select the first speech recognition result or to wait for the second speech recognition result based at least in part on the first speech recognition result and the first plurality of features and selecting the first speech recognition result based on the determining.

15. Software stored on a non-transitory, computer-readable medium, the software including instructions for causing one or more processors to perform steps to select a speech recognition result on a computing device, the steps including:

acquiring first speech data for a first utterance;

soliciting a first speech recognition result from the computing device including sending the first speech data to the computing device;

soliciting a second speech recognition result from a cloud computing service including sending the first speech data to the cloud computing service;

receiving the first speech recognition result determined by the computing device;

receiving a first plurality of features, at least some of the features being determined using the first speech recognition result;

determining, prior to receiving the second speech recognition result from the cloud computing service, whether to select the first speech recognition result or to wait for the second speech recognition result based at least in part on the first speech recognition result and the first plurality of features; and

selecting the first speech recognition result based on the determining.

Assignments (4)
RELEASE (REEL 067417 / FRAME 0303) Recorded Jan 2, 2025
From: WELLS FARGO BANK, NATIONAL ASSOCIATION
To: CERENCE OPERATING COMPANY
Reel/Frame 069797/0422 →
SECURITY AGREEMENT Recorded Apr 15, 2024
From: CERENCE OPERATING COMPANY
To: WELLS FARGO BANK, N.A., AS COLLATERAL AGENT
Reel/Frame 067417/0303 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2022
From: TANG, MIN
To: CERENCE OPERATING COMPANY
Reel/Frame 061347/0632 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2022
From: TANG, MIN
To: CERENCE OPERATING COMPANY
Reel/Frame 061288/0140 →
Continuity (3)
Continuation 16830638 · Mar 26, 2020
Provisional Application 62825391 · Mar 28, 2019
Related Publication 20230169971A1 · Jun 1, 2023