IP Library Granted Patent US 9,384,736
Granted Patent B2
US 9,384,736 · App. 13/590,699 · Granted Jul 5, 2016

Method to provide incremental UI response based on multiple asynchronous evidence about user input

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,384,736
App. No.
13/590,699
Granted
Jul 5, 2016
Kind
B2
Abstract

Techniques disclosed herein include systems and methods for managing user interface responses to user input including spoken queries and commands. This includes providing incremental user interface (UI) response based on multiple recognition results about user input that are received with different delays. Such techniques include providing an initial response to a user at an early time, before remote recognition results are available. Systems herein can respond incrementally by initiating an initial UI response based on first recognition results, and then modify the initial UI response after receiving secondary recognition results. Since an initial response begins immediately, instead of waiting for results from all recognizers, it reduces the perceived delay by the user before complete results get rendered to the user.

Claims (38)

1. A computer-implemented method for managing speech recognition response, the computer-implemented method comprising:

receiving a spoken utterance at a client electronic device, the client electronic device having a local automated speech recognizer;

analyzing the spoken utterance using the local automated speech recognizer;

transmitting at least a portion of the spoken utterance over a communication network to a remote automated speech recognizer that analyzes spoken utterances and returns remote speech recognition results;

prior to receiving a remote speech recognition result from the remote automated speech recognizer, initiating a response via a user interface of the client electronic device, the response corresponding to the spoken utterance, wherein at least an initial portion of the response is based on a local speech recognition result from the local automated speech recognizer, wherein the local speech recognition result is classified into one of multiple reliability classes based on a confidence value representing accuracy of the local speech recognition result, and wherein the initiated response is selected based on the reliability class assigned to the local speech recognition result; and

modifying the response after the response has been initiated and prior to completing delivery of the response via the user interface such that modifications to the response are delivered via the user interface as a portion of the response, the modifications being based on the remote speech recognition result.

2. The computer-implemented method of claim 1 , wherein initiating the response includes using a text-to-speech system that begins producing audible speech.

3. The computer-implemented method of claim 2 , wherein producing the audible speech includes using one or more filler words that convey possession of the speech recognition results.

4. The computer-implemented method of claim 3 , wherein modifying the response includes adding words to the response that are produced audibly by the text-to-speech system after the one or more filler words have been audibly produced.

5. The computer-implemented method of claim 1 , wherein initiating the response includes using a graphical user interface that begins displaying graphics corresponding to the spoken utterance.

6. The computer-implemented method of claim 5 , wherein displaying the graphics includes using a sequence of graphics that appears as a commencement of search results.

7. The computer-implemented method of claim 1 , wherein the remote speech recognition result is received during delivery of the response via the user interface.

8. The computer-implemented method of claim 1 , further comprising:

identifying a predicted delay in receiving remote recognition results; and

basing the initiated response on the predicted delay.

9. The computer-implemented method of claim 8 , wherein basing the initiated response on the predicted delay includes selecting an initial response span based on the predicted delay.

10. The computer-implemented method of claim 9 , wherein selecting the initial response span includes selecting a number of words to produce that provides sufficient time to receive the remote speech recognition results.

11. The computer-implemented method of claim 9 , wherein basing the initiated response on the predicted delay includes adjusting a speed of speech production based on the predicted delay.

12. The computer-implemented method of claim 8 , wherein identifying the predicted delay includes recalculating predicted delays after producing each respective word of an initial response via the user interface.

13. The computer-implemented method of claim 1 , further comprising identifying that the initiated response conveyed via text-to-speech is incorrect based on remote speech recognition results, and correcting the initiated response using an audible excuse transitions.

14. The computer-implemented method of claim 1 , wherein the remote automated speech recognizer is a general-purpose open-domain recognizer.

15. The computer-implemented method of claim 1 , wherein the remote automated speech recognizer is a content-specific recognizer.

16. The computer-implemented method of claim 1 , wherein the spoken utterance is either a voice query or a voice command.

17. The computer-implemented method of claim 1 , wherein the client electronic device is a mobile telephone.

18. A system for managing speech recognition response, the system comprising:

a processor; and

a memory coupled to the processor, the memory storing instructions that, when executed by the processor, causes the system to perform:

receiving a spoken utterance at a client electronic device, the client electronic device having a local automated speech recognizer;

analyzing the spoken utterance using the local automated speech recognizer;

transmitting at least a portion of the spoken utterance over a communication network to a remote automated speech recognizer that analyzes spoken utterances and returns remote speech recognition results;

prior to receiving a remote speech recognition result from the remote automated speech recognizer, initiating a response via a user interface of the client electronic device, the response corresponding to the spoken utterance, wherein at least an initial portion of the response is based on a local speech recognition result from the local automated speech recognizer, wherein the local speech recognition result is classified into one of multiple reliability classes based on a confidence value representing accuracy of the local speech recognition result, and wherein the initiated response is selected based on the reliability class assigned to the local speech recognition result; and

modifying the response after the response has been initiated and prior to completing delivery of the response via the user interface such that modifications to the response are delivered via the user interface as a portion of the response, the modifications being based on the remote speech recognition result.

19. A computer program product including a non-transitory computer-storage medium having instructions stored thereon for processing data information, such that the instructions, when carried out by a processing device, cause the processing device to perform:

receiving a spoken utterance at a client electronic device, the client electronic device having a local automated speech recognizer;

analyzing the spoken utterance using the local automated speech recognizer;

transmitting at least a portion of the spoken utterance over a communication network to a remote automated speech recognizer that analyzes spoken utterances and returns remote speech recognition results;

prior to receiving a remote speech recognition result from the remote automated speech recognizer, initiating a response via a user interface of the client electronic device, the response corresponding to the spoken utterance, wherein at least an initial portion of the response is based on a local speech recognition result from the local automated speech recognizer, wherein the local speech recognition result is classified into one of multiple reliability classes based on a confidence value representing accuracy of the local speech recognition result, and wherein the initiated response is selected based on the reliability class assigned to the local speech recognition result; and

modifying the response after the response has been initiated and prior to completing delivery of the response via the user interface such that modifications to the response are delivered via the user interface as a portion of the response, the modifications being based on the remote speech recognition result.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2023
From: NUANCE COMMUNICATIONS, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 065578/0676 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 21, 2012
From: LABSKY, MARTIN; MACEK, TOMAS; KUNC, LADISLAV; KLEINDIENST, JAN
To: NUANCE COMMUNICATIONS, INC.
Reel/Frame 028821/0691 →