IP Library Granted Patent US 11,935,521
Granted Patent B2
US 11,935,521 · App. 16/941,051 · Granted Mar 19, 2024

Real-time feedback for efficient dialog processing

Inventor: Michael Richard Kennewick (Bellevue, WA)
Assignee: Oracle International Corporation
G10L15/1822G10L13/08G10L15/1815G10L15/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,935,521
App. No.
16/941,051
Granted
Mar 19, 2024
Kind
B2
Abstract

Techniques are described for improving the efficiency of dialog processing by prompting and processing user feedback substantially in real time. A dialog system receives speech input from a user and processes a first portion of the speech input to determine an initial discerned intent. The dialog system causes display of a visual indication of the initial discerned intent. The visual indication of the discerned intent is used to guide the dialog so that the user can correct or confirm the initial discerned intent in a natural fashion. If the initial discerned intent is inaccurate, the user can provide feedback correcting the dialog system, and the dialog system processes a second portion of the speech input to determine a modified discerned intent. Thus, the dialog system can use the feedback to correct misunderstandings on the fly.

Claims (52)

1. A method comprising:

receiving, by a dialog system, speech input from a user;

processing, by the dialog system using a trained neural network, wherein the neural network has been trained to predict a likely domain based upon sample domain to utterance pairs, a first portion of the speech input to determine an initial discerned domain;

responsive to determining the initial discerned domain and while receiving the speech input, causing display, by the dialog system, of a visual indication of the initial discerned domain, wherein the visual indication of the initial discerned domain comprises an image indicating the initial discerned domain, the initial discerned domain comprising a category corresponding to multiple intents; and

processing, by the dialog system, a second portion of the speech input to determine a modified discerned domain, wherein the second portion of the speech input completes the first portion of the speech input and clarifies the initial discerned domain responsive to the visual indication of the initial discerned domain, wherein the modified discerned domain is different from the initial discerned domain, and wherein the first portion of speech input and the second portion of speech input are processed as a continuous stream of input.

2. The method of claim 1 , wherein the dialog system processes the speech input substantially in real time.

3. The method of claim 1 , further comprising executing, by the dialog system, a task corresponding to the modified discerned domain.

4. The method of claim 1 , further comprising:

determining, by the dialog system, a response based upon the modified discerned domain; and

providing, by the dialog system, the response to the user.

5. The method of claim 1 , wherein determining the initial discerned domain comprises:

computing, by the dialog system, a plurality of scores for a respective plurality of potential domains and the first portion of the speech input; and

determining, by the dialog system, that the score for a particular domain, of the plurality of potential domains, exceeds a threshold value, thereby setting the particular domain to the initial discerned domain,

wherein causing display of the visual indication of the initial discerned domain is triggered by the determination that the score exceeds the threshold value.

6. The method of claim 1 , wherein determining the initial discerned domain comprises:

determining, by an automated speech recognition subsystem of the dialog system, a first text utterance corresponding to the first portion of the speech input;

providing, by the automated speech recognition subsystem to a natural language understanding subsystem of the dialog system, the first text utterance; and

determining, by the natural language understanding subsystem based upon the first text utterance, the initial discerned domain.

7. A non-transitory computer-readable memory storing a plurality of instructions executable by one or more processors, the plurality of instructions comprising instructions that when executed by the one or more processors cause the one or more processors to perform processing comprising:

receiving speech input from a user;

processing, using a trained neural network, wherein the neural network has been trained to predict a likely domain based upon sample domain to utterance pairs, a first portion of the speech input to determine an initial discerned domain;

responsive to determining the initial discerned domain and while receiving the speech input, causing display of a visual indication of the initial discerned domain, wherein the visual indication of the initial discerned domain comprises an image indicating the initial discerned domain, the initial discerned domain comprising a category corresponding to multiple intents; and

processing a second portion of the speech input to determine a modified discerned domain, wherein the second portion of the speech input completes the first portion of the speech input and clarifies the initial discerned domain responsive to the visual indication of the initial discerned domain, wherein the modified discerned domain is different from the initial discerned domain, and wherein the first portion of speech input and the second portion of speech input are processed as a continuous stream of input.

8. The non-transitory computer-readable memory of claim 7 , wherein the speech input is processed substantially in real time.

9. The non-transitory computer-readable memory of claim 7 , the processing further comprising executing a task corresponding to the modified discerned domain.

10. The non-transitory computer-readable memory of claim 7 , the processing further comprising:

determining a response based upon the modified discerned domain; and

providing the response to the user.

11. The non-transitory computer-readable memory of claim 7 , wherein determining the initial discerned domain comprises:

computing a plurality of scores for a respective plurality of potential domains and the first portion of the speech input; and

determining that the score for a particular domain, of the plurality of potential domains, exceeds a threshold value, thereby setting the particular domain to the initial discerned domain,

wherein causing display of the visual indication of the initial discerned domain is triggered by the determination that the score exceeds the threshold value.

12. The non-transitory computer-readable memory of claim 7 , wherein determining the initial discerned domain comprises:

determining, by an automated speech recognition subsystem, a first text utterance corresponding to the first portion of the speech input;

providing, by the automated speech recognition subsystem to a natural language understanding subsystem, the first text utterance; and

determining, by the natural language understanding subsystem based upon the first text utterance, the initial discerned domain.

13. A system comprising:

one or more processors;

a memory coupled to the one or more processors, the memory storing a plurality of instructions executable by the one or more processors, the plurality of instructions comprising instructions that when executed by the one or more processors cause the one or more processors to perform processing comprising:

receiving speech input from a user;

processing, using a trained neural network, wherein the neural network has been trained to predict a likely domain based upon sample domain to utterance pairs, a first portion of the speech input to determine an initial discerned domain;

responsive to determining the initial discerned domain and while receiving the speech input, causing display of a visual indication of the initial discerned domain, wherein the visual indication of the initial discerned domain comprises an image indicating the initial discerned domain, the initial discerned domain comprising a category corresponding to multiple intents; and

processing a second portion of the speech input to determine a modified discerned domain, wherein the second portion of the speech input completes the first portion of the speech input and clarifies the initial discerned domain responsive to the visual indication of the initial discerned domain, wherein the modified discerned domain is different from the initial discerned domain, and wherein the first portion of speech input and the second portion of speech input are processed as a continuous stream of input.

14. The system of claim 13 , wherein the speech input is processed substantially in real time.

15. The system of claim 13 , the processing further comprising executing a task corresponding to the modified discerned domain.

16. The system of claim 13 , the processing further comprising:

determining a response based upon the modified discerned domain; and

providing the response to the user.

17. The system of claim 13 , wherein determining the initial discerned domain comprises:

computing a plurality of scores for a respective plurality of potential domains and the first portion of the speech input; and

determining that the score for the initial discerned domain exceeds a threshold value,

wherein causing display of the visual indication of the initial discerned domain is triggered by the determination that the score for the initial discerned domain exceeds the threshold value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 29, 2020
From: KENNEWICK, MICHAEL RICHARD
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 053345/0218 →
Continuity (2)
Provisional Application 62899641 · Sep 12, 2019
Related Publication 20210082412A1 · Mar 18, 2021