IP Library Granted Patent US 11,790,899
Granted Patent B2
US 11,790,899 · App. 16/952,413 · Granted Oct 17, 2023

Determining state of automated assistant dialog

Inventors: Abhinav Rastogi (Santa Clara, CA); Larry Paul Heck (Los Altos, CA); Dilek Hakkani-Tur (Los Altos, CA)
Assignee: GOOGLE LLC
G10L15/197G06N3/08G10L15/16G10L15/1815G10L15/22G10L15/30G06N3/044G10L15/1822G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,790,899
App. No.
16/952,413
Granted
Oct 17, 2023
Kind
B2
Abstract

Determining a dialog state of an electronic dialog that includes an automated assistant and at least one user, and performing action(s) based on the determined dialog state. The dialog state can be represented as one or more slots and, for each of the slots, one or more candidate values for the slot and a corresponding score (e.g., a probability) for each of the candidate values. Candidate values for a slot can be determined based on language processing of user utterance(s) and/or system utterance(s) during the dialog. In generating scores for candidate value(s) of a given slot at a given turn of an electronic dialog, various features are determined based on processing of the user utterance and the system utterance using a memory network. The various generated features can be processed using a scoring model to generate scores for candidate value(s) of the given slot at the given turn.

Claims (56)

1. A method implemented by one or more processors, comprising:

identifying a conversation context of an electronic dialog that includes an automated assistant and a user, the conversation context based at least in part on a system utterance of the automated assistant, and a user utterance of the user, the system utterance and the user utterance provided during a turn of the electronic dialog;

determining, based on the conversation context, one or more candidate values for a slot;

identifying a textual descriptor, for the slot, that describes the parameters that can be defined by the candidate values for the slot;

generating, based on processing the conversation context using one or more memory networks:

one or more representations for the system utterance and the user utterance, and

candidate value features for each of the candidate values for the slot, wherein generating the candidate value features for each of the candidate values for the slot comprises processing the textual descriptor, for the slot, using one or more of the memory networks;

generating, based on processing the one or more representations and the candidate value features, a score for each of the candidate values for the slot;

selecting a given value, of the candidate values for the slot, based on the scores for the candidate values; and

performing a further action based on the selected given value for the slot.

2. The method of claim 1 , wherein generating the one or more representations for the system utterance and the user utterance comprises:

generating the one or more representations based on final and/or hidden states of the memory network after the processing of the system utterance and the user utterance.

3. The method of claim 1 , wherein determining the candidate values for the slot comprises determining the given value based on one or more given terms of the user utterance.

4. The method of claim 1 , wherein the one or more candidate values include the given value and an additional value.

5. The method of claim 4 , wherein generating the score for the given value is based on processing, using a trained scoring model: the one or more representations and the candidate value features for the given value; and wherein generating the score for the additional value is based on processing, using the trained scoring model: the one or more representations and the candidate value features for the additional value.

6. The method of claim 1 , wherein the one or more candidate values further comprise an indifferent value, and wherein generating the score for the indifferent value is based on the one or more representations, and a score for the indifferent value in an immediately preceding turn of the electronic dialog.

7. The method of claim 1 , wherein each of the scores is a probability.

8. The method of claim 1 , further comprising:

selecting a domain based on the electronic dialog;

selecting the slot based on it being assigned to the domain.

9. The method of claim 1 , wherein performing the further action based on the selected given value for the slot comprises:

generating an agent command that includes the selected given value for the slot; and

transmitting the agent command to an agent over one or more networks, wherein the agent command causes the agent to generate responsive content and transmit the responsive content over one or more networks.

10. The method of claim 9 , further comprising:

receiving the responsive content generated by the agent; and

transmitting, to a client device at which the user utterance was provided, output that is based on the responsive content generated by the agent.

11. The method of claim 1 , wherein performing the further action based on the selected given value for the slot comprises:

generating an additional system utterance based on the selected given value; and

incorporating the additional system utterance in a following turn of the electronic dialog for presentation to the user, the following turn immediately following the turn in the electronic dialog.

12. An apparatus, comprising:

memory storing instructions;

one or more processors configured to execute the instructions stored in the memory to perform a method comprising:

identifying a conversation context of an electronic dialog that includes an automated assistant and a user, the conversation context based at least in part on a system utterance of the automated assistant, and a user utterance of the user, the system utterance and the user utterance provided during a turn of the electronic dialog;

determining, based on the conversation context, one or more candidate values for a slot;

identifying a textual descriptor, for the slot, that describes the parameters that can be defined by the candidate values for the slot;

generating, based on processing the conversation context using one or more memory networks:

one or more representations for the system utterance and the user utterance, and

candidate value features for each of the candidate values for the slot, wherein generating the candidate value features for each of the candidate values for the slot comprises processing the textual descriptor, for the slot, using one or more of the memory networks;

generating, based on processing the one or more representations and the candidate value features, a score for each of the candidate values for the slot;

selecting a given value, of the candidate values for the slot, based on the scores for the candidate values; and

performing a further action based on the selected given value for the slot.

13. The apparatus of claim 12 , wherein generating the one or more representations for the system utterance and the user utterance comprises:

generating the one or more representations based on final and/or hidden states of the memory network after the processing of the system utterance and the user utterance.

14. The apparatus of claim 12 , wherein determining the candidate values for the slot comprises determining the given value based on one or more given terms of the user utterance.

15. The apparatus of claim 12 , wherein the one or more candidate values include the given value and an additional value.

16. The apparatus of claim 15 , wherein generating the score for the given value is based on processing, using a trained scoring model: the one or more representations and the candidate value features for the given value; and wherein generating the score for the additional value is based on processing, using the trained scoring model: the one or more representations and the candidate value features for the additional value.

17. The apparatus of claim 12 , wherein the one or more candidate values further comprise an indifferent value, and wherein generating the score for the indifferent value is based on the one or more representations, and a score for the indifferent value in an immediately preceding turn of the electronic dialog.

18. The apparatus of claim 12 , wherein the method further comprises:

selecting a domain based on the electronic dialog;

selecting the slot based on it being assigned to the domain.

19. The apparatus of claim 12 , wherein performing the further action based on the selected given value for the slot comprises:

generating an agent command that includes the selected given value for the slot; and

transmitting the agent command to an agent over one or more networks, wherein the agent command causes the agent to generate responsive content and transmit the responsive content over one or more networks.

20. The apparatus of claim 12 , wherein performing the further action based on the selected given value for the slot comprises:

generating an additional system utterance based on the selected given value; and

incorporating the additional system utterance in a following turn of the electronic dialog for presentation to the user, the following turn immediately following the turn in the electronic dialog.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 29, 2020
From: RASTOGI, ABHINAV; HECK, LARRY PAUL; HAKKANI-TUR, DILEK
To: GOOGLE LLC
Reel/Frame 054764/0282 →
Continuity (2)
Continuation 16321294
Related Publication 20210074279A1 · Mar 11, 2021
Cited By (2)
US 12,254,875 US 12,361,931