IP Library Granted Patent US 10,713,317
Granted Patent B2
US 10,713,317 · App. 15/419,497 · Granted Jul 14, 2020

Conversational agent for search

Inventors: Balaji Krishnamurthy (Kailash Dham, IN); Shagun Sodhani (Vasundhra Enclave, IN); Aarushi Arora (Faridabad, IN); Milan Aggarwal (Pitampura, IN)
Assignee: ADOBE INC.
G06F16/9535G06F16/90332G06F40/30G06F40/35G06N3/006G06N3/08G06N20/00G06N7/005
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,713,317
App. No.
15/419,497
Granted
Jul 14, 2020
Kind
B2
Abstract

A conversational agent facilitates conversational searches for users. The conversational agent is a reinforcement learning (RL) agent trained using a user model generated from existing session logs from a search engine. The user model is generated from the session logs by mapping entries from the session logs to user actions understandable by the RL agent and computing conditional probabilities of user actions occurring given previous user actions in the session logs. The RL agent is trained by conducting conversations with the user model in which the RL agent selects agent actions in response to user actions sampled using the conditional probabilities from the user model.

Claims (40)

1. A computer system comprising:

one or more processors; and

one or more computer storage media storing computer-useable instructions that, when used by the one or more processors, cause the one or more processors to:

generate a user model using session logs from a search engine, the session logs comprising user activities during search sessions that do not include conversation data, the user model generated by:

employing a set of rules with mappings between certain types of session log entries to specific user actions defined by a user action space understandable by a reinforcement learning agent to represent entries from the session logs as user actions from the user action space and each search session within the session logs as a series of user actions from the user action space, and

using the mapped entries from the session logs to compute conditional probabilities for user actions from the user action space for each of a plurality of different sets of previous user actions;

train the reinforcement learning agent using the user model by iteratively performing dialog turns of selecting an agent action and selecting a user action, wherein the user actions are selected during the dialog turns based at least in part on the conditional probabilities from the user model; and

employing the reinforcement learning agent as a conversational agent in a conversational search system.

2. The system of claim 1 , wherein the user model further comprises predefined probabilities for user actions defined by the user action space given various sets of previous agent actions by the reinforcement learning agent, and wherein the user actions are selected during the dialog turns based at least in part on the predefined probabilities.

3. The system of claim 1 , wherein the user actions defined by the user action space include one or more selected from the following: submitting a new search query, refining a search query, requesting more search result, selecting a search result, adding a search result to a cart, bookmarking a search result, selecting a cluster category, and searching for assets similar to a selected search result.

4. The system of claim 1 , wherein the agent action selected at each dialog turn is selected from an agent action space setting forth available agent actions for the reinforcement learning agent, wherein the agent action space includes one or more selected from the following: providing search results, asking for a use case, asking to refine a search query, asking for feedback, providing cluster categories, providing an offer to sign up, asking to add a search result item to a cart, providing a salutation, and providing help information.

5. The system of claim 1 , wherein the agent action selected at each dialog turn is based on a current system state and expected reward values.

6. The system of claim 5 , wherein the current system state is a vector comprising a value for at least one or more of the following parameters: a current user action, one or more previous user actions, one or more previous agent actions, search result scores, and length of conversation.

7. The system of claim 1 , wherein reinforcement learning agent is trained using Q-learning.

8. The system of claim 1 , wherein the reward provided at each dialog turn is based on a current system state by looking up a reward value in predefined data setting a reward value for each of a plurality of system states.

9. One or more computer storage media storing computer-useable instructions that, when executed by a computing device, cause the computing device to perform operations, the operations comprising:

generating a user model from entries in session logs from a search engine, the entries representing user activity during search sessions with the search engine that do not include conversation data, the user model generated by employing a set of rules with mapping between certain types of session log entries to specific user actions defined by a user action space understandable by a reinforcement learning agent to represent entries from the session logs as user actions from the user action space and each search session within the session logs as a series of user actions from the user action space;

training the reinforcement learning agent by using the user model to select user actions during training dialog turns; and

employing the reinforcement learning agent as a conversational agent in a conversational search system.

10. The one or more computer storage media of claim 9 , wherein the user model comprises conditional probabilities for user actions given various sets of previous user actions, the conditional probabilities calculated from the entries in the session logs, and wherein the user actions are selected during the training dialog turns based at least in part on the conditional probabilities.

11. The one or more computer storage media of claim 10 , wherein the user model further comprises predefined probabilities for user actions given various sets of previous agent actions by the reinforcement learning agent, and wherein the user actions are selected during the training dialog turns based at least in part on the predefined probabilities.

12. The one or more computer storage media of claim 9 , wherein generating the user model comprises:

computing conditional probabilities for user actions from the user action space given different sets of previous user actions, the conditional probabilities being calculated from the mapped entries from the session logs; and

generating the user model as a finite state machine using the conditional probabilities.

13. The one or more computer storage media of claim 9 , wherein the user actions defined by the user action space include one or more selected from the following: submitting a new search query, refining a search query, requesting more search result, selecting a search result, adding a search result to a cart, bookmarking a search result, selecting a cluster category, and searching for assets similar to a selected search result.

14. The one or more computer storage media of claim 9 , wherein training the reinforcement learning agent comprises performing a reinforcement learning conversation with the reinforcement learning agent and the user model, wherein the reinforcement learning conversation comprises a series of dialog turns, each dialog turn including an interaction involving a selected agent action and a selected user action, wherein a reward provided for each dialog turn, and wherein the user action is selected at each dialog turn based on the user model and the agent action is selected at each dialog turn to maximize an overall reward provided from the dialog turns of the reinforcement learning conversation.

15. The one or more computer storage media of claim 9 , wherein training the reinforcement learning agent comprises:

initializing a Q-value for each agent action defined by an agent action space;

selecting a first agent action based on the initialized Q-values; and

iteratively:

employing the user model to select a current user action;

measuring a reward based on the current system state;

updating the Q-value for each agent action based on the current system state; and

selecting a next agent action based on the updated Q-values.

16. The one or more computer storage media of claim 15 , the current system state is a vector comprising a value for at least one or more of the following parameters: the current user action, one or more previous user actions, one or more previous agent actions, search result scores, and length of conversation.

17. The one or more computer storage media of claim 15 , wherein measuring a reward based on the current system state comprises looking up a reward value in predefined data setting a reward value for each of a plurality of system states.

18. The one or more computer storage media of claim 9 , wherein the operations further comprise retraining the reinforcement learning agent using interactions with humans.

19. A computer system comprising:

means for generating a user model having conditional probabilities for user actions based on entries in session logs from a search engine, the entries in the session logs comprising user activities during search sessions that do not include conversation data, the user model generated using a set of rules with mappings between certain types of session log entries to specific user actions defined by a user action space understandable by a reinforcement learning agent to represent entries from the session logs as user actions from the user action space and each search session within the session logs as a series of user actions from the user action space; and

means for training a reinforcement learning agent by using the user model to select user actions during training dialog turns based at least in part on the conditional probabilities of the user model.

Assignments (2)
CHANGE OF NAME Recorded Nov 29, 2018
From: ADOBE SYSTEMS INCORPORATED
To: ADOBE INC.
Reel/Frame 047687/0115 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 14, 2017
From: KRISHNAMURTHY, BALAJI; SODHANI, SHAGUN; ARORA, AARUSHI; AGGARWAL, MILAN
To: ADOBE SYSTEMS INCORPORATED
Reel/Frame 041250/0771 →
Cited By (1)
US 12,363,158