IP Library › Granted Patent US 10,824,798
Granted Patent B2
US 10,824,798 · App. 15/804,093 · Granted Nov 3, 2020

Data collection for a new conversational dialogue system

Inventors: Percy Shuo Liang (Palo Alto, CA); Daniel Klein (Orinda, CA); Laurence Steven Gillick (Newton, MA); Jordan Rian Cohen (Kure Beach, NC); Linda Kathleen Arsenault (Chelmsford, MA); Joshua James Clausman (Somerville, MA); Adam David Pauls (Berkeley, CA); David Leo Wright Hall (Berkeley, CA)
Assignee: Semantic Machines, Inc.
G06F40/169G06F3/0482G06F7/08G06F40/20G06F40/205G06F40/30G06N20/00G10L15/063G10L15/22G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,824,798
App. No.
15/804,093
Granted
Nov 3, 2020
Kind
B2
Abstract

A data collection system is based on a general set of dialogue acts which are derived from a database schema. Crowd workers perform two types of tasks: (i) identification of sensical dialogue paths and (ii) performing context-dependent paraphrasing of these dialogue paths into real dialogues. The end output of the system is a set of training examples of real dialogues which have been annotated with their logical forms. This data can be used to train all three components of the dialogue system: (i) the semantic parser for understanding context-dependent utterances, (ii) the dialogue policy for generating new dialogue acts given the current state, and (iii) the generation system for both deciding what to say and how to render it in natural language.

Claims (47)

1. A method for training an annotated dialogue system, comprising:

generating, by a data generation application executing on a machine, a first list of canonical utterances in a dialogue tree of possible dialogues, the first list of canonical utterances generated in response to an input received from an annotating user;

receiving from the annotating user a first selection of a first canonical utterance from the first list, the selection indicating a first step in a multi-step dialogue;

generating a second list of canonical utterances, the second list of canonical utterances generated in response to the first selection from the first list;

receiving from the user a second selection of a second canonical utterance from the second list, the second selection indicating a second step in the multi-step dialogue that also includes the first step;

presenting a user interface indicating a dialogue path including both the first step and the second step;

receiving from the annotating user via the user interface a compound paraphrase for the dialogue path including both the first step and the second step; and

training the annotated dialogue system on annotated data including the compound paraphrase, the first step, and the second step.

2. The method of claim 1 , wherein the compound paraphrase represents paraphrasing for the whole dialogue path, wherein the first selection from the first list and the second selection from the second list are two canonical utterances of a plurality of three or more selected canonical utterances in the dialogue path.

3. The method of claim 1 , further comprising receiving, through the user interface, a ranking from the annotating user for one or more canonical utterances of the first list and the second list.

4. The method of claim 1 , wherein further comprising storing the annotated training data in a format suitable for training the annotated dialogue system.

5. The method of claim 1 , wherein generating a first list of a plurality of canonical utterances includes:

querying a database of possible actions,

generating a logical form at least in part from each action; and

generating a canonical utterance from each logical form.

6. The method of claim 5 , wherein a logical form generated from an action includes computer code associated with the action, and wherein a canonical utterance generated from the logical form is a pseudocode representation of the computer code associated with the action.

7. The method of claim 1 , wherein the first list and second list are each ranked according to a ranking model.

8. The method of claim 7 , wherein the ranking model is based at least in part from user rankings of previous canonical utterances.

9. The method of claim 1 , wherein the first list and the second list are provided via the user interface.

10. The method of claim 9 , wherein the user interface includes state information for the current dialogue.

11. The method of claim 9 , wherein the user interface includes a search box.

12. The method of claim 9 , further comprising, in response to receiving the first input from the list, removing the one or more canonical utterances of the list that were not selected by the annotating user from the user interface.

13. The method of claim 9 , further comprising, in response to receive the first input, adding the selected canonical utterance to a list of selected canonical utterances.

14. A system for training an annotated dialogue system, comprising:

a processor;

memory;

one or more modules stored in memory and executable by the processor to:

generate, by a data generation application, a first list of canonical utterances in a dialogue tree of possible dialogues, the first list of canonical utterances generated in response to an input received from an annotating user,

receive from the annotating user a first selection of a first canonical utterance from the first list, the first selection indicating a first step in a multi-step dialogue,

generate a second list of canonical utterances, the second list of canonical utterances generated in response to the first selection from the first list,

receive from the annotating user a second selection of a second canonical utterance from the second list, the second selection indicating a second step in the multi-step dialogue that also includes the first step,

present a user interface indicating a dialogue path including both the first step and the second step,

receive from the annotating user via the user interface a compound paraphrase for the dialogue path including both the first step and the second step, and

train the annotated dialogue system on annotated data including the compound paraphrase, the first step, and the second step.

15. The system of claim 14 , wherein the one or more modules stored in memory are further executable by the processor to receive, through the user interface, a ranking from the annotating user for one or more canonical utterances of the first list and the second list.

16. The system of claim 14 , further comprising storing the annotated data in a format suitable for training a dialogue system.

17. The system of claim 14 , wherein the first list and second list are each ranked according to a ranking model.

18. The system of claim 17 , wherein the ranking model is based at least in part on user rankings of previous canonical utterances.

19. The system of claim 18 , wherein the first list and second list are provided within the interface.

20. A computer system including a processor and memory holding instructions executable by the processor to perform a method for training an annotated dialogue system, the method comprising:

generating a first list of canonical utterances in a dialogue tree of possible dialogues, the first list of canonical utterances generated in response to an input received from an annotating user;

receiving from the annotating user a first selection of a first canonical utterance from the first list, the first selection indicating a first step in a multi-step dialogue;

generating a second list of canonical utterances, the second list of canonical utterances generated in response to the first selection from the first list;

receiving from the annotating user a second selection of a second canonical utterance from the second list, the second selection indicating a second step in the multi-step dialogue that also includes the first step;

presenting a user interface indicating a dialogue path including both the first step and the second step and an input box configured for receiving a compound paraphrase for the dialogue path including both the first step and the second step;

receiving from the annotating user via the input box of the user interface a compound paraphrase for the dialogue path including both the first step and the second step; and

training the annotated dialogue system on annotated data including the compound paraphrase, the first step, and the second step.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2020
From: SEMANTIC MACHINES, INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 053904/0601 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 29, 2018
From: LIANG, PERCY SHUO; HALL, DAVID LEO WRIGHT; KLEIN, DANIEL; GILLICK, LAURENCE STEVEN; COHEN, JORDAN RIAN; CLAUSMAN, JOSHUA JAMES; PAULS, ADAM DAVID; ARSENAULT, LINDA KATHLEEN
To: SEMANTIC MACHINES, INC.
Reel/Frame 045380/0289 →
Continuity (2)
Provisional Application 62418035 · Nov 4, 2016
Related Publication 20180203833A1 · Jul 19, 2018