IP Library Granted Patent US 10,395,646
Granted Patent B2
US 10,395,646 · App. 15/594,308 · Granted Aug 27, 2019

Two-stage training of a spoken dialogue system

Inventors: Seyed Mehdi Fatemi Booshehri (Montreal, CA); Layla El Asri (Montreal, CA); Hannes Schulz (Karlsruhe, DE); Jing He (Toronto, CA); Kaheer Suleman (Cambridge, CA)
Assignee: Microsoft Technology Licensing, LLC
G10L15/063G06F3/167G06N3/0427G10L15/14G10L15/16G10L25/51G06N3/0445G10L15/22G10L2015/0638
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,395,646
App. No.
15/594,308
Granted
Aug 27, 2019
Kind
B2
Abstract

Described herein are systems and methods for two-stage training of a spoken dialog system. The first stage trains a policy network using external data to produce a semi-trained policy network. The external data includes one or more known fixed dialogs. The second stage trains the semi-trained policy network through interaction to produce a trained policy network. The interaction may be interaction with a user simulator.

Claims (37)

1. A system, comprising:

a spoken dialogue system, comprising:

a policy network for producing a probability distribution over all possible actions performable in response to a given state of a dialogue; and

a value network operably connected to the policy network for estimating the given state of the dialogue and providing an advantage signal to the policy network that indicates a success level of the policy network;

a storage device operably connected to the policy network and storing one or more fixed dialogues used to train the policy network in a first stage of training; and

a user simulator operably connected to the policy network and to the value network and used to simulate one or more user dialogues to train the policy network in a second stage of training.

2. The system of claim 1 , wherein the first stage of training produces a semi-trained policy network and the second stage of training produces a trained policy network.

3. The system of claim 1 , wherein the policy network and the value network each comprises a neural network.

4. The system of claim 1 , wherein the trained spoken dialogue system is accessed by a client-computing device.

5. A method, comprising:

training a policy network in a spoken dialogue system using external data comprising one or more fixed dialogues in which a machine action at each turn of a dialogue is known to produce a semi-trained policy network that has a first level of training; and

training the semi-trained policy network through interaction with one or more dialogues in which a machine action at each turn of a dialogue is not known to produce a trained policy network that has a second level of training that is greater than the first level of training.

6. The method of claim 5 , wherein training the policy network using one or more fixed dialogues comprises:

receiving, from a storage device, a state of a fixed dialogue in the one or more fixed dialogues;

producing a predicted output, the predicted output comprising a predicted probability distribution over all possible actions; and

comparing the predicted output to an expected output, the expected output comprising a known probability distribution over all of the possible actions.

7. The method of claim 6 , further comprising repeating the operations of receiving, producing, and comparing to reduce a difference between the predicted output and the expected output.

8. The method of claim 7 , wherein the operations of receiving, producing, and comparing are repeated until the difference between the predicted output and the expected output is below a threshold value.

9. The method of claim 7 , wherein the operations of receiving, producing, and comparing are repeated until a categorical cross-entropy between the predicted output and the expected output is minimized.

10. The method of claim 5 , wherein training the semi-trained policy network through interaction with the one or more dialogues in which the machine action at each turn of the dialogue is not known comprises training the semi-trained policy network using a user simulator that simulates the one or more dialogues in which the machine action at each turn of the dialogue is not known.

11. The method of claim 10 , wherein training the semi-trained policy network using the user simulator comprises:

receiving, from the user simulator, a user turn in a dialogue;

in response to receiving the user turn, determining a state of the dialogue;

producing a predicted output based on the determined state of the dialogue, the predicted output comprising a predicted probability distribution over all possible actions or a probability associated with one possible action;

receiving, from a value network, an advantage signal representing a success level of the policy network associated with the predicted output.

12. The method of claim 11 , further comprising repeating the operations of receiving, producing, and receiving until the semi-trained policy network achieves a respective convergence.

13. The method of claim 11 , wherein the probability associated with the one possible action is included in a sequence of probabilities associated with a sequence of possible actions.

14. A spoken dialogue system, comprising:

a policy network configured to produce a probability distribution over one or more possible actions performable in response to a given state of a dialogue; and

a value network connected to the policy network and configured to receive the given state of the dialogue and provide an advantage signal to the policy network that indicates an accuracy of the probability distribution,

wherein the policy network is trained using one or more fixed dialogues in which a machine action at each turn of a dialogue is known and one or more simulated dialogues in which a machine action at each turn of a dialogue is not known.

15. The spoken dialogue system of claim 14 , wherein the one or more dialogues in which the machine action at each turn is not known comprises one or more simulated dialogues received from a user simulator.

16. The spoken dialogue system of claim 14 , wherein the policy network uses a policy gradient algorithm to learn to produce the probability distribution, the policy gradient algorithm comprising an advantage function.

17. The spoken dialogue system of claim 16 , wherein the value network produces the advantage function using a regression algorithm.

18. The spoken dialogue system of claim 14 , wherein the probability distribution over the one or more possible actions comprises a probability distribution over all possible actions.

19. The spoken dialogue system of claim 14 , wherein the probability distribution over the one or more possible actions comprises a probability distribution over a sequence of possible actions.

20. The spoken dialogue system of claim 14 , wherein the policy network and the value network each comprises a neural network.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 1, 2020
From: MALUUBA INC.
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 053116/0878 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 12, 2017
From: FATEMI BOOSHEHRI, SEYED MEHDI; EL ASRI, LAYLA; SCHULZ, HANNES; HE, JING; SULEMAN, KAHEER
To: MALUUBA INC.
Reel/Frame 042363/0974 →
Continuity (2)
Provisional Application 62336163 · May 13, 2016
Related Publication 20170330556A1 · Nov 16, 2017
Cited By (1)
US 12,518,140