IP Library › Granted Patent US 12,112,138
Granted Patent B2
US 12,112,138 · App. 17/830,889 · Granted Oct 8, 2024

Systems and methods for an end-to-end evaluation and testing framework for task-oriented dialog systems

Inventors: Guangsen Wang (Singapore, SG); Samson Min Rong Tan (Singapore, SG); Shafiq Rayhan Joty (Singapore, SG); Gang Wu (Santa Clara, CA); Chu Hong Hoi (Singapore, SG); Ka Chun Au (San Francisco, CA)
Assignee: Salesforce, Inc.
G06F40/35G06F40/186G06F40/40H04L51/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,112,138
App. No.
17/830,889
Granted
Oct 8, 2024
Kind
B2
Abstract

Embodiments provide a software framework for evaluating and troubleshooting real-world task-oriented bot systems. Specifically, the evaluation framework includes a generator that infers dialog acts and entities from bot definitions and generates test cases for the system via model-based paraphrasing. The framework may also include a simulator for task-oriented dialog user simulation that supports both regression testing and end-to-end evaluation. The framework may also include a remediator to analyze and visualize the simulation results, remedy some of the identified issues, and provide actionable suggestions for improving the task-oriented dialog system.

Claims (54)

1. A method of simulating a dialog, the method comprising:

receiving, via a communication interface, a plurality of task-oriented dialog data generated from a dialog agent;

determining a plurality of natural language understanding pairs including bot dialog acts and respective bot messages based on the plurality of task-oriented dialog data;

determining a plurality of goal pairs including goal entity slots and respective goal entity slot values based on the plurality of task-oriented dialog data;

determining a plurality of natural language generation templates based on the plurality of task-oriented dialog data;

generating a simulated task-oriented dialog based on the plurality of natural language understanding pairs, the plurality of goal pairs, and the plurality of natural language generation templates;

generating simulation results from the simulated task-oriented dialog; and

generating, based on simulation results, an actionable suggestion relating to an adjustment to the dialog agent.

2. The method of claim 1 , wherein the plurality of task-oriented dialog data comprises responses to application programming interface (API) calls to the dialog agent.

3. The method of claim 1 , wherein the plurality of task-oriented dialog data comprises metadata associated with the dialog agent.

4. The method of claim 1 , further comprising:

determining an intent training utterance based on the plurality of task-oriented dialog data.

5. The method of claim 4 , further comprising:

generating, by a plurality of models, a plurality of paraphrases based on the intent training utterance.

6. The method of claim 5 , further comprising:

filtering the plurality of paraphrases based on a similarity metric to produce a subset of the plurality of paraphrases.

7. The method of claim 6 , further comprising:

generating the simulated task-oriented dialog based on the subset of the plurality of paraphrases.

8. A system for dialog simulation, the system comprising:

a communication interface that receives a plurality of task-oriented dialog data generated from a dialog agent; and

one or more hardware processors that:

determines a plurality of natural language understanding pairs including bot dialog acts and respective bot messages based on the plurality of task-oriented dialog data;

determines a plurality of goal pairs including goal entity slots and respective goal entity slot values based on the plurality of task-oriented dialog data;

determines a plurality of natural language generation templates based on the plurality of task-oriented dialog data;

generates a simulated task-oriented dialog based on the plurality of natural language understanding pairs, the plurality of goal pairs, and the plurality of natural language generation templates;

generates simulation results from the simulated task-oriented dialog; and

generates, based on simulation results, an actionable suggestion relating to an adjustment to the dialog agent.

9. The system of claim 8 , wherein the plurality of task-oriented dialog data comprises responses to application programming interface (API) calls to the dialog agent.

10. The system of claim 8 , wherein the plurality of task-oriented dialog data comprises metadata associated with the dialog agent.

11. The system of claim 8 , wherein the one or more hardware processors further:

determines an intent training utterance based on the plurality of task-oriented dialog data.

12. The system of claim 11 , wherein the one or more hardware processors further:

generates, by a plurality of models, a plurality of paraphrases based on the intent training utterance.

13. The system of claim 12 , wherein the one or more hardware processors further:

filters the plurality of paraphrases based on a similarity metric to produce a subset of the plurality of paraphrases.

14. The system of claim 13 , wherein the one or more hardware processors further:

generates the simulated task-oriented dialog based on the subset of the plurality of paraphrases.

15. A processor-readable non-transitory storage medium storing a plurality of processor-executable instructions, the instructions being executed by a processor to perform operations comprising:

receiving, via a communication interface, a plurality of task-oriented dialog data generated from a dialog agent;

determining a plurality of natural language understanding pairs including bot dialog acts and respective bot messages based on the plurality of task-oriented dialog data;

determining a plurality of goal pairs including goal entity slots and respective goal entity slot values based on the plurality of task-oriented dialog data;

determining a plurality of natural language generation templates based on the plurality of task-oriented dialog data;

generating a simulated task-oriented dialog based on the plurality of natural language understanding pairs, the plurality of goal pairs, and the plurality of natural language generation templates;

generating simulation results from the simulated task-oriented dialog; and

generating, based on simulation results, an actionable suggestion relating to an adjustment to the dialog agent.

16. The processor-readable non-transitory storage medium of claim 15 , wherein the plurality of task-oriented dialog data comprises metadata associated with the dialog agent.

17. The processor-readable non-transitory storage medium of claim 15 , further comprising:

determining an intent training utterance based on the plurality of task-oriented dialog data.

18. The processor-readable non-transitory storage medium of claim 17 , further comprising:

generating, by a plurality of models, a plurality of paraphrases based on the intent training utterance.

19. The processor-readable non-transitory storage medium of claim 18 , further comprising:

filtering the plurality of paraphrases based on a similarity metric to produce a subset of the plurality of paraphrases.

20. The processor-readable non-transitory storage medium of claim 19 , further comprising:

generating the simulated task-oriented dialog based on the subset of the plurality of paraphrases.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 14, 2022
From: WANG, GUANGSEN; TAN, SAMSON; JOTY, SHAFIQ RAYHAN; WU, GANG; HOI, CHU HONG; AU, KA CHUN
To: SALESFORCE.COM, INC.
Reel/Frame 060188/0844 →
Continuity (2)
Provisional Application 63303850 · Jan 27, 2022
Related Publication 20230237275A1 · Jul 27, 2023