IP Library › Granted Patent US 12,688,847
Granted Patent B2
US 12,688,847 · App. 17/783,083 · Granted Jul 21, 2026

Error-correction and extraction in request dialogs

Inventors: Stefan Constantin (Karlsruhe, DE); Alexander Waibel (Sammamish, WA)
Assignee: Zoom Communications, Inc.
G10L15/01G10L15/063G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,847
App. No.
17/783,083
Filed
Jun 7, 2022
Granted
Jul 21, 2026
Kind
B2
Art Unit
2656
USPC
704/235
Abstract

A system comprises a machine that is configured to act upon requests from a user and sensing means for sensing an operational-mode dialog stream from the user for the machine. The system also comprises a computing system that is configured to train a neural network through machine learning to output, for each training example in a training dialog stream dataset, a corrected request for the machine. The computing system is also configure to, in an operational mode, using the trained neural network, generate a corrected, operational-mode request for the machine based on the operational-mode dialog stream from the user for the machine, wherein the operational-mode dialog stream is sensed by the sensing means.

Claims (54)

1 . A system comprising:

a non-transitory computer-readable medium; and

one or more processor communicatively coupled to the non-transitory computer-readable medium, the one or more processors configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:

receive, from a sensing means for sensing an operational-mode communication from a user,

a request indicating a user intention;

use a trained neural network to:

label every word token in the request;

determine one or more reparandums in the request based on the labels; and

determine one or more repair phrases in the request based on the labels;

use a second trained neural network to determine additional semantic information based on the request;

use the trained neural network to determine a corrected user intention based on the one or more reparandums and the one or more repair phrases and the additional semantic information;

update the second trained neural network based on first semantic information within the request;

use the trained neural network to generate one or more actions based on the corrected user intention; and

output one or more commands to a machine to cause the machine to perform the one or more actions.

2 . The system of claim 1 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to:

train the trained neural network through machine learning to output, for each training example in a training dataset, one or more updated actions; and wherein the training dataset comprises training communications, wherein the training communications each comprise a training user-intention and a training correction for the training user-intention.

3 . The system of claim 2 , wherein the trained neural network:

is trained to identify reparandums in the training user-intentions in the training dataset and to identify training repairs in corrections to the training user-intentions in the training dataset, and to generate, based on the identified reparandums and repairs in the training dataset, the corrected user intention; and

is configured to generate the one or more actions based on a reparandum identified in the request.

4 . The system of claim 3 , wherein the trained neural network is configured to generate the one or more actions based on a reparandum identified in the request and a repair.

5 . The system of claim 3 , wherein the one or more processors are configured to execute further processor-executable instructions stored in the non-transitory computer-readable medium to use a second neural network to determine relations between reparandums and the intention within the request.

6 . The system of claim 1 , wherein the trained neural network has a fixed size output vocabulary.

7 . The system of claim 3 , wherein the trained neural network is trained to assign a label to word tokens in communications in the training dataset and determine one or more corrected user-intention for each training example in the training dataset based on the assigned labels.

8 . The system of claim 1 , wherein the trained neural network does not have a fixed size output vocabulary.

9 . The system of claim 1 , wherein the machine comprises a processor-based device selected from the group consisting of a robot, a computer, a mobile device, an appliance, a home entertainment system, a personal assistant, an automotive system, a healthcare system, and a medical device.

10 . The system of claim 1 , wherein:

the request comprises audio from the user; and

the sensing means comprises a microphone and a Natural Language Processor (NLP).

11 . The system of claim 1 , wherein:

the request comprises an electronic message comprising text; and

the sensing means comprises a Natural Language Processor (NLP) for processing the text in the electronic message.

12 . The system of claim 1 , wherein the sensing means comprises a sensor selected from the group consisting of a motion sensor, a camera, a pressure sensor, a proximity sensor, a humidity sensor, an ambient light sensor, a GPS receiver, and a touch-sensitive display.

13 . The system of claim 1 , wherein the sensing means is part of the machine.

14 . The system of claim 1 , wherein the system is part of the machine.

15 . The system of claim 1 , wherein the request from the user comprises an imperative request from the user for the machine.

16 . The system of claim 1 , wherein the request sensed by the sensing means comprises a communication modality selected from the group consisting of text, speech, physical things like gestures, head movements, actions.

17 . The system of claim 1 , wherein the request sensed by the sensing means comprises a dialog stream, wherein the dialog stream comprises dialog from the user.

18 . A method comprising:

receiving, from a sensing means for sensing an operational-mode communication from a user,

a request from a user, the request indicating a user intention;

using a trained neural network to:

label every word token in the request;

determine one or more reparandums in the request based on the labels; and

determine one or more repair phrases in the request based on the labels;

using a second trained neural network to determine additional semantic information based on the request;

using the trained neural network to determine a corrected user intention based on one or more reparandums and the one or more repair phrases and the additional semantic information;

updating the second trained neural network based on first semantic information within the request;

generating, using the trained neural network, one or more actions based on the corrected user intention; and

outputting one or more commands to a robotic system to cause the robotic system to perform the one or more actions.

19 . The method of claim 18 , further comprising training the trained neural network through machine learning to output, for each training example in a training dataset, one or more corrected user-intention for the machine, wherein the machine is configured to act upon user-intentions from a user, wherein:

the training dataset comprises training communications, wherein the training communications each comprise a training user-intention for the machine and a training correction for the training user-intention.

20 . The method of claim 19 , wherein:

training the trained neural network comprises training the trained neural network to identify reparandums in the training user-intentions in the training dataset and to identify training repairs in corrections to the training user-intentions in the training dataset, and to generate, based on the identified reparandums and repairs in the training dataset, one or more corrected user-intentions for the machine; and

generating the corrected user intention comprises generating the corrected user intention based on a reparandum identified in the request and an operational-mode repair identified in the request.

Assignments (2)
CHANGE OF NAME Recorded Jun 17, 2026
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 075767/0910 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 7, 2022
From: INTERACTIVE-AI LLC
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 060121/0079 →
Continuity (2)
Provisional Application 62947946 · Dec 13, 2019
Related Publication 20230013768A1 · Jan 19, 2023
References Cited (33)
US 7860719B2 · Maskey et al. · 2010 [cited by applicant]
US 9514098B1 · Subramanya · 2016 [cited by examiner]
US 10431207B2 · Angkititrakul · 2019 [cited by examiner]
US 20080004865A1 · Weng · 2008 [cited by examiner]
US 20080046229A1 · Maskey et al. · 2008 [cited by applicant]
US 20110040554A1 · Audhkhasi · 2011 [cited by examiner]
US 20170116185A1 · Erickson · 2017 [cited by examiner]
US 20170161256A1 · Hori · 2017 [cited by examiner]
US 20170213546A1 · Gilbert · 2017 [cited by examiner]
US 20180090140A1 · Georges · 2018 [cited by examiner]
US 20180268813A1 · Georges · 2018 [cited by examiner]
US 20180370029A1 · Hall · 2018 [cited by examiner]
US 20180370032A1 · Ichikawa · 2018 [cited by examiner]
US 20190001489A1 · Hudson · 2019 [cited by examiner]
US 20190197109A1 · Peters · 2019 [cited by examiner]
US 20190197185A1 · Miseldine · 2019 [cited by examiner]
US 20190224849A1 · Tan et al. · 2019 [cited by applicant]
US 20190244603A1 · Angkititrakul · 2019 [cited by examiner]
US 20190295546A1 · Sugiyama et al. · 2019 [cited by applicant]
US 20190340485A1 · Ngo et al. · 2019 [cited by applicant]
US 20190384815A1 · Patel · 2019 [cited by examiner]
JP 2018513405A · 2018 [cited by applicant]
WO 2019142427A1 · 2019 [cited by applicant]
Li, X., Chen, Y. N., Li, L., Gao, J., & Celikyilmaz, A. (2017). End-to-end task-completion neural dialogue systems. arXiv preprint arXiv:1703.01008. (Year: 2017). [cited by examiner]
Broad, A., Arkin, J., Ratliff, N., Howard, T., & Argall, B. (2017). Real-time natural language corrections for assistive robotic manipulators. The International Journal of Robotics Research, 36(5-7), 684-698. (Year: 201… [cited by examiner]
Matuszek, C., Herbst, E., Zettlemoyer, L., & Fox, D. (2013). Learning to parse natural language commands to a robot control system. In Experimental robotics: the 13th international symposium on experimental robotics (pp… [cited by examiner]
EP Exended Search Report and Opinion for EP20898102.7 mailed Dec. 6, 2023. [cited by applicant]
Shalyminov et al., “Multi-Task Learning for Domain-General Spoken Disfluency Detection in Dialog Systems”, arxiv.org, Oct. 8, 2018; pp. 1-9. [cited by applicant]
Wu et al., “ScratchThat: Supporting Command-Agnostic Speech Repair in Voice-Driven Assistants”, Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, ACMPUB27, New York, NY, vol. 3, No. 2,… [cited by applicant]
Dong et al., “Adapting Translation Models for Transcript Disfluency Detection”, Proceedings of the Association for the Advancement of Artificial Intelligence Conference on Artificial Intelligence, vol. 33, No. 1, Jul. 1… [cited by applicant]
Application No. JP2022-535208 , Office Action, Mailed On Feb. 14, 2025, 13 pages. [cited by applicant]
Application No. EP20898102.7 , Office Action, Mailed On Nov. 11, 2025, 14 pages. [cited by applicant]
Application No. JP2022-535208 , Office Action, Mailed On Oct. 21, 2025, 6 pages. [cited by applicant]