IP Library Granted Patent US 11,222,632
Granted Patent B2
US 11,222,632 · App. 16/233,678 · Granted Jan 11, 2022

System and method for intelligent initiation of a man-machine dialogue based on multi-modal sensory inputs

Inventors: Changsong Liu (Los Angeles, CA); Rui Fang (Los Angeles, CA)
Assignee: DMAI, INC.
G10L15/22G06K9/00302G06K9/00664G10L15/24G10L15/1815G10L25/63G10L2015/225G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,222,632
App. No.
16/233,678
Granted
Jan 11, 2022
Kind
B2
Abstract

The present teaching relates to method, system, medium, and implementations for enabling communication with a user. Information representing surrounding of a user to be engaged in a new dialogue is received via the communication platform, wherein the information is acquired from a scene in which the user is present and captures characteristics of the user and the scene. Relevant features are extracted from the information. A state of the user is estimated based on the relevant features, and a dialogue context surrounding the scene is determined based on the relevant features. A topic for the new dialogue is determined based on the user, and a feedback is generated to initiate the new dialogue with the user based on the topic, the state of the user, and the dialogue context.

Claims (77)

1. A method, implemented on a machine having at least one processor, storage, and a communication platform for enabling communication between a machine dialogue agent and a user, comprising:

receiving, by the dialogue agent via the communication platform, information representing surrounding of a user to be engaged in a new dialogue, wherein the information is acquired from a scene in which the user is present and captures characteristics of the user and the scene;

extracting relevant features from the information;

estimating a state of the user based on the relevant features;

determining a dialogue context surrounding the scene based on the relevant features;

determining a topic for the new dialogue that is considered appropriate based on the characteristics of the user and the scene;

replacing a predetermined starting sentence, of the new dialogue, indicated by an initiation node in a dialogue tree with a new starting sentence determined based on the state of user and the dialogue context, wherein the dialogue tree is for governing the new dialogue with the user on the topic; and

initiating the new dialogue between the dialogue agent and the user on the topic by starting the new dialogue with the new starting sentence.

2. The method of claim 1 , wherein the information includes multimodal sensor data in at least one of audio, visual, textual, and haptic modalities, where the audio sensor data record acoustic sound from the scene and/or a speech from the user.

3. The method of claim 1 , wherein the state of the user characterizes at least one of:

an appearance of the user observed in the scene;

an expression of the user estimated based on the information acquired from the scene;

one or more emotions of the user inferred based on the expression of the user; and

an intent of the user inferred based on at least one of the expression and the one or more emotions.

4. The method of claim 1 , wherein the dialogue context includes at least one of:

at least one object present in the scene and a characterization thereof;

an estimated classification of the scene;

a characterization of the scene; and

a sound heard from the environment of the scene.

5. The method of claim 1 , wherein the topic is one of

a subject matter related to a program that the user previously signed up;

a subject matter dynamically determined based on the state of the user and/or the dialogue context related to the scene; and

a combination thereof.

6. The method of claim 1 , wherein the feedback is to be conveyed to the user to initiate the new dialogue via at least one of speech and visual means.

7. The method of claim 1 , further comprising generating one or more instructions to be used for rendering the feedback to the user.

8. The system of claim 1 , wherein the dialogue context includes at least one of:

at least one object present in the scene and a characterization thereof;

an estimated classification of the scene;

a characterization of the scene; and

a sound heard from the environment of the scene.

9. The system of claim 1 , wherein the topic is one of

a subject matter related to a program that the user previously signed up;

a subject matter dynamically determined based on the state of the user and/or the dialogue context related to the scene; and

a combination thereof.

10. Machine readable and non-transitory medium coded with information for enabling communication between a dialogue agent and a user, wherein the information, once read by the machine, causes the machine to perform:

receiving, via the communication platform, information representing surrounding of a user to be engaged in a new dialogue, wherein the information is acquired from a scene in which the user is present and captures characteristics of the user and the scene;

extracting relevant features from the information;

estimating a state of the user based on the relevant features;

determining a dialogue context surrounding the scene based on the relevant features;

determining a topic for the new dialogue that is considered appropriate based on the characteristics of the user and the scene;

replacing a predetermined starting sentence, of the new dialogue, indicated by an initiation node in a dialogue tree with a new starting sentence determined based on the state of user and the dialogue context, wherein the dialogue tree is for governing the new dialogue with the user on the topic; and

initiating the new dialogue between the dialogue agent and the user on the topic by starting the new dialogue with the new starting sentence.

11. The medium of claim 10 , wherein the information includes multimodal sensor data in at least one of audio, visual, textual, and haptic modalities, where the audio sensor data record acoustic sound from the scene and/or a speech from the user.

12. The medium of claim 10 , wherein the state of the user characterizes at least one of:

an appearance of the user observed in the scene;

an expression of the user estimated based on the information acquired from the scene;

one or more emotions of the user inferred based on the expression of the user; and

an intent of the user inferred based on at least one of the expression and the one or more emotions.

13. The medium of claim 10 , wherein the dialogue context includes at least one of:

at least one object present in the scene and a characterization thereof;

an estimated classification of the scene;

a characterization of the scene; and

a sound heard from the environment of the scene.

14. The medium of claim 10 , wherein the topic is one of

a subject matter related to a program that the user previously signed up;

a subject matter dynamically determined based on the state of the user and/or the dialogue context related to the scene; and

a combination thereof.

15. The medium of claim 10 , wherein the feedback is to be conveyed to the user to initiate the new dialogue via at least one of speech and visual means.

16. The medium of claim 10 , wherein the information, when read by the machine, further causes the machine to perform generating one or more instructions to be used for rendering the feedback to the user.

17. A system for enabling communication between a dialogue agent and a user, comprising:

a multimodal data analysis unit implemented on a processor and configured for

receiving information representing surrounding of a user to be engaged in a new dialogue, wherein the information is acquired from a scene in which the user is present and captures characteristics of the user and the scene, and

extracting relevant features from the information;

a user state estimator implemented on a processor and configured for estimating a state of the user based on the relevant features;

a dialogue contextual info determiner implemented on a processor and configured for determining a dialogue context surrounding the scene based on the relevant features; and

a dialogue controller implemented on a processor and configured for

determining a topic for the new dialogue that is considered appropriate based on the characteristics of the user and the scene,

replacing a predetermined starting sentence, of the new dialogue, indicated by an initiation node in a dialogue with a new starting sentence determined based on the state of user and the dialogue context, wherein the dialogue tree is for governing the new dialogue with the user on the topic, and

initiating the new dialogue between the dialogue agent and the user on the topic by starting the new dialogue with the new starting sentence.

18. The system of claim 17 , wherein the information includes multimodal sensor data in at least one of audio, visual, textual, and haptic modalities, where the audio sensor data record acoustic sound from the scene and/or a speech from the user.

19. The system of claim 17 , wherein the state of the user characterizes at least one of:

an appearance of the user observed in the scene;

an expression of the user estimated based on the information acquired from the scene;

one or more emotions of the user inferred based on the expression of the user; and

an intent of the user inferred based on at least one of the expression and the one or more emotions.

20. The system of claim 17 , wherein the feedback is to be conveyed to the user to initiate the new dialogue via at least one of speech and visual means.

21. The system of claim 17 , further comprising a feedback instruction generator implemented on a processor and configured for generating one or more instructions to be used for rendering the feedback to the user.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2023
From: DMAL, INC.
To: DMAI (GUANGZHOU) CO.,LTD.
Reel/Frame 065489/0038 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 27, 2018
From: LIU, CHANGSONG; FANG, RUI
To: DMAI, INC.
Reel/Frame 047860/0773 →
Continuity (2)
Provisional Application 62612163 · Dec 29, 2017
Related Publication 20190206401A1 · Jul 4, 2019
Cited By (1)
US 12,602,901