IP Library › Granted Patent US 10,831,442
Granted Patent B2
US 10,831,442 · App. 16/165,777 · Granted Nov 10, 2020

Digital assistant user interface amalgamation

Inventors: Jeremy R. Fox (Georgetown, TX); Gregory J. Boss (Saginaw, MI); Kelley Anders (East New Market, MD); Sarbajit K. Rakshit (Kolkata, IN)
Assignee: International Business Machines Corporation
G06F3/167G06F3/017G06K9/00302G06K9/00335G06N20/00G10L15/063G10L15/065G10L15/24G10L15/25G10L15/28H04R1/04G10L2015/223G10L2015/225G10L2015/226
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,831,442
App. No.
16/165,777
Granted
Nov 10, 2020
Kind
B2
Abstract

An approach is provided that receives, from a user, an amalgamation at a digital assistant. The amalgamation includes one or more words spoken by the user that are captured by a digital microphone and a set of digital images corresponding to one or more gestures that are performed by the user with the digital images captured by a digital camera. The system then determines an action that is responsive to the amalgamation and then performs the determined action.

Claims (67)

1. A method implemented by an information handling system that includes a processor and a memory accessible by the processor, the method comprising:

receiving, from a user, an amalgamation at a digital assistant, wherein the amalgamation includes a first set of words spoken by the user and captured by a microphone and a set of digital images corresponding to one or more gestures performed by the user captured by a digital camera;

determining an action responsive to the amalgamation;

performing the action by the digital assistant;

determining that the action is incorrect based on a facial expression received from the user and responsively:

receiving, from the user, a set of further amalgamations at the digital assistant;

indicating a set of responsive actions corresponding to the set of further amalgamations;

receiving a set of user feedback to the set of responsive actions; and

selecting one amalgamation from the set of further amalgamations based on the set of user feedback; and

storing the selected amalgamation and a corresponding responsive action from the set of responsive actions in a data store.

2. The method of claim 1 further comprising:

training a machine learning system, wherein the training includes the determined action and a corresponding amalgamation.

3. The method of claim 1 wherein the set of user feedback comprises a subsequent facial expression of the user captured by the digital camera.

4. The method of claim 1 further comprising:

receiving, from the user, a second amalgamation at the digital assistant, wherein the second amalgamation includes a second set of words spoken by the user and captured by the microphone and a second set of digital images corresponding to a second set of one or more gestures performed by the user captured by a digital camera;

identifying that the second amalgamation matches the stored selected amalgamation and responsively receiving the stored corresponding responsive action from the data store; and

performing, by the digital assistant, the stored corresponding responsive action.

5. The method of claim 1 wherein the selected amalgamation and the stored corresponding responsive action are stored in a question-answering (QA) system.

6. The method of claim 1 further comprising:

determining a meaning of at least one of the one or more gestures;

determining a set of ingested words related to the first set of words, wherein the set of ingested words correspond to the determined meaning of the at least one of the one or more gestures; and

identify the determined action based on the set of ingested words and the meaning of the at least one of the one or more gestures.

7. An information handling system comprising:

one or more processors;

a memory coupled to at least one of the one or more processors;

a digital microphone accessible by at least one of the one or more processors;

a digital camera accessible by at least one of the one or more processors; and

a set of computer program instructions stored in the memory and executed by at least one of the one or more processors in order to perform actions comprising:

receiving, from a user, an amalgamation at a digital assistant, wherein the amalgamation includes a first set of words spoken by the user and captured by the digital microphone and a set of digital images corresponding to one or more gestures performed by the user captured by the digital camera;

determining an action responsive to the amalgamation;

performing the action by the digital assistant;

determining that the action is incorrect based on a facial expression received from the user and responsively:

receiving, from the user, a set of further amalgamations at the digital assistant;

indicating a set of responsive actions corresponding to the set of further amalgamations;

receiving a set of user feedback to the set of responsive actions; and

selecting one amalgamation from the set of further amalgamations based on the set of user feedback; and

storing the selected amalgamation and a corresponding responsive action from the set of responsive actions in a data store.

8. The information handling system of claim 7 wherein the actions further comprise:

training a machine learning system, wherein the training includes the determined action and a corresponding amalgamation.

9. The information handling system of claim 7 wherein the set of user feedback comprises a subsequent facial expression of the user captured by the digital camera.

10. The information handling system of claim 7 wherein the actions further comprise:

receiving, from the user, a second amalgamation at the digital assistant, wherein the second amalgamation includes a second set of words spoken by the user and captured by the microphone and a second set of digital images corresponding to a second set of one or more gestures performed by the user captured by a digital camera;

identifying that the second amalgamation matches the stored selected amalgamation and responsively receiving the stored corresponding responsive action from the data store; and

performing, by the digital assistant, the stored corresponding responsive action.

11. The information handling system of claim 7 wherein the selected amalgamation and the stored corresponding responsive action are stored in a question-answering (QA) system.

12. The information handling system of claim 7 wherein the actions further comprise:

determining a meaning of at least one of the one or more gestures;

determining a set of ingested words related to the first set of words, wherein the set of ingested words correspond to the determined meaning of the at least one of the one or more gestures; and

identify the determined action based on the set of ingested words and the meaning of the at least one of the one or more gestures.

13. A computer program product stored in a computer readable storage medium, comprising computer program code that, when executed by an information handling system, performs actions comprising:

receiving, from a user, an amalgamation at a digital assistant, wherein the amalgamation includes a first set of words spoken by the user and captured by a microphone and a set of digital images corresponding to one or more gestures performed by the user captured by a digital camera;

determining an action responsive to the amalgamation;

performing the action by the digital assistant;

determining that the action is incorrect based on a facial expression received from the user and responsively:

receiving, from the user, a set of further amalgamations at the digital assistant;

indicating a set of responsive actions corresponding to the set of further amalgamations;

receiving a set of user feedback to the set of responsive actions; and

selecting one amalgamation from the set of further amalgamations based on the set of user feedback; and

storing the selected amalgamation and a corresponding responsive action from the set of responsive actions in a data store.

14. The computer program product of claim 13 wherein the actions further comprise:

training a machine learning system, wherein the training includes the determined action and a corresponding amalgamation.

15. The computer program product of claim 13 wherein the set of user feedback comprises a subsequent facial expression of the user captured by the digital camera.

16. The computer program product of claim 13 wherein the actions further comprise:

receiving, from the user, a second amalgamation at the digital assistant, wherein the second amalgamation includes a second set of words spoken by the user and captured by the microphone and a second set of digital images corresponding to a second set of one or more gestures performed by the user captured by a digital camera;

identifying that the second amalgamation matches the stored selected amalgamation and responsively receiving the stored corresponding responsive action from the data store; and

performing, by the digital assistant, the stored corresponding responsive action.

17. The computer program product of claim 13 wherein the selected amalgamation and the stored corresponding responsive action are stored in a question-answering (QA) system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 19, 2018
From: FOX, JEREMY R.; BOSS, GREGORY J.; ANDERS, KELLEY; RAKSHIT, SARBAJIT K.
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 047238/0035 →
Continuity (1)
Related Publication 20200125321A1 · Apr 23, 2020