IP Library › Granted Patent US 10,884,503
Granted Patent B2
US 10,884,503 · App. 15/332,494 · Granted Jan 5, 2021

VPA with integrated object recognition and facial expression recognition

Inventors: Ajay Divakaran (Monmouth Junction, NJ); Amir Tamrakar (Philadelphia, PA); Girish Acharya (Redwood City, CA); William Mark (San Mateo, CA); Greg Ho (South Brunswick, NJ); Jihua Huang (Philadelphia, PA); David Salter (Bensalem, PA); Edgar Kalns (San Jose, CA); Michael Wessel (Palo Alto, CA); Min Yin (San Jose, CA); James Carpenter (Mountain View, CA); Brent Mombourquette (Menlo Park, CA); Kenneth Nitz (Redwood City, CA); Elizabeth Shriberg (Berkeley, CA); Eric Law (Hayward, CA); Michael Frandsen (Helena, MT); Hyong-Gyun Kim (Santa Clara, CA); Cory Albright (Helena, MT); Andreas Tsiartas (Santa Clara, CA)
Assignee: SRI International
G06F3/017G06F3/0304G06F3/167G06K9/00221G06K9/00335G06N3/006G06N5/022G06N20/00G10L15/1815G10L15/22G10L25/63G06N7/005G10L15/1822G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,884,503
App. No.
15/332,494
Filed
Oct 24, 2016
Granted
Jan 5, 2021
Kind
B2
Art Unit
2692
USPC
345/156
Abstract

Methods, computing devices, and computer-program products are provided for implementing a virtual personal assistant. In various implementations, a virtual personal assistant can be configured to receive sensory input, including at least two different types of information. The virtual personal assistant can further be configured to determine semantic information from the sensory input, and to identify a context-specific framework. The virtual personal assistant can further be configured to determine a current intent. Determining the current intent can include using the semantic information and the context-specific framework. The virtual personal assistant can further be configured to determine a current input state. Determining the current input state can include using the semantic information and one or more behavioral models. The behavioral models can include one or more interpretations of previously-provided semantic information. The virtual personal assistant can further be configured to determine an action using the current intent and the current input state.

Claims (47)

1. A method, comprising:

receiving, by a computing device, sensory input, wherein the sensory input includes at least two different types of information;

determining semantic information from the sensory input, wherein the semantic information provides an interpretation of the sensory input;

determining an action, wherein determining the action includes using the semantic information;

identifying a context-specific framework, wherein the context-specific framework includes a cumulative sequence of one or more previous intents;

determining a current intent, wherein determining the current intent includes using the semantic information and the context-specific framework;

determining a current input state, wherein determining the current input state includes using the semantic information and one or more behavioral models, and wherein the behavioral models include one or more interpretations of previously-provided semantic information;

using one or more of the current intent and the current input state to modify the action prior to outputting the action to a user of the computing device, wherein to modify the action comprises changing a question type from a first type designed to elicit a first type of answer to a second type designed to elicit a second type of answer.

2. The method of claim 1 , wherein types of information include speech information, graphical information, audio input, image input, or tactile input.

3. The method of claim 1 , wherein an emotional state is capable of being derived from audio input.

4. The method of claim 1 , wherein physical gestures are capable of being derived from image input.

5. The method of claim 1 , wherein an emotional state is capable of being derived from image input.

6. The method of claim 1 , wherein behavioral models include a configurable preference model, and wherein the configurable preference model includes one or more previous input states and semantic information associated with the one or more previous input states.

7. The method of claim 1 , wherein the context-specific framework includes a dynamic ontology, wherein the dynamic ontology includes a cumulative sequence of one or more previous intents derived from a context-specific ontology.

8. The method of claim 1 , wherein to modify the action comprises increasing a volume of system output or decreasing the volume of system output or increasing a rate of system output or decreasing the rate of system output.

9. A virtual personal assistant device, comprising:

one or more processors; and

a non-transitory computer-readable medium including instructions that, when executed by the one or more processors, cause the one or more processors to perform operations including:

receiving, by a computing device, sensory input, wherein the sensory input includes at least two different types of information;

determining semantic information from the sensory input, wherein the semantic information provides an interpretation of the sensory input;

determining an action, wherein determining the action includes using the semantic information;

identifying a context-specific framework, wherein the context-specific framework includes a cumulative sequence of one or more previous intents;

determining a current intent, wherein determining the current intent includes using the semantic information and the context-specific framework;

determining a current input state, wherein determining the current input state includes using the semantic information and one or more behavioral models, and wherein the behavioral models include one or more interpretations of previously-provided semantic information; and

using one or more of the current intent and the current input state to modify the action prior to outputting the action to a user of the computing device, wherein to modify the action comprises changing a question type from a first type designed to elicit a first type of answer to a second type designed to elicit a second type of answer.

10. The virtual personal assistant device of claim 9 , wherein types of information include speech information, graphical information, audio input, image input, or tactile input.

11. The virtual personal assistant device of claim 9 , wherein an emotional state is capable of being derived from audio input.

12. The virtual personal assistant device of claim 9 , wherein physical gestures are capable of being derived from image input.

13. The virtual personal assistant device of claim 9 , wherein an emotional state is capable of being derived from image input.

14. The virtual personal assistant device of claim 9 , wherein behavioral models include a configurable preference model, and wherein the configurable preference model includes one or more previous input states and semantic information associated with the one or more previous input states.

15. The virtual personal assistant device of claim 9 , wherein the context-specific framework includes a dynamic ontology, wherein the dynamic ontology includes a cumulative sequence of one or more previous intents derived from a context-specific ontology.

16. The virtual personal assistant device of claim 9 , wherein to modify the action comprises causing the action to change from an open-ended dialog approach to a yes/no questioning approach or from a yes/no questioning approach to an open-ended dialog approach, or increasing a volume of system output or decreasing the volume of system output, or increasing a rate of system output or decreasing the rate of system output.

17. A computer-program product tangibly embodied in a non-transitory machine-readable storage medium, including instructions that, when executed by one or more processors, cause the one or more processors to:

receive, by a computing device, sensory input, wherein the sensory input includes at least two different types of information;

determine semantic information from the sensory input, wherein the semantic information provides an interpretation of the sensory input;

determine an action, wherein determining the action includes using the semantic information;

determine an action, wherein determining the action includes using the semantic information;

identify a context-specific framework, wherein the context-specific framework includes a cumulative sequence of one or more previous intents;

determine a current intent, wherein determining the current intent includes using the semantic information and the context-specific framework;

determine a current input state, wherein determining the current input state includes using the semantic information and one or more behavioral models, and wherein the behavioral models include one or more interpretations of previously-provided semantic information; and

use one or more of the current intent and the current input state to modify the action prior to outputting the action to a user of the computing device, wherein to modify the action comprises changing the action from an open-ended dialog approach to a yes/no questioning approach or from a yes/no questioning approach to an open-ended dialog approach or from a first question not specifically designed to elicit easy to remember information to a second question designed to elicit easy to remember information.

18. The computer-program product of claim 17 , wherein types of information include speech information, graphical information, audio input, image input, or tactile input.

19. The computer-program product of claim 17 , wherein an emotional state is capable of being derived from audio input.

20. The computer-program product of claim 17 , wherein physical gestures are capable of being derived from image input.

21. The computer-program product of claim 17 , wherein an emotional state is capable of being derived from image input.

22. The computer-program product of claim 17 , wherein behavioral models include a configurable preference model, and wherein the configurable preference model includes one or more previous input states and semantic information associated with the one or more previous input states.

23. The computer-program product of claim 17 , wherein to modify the action comprises causing the action to change from an open-ended dialog approach to a yes/no questioning approach or from a yes/no questioning approach to an open-ended dialog approach, or increasing a volume of system output or decreasing the volume of system output, or increasing a rate of system output or decreasing the rate of system output.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2016
From: DIVAKARAN, AJAY; TAMRAKAR, AMIR; ACHARYA, GIRISH; MARK, WILLIAM; HO, GREG; HUANG, JIHUA; SALTER, DAVID; KALNS, EDGAR; WESSEL, MICHAEL; YIN, MIN; CARPENTER, JAMES; MOMBOURQUETTE, BRENT; NITZ, KENNETH; SHRIBERG, ELIZABETH; LAW, ERIC; FRANDSEN, MICHAEL; KIM, HYONG-GYUN; ALBRIGHT, CORY; TSIARTAS, ANDREAS
To: SRI INTERNATIONAL
Reel/Frame 040263/0001 →
Continuity (4)
Provisional Application 62264228 · Dec 7, 2015
Provisional Application 62329055 · Apr 28, 2016
Provisional Application 62339547 · May 20, 2016
Related Publication 20170160813A1 · Jun 8, 2017
Cited By (4)
US 12,282,606 US 12,505,503 US 12,505,852 US 12,705,286