IP Library › Granted Patent US 12,282,606
Granted Patent B2
US 12,282,606 · App. 17/107,958 · Granted Apr 22, 2025

VPA with integrated object recognition and facial expression recognition

Inventors: Ajay Divakaran (Monmouth Junction, NJ); Amir Tamrakar (Philadelphia, PA); Girish Acharya (Redwood City, CA); William Mark (San Mateo, CA); Greg Ho (South Brunswick, NJ); Jihua Huang (Philadelphia, PA); David Salter (Bensalem, PA); Edgar Kalns (San Jose, CA); Michael Wessel (Palo Alto, CA); Min Yin (San Jose, CA); James Carpenter (Mountain View, CA); Brent Mombourquette (Menlo Park, CA); Kenneth Nitz (Redwood City, CA); Elizabeth Shriberg (Berkeley, CA); Eric Law (Hayward, CA); Michael Frandsen (Helena, MT); Hyong-Gyun Kim (Santa Clara, CA); Cory Albright (Helena, MT); Andreas Tsiartas (Santa Clara, CA)
Assignee: SRI International
G06F3/017G06F3/0304G06F3/167G06N3/006G06N5/022G06N20/00G06N20/10G06V40/16G06V40/20G10L15/1815G10L15/22G10L25/63G06N7/01G10L15/1822G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,282,606
App. No.
17/107,958
Granted
Apr 22, 2025
Kind
B2
Abstract

Methods, computing devices, and computer-program products are provided for implementing a virtual personal assistant. In various implementations, a virtual personal assistant can be configured to receive sensory input, including at least two different types of information. The virtual personal assistant can further be configured to determine semantic information from the sensory input, and to identify a context-specific framework. The virtual personal assistant can further be configured to determine a current intent. Determining the current intent can include using the semantic information and the context-specific framework. The virtual personal assistant can further be configured to determine a current input state. Determining the current input state can include using the semantic information and one or more behavioral models. The behavioral models can include one or more interpretations of previously-provided semantic information. The virtual personal assistant can further be configured to determine an action using the current intent and the current input state.

Claims (47)

1. A method comprising:

receiving, by a computing device, sensory input comprising at least audio input and visual input;

determining semantic information from the sensory input;

determining a first questioning style to question a user using the semantic information;

determining scene information from the visual input;

determining, after waiting for audio input from the user, an input state associated with the user using the audio input and the scene information;

changing the first questioning style to a second questioning style using the determined input state, the second questioning style being different from the first questioning style; and

outputting a question to the user using the second questioning style.

2. The method of claim 1 , further comprising extracting the scene information from at least one image input or at least one video input.

3. The method of claim 1 , further comprising determining object information from the scene information and using the object information to determine the input state.

4. The method of claim 1 , further comprising determining event information from the scene information and using the event information to determine the input state.

5. The method of claim 1 , further comprising determining motion information from the scene information and using the motion information to determine the input state.

6. The method of claim 1 , further comprising determining the scene information using at least one machine learned model trained using domain-specific training samples including at least one image sample or at least one video sample.

7. The method of claim 1 , further comprising, using the scene information, determining a temporal sequence and using the temporal sequence to determine the input state, wherein the temporal sequence comprises (i) an object, (ii) an event, or (iii) a combination of (i) and (ii).

8. The method of claim 1 , further comprising extracting an image of an iris of a human eye from the visual input, using the image to generate a code, comparing the code to a template, determining biometric information based on a comparison of the code to the template, and using the biometric information to determine the input state.

9. The method of claim 1 , wherein the input state is at least one of an emotional state, cognitive state, or mental state of the user.

10. The method of claim 1 , wherein the sensory input is generated by the user.

11. A device, comprising: at least one processor;

at least one sensor coupled to the at least one processor; and memory coupled to the at least one processor;

the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations comprising:

receiving, by the at least one sensor, audio input;

receiving, by the at least one sensor, visual input;

determining semantic information from the audio input;

determining a first questioning style to question a user using the semantic information;

determining scene information from the visual input;

determining, after waiting for audio input from the user, an input state associated with the user using the audio input and the scene information;

changing the first questioning style to a second questioning style using the determined input state, the second questioning style being different from the first questioning style; and

outputting a question to the user using the second questioning style.

12. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising extracting the scene information from at least one image input or at least one video input.

13. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising determining object information from the scene information and using the object information to determine the input state.

14. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising determining event information from the scene information and using the event information to determine the input state.

15. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising determining motion information from the scene information and using the motion information to determine the input state.

16. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising determining the input state using at least one machine-learned model trained using domain-specific training samples including at least one image sample or at least one video sample.

17. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising, using the scene information, determining a temporal sequence and using the temporal sequence to determine the input state, wherein the temporal sequence comprises (i) at least one object, (ii) at least one event, or (iii) a combination of (i) and (ii).

18. The device of claim 11 , the memory cooperating with the at least one processor to cause the at least one processor to be capable of performing operations further comprising extracting an image of an iris of a human eye from the visual input, using the image to generate a code, comparing the code to a template, determining biometric information based on a comparison of the code to the template, and using the biometric information to determine the input state.

19. At least one memory device storing instructions which, when executed by at least one processor, cause the at least one processor to be capable of performing operations comprising:

receiving an audio input;

receiving a visual input;

determining semantic information from the audio input;

determining a first questioning style to question a user using the semantic information;

determining scene information from the visual input;

determining, after waiting for audio input from the user, an input state associated with the user using the audio input and the scene information;

changing the first questioning style to a second questioning style using the determined input state, the second questioning style being different from the first questioning style; and

outputting a question to the user using the second questioning style.

20. The at least one memory device of claim 19 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to be capable of performing operations further comprising extracting the scene information from at least one image input or at least one video input and determining the input state using at least one machine-learned model trained using domain-specific training samples including at least one image sample or at least one video sample.

21. The at least one memory device of claim 19 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to be capable of performing operations further comprising determining object information from the scene information and using the object information to determine the input state.

22. The at least one memory device of claim 19 , wherein the instructions, when executed by the at least one processor, cause the at least one processor to be capable of performing operations further comprising, using the scene information, determining a temporal sequence and using the temporal sequence to determine the input state, wherein the temporal sequence comprises (i) at least one object, (ii) at least one event, or (iii) a combination of (i) and (ii).

Continuity (5)
Continuation 15332494 · Oct 24, 2016
Provisional Application 62264228 · Dec 7, 2015
Provisional Application 62329055 · Apr 28, 2016
Provisional Application 62339547 · May 20, 2016
Related Publication 20210081056A1 · Mar 18, 2021
References Cited (35)
US 5357596A · Takebayashi · 1994 [cited by examiner]
US 8407055B2 · Asano · 2013 [cited by examiner]
US 8781802B2 · Hagelin · 2014 [cited by examiner]
US 8914290B2 · Hendrickson · 2014 [cited by examiner]
US 9613025B2 · Heo · 2017 [cited by examiner]
US 9898850B2 · Furukawa · 2018 [cited by examiner]
US 9916628B1 · Wang · 2018 [cited by examiner]
US 10884503B2 · Divakaran · 2021 [cited by examiner]
US 20030167167A1 · Gong · 2003 [cited by applicant]
US 20060120570A1 · Azuma · 2006 [cited by examiner]
US 20060192868A1 · Wakamori · 2006 [cited by examiner]
US 20090231441A1 · Walker · 2009 [cited by examiner]
US 20100014718A1 · Savvides · 2010 [cited by examiner]
US 20110087483A1 · Hsieh · 2011 [cited by applicant]
US 20110112826A1 · Wang · 2011 [cited by examiner]
US 20130110520A1 · Cheyer et al. · 2013 [cited by applicant]
US 20130185066A1 · Tzirkel-Hancock · 2013 [cited by examiner]
US 20130185078A1 · Tzirkel-Hancock · 2013 [cited by applicant]
US 20140195221A1 · Frank · 2014 [cited by examiner]
US 20140272847A1 · Grimes · 2014 [cited by examiner]
US 20150007307A1 · Grimes · 2015 [cited by examiner]
US 20150340031A1 · Kim · 2015 [cited by examiner]
US 20160163332A1 · Un · 2016 [cited by examiner]
US 20170125008A1 · Maisonnier · 2017 [cited by examiner]
US 20180033432A1 · Ikeno · 2018 [cited by examiner]
US 20180075142A1 · Fridental · 2018 [cited by examiner]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Office Action, Feb. 27, 2020. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Office Action, Jan. 17, 2019. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Advisory Action, Feb. 12, 2020. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Notice of Allowance, Sep. 4, 2020. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Interview Summary, Jun. 9, 2020. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, First Office Action Interview, Mar. 26, 2019. [cited by applicant]
Divakaran, U.S. Appl. No. 15/332,494, filed Oct. 24, 2016, Final Office Action, Sep. 9, 2019. [cited by applicant]
The International Bureau of WIPO, “Preliminary Report on Patentability”, in application No. PCT/US2016/065400, dated Jun. 21, 2018, 13 pages. [cited by applicant]
Current Claims in application No. PCT/US2016/065400, dated Jun. 2018, 8 pages. [cited by applicant]