IP Library › Granted Patent US 11,561,621
Granted Patent B2
US 11,561,621 · App. 16/600,830 · Granted Jan 24, 2023

Multi media computing or entertainment system for responding to user presence and activity

Inventors: Feng Tang (Cupertino, CA); Chong Chen (Sunnyvale, CA); Haitao Guo (Cupertino, CA); Xiaojin Shi (Cupertino, CA); Thorsten Gernoth (San Francisco, CA)
Assignee: Apple Inc.
G06F3/017G06F3/011G06F3/012G06F3/0304G06F3/16G06F3/165G06V40/107G06V40/113G06K9/6269G06K9/6282G06V2201/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,561,621
App. No.
16/600,830
Granted
Jan 24, 2023
Kind
B2
Abstract

Intelligent systems are disclosed that respond to user intent and desires based upon activity that may or may not be expressly directed at the intelligent system. In some embodiments, the intelligent system acquires a depth image of a scene surrounding the system. A scene geometry may be extracted from the depth image and elements of the scene may be monitored. In certain embodiments, user activity in the scene is monitored and analyzed to infer user desires or intent with respect to the system. The interpretation of the user's intent as well as the system's response may be affected by the scene geometry surrounding the user and/or the system. In some embodiments, techniques and systems are disclosed for interpreting express user communication, e.g., expressed through hand gesture movements. In some embodiments, such gesture movements may be interpreted based on real-time depth information obtained from, e.g., optical or non-optical type depth sensors.

Claims (21)

1. A non-transitory program storage device, readable by a processor and comprising instructions stored thereon to cause one or more processors to: acquire a depth image of a scene in a vicinity of a first device; acquire an image of the scene; store the depth image and the image in a memory; develop a scene geometry based upon the depth image; monitor the activity of one or more humans present in the scene geometry, wherein one of the one or more humans comprises a user of the first device; detect a human face in the acquired image corresponding to the user of the first device; employ the depth image to detect a head orientation associated with the detected human face; correlate the detected human face with the head orientation; determine whether the user is engaged in conversation with at least one other of the one or more humans based, at least in part, upon the correlated detected human face and the head orientation; and in response to a determination that the user is engaged in conversation with the at least one other of the one or more humans, adjust an audio output of the first device based, at least in part, upon a characteristic of the determined conversation.

2. The non-transitory program storage device of claim 1 , wherein the characteristic of the determined conversation comprises an indication that the determined conversation is at least one of: a phone conversation, a conversation wherein the user is not facing the first device, a conversation wherein the user is facing the first device, or a conversation wherein the user is whispering.

3. The non-transitory program storage device of claim 1 , wherein the instructions stored thereon to adjust an output of the first device further cause the one or more processors to perform at least one of the following actions: turn down the volume of the first device, pause media being rendered by the first device, or cause the first device to enter a power-saving mode.

4. The non-transitory program storage device of claim 1 , wherein the instructions stored thereon further cause the one or more processors to:

identify the user of the first device; and

determine preferences associated with the user,

wherein the determination of whether the user is engaged in conversation with at least one other of the one or more humans is further based, at least in part, upon the determined preferences.

5. A method, comprising: acquiring a depth image of a scene in a vicinity of a first device; acquiring an image of the scene; storing the depth image and the image in a memory; developing a scene geometry based upon the depth image; monitoring the activity of one or more humans present in the scene geometry, wherein one of the one or more humans comprises a user of the first device; detecting a human face in the acquired image corresponding to the user of the first device; employing the depth image to detect a head orientation associated with the detected human face; correlating the detected human face with the head orientation; determining whether the user is engaged in conversation with at least one other of the one or more humans based, at least in part, upon the correlated detected human face and the head orientation; and in response to a determination that the user is engaged in conversation with the at least one other of the one or more humans, adjusting an audio output of the first device based, at least in part, upon a characteristic of the determined conversation.

6. The method of claim 5 , wherein the characteristic of the determined conversation comprises an indication that the determined conversation is at least one of: a phone conversation, a conversation wherein the user is not facing the first device, a conversation wherein the user is facing the first device, or a conversation wherein the user is whispering.

7. The method of claim 5 , wherein adjusting an output of the first device further comprises performing at least one of the following actions: turning down the volume of the first device, pausing media being rendered by the first device, or causing the first device to enter a power-saving mode.

8. The method of claim 5 , further comprising:

identifying the user of the first device; and

determining preferences associated with the user,

wherein the determination of whether the user is engaged in conversation with at least one other of the one or more humans is further based, at least in part, upon the determined preferences.

9. An electronic device comprising: a memory; a depth sensor; an image capture unit; one or more processors, communicatively coupled to the memory, wherein the memory stores instructions to cause the one or more processors to: acquire, from the depth sensor, a depth image of a scene in a vicinity of the electronic device; acquire, from the image capture unit, an image of the scene; store the depth image and the image in the memory; develop a scene geometry based upon the depth image; monitor the activity of one or more humans present in the scene geometry, wherein one of the one or more humans comprises a user of the electronic device; detect a human face in the acquired image corresponding to the user of the first device; employ the depth image to detect a head orientation associated with the detected human face; correlate the detected human face with the head orientation; determine whether the user is engaged in conversation with at least one other of the one or more humans based, at least in part, upon the correlated detected human face and the head orientation; and in response to a determination that the user is engaged in conversation with the at least one other of the one or more humans, adjust an audio output of the electronic device based, at least in part, upon a characteristic of the determined conversation.

10. The electronic device of claim 9 , wherein the characteristic of the determined conversation comprises an indication that the determined conversation is at least one of: a phone conversation, a conversation wherein the user is not facing the electronic device, a conversation wherein the user is facing the electronic device, or a conversation wherein the user is whispering.

11. The electronic device of claim 9 , wherein the instructions stored in the memory to adjust an output of the electronic device further cause the one or more processors to perform at least one of the following actions: turn down the volume of the electronic device, pause media being rendered by the electronic device, or cause the electronic device to enter a power-saving mode.

12. The electronic device of claim 9 , wherein the instructions stored in the memory further cause the one or more processors to:

identify the user of the first device; and;

determine preferences associated with the user,

wherein the determination of whether the user is engaged in conversation with at least one other of the one or more humans is further based, at least in part, upon the determined preferences.

Continuity (3)
Continuation 16055994 · Aug 6, 2018
Continuation 14865850 · Sep 25, 2015
Related Publication 20200042096A1 · Feb 6, 2020
Cited By (2)
US 12,424,222 US 12,443,272