IP Library Granted Patent US 12672068
Granted Patent B2
US 12672068 · App. 18/535,352 · Granted Jun 30, 2026

Providing safety and environmental features using human presence detection

Inventors: Jan Neerbek (Beder, DK); Rafal Krzysztof Malewski (Aalborg, DK); Brian Thoft Moth Møller (Hojberg, DK); Paul Nangeroni (San Francisco, CA); Amalavoyal Narasimha Chari (Palo Alto, CA)
Assignee: ROKU, INC.
H04W52/0254G06F18/251G06V40/20G08B13/24G08B29/188H04N5/57H04Q9/00G16Y10/65G16Y20/40G16Y40/50H04Q2213/002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12672068
App. No.
18/535,352
Granted
Jun 30, 2026
Kind
B2
Abstract

Disclosed herein are system, method, and computer program product embodiments for the detection of human presence in front of a plurality of sensors such as those of speakers and a device with a processor, such as a television. Data gathered from the plurality of sensors may be analyzed by the processor to determine if one or more humans are present proximate to the device. Based on the determined presence or absence of one or more humans, further actions including, inter alia, activating a sleep mode for the one or more humans, shutting off the device in a green mode, or alerting an owner-user to the presence of an intruder can be taken.

Claims (71)

1 . A computer-implemented method comprising:

receiving, by at least one computer processor, a plurality of received signal strength indications (RSSIs) respectively associated with a plurality of Wi-Fi signals, wherein each Wi-Fi signal of the plurality of Wi-Fi signals is transmitted by a respective module in a plurality of modules that are arranged to form a detection zone around a user space, wherein a portion of the plurality of modules each comprise at least one speaker of a home entertainment system, and the user space comprises a smart television of the home entertainment system;

receiving ambient sound recorded by a microphone incorporated into a module of the plurality of modules;

forming a Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound;

recognizing an intruder-indicative sound in the ambient sound;

providing at least the Wi-Fi signature as input to a machine learning model that is configured to determine whether an intruder is present based on at least the Wi-Fi signature and the recognized intruder-indicative sound;

based on determining that a presence of a user is not detected in the detection zone, and in response to receiving a determination from the machine learning model that the intruder is present and determining that the intruder is within the detection zone, sending a message to a user device of the user; and

based on receiving a user response to the message;

saving metadata based on a type of the user response to an in-memory database of the smart television;

playing a sound on at least one speaker of the portion of the plurality of modules via the smart television; and

transmitting the metadata over a network to a cloud computing environment configured to train the machine learning model based on the metadata and on metadata from a plurality of other smart televisions corresponding to a plurality of other users.

2 . The computer-implemented method of claim 1 , wherein the machine learning model comprises a neural network that includes an input layer having nodes corresponding to the Wi-Fi signature and an output layer including a first node representing that the intruder is present and a second node representing that the intruder is not present, wherein the neural network determines that the intruder is present in response to determining that a value of the first node is greater than a value of the second node.

3 . The computer-implemented method of claim 1 , wherein the machine learning model comprises at least one support vector machine that is configured to construct a hyper plane in multiple dimensions between a first class that corresponds to the intruder being present and a second class that corresponds to the intruder not being present.

4 . The computer-implemented method of claim 1 , further comprising:

receiving a plurality of received signal destinations respectively associated with the plurality of Wi-Fi signals;

wherein forming the Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound comprises:

forming the Wi-Fi signature based at least on the plurality of RSSIs, the plurality of received signal destinations, and the recorded ambient sound.

5 . The computer-implemented method of claim 1 , wherein providing at least the Wi-Fi signature as input to the machine learning model comprises:

providing multiple Wi-Fi signatures corresponding to different times as input to the machine learning model.

6 . The computer-implemented method of claim 2 , further comprising:

in response to receiving the determination that the intruder is present from the machine learning model:

determining that the intruder is not within the detection zone; and

in response to determining that the intruder is not within the detection zone, causing the smart television to turn on and play content.

7 . The computer-implemented method of claim 1 , wherein:

the machine learning model is executed by the smart television to determine whether the intruder is present.

8 . A system, comprising:

a memory; and

at least one processor coupled to the memory and configured to perform operations comprising:

receiving a plurality of received signal strength indications (RSSIs) respectively associated with a plurality of Wi-Fi signals, wherein each Wi-Fi signal of the plurality of Wi-Fi signals is transmitted by a respective module in a plurality of modules that are arranged to form a detection zone around a user space, wherein a portion of the plurality of modules each comprise at least one speaker of a home entertainment system, and the user space comprises a smart television of the home entertainment system;

receiving ambient sound recorded by a microphone incorporated into a module of the plurality of modules;

forming a Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound;

recognizing an intruder-indicative sound in the ambient sound;

providing at least the Wi-Fi signature as input to a machine learning model that is configured to determine whether an intruder is present based on at least the Wi-Fi signature and the recognized intruder-indicative sound;

based on determining that a presence of a user is not detected in the detection zone, and in response to receiving a determination from the machine learning model that the intruder is present and determining that the intruder is within the detection zone, sending a message to a user device of the user; and

based on receiving a user response to the message;

saving metadata based on a type of the user response to an in-memory database of the smart television;

playing a sound on at least one speaker of the portion of the plurality of modules via the smart television; and

transmitting the metadata over a network to a cloud computing environment configured to train the machine learning model based on the metadata and on metadata from a plurality of other smart televisions corresponding to a plurality of other users.

9 . The system of claim 8 , wherein the machine learning model comprises a neural network that includes an input layer having nodes corresponding to the Wi-Fi signature and an output layer including a first node representing that the intruder is present and a second node representing that the intruder is not present, wherein the neural network determines that the intruder is present in response to determining that a value of the first node is greater than a value of the second node.

10 . The system of claim 9 , wherein the machine learning model comprises at least one support vector machine that is configured to construct a hyper plane in multiple dimensions between a first class that corresponds to the intruder being present and a second class that corresponds to the intruder not being present.

11 . The system of claim 8 , wherein the operations further comprise:

receiving a plurality of received signal destinations respectively associated with the plurality of Wi-Fi signals;

wherein forming the Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound comprises:

forming the Wi-Fi signature based at least on the plurality of RSSIs, the plurality of received signal destinations, and the recorded ambient sound.

12 . The system of claim 8 , wherein providing at least the Wi-Fi signature as input to the machine learning model comprises:

providing multiple Wi-Fi signatures corresponding to different times as input to the machine learning model.

13 . The system of claim 9 , wherein the operations further comprise:

in response to receiving the determination that the intruder is present from the machine learning model:

determining that the intruder is not within the detection zone; and

in response to determining that the intruder is not within the detection zone, causing, by the at least one processor, the smart television to turn on and play content.

14 . A non-transitory computer-readable medium having instructions stored thereon that, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

receiving a plurality of received signal strength indications (RSSIs) respectively associated with a plurality of Wi-Fi signals, wherein each Wi-Fi signal of the plurality of Wi-Fi signals is transmitted by a respective module in a plurality of modules that are arranged to form a detection zone around a user space, wherein a portion of the plurality of modules each comprise at least one speaker of a home entertainment system, and the user space comprises a smart television of the home entertainment system;

receiving ambient sound recorded by a microphone incorporated into a module of the plurality of modules;

forming a Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound;

recognizing an intruder-indicative sound in the ambient sound;

providing at least the Wi-Fi signature as input to a machine learning model that is configured to determine whether an intruder is present based on at least the Wi-Fi signature and the recognized intruder-indicative sound;

based on determining that a presence of a user is not detected in the detection zone, and in response to receiving a determination from the machine learning model that the intruder is present and determining that the intruder is within the detection zone, sending a message to a user device of the user; and

based on receiving a user response to the message:

saving metadata based on a type of the user response to an in-memory database of the smart television;

playing a sound on at least one speaker of the portion of the plurality of modules via the smart television; and

transmitting the metadata over a network to a cloud computing environment configured to train the machine learning model based on the metadata and on metadata from a plurality of other smart televisions corresponding to a plurality of other users.

15 . The non-transitory computer-readable medium of claim 14 , wherein the machine learning model comprises a neural network that includes an input layer having nodes corresponding to the Wi-Fi signature and an output layer including a first node representing that the intruder is present and a second node representing that the intruder is not present, wherein the neural network determines that the intruder is present in response to determining that a value of the first node is greater than a value of the second node.

16 . The non-transitory computer-readable medium of claim 14 , wherein the machine learning model comprises at least one support vector machine that is configured to construct a hyper plane in multiple dimensions between a first class that corresponds to the intruder being present and a second class that corresponds to the intruder not being present.

17 . The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise:

receiving a plurality of received signal destinations respectively associated with the plurality of Wi-Fi signals;

wherein forming the Wi-Fi signature based at least on the plurality of RSSIs and the recorded ambient sound comprises:

forming the Wi-Fi signature based at least on the plurality of RSSIs, the plurality of received signal destinations, and the recorded ambient sound.

18 . The non-transitory computer-readable medium of claim 14 , wherein the operations further comprise:

in response to receiving the determination that the intruder is present from the machine learning model:

determining that the intruder is not within the detection zone; and

in response to determining that the intruder is not within the detection zone, causing the smart television to turn on and play content.