IP Library Granted Patent US 11,295,139
Granted Patent B2
US 11,295,139 · App. 16/279,512 · Granted Apr 5, 2022

Human presence detection in edge devices

Inventors: Krishna Khadloya (San Jose, CA); Maxim Sokolov (Nizhny Novgorod, RU); Vaidhi Nathan (San Jose, CA); Chandan Gope (Cupertino, CA)
Assignee: INTELLIVISION TECHNOLOGIES CORP.
G06K9/00771G06K9/00369G06K9/00711G06T7/215G06T7/246G08B13/19608
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,295,139
App. No.
16/279,512
Granted
Apr 5, 2022
Kind
B2
Abstract

A system and method for detecting human presence in or absence from a field-of-view of a camera by analyzing camera data using a processor inside of or adjacent to the camera itself. In an example, the camera can be integrated with or embedded in another edge-based sensor device. In an example, a video signal processing system receives image data from one or more image sensors and uses a local processing circuit to process the image data and determine if a human being is or is not present during a particular time, interval, or sequence of frames. In an example, the human being identification technique can be used in security or surveillance applications such as for home, business, or other monitoring cameras.

Claims (55)

1. A method for performing a multiple-factor human recognition and validation routine that includes determining whether a human being is present in or absent from an environment monitored by a camera using image information from the camera, the method comprising:

receiving, at a first processor circuit, multiple frames from the camera, the multiple frames corresponding to a sequence of substantially adjacent time instances;

using the first processor circuit, determining whether a frame difference between a portion of at least first and second frames from among the multiple frames indicates movement by an object in the environment monitored by the camera; and

in response to the frame difference indicating movement by the object, continuing the human recognition and validation routine by:

using the first processor circuit, selecting a third frame from among the multiple frames for full-frame analysis by a first neural network;

using the first processor circuit, applying the third frame as an input to the first neural network and, in response, receiving a first indication of a likelihood that the third frame includes at least a portion of an image of a human being; and

using the first processor circuit, providing an indication that a human being is present in or absent from the environment monitored by the camera based on the identified frame difference and on the received first indication of the likelihood that the third frame includes at least a portion of an image of a human being, wherein the indication that the human being is present includes an activity report with information about whether a particular activity was performed by the human being in the environment monitored by the camera.

2. The method of claim 1 , wherein the selecting the third frame from among the multiple frames includes selecting a frame that includes information about the object in the environment.

3. The method of claim 1 , wherein the selecting the third frame from among the multiple frames includes selecting one of the first and second frames.

4. The method of claim 1 , wherein the identifying the change in the one or more portions of the at least first and second frames includes applying the first frame or second frame as an input to a second neural network and, in response, receiving a second indication of a likelihood that the first or second frame includes at least a portion of an image of a human being; and

wherein the providing the indication that a human being is present or absent includes using the first indication of a likelihood that the third frame includes at least a portion of an image of a human being, and using the second indication of a likelihood that the first or second frame includes at least a portion of an image of a human being.

5. The method of claim 4 , further comprising differently weighting the likelihood indications from the first and second neural networks.

6. The method of claim 4 , wherein the applying the first frame or the second frame as an input to the second neural network includes applying information corresponding to the portion in which the change was identified and excluding information not corresponding to the portion in which the change was identified.

7. The method of claim 1 , wherein when the processor circuit provides an indication that a human being is present in the environment, the method further comprises determining whether the human being is in a permitted or unpermitted location in the environment.

8. The method of claim 1 , further comprising generating a dashboard of information about the indication that a human being is present in or absent from the environment monitored by the camera, wherein the dashboard comprises the activity report and one or more of a dwell time indicator, a heat map indicating a location of the human being in the environment, or demographic indicator that includes demographic information about the human being.

9. The method of claim 1 , further comprising generating a dashboard of information about the indication that the human being is present in the environment monitored by the camera, the dashboard comprising dwell time information about the human being at a particular location in the environment.

10. The method of claim 9 , wherein the dashboard further comprises occupancy information about the human being, and about one or more other human beings over time, for the particular location in the environment.

11. The method of claim 10 , wherein the dashboard further comprises demographic information about the human being and about the one or more other human beings, and wherein the demographic information is provided at least in part using the first processor circuit and the first neural network.

12. A machine learning-based multiple-factor image recognition system for determining when human beings are present in or absent from an environment and reporting information about an occupancy of the environment over time, the system comprising:

a camera configured to receive a series of images of the environment, wherein each of the images is a different frame acquired at a different time; and

an image processor circuit configured to:

determine whether a frame difference, between a portion of at least first and second frames acquired by the camera, indicates movement by one or more objects in the environment monitored by the camera;

when the frame difference indicates movement by the object, configuring the image processor circuit to:

select a third frame from among the multiple frames for full-frame analysis by a first neural network;

apply the third frame as an input to the first neural network and in response receive a first indication that the third frame includes an image of at least a portion of a first human being;

determine whether the first human being is present in or absent from the environment based on the identified frame difference and on the received first indication;

determine whether an activity of interest was performed by the first human being in the environment; and

store information about the first human being, including information about whether the activity of interest was performed by the first human being, in a memory circuit, the stored information further including at least one of demographic information, dwell time information, or location information about the first human being.

13. The system of claim 12 , further comprising a second processor circuit configured to:

receive information from the memory circuit about the first human being and other human beings detected in the environment using information from the camera; and

generate a pictorial dashboard for presenting to a user the demographic information, dwell time information, and location information for the first human being and for the other human beings.

14. The system of claim 12 , further comprising a second camera configured to receive a series of images of a second environment, wherein each of the images is a different frame acquired at a different time; and

wherein the image processor circuit is configured to determine whether the first human being is present in or absent from the second environment based on information from the second camera.

15. The system of claim 12 , further comprising:

a second camera configured to receive a series of images of a second environment, wherein each of the images is a different frame acquired at a different time; and

a second processor circuit configured to:

receive information from the memory circuit about the first human being and receive information about one or more other human beings detected in images from the second camera;

perform facial recognition to determine if the first human being is a recognized individual; and

generate a dashboard for presenting to a user information about the first human being together with information about the one or more other human beings.

16. The system of claim 12 , wherein the image processor circuit is configured to use the first neural network to determine the demographic information about the first human being, and wherein the image processor circuit is configured to store, in the memory circuit, the demographic information, dwell time information, and location information about the first human being.

17. A machine learning-based image classifier system for determining whether a human being is present in or absent from an environment using neural network processing, the system comprising:

a first camera configured to receive a series of images of the environment, wherein each of the images corresponds to a different frame acquired at a different time, and wherein the series of images comprises image information in a YUV color space;

an image processor circuit configured to:

identify, using at least the Y information from the YUV color space, a difference between a portion of at least first and second frames acquired by the first camera, wherein when the difference indicates movement by the object in the environment monitored by the first camera then further configuring the processor to:

select a third frame from among the multiple frames for full-frame analysis using a first neural network;

apply the third frame as an input to the first neural network and in response determine a first indication of a likelihood that the third frame includes at least a portion of an image of a first human being; and

provide an indication that a human being is present in or absent from the environment based on the identified difference and on the determined first indication of the likelihood that the third frame includes at least a portion of an image of the first human being; and

a second processor circuit configured to generate a visual dashboard of information for presentation to a user about the first human being, the dashboard comprising dwell time information about the first human being for a particular location in the environment and dwell time information about one or more other human beings over time for the same particular location in the environment.

18. The system of claim 17 , further comprising:

a second camera configured to receive a series of images of the same or different environment, wherein each of the images corresponds to a different frame; and

wherein the second processor circuit is configured to generate the visual dashboard of information using the series of images from the second camera.

19. The system of claim 17 ,

wherein the image processor circuit is configured to identify the difference between a portion of at least first and second frames acquired by the first camera using only the Y information from the YUV color space; and

wherein the image processor circuit is configured to apply the Y information about the third frame from the YUV color space as the input to the first neural network.

20. The system of claim 17 , wherein the image processor circuit is configured to identify the difference between a portion of at least first and second frames acquired by the first camera using Y, U, and V information from the YUV color space.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 7, 2024
From: INTELLIVISION TECHNOLOGIES CORP.
To: NICE NORTH AMERICA LLC
Reel/Frame 068815/0671 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 6, 2020
From: KHADLOYA, KRISHNA; SOKOLOV, MAXIM; NATHAN, VAIDHI; GOPE, CHANDAN
To: INTELLIVISION TECHNOLOGIES CORP.
Reel/Frame 052322/0227 →
Continuity (3)
Provisional Application 62632416 · Feb 19, 2018
Provisional Application 62632417 · Feb 19, 2018
Related Publication 20190258866A1 · Aug 22, 2019
Cited By (4)
US 12,314,889 US 12,475,707 US 12,518,605 US 12,637,896