IP Library › Granted Patent US 9,430,701
Granted Patent B2
US 9,430,701 · App. 14/614,891 · Granted Aug 30, 2016

Object detection system and method

Inventors: Sangheeta Roy (West Bengal, IN); Tanushyam Chattopadhyay (West Bengal, IN); Dipti Prasad Mukherjee (West Bengal, IN)
Assignee: Tata Consultancy Services Limited
G06K9/00369G06K9/00201G06K9/00771G06K9/4638G06T7/0081G06T7/0093G06T7/0097G06T7/408G06K2009/485G06T2207/10024G06T2207/10028G06T2207/20021G06T2207/20072G06T2207/30196G06T2207/30232
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,430,701
App. No.
14/614,891
Granted
Aug 30, 2016
Kind
B2
Abstract

Disclosed is a system and method for detecting a human in an image, and a corresponding activity. The image is captured, wherein the image comprises a plurality of pixels having gray scale information and a depth information. The image is segmented into a plurality of segments based upon the depth information. A connected component analysis is performed on a segment in order to segregate the one or more objects into noisy objects and candidate objects, the noisy objects are eliminated from the segment. A plurality of features are extracted from the candidate objects, and are evaluated using a Hidden Markov Model (HMM) model in order to determine the candidate objects as one of the human or non-human. The corresponding activity associated with the human is detected based on a depth value associated with each pixel corresponding to the candidate object in the image.

Claims (67)

1. A method for detecting a human in at least one image, the method comprising:

capturing at least one image using a motion sensing device, wherein the at least one image comprises a plurality of pixels having gray scale information and a depth information, and wherein the gray scale information comprises intensity of each pixel corresponding to a plurality of objects in the at least one image, and wherein the depth information comprises a distance information between each object and the motion sensing device;

segmenting, by a processor, the at least one image into a plurality of segments based on the depth information of the plurality of objects, wherein each segment of the plurality of segments comprises a subset of the plurality of pixels, and each segment corresponds to one or more objects in the at least one image;

performing a connected component analysis on a segment of the plurality of segments to obtain one or more noisy objects and one or more candidate objects;

eliminating the one or more noisy objects from the segment using a vertical pixel projection technique;

extracting a plurality of features from the one or more candidate objects present in the segment, wherein the plurality of features are extracted by,

applying a windowing technique on the segment in order to divide the segment into one or more blocks, wherein each block of the one or more blocks comprises one or more sub-blocks,

calculating a local gradient histogram (LGH) corresponding to each sub-block of the one or more sub-blocks, and

generating a vector comprising the plurality of features by concatenating the LGH of each sub-block; and

detecting a human or a non-human from the one or more candidate objects based on the plurality of features.

2. The method of claim 1 , wherein the segmenting further comprises:

maintaining the gray scale information of the subset of the plurality of pixels; and

transforming the gray scale information of remaining pixels other than the subset into a black color.

3. The method of claim 1 , wherein the human or the non-human are detected from the one or more candidate objects based on an evaluation of a state transition sequence of the plurality of features with a sequence of features pre-stored in a database.

4. The method of claim 1 , wherein the plurality of features are evaluated by using a Hidden Markov Model (HMM) to detect the human or the non-human from the one or more candidate objects.

5. The method of claim 1 , further comprising detecting one or more activities associated with the human in the at least one image by:

analyzing each pixel of the one or more candidate objects in the at least one image by,

comparing the gray scale value of each pixel with a pre-defined gray scale value,

replacing a subset of the pixels having the gray scale value less than the pre-defined gray scale value with 0 and a remaining subset of the pixels with 1 in order to derive a binary image corresponding to the at least one image, and

determining the subset as the one or more candidate objects in the binary image;

performing a connected component analysis on the binary image in order to detect a candidate object of the one or more candidate objects as the human;

retrieving the depth value associated with each pixel corresponding to the candidate object from a look-up table; and

detecting an activity of the candidate object by using the depth value or a floor map algorithm.

6. The method of claim 5 , further comprising de-noising each pixel in order to retain the depth value associated with each pixel using a nearest neighbor interpolation algorithm, wherein each pixel is de-noised when the depth value of each pixel bounded by the one or more pixels are not within a pre-defined depth value of the one or more pixels.

7. The method of claim 5 , wherein the activity comprises any of a walking, a standing, a sitting, a sleeping, or combinations thereof.

8. A system for detecting a human in at least one image, the system comprising:

a processor; and

a memory coupled to the processor, wherein the processor is capable of executing a plurality of modules stored in the memory, and wherein the plurality of modules comprising:

an image capturing module configured to:

capture at least one image using a motion sensing device, wherein the at least one image comprises a plurality of pixels having gray scale information and a depth information, and wherein the gray scale information comprises intensity of each pixel corresponding to a plurality of objects in the at least one image, and wherein the depth information comprises a distance information between each object and the motion sensing device,

segment the at least one image into a plurality of segments based on the depth information of the plurality of objects, wherein each segment of the plurality of segments comprises a subset of the plurality of pixels, and each segment corresponds to one or more objects in the at least one image,

perform a connected component analysis on a segment of the plurality of segments to obtain one or more noisy objects and one or more candidate objects, and

eliminate the one or more noisy objects from the segment using a vertical pixel projection technique,

a feature extraction module that is configured to extract a plurality of features from the one or more candidate objects present in the segment, wherein the plurality of features are extracted by,

applying a windowing technique on the segment in order to divide the segment into one or more blocks, wherein each block of the one or more blocks comprises one or more sub-blocks,

calculating a local gradient histogram (LGH) corresponding to each sub-block of the one or more sub-blocks, and

generating a vector comprising the plurality of features by concatenating the LGH of each sub-block; and

an object determination module that is configured to detect a human or a non-human from the one or more candidate objects based on the plurality of features.

9. The system of claim 8 , further comprising

an image processing module that is configured to analyze each pixel of the one or more candidate objects in the at least one image by,

comparing the gray scale value of each pixel with a pre-defined gray scale value,

replacing a subset of the pixels having the gray scale value less than the pre-defined gray scale value with 0 and a remaining subset of the pixels with 1 in order to derive a binary image corresponding to the at least one image, and

determining the subset as the one or more candidate objects in the binary image;

performing the connected component analysis on the binary image in order to detect a candidate object of the one or more candidate objects as the human; and

retrieving the depth value associated with each pixel corresponding to the candidate object from a look-up table; and

an activity detection module that is configured to detect one or more activities of the candidate object by using the depth value or a floor map algorithm.

10. A non-transitory computer program product having embodied thereon a computer program for detecting a human in at least one image by performing the step of:

capturing at least one image using a motion sensing device, wherein the at least one image comprises a plurality of pixels having gray scale information and a depth information, and wherein the gray scale information comprises intensity of each pixel corresponding to a plurality of objects in the at least one image, and wherein the depth information comprises a distance information between each object and the motion sensing device;

segmenting, by a processor, the at least one image into a plurality of segments based on the depth information of the plurality of objects, wherein each segment of the plurality of segments comprises a subset of the plurality of pixels, and each segment corresponds to one or more objects in the at least one image;

performing a connected component analysis on a segment of the plurality of segments to obtain one or more noisy objects and one or more candidate objects;

eliminating the one or more noisy objects from the segment using a vertical pixel projection technique;

extracting a plurality of features from the one or more candidate objects present in the segment, wherein the plurality of features are extracted by,

applying a windowing technique on the segment in order to divide the segment into one or more blocks, wherein each block of the one or more blocks comprises one or more sub-blocks,

calculating a local gradient histogram (LGH) corresponding to each sub-block of the one or more sub-blocks, and

generating a vector comprising the plurality of features by concatenating the LGH of each sub-block; and

detecting a human or a non-human from the one or more candidate objects based on the plurality of features.

11. The non-transitory computer program product of claim 10 , wherein the one or more instructions, which when executed by the one or more processors further causes

detecting one or more activities associated with the human in the at least one image

by: analyzing each pixel of the one or more candidate objects in the at least one image by,

comparing the gray scale value of each pixel with a pre-defined gray scale value,

replacing a subset of the pixels having the gray scale value less than the pre-defined gray scale value with 0 and a remaining subset of the pixels with 1 in order to derive a binary image corresponding to the at least one image, and

determining the subset as the one or more candidate objects in the binary image;

performing a connected component analysis on the binary image in order to detect a candidate object of the one or more candidate objects as the human;

retrieving the depth value associated with each pixel corresponding to the candidate object from a look-up table; and

detecting an activity of the candidate object by using the depth value or a floor map algorithm.

12. The non-transitory computer program product of claim 11 , wherein the one or more instructions, which when executed by the one or more processors further causes de-noising each pixel in order to retain the depth value associated with each pixel using a nearest neighbor interpolation algorithm, wherein each pixel is de-noised when the depth value of each pixel bounded by the one or more pixels are not within a pre-defined depth value of the one or more pixels.

13. The non-transitory computer program product of claim 11 , wherein the activity comprises any of a walking, a standing, a sitting, a sleeping, or combinations thereof.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 10, 2015
From: ROY, SANGHEETA; CHATTOPADHYAY, TANSUHYAM; MUKHERJEE, DIPTI PRASAD
To: TATA CONSULTANCY SERVICES LIMITED
Reel/Frame 034929/0311 →
Priority Claims (2)
IN 456/MUM/2014 · Feb 7, 2014 · national
IN 949/MUM/2014 · Mar 22, 2014 · national
Continuity (1)
Related Publication 20150227784A1 · Aug 13, 2015