IP Library › Granted Patent US 12,456,221
Granted Patent B2
US 12,456,221 · App. 18/075,875 · Granted Oct 28, 2025

Electronic apparatus for real-time human detection and tracking system and controlling method thereof

Inventor: Yusun Lim (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06T7/70G06T7/20G06T7/50G06V10/25G06V10/761G06V10/82G06T2207/10024G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,456,221
App. No.
18/075,875
Granted
Oct 28, 2025
Kind
B2
Abstract

An electronic apparatus is provided. The electronic apparatus includes a first sensor configured to obtain a color image, a second sensor configured to obtain a depth image, a memory storing a neural network model, and a processor configured to, based on a first color image being received from the first sensor, obtain a first region of interest by inputting the first color image to the neural network model, and identify whether a distance between an object included in the first region of interest and the electronic apparatus is less than a threshold distance.

Claims (71)

1. An electronic apparatus comprising:

a first sensor configured to obtain a color image;

a second sensor configured to obtain a depth image;

memory storing a neural network model and instructions; and

one or more processors communicatively coupled to the first sensor, second sensor and the memory,

wherein the instructions, when executed by the one or more processors individually or collectively, cause the electronic apparatus to:

based on a first color image being obtained via the first sensor, obtain a first region of interest from the first color image by inputting the first color image to the neural network model,

identify whether a distance between an object included in the first region of interest and the electronic apparatus is less than a threshold distance,

based on the identified distance being less than the threshold distance, identify a first region including pixels including depth information less than the threshold distance among a plurality of pixels included in a first depth image corresponding to the first color image,

obtain intersection information between the first region of interest and the first region,

based on the obtained intersection information being a threshold value or more, identify that the first region includes the object and merge the first region of interest and the first region to obtain a merged region, and

track a position of the object based on the merged region.

2. The electronic apparatus of claim 1 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to:

identify a second region including pixels including depth information less than the threshold distance in a second depth image corresponding to a second color image received from the second sensor,

obtain second intersection information between the first region and the second region,

based on the obtained second intersection information being the threshold value or more, identify that the second region includes the object, and

track the position of the object based on the first region and the second region.

3. The electronic apparatus of claim 2 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to:

input a second color image received from the first sensor to the neural network model, and

based on a region of interest being not identified in the second color image based on an output of the neural network model, identify the second region including the pixels including the depth information less than the threshold distance in the second depth image corresponding to the second color image.

4. The electronic apparatus of claim 2 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to:

based on a proportion of pixels including the depth information less than the threshold distance in the second depth image being less than a threshold proportion, obtain a second region of interest including the object from the first color image by inputting the second color image to the neural network model,

identify whether a distance between the object included in the second region of interest and the electronic apparatus is less than a threshold distance,

based on the identified distance being less than the threshold distance, identify the second region including the pixels including the depth information less than the threshold distance among a plurality of pixels included in the second depth image,

obtain third intersection information between the second region of interest and the second region, and

based on the obtained third intersection information being a threshold value or more, identify that the second region includes the object.

5. The electronic apparatus of claim 4 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to, based on the identified distance being the threshold distance or more, track the position of the object based on the first region of interest, the first region, and the second region of interest.

6. The electronic apparatus of claim 1 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to:

based on the obtained intersection information being the threshold value or more, obtain a first merged region in which the first region of interest and the first region are merged,

identify a second region including pixels including depth information less than the threshold distance in a second depth image corresponding to a second color image received from the second sensor,

obtain second intersection information between the first merged region and the second region,

based on the obtained second intersection information being the threshold value or more, identify that the second region includes the object, and

track the position of the object based on the first merged region and the second region.

7. The electronic apparatus of claim 6 , wherein the instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to, based on a proportion of pixels including depth information less than the threshold distance in the second depth image being a threshold proportion or more, identify the second region including pixels including depth information less than the threshold distance among a plurality of pixels included in the second depth image.

8. The electronic apparatus of claim 1 , wherein the processor is instructions that, when executed by the one or more processors individually or collectively, further cause the electronic apparatus to, based on the identified distance being the threshold distance or more, identify the position of the object based on the first region of interest.

9. The electronic apparatus of claim 1 ,

wherein the first sensor comprises at least one of a camera or a red, green, and blue (RGB) color sensor, and

wherein the second sensor comprises at least one of a stereo vision sensor, a Time-of-Flight (ToF) sensor, or a light detection and ranging (LiDAR) sensor.

10. A method an electronic apparatus, the method comprising:

based on a first color image being obtained via a first sensor, obtaining, by the electronic apparatus, a first region of interest from the first color image by inputting the first color image to a neural network model; and

identifying, by the electronic apparatus, whether a distance between an object included in the first region of interest and the electronic apparatus is less than a threshold distance,

wherein the identifying comprises:

based on the identified distance being less than the threshold distance, identifying, by the electronic apparatus, a first region including pixels including depth information less than the threshold distance among a plurality of pixels included in a first depth image corresponding to the first color image,

obtaining, by the electronic apparatus, intersection information between the first region of interest and the first region,

based on the obtained intersection information being a threshold value or more, identifying, by the electronic apparatus, that the first region includes the object and merging the first region of interest and the first region to obtain a merged region, and

tracking a position of the object based on the merged region.

11. The method of claim 10 , further comprising:

identifying a second region including pixels including depth information less than the threshold distance in a second depth image corresponding to a second color image received from a second sensor;

obtaining second intersection information between the first region and the second region;

based on the obtained second intersection information being the threshold value or more, identifying that the second region includes the object; and

tracking the position of the object based on the first region and the second region.

12. The method of claim 11 , further comprising:

inputting a second color image received from the first sensor to the neural network model,

wherein the identifying of the second region comprises, based on a region of interest being not identified in the second color image based on an output of the neural network model, identifying the second region including the pixels including the depth information less than the threshold distance in the second depth image corresponding to the second color image.

13. The method of claim 11 , further comprising:

based on a proportion of pixels including the depth information less than the threshold distance in the second depth image being less than a threshold proportion, obtaining a second region of interest including the object from the first color image by inputting the second color image to the neural network model; and

identifying whether a distance between the object included in the second region of interest and the electronic apparatus is less than a threshold distance,

wherein the identifying of the second region comprises, based on the identified distance being less than the threshold distance, identifying the second region including the pixels including the depth information less than the threshold distance among a plurality of pixels included in the second depth image, and

wherein the method further comprises:

obtaining intersection information between the second region of interest and the second region, and

based on the obtained intersection information being a threshold value or more, identifying that the second region includes the object.

14. The method of claim 11 , wherein the neural network model comprises at least one of: a trained convolution neural network (CNN), a deep neural network (DNN) model of a Recurrent Neural Network (RNN), a Long Short-Term Memory Network (LSTM), a Gated Recurrent Units (GRU), or a Generative Adversarial Networks (GAN).

15. The method of claim 11 , wherein the neural network model comprises a neural network model trained to identify, based on the first color image, a region predicted to include the object in the first color image as the first region of interest.

16. The method of claim 15 , wherein the neural network model comprises a neural network model trained to identify a region in which a pixel value is changed by comparing the first color image preceding in chronological sequence and a following second color image, and determine the identified region as the first region of interest predicted to include the object.

17. The method of claim 10 , further comprising:

based on the obtained intersection information being the threshold value or more, obtaining a first merged region in which the first region of interest and the first region are merged;

identifying a second region including pixels including depth information less than the threshold distance in a second depth image corresponding to a second color image received from a second sensor;

obtaining third intersection information between the first merged region and the second region;

based on the obtained third intersection information being the threshold value or more, identifying that the second region includes the object; and

tracking the position of the object based on the first merged region and the second region.

18. The method of claim 17 , wherein the identifying of the second region comprises, based on a proportion of pixels including depth information less than the threshold distance in the second depth image being a threshold proportion or more, identifying the second region including pixels including depth information less than the threshold distance among a plurality of pixels included in the second depth image.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2022
From: LIM, YUSUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 061997/0549 →
Priority Claims (1)
KR 10-2021-0142133 · Oct 22, 2021 · national
Continuity (2)
Continuation PCTKR2022012226 · Aug 17, 2022
Related Publication 20230131404A1 · Apr 27, 2023
References Cited (31)
US 7590262B2 · Fujimura et al. · 2009 [cited by applicant]
US 9201425B2 · Yoon et al. · 2015 [cited by applicant]
US 9247211B2 · Zhang et al. · 2016 [cited by applicant]
US 9460339B2 · Litvak et al. · 2016 [cited by applicant]
US 9639951B2 · Salahat et al. · 2017 [cited by applicant]
US 9729865B1 · Kuo et al. · 2017 [cited by applicant]
US 10761538B2 · Ball et al. · 2020 [cited by applicant]
US 10984543B1 · Srinivasan · 2021 [cited by applicant]
US 11417007B2 · Kang et al. · 2022 [cited by applicant]
US 11605244B2 · Lee et al. · 2023 [cited by applicant]
US 11734910B2 · Mithun et al. · 2023 [cited by applicant]
US 20110211754A1 · Litvak · 2011 [cited by examiner]
US 20170154471A1 · Woo · 2017 [cited by examiner]
US 20170200044A1 · Lee et al. · 2017 [cited by applicant]
US 20170372479A1 · Somanath · 2017 [cited by examiner]
US 20200050839A1 · Wolf et al. · 2020 [cited by applicant]
US 20200118988A1 · Yueh et al. · 2020 [cited by applicant]
US 20200380251A1 · Lee et al. · 2020 [cited by applicant]
US 20210150746A1 · Kang et al. · 2021 [cited by applicant]
US 20220335252A1 · Georgievskaya · 2022 [cited by examiner]
US 20230251744A1 · Peuhkurinen · 2023 [cited by examiner]
CN 108830179A · 2018 [cited by applicant]
KR 1020160078082A · 2016 [cited by applicant]
KR 1020180108123A · 2018 [cited by applicant]
KR 1020210061839A · 2021 [cited by applicant]
KR 102565823B1 · 2023 [cited by applicant]
Tang et al., “Hand tracking and pose recognition via depth and color information” (Year: 2012). [cited by examiner]
International Search Report and Written Opinion dated Dec. 13, 2022, issued in International Application No. PCT/KR2022/012226. [cited by applicant]
XP 034051341, Pallet detection and docking strategy for autonomous pallet truck AGV operation, Sep. 27, 2021. [cited by applicant]
XP 033251277, Deep Detection of People and their Mobility Aids for a Hospital Robot, Sep. 6, 2017. [cited by applicant]
European Search Report dated Oct. 4, 2024, issued in European Application No. 22883739.9. [cited by applicant]