IP Library Granted Patent US 12,494,085
Granted Patent B2
US 12,494,085 · App. 17/926,322 · Granted Dec 9, 2025

Technologies for analyzing behaviors of objects or with respect to objects based on stereo imageries thereof

Inventors: Aluisio Figueiredo (Freehold, NJ); Roman Jarkoi (Sunrise, FL); Oleg Vladimirovich Stepanenko (Balashikha, RU); Valery Arzumanov (Moscow, RU)
Assignee: Intelligent Security Systems Corporation
G06V40/20G06T7/251G06T7/285G06V10/82G06V20/52G06V20/64G06V40/28G08B21/245G06T2207/30196
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,494,085
App. No.
17/926,322
Granted
Dec 9, 2025
Kind
B2
Abstract

This disclosure enables various technologies for analyzing behaviors of objects or with respect to objects based on stereo imageries thereof. For example, such analysis may be useful in enforcement of certain actions by objects or with respect to objects, surveillance of objects or with respect to objects, or other situations involving analyzing behaviors of objects or with respect to objects.

Claims (47)

1 . A device, comprising:

a processor programmed to:

access, in real-time, a stereo imagery of an area including a first object and a second object engaging with the first object;

form, in real-time, a reconstruction of the second object in the area based on the stereo imagery, wherein the reconstruction including a 3D area model and a 3D skeletal model within the 3D area model, wherein the 3D area model simulating the area, wherein the 3D skeletal model simulating the second object in the area;

identify, in real-time, a set of virtual movements of the 3D skeletal model in the 3D area model, wherein the set of virtual movements simulating the second object engaging with the first object;

identify, in real-time, a set of atomic movements of the 3D skeletal model corresponding to the set of virtual movements;

identify, in real-time, an event semantically defined by an expert system as a combination of at least two atomic movements of the set of atomic movements; and

take, in real-time, an action responsive to the event being identified.

2 . The device of claim 1 , wherein the second object including a real limb, wherein the 3D skeletal model including a virtual limb simulating the real limb, wherein the set of virtual movements including a virtual movement of the virtual limb, wherein the set of atomic movements including an atomic movement of the virtual limb, wherein the processor is further programmed to identify, in real-time, the set of atomic movements based on the atomic movement of the virtual limb corresponding the virtual movement of the virtual limb.

3 . The device of claim 2 , wherein the atomic movement of the virtual limb including a bending of the virtual limb according to a preset rule.

4 . The device of claim 3 , wherein the real limb is a real hand, wherein the first object is a sanitizing station, wherein the second object engaging with the first object includes the second object sanitizing the real hand at the sanitizing station, wherein the event is a hand sanitizing event.

5 . The device of claim 3 , wherein the first object is a first person having a first real hand, wherein the second object is a second person having a second real hand corresponding to the real limb, wherein the second object engaging with the first object including the second real hand shaking the first real hand, wherein the event is a hand shaking event.

6 . The device of claim 1 , wherein the set of atomic movements is a first set of atomic movements, wherein the event is further defined by a second set of atomic movements of the 3D skeletal model corresponding to the set of virtual movements, wherein the processor is further programmed to identify, in real-time, the event based on the first set of atomic movements and the second set of atomic movements.

7 . The device of claim 1 , wherein the second object having a pose in the area, wherein the second object including a real limb in the area, wherein the processor is further programmed to form, in real-time, the 3D skeletal model from the pose based on the stereo imagery and from how the real limb is positioned in the area based on the stereo imagery.

8 . The device of claim 7 , wherein the processor is further programmed to identify, in real-time, the set of atomic movements based on the 3D skeletal model simulating the pose and how the real limb is positioned in the area.

9 . The device of claim 1 , wherein the second object has a real face, wherein the stereo imagery depicts the real face, wherein the processor is further programmed to perform, in real-time, a recognition of the real face based on at least an image of the stereo imagery, wherein the action is based on the recognition, wherein the image depicts the real face.

10 . The device of claim 9 , wherein the action is a first action, wherein the processor is further programmed to take a second action responsive to the event not being identified, wherein the second action is based on the recognition.

11 . The device of claim 10 , wherein the second action includes adding an identifier of the second object associated with the real face to a list of identifiers not associated with the event.

12 . The device of claim 1 , wherein the processor is further programmed to modify, in real-time, a log for each occurrence of the event being identified and the event not being identified, wherein each of the occurrences includes an identifier of the second object.

13 . The device of claim 1 , wherein the processor is further programmed to retrieve a record with a set of information for the event not being identified responsive to a request from a user input device.

14 . The device of claim 13 , wherein the set of information includes at least a user interface element programmed to retrieve or play a portion of the stereo imagery.

15 . The device of claim 1 , wherein the processor is further programmed to access a log containing a set of information for the event being identified and the event not being identified responsive to a request from a user input device, and export a subset of the set of information for a preset time period responsive to the request from the user input device.

16 . The device of claim 1 , wherein the processor is further programmed to generate a dashboard based on a set of information for the event being identified and the event not being identified responsive a request from a user input device, and instruct a display to output the dashboard responsive to the request from the user input device.

17 . The device of claim 1 , further comprising:

a next unit of computing (NUC) apparatus including the processor; and

a housing including a pair of video cameras forming a stereo pair sourcing the stereo imagery.

18 . The device of claim 1 , wherein the action is a first action, wherein the processor is programmed to take a second action responsive to the event not being identified.

19 . The device of claim 18 , wherein the processor is further programmed to access an identifier for the second object, wherein each of the first action and the second action involves the identifier.

20 . The device of claim 18 , wherein the second action includes requesting an output device to output an alert, an alarm, or a notice indicative of the event not being identified.

21 . The device of claim 20 , wherein the output device is a display.

22 . The device of claim 20 , wherein the output device is a speaker.

23 . The device of claim 20 , wherein the stereo imagery is sourced from a source, wherein the output device is remote from the source.

24 . The device of claim 20 , wherein the stereo imagery is sourced from a source, wherein the output device is local to the source.

25 . The device of claim 20 , wherein the first object is a static object and the second object is a dynamic object.

26 . The device of claim 20 , wherein the first object is a dynamic object and the second object is a dynamic object.

27 . The device of claim 1 , wherein the stereo imagery is markerless motion capture.

28 . The device of claim 1 , wherein the expert system is programmed to employ a natural language to define the event as the combination of the at least two atomic movements of the set of atomic movements.

29 . The device of claim 1 , wherein the stereo imagery includes a set of depth metadata, wherein the expert system is programmed to process the set of depth metadata to define the event.

30 . A method, comprising:

accessing, via a processor, in real-time, a stereo imagery of an area including a first object and a second object engaging with the first object;

forming, via the processor, in real-time, a reconstruction of the second object in the area based on the stereo imagery, wherein the reconstruction including a 3D area model and a 3D skeletal model within the 3D area model, wherein the 3D area model simulating the area, wherein the 3D skeletal model simulating the second object in the area;

identifying, via the processor, in real-time, a set of virtual movements of the 3D skeletal model in the 3D area model, wherein the set of virtual movements simulating the second object engaging with the first object;

identifying, via the processor, in real-time, a set of atomic movements of the 3D skeletal model corresponding to the set of virtual movements;

identifying, via the processor, in real-time, an event semantically defined by an expert system as a combination of at least two atomic movements of the set of atomic movements; and

taking, via the processor, in real-time, an action responsive to the event being identified.

31 . The method of claim 30 , wherein the expert system is programmed to employ a natural language to define the event as the combination of the at least two atomic movements of the set of atomic movements.

32 . The method of claim 30 , wherein the stereo imagery includes a set of depth metadata, wherein the expert system is programmed to process the set of depth metadata to define the event.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2023
From: FIGUEIREDO, ALUISIO; JARKOI, ROMAN; STEPANANKO, OLEG VLADIMIROVICH; ARZUMANOV, VALERY
To: INTELLIGENT SECURITY SYSTEMS CORPORATION
Reel/Frame 062612/0858 →
Continuity (2)
Provisional Application 63027215 · May 19, 2020
Related Publication 20230237850A1 · Jul 27, 2023
References Cited (44)
US 6847957B1 · Morley · 2005 [cited by examiner]
US 9244924B2 · Cheng · 2016 [cited by examiner]
US 9507768B2 · Cobb · 2016 [cited by examiner]
US 10482181B1 · Rubin · 2019 [cited by examiner]
US 11182924B1 · Akbas · 2021 [cited by examiner]
US 11238401B1 · Guan · 2022 [cited by examiner]
US 11373331B2 · Huelsdunk · 2022 [cited by examiner]
US 11826636B2 · Argiro · 2023 [cited by examiner]
US 11868956B1 · Dillon · 2024 [cited by examiner]
US 12250362B1 · Berme · 2025 [cited by examiner]
US 20060149553A1 · Begeja · 2006 [cited by examiner]
US 20070285419A1 · Givon · 2007 [cited by applicant]
US 20100201783A1 · Ueda · 2010 [cited by applicant]
US 20120287243A1 · Parulski · 2012 [cited by applicant]
US 20120301013A1 · Gu · 2012 [cited by applicant]
US 20150000026A1 · Clements · 2015 [cited by examiner]
US 20150248917A1 · Chang · 2015 [cited by examiner]
US 20170035330A1 · Bunn · 2017 [cited by examiner]
US 20170262584A1 · Gallix · 2017 [cited by examiner]
US 20170275023A1 · Harris · 2017 [cited by applicant]
US 20180101966A1 · Lee · 2018 [cited by examiner]
US 20200129109A1 · Bunn · 2020 [cited by examiner]
US 20200193615A1 · Goncharov · 2020 [cited by examiner]
US 20200296249A1 · Parian · 2020 [cited by examiner]
US 20200311395A1 · Kim · 2020 [cited by examiner]
US 20200394364A1 · Venkateshwaran · 2020 [cited by examiner]
US 20200401617A1 · Spiegel · 2020 [cited by examiner]
US 20210209169A1 · Sundararajan · 2021 [cited by examiner]
US 20210279475A1 · Tusch · 2021 [cited by examiner]
US 20220222849A1 · Zhang · 2022 [cited by examiner]
US 20220308359A1 · Karafin · 2022 [cited by examiner]
US 20220335720A1 · Chang · 2022 [cited by examiner]
US 20240148464A1 · Freeman · 2024 [cited by examiner]
US 20240225579A1 · Nandi · 2024 [cited by examiner]
US 20240232539A1 · Venkateshwaran · 2024 [cited by examiner]
CN 104461525A · 2015 [cited by examiner]
CN 118821792A · 2024 [cited by examiner]
Singh, Alok et al. “A Comprehensive Review on Recent Methods and Challenges of Video Description”, Cornell University, Dec. 1, 2020 (Year: 2020). [cited by examiner]
Extended European Search Report Apr. 25, 2024 in related EP Application No. 21809882.0 filed May 16, 2021 (10 pages). [cited by applicant]
Troung Anh Minh et al., “Structured RNN for human interaction”, IET Computer Vision, The Institution of Engineering and Technology, Michael Faraday House, Six Fllls Way, Stevenage, Herts. SG1 2AY, UK, vol. 12, No. 6 Sep… [cited by applicant]
Anonymous: “Human-Object Interaction Detection I Paper with Code”, Apr. 11, 2024, XP093150665, retrieved from the internet: URL:https://paperswithcode.com/task/huma-object-interaction-detection, Apr. 11, 2024 (9 pages). [cited by applicant]
Nina Wiedemann et al., “A Tracking System for Baseball Game Reconstruction”, ARXIV.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Mar. 8, 2020, XP081617369 (27 pages). [cited by applicant]
Hand, A survey of 3D interaction techniques. Computer graphics forum. Vol. 16. No. 5. Oxford, UK and Boston, USA: Blackwell Publishers, 1997. Jun. 28, 2008 (Jun. 28, 2008) Retrieved on Jul. 12, 2021 (Jul. 12, 2021) from… [cited by applicant]
International Search Report and Written Opinion mailed on Aug. 17, 2021 in related International Application PCT/US2021/032649 filed May 16, 2021 (14 pages). [cited by applicant]