IP Library › Granted Patent US 12,536,698
Granted Patent B2
US 12,536,698 · App. 18/213,688 · Granted Jan 27, 2026

Intelligent projection point prediction for overhead objects

Inventors: Poyraz Umut Hatipoglu (Istanbul, TR); Sercan Esen (New York, NY); Serhat Çillidag (Berlin, DE); Okan Ulusoy (Istanbul, TR); Ali Ufuk Yaman (Istanbul, TR)
Assignee: Intenseye, Inc.
G06T7/74B66C13/46B66C15/065G06T7/11G06T7/248G06T7/55B66C2700/084G06T2207/20068G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,536,698
App. No.
18/213,688
Filed
Jun 23, 2023
Granted
Jan 27, 2026
Kind
B2
Art Unit
2672
USPC
382/103
Abstract

Methods and systems provide for the intelligent prediction of projection points of overhead objects in a workplace or similar environment. In one embodiment, the system receives one or more two-dimensional (hereinafter “2D”) training images of unique overhead objects within an environment; processes the 2D training images to determine training positions and projection points of each overhead object; trains one or more artificial intelligence (hereinafter “AI”) models to predict 2D positions of projection points of the overhead object; receives one or more 2D inference images of the overhead objects; processes the 2D inference images to determine one or more inferred positions of the overhead objects; predicts, via the trained AI models, 2D inferred positions of the projection points of the overhead objects; and provides them to one or more client devices. In some embodiments, the system alerts a person based on their proximity to the projection point of an overhead object.

Claims (65)

1 . A method, comprising:

receiving one or more two-dimensional (2D) training images of one or more unique overhead objects within an environment;

processing the 2D training images to determine one or more training positions of each overhead object and one or more training positions of projection points of each overhead object;

based on at least the training positions of the overhead objects, training one or more artificial intelligence (AI) models to predict 2D positions of projection points of the overhead object;

receiving one or more 2D inference images of the overhead objects captured by one or more cameras within the environment;

processing the 2D inference images to determine one or more inferred positions of the overhead objects by:

identifying an overhead object in the one or more 2D inference images; and

placing a bounding box about the identified overhead object, the bounding box comprising the inferred positions;

predicting, via the one or more trained AI models and using the inferred positions of the overhead objects as input, 2D inferred positions of the projection points of the overhead objects; and

providing, to one or more client devices, the predicted 2D inferred positions of the projection points of the overhead objects.

2 . The method of claim 1 , wherein the 2D inference images further comprise one or more people in the environment, the method further comprising:

processing the 2D inference images to determine one or more inferred positions of the people in the environment.

3 . The method of claim 2 , wherein processing the 2D inference images to determine the inferred positions of the people in the environment is performed via one of:

amodal object detection, or instance segmentation.

4 . The method of claim 2 , further comprising:

determining that an inferred position of at least one of the people in the environment is within a specified proximity of at least one of the predicted 2D inferred positions of the projection points of the overhead objects; and

in response to the determination, providing an alert to at least one of the people in the environment warning of dangerous proximity to at least one of the overhead objects.

5 . The method of claim 4 , wherein the determination comprises determining that the distance between at least one of the predicted 2D inferred positions of the projection points of the overhead objects and at least one of the positions of the people in the environment is lower than a threshold distance.

6 . The method of claim 1 , further comprising:

normalizing the inferred positions of the people in the environment to a range between 0 and 1.

7 . The method of claim 1 , wherein the training positions of the overhead objects, the training positions of the projection points of the overhead objects, and the inferred positions of the overhead objects are configured to be pixel coordinates.

8 . The method of claim 1 , wherein training the one or more AI models comprises normalizing the training positions of the overhead objects and the training positions of the projection points of the overhead objects to a range between 0 and 1.

9 . The method of claim 8 , wherein training the one or more AI models further comprises applying coordinate transformation to the normalized training positions of the overhead objects and the normalized training positions of the projection points of the overhead objects, and wherein applying coordinate transformation comprises defining the normalized training positions of the overhead objects and the normalized training positions of the projection points of the overhead objects as relative to the bottom center of the bounding box for the overhead objects.

10 . The method of claim 1 , wherein the processing of the 2D training images and determining of the training positions of the overhead objects is performed via one or more of: amodal object detection, instance segmentation, and manual annotation.

11 . The method of claim 1 , wherein the processing of the 2D inference images and determining of the inferred positions of the overhead objects is performed via one of: amodal object detection, and instance segmentation.

12 . The method of claim 1 , wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises normalizing the inferred positions of the overhead objects to a range between 0 and 1.

13 . The method of claim 12 , wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises applying coordinate transformation to the normalized actual positions of the overhead objects, and wherein applying coordinate transformation comprises defining the normalized inferred positions of the overhead objects as relative to the bottom center of the bounding box for the overhead objects.

14 . The method of claim 1 , wherein the determined training positions of the overhead objects are configured to be 4 dimensional vectors, and wherein the determined training positions of the projection points of the overhead objects are configured to be 2 dimensional vectors.

15 . The method of claim 1 , wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises using the inferred positions of the overhead objects as input to the one or more trained AI models, wherein the inferred positions of the overhead objects are configured to be 4 dimensional vectors, and wherein the predicted 2D inferred positions of the projection points of the overhead objects are configured to be 2 dimensional vectors.

16 . The method of claim 1 , wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises converting a pixel representation of bounding boxes from a standard bounding box format to a format comprising a bottom-center x-axis, a bottom-center y-axis, a width, and a height.

17 . The method of claim 1 , wherein training the one or more AI models comprises comparing, via a loss function, the training positions of projection points of each overhead object to one or more specified actual positions of projection points of each overhead object.

18 . The method of claim 17 , wherein the loss function is a regression loss function comprising one of: a mean squared error (MSE) function, a mean absolute error (MAE) function, and a mean high-order error loss (MHOE) function.

19 . The method of claim 1 , wherein training the one or more AI models comprises:

feeding back one or more specified faulty predicted positions of projection points of the overhead objects to the one or more AI models to refine future output of the AI models.

20 . The method of claim 1 , wherein training the one or more AI models comprises:

obtaining an optimally performing AI model state by determining a state minimizing the specified faulty predicted positions.

21 . The method of claim 1 , wherein at least a subset of the received 2D training images is generated via a simulation, the simulation comprising at least a model of the overhead object.

22 . The method of claim 21 , wherein the simulation further comprises a model of each of the one or more cameras and one or more lens characteristics pertaining to the one or more cameras.

23 . The method of claim 1 , wherein the 2D inferred positions of the projection points of the overhead objects comprise a 2D inferred position of a center projection point of each of the overhead objects.

24 . The method of claim 23 , wherein the method is further configured to predict one or more additional 2D inferred positions of additional projection points of the overhead objects other than the center projection points of the overhead objects.

25 . The method of claim 1 , wherein each of the one or more cameras are configured to operate at any height and at any angle as long as the overhead object is within view of the camera, and wherein the ground level of the environment is visible within view of the camera.

26 . The method of claim 1 , further comprising:

wherein the inferred positions of the people are projection points of people in the environment;

inputting the determined position into a proximity based alert analyzer; and

determining, by the proximity based alert analyzer, based on a distance threshold limit, whether there is a close proximity between the people within the environment and the overhead objects within the environment.

27 . The method of claim 26 , wherein the proximity based alert analyzer measures the distance between each projection point of each overhead objects and each projection point of each person in the area.

28 . The method of claim 27 , wherein each projection point of each person are feet projection points of each person.

29 . A method, comprising:

receiving one or more two-dimensional (2D) training images of one or more unique overhead objects within an environment;

processing the 2D training images to determine one or more training positions of each overhead object and one or more training positions of projection points of each overhead object;

based on at least the training positions of the overhead objects, training one or more artificial intelligence (AI) models to predict 2D positions of projection points of the overhead object;

receiving one or more 2D inference images of the overhead objects captured by one or more cameras within the environment;

processing the 2D inference images to determine one or more inferred positions of the overhead objects;

predicting, via the one or more trained AI models and using the inferred positions of the overhead objects as input, 2D inferred positions of the projection points of the overhead objects; and

providing, to one or more client devices, the predicted 2D inferred positions of the projection points of the overhead objects;

wherein training the one or more AI models comprises normalizing the training positions of the overhead objects and the training positions of the projection points of the overhead objects to a range between 0 and 1, and wherein training the one or more AI models further comprises applying coordinate transformation to the normalized training positions of the overhead objects and the normalized training positions of the projection points of the overhead objects, and wherein applying coordinate transformation comprises defining the normalized training positions of the overhead objects and the normalized training positions of the projection points of the overhead objects as relative to the bottom center of a bounding box for the overhead objects.

30 . A method, comprising:

receiving one or more two-dimensional (2D) training images of one or more unique overhead objects within an environment;

processing the 2D training images to determine one or more training positions of each overhead object and one or more training positions of projection points of each overhead object;

based on at least the training positions of the overhead objects, training one or more artificial intelligence (AI) models to predict 2D positions of projection points of the overhead object;

receiving one or more 2D inference images of the overhead objects captured by one or more cameras within the environment;

processing the 2D inference images to determine one or more inferred positions of the overhead objects;

predicting, via the one or more trained AI models and using the inferred positions of the overhead objects as input, 2D inferred positions of the projection points of the overhead objects; and

providing, to one or more client devices, the predicted 2D inferred positions of the projection points of the overhead objects;

wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises normalizing the inferred positions of the overhead objects to a range between 0 and 1, and wherein predicting the 2D inferred positions of the projection points of the overhead objects comprises applying coordinate transformation to the normalized actual positions of the overhead objects, and wherein applying coordinate transformation comprises defining the normalized inferred positions of the overhead objects as relative to the bottom center of a bounding box for the overhead objects.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2023
From: ESEN, SERCAN; ÇILLIDAG, SERHAT; YAMAN, ALI UFUK; ULUSOY, OKAN; HATIPOGLU, POYRAZ UMUT
To: INTENSEYE, INC.
Reel/Frame 064138/0956 →
Continuity (1)
Related Publication 20240428446A1 · Dec 26, 2024
References Cited (28)
US 10482607B1 · Walters · 2019 [cited by examiner]
US 10528812B1 · Brouard · 2020 [cited by examiner]
US 11003956B2 · Weinzaepfel · 2021 [cited by examiner]
US 20120288189A1 · Hu · 2012 [cited by examiner]
US 20200140239A1 · Schoonmaker et al. · 2020 [cited by applicant]
US 20200302168A1 · Vo · 2020 [cited by examiner]
US 20210027598A1 · Tran et al. · 2021 [cited by applicant]
US 20210049780A1 · Westmacot · 2021 [cited by examiner]
US 20210407125A1 · Mahendran · 2021 [cited by examiner]
US 20220324679A1 · Benzing · 2022 [cited by examiner]
US 20230061389A1 · Bartek et al. · 2023 [cited by applicant]
US 20230153962A1 · Tang · 2023 [cited by examiner]
US 20230196188A1 · Horowitz · 2023 [cited by examiner]
US 20230326183A1 · Kah · 2023 [cited by examiner]
US 20230368414A1 · Afrooze · 2023 [cited by examiner]
US 20240281954A1 · Osman · 2024 [cited by examiner]
US 20240290027A1 · Bu · 2024 [cited by examiner]
US 20240428454A1 · Benkert · 2024 [cited by examiner]
CN 111461079A · 2020 [cited by examiner]
CN 113128346A · 2021 [cited by applicant]
CN 113656628A · 2021 [cited by applicant]
Chian et al. “Dynamic identification of crane load fall zone: A computer vision approach.” Safety science 156 (2022): 105904. (Year: 2022). [cited by examiner]
Ren et al. “Faster R-CNN: Towards real-time object detection with region proposal networks.” IEEE transactions on pattern analysis and machine intelligence 39.6 (2016): 1137-1149. (Year: 2016). [cited by examiner]
Yang et al. “Safety distance identification for crane drivers based on mask R-CNN.” Sensors 19.12 (2019): 2789. (Year: 2019). [cited by examiner]
Jeelani et al. “Real-time vision-based worker localization & hazard detection for construction.” Automation in Construction 121 (2021): 103448. (Year: 2021). [cited by examiner]
International Search Report and Written Opinion of the International Search Authority in international application No. PCT/US2024/035271, mailed on Sep. 23, 2024. [cited by applicant]
Pateraki et al. Crane Spreader Pose Estimation from a Single View. InVISIGRAPP (5: VISAPP) Jan. 2023 (pp. 796-805). https://www.researchgate.net/publication/367051971_Crane_Spreader_Pose_Estimation_from_a_Single_View. [cited by applicant]
Aaltonen. Evaluation of a computer vision model for a crane's target position measurement system (Master's thesis). published Apr. 2021. [cited by applicant]