IP Library › Granted Patent US 12,361,570
Granted Patent B2
US 12,361,570 · App. 18/641,880 · Granted Jul 15, 2025

Three-dimensional object tracking using unverified detections registered by one or more sensors

Inventors: Daniel Forsgren (Enebyberg, SE); Anton Mikael Jansson (Stockholm, SE); Stein Norheim (Spånga, SE)
Assignee: Topgolf Sweden AB
G06T7/292G06T7/251G06T7/77G06V10/255G06V20/52G06V20/647G06T2207/10021G06T2207/10028G06T2207/20084G06T2207/30224G06T2207/30241
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,570
App. No.
18/641,880
Granted
Jul 15, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including medium-encoded computer program products, for three-dimensional object tracking includes, in at least one aspect, a method including: obtaining three-dimensional positions of objects registered by a detection system configured to allow more false positives so as to minimize false negatives, forming hypotheses using a filter that allows connections between registered objects when estimated three-dimensional velocity vectors roughly correspond to an object in motion in three-dimensional space, eliminating a proper subset of the hypotheses that are not further extended during the forming, specifying at least one three-dimensional track of at least one ball in motion in three-dimensional space by applying a full three-dimensional physics model to data for the three-dimensional positions used in the forming of at least one hypothesis that survives the eliminating, and outputting for display the at least one three-dimensional track of the at least one ball in motion in three-dimensional space.

Claims (103)

1. A method comprising:

obtaining three-dimensional positions of objects of interest registered by an object detection system configured to allow more false positives for the registered objects of interest so as to minimize false negatives;

forming hypotheses of objects in motion in three-dimensional space using a filter applied to the three-dimensional positions of the registered objects of interest, wherein the filter allows connections between specific ones of the registered objects of interest when estimated three-dimensional velocity vectors for the specific objects of interest roughly correspond to an object in motion in three-dimensional space across time;

eliminating a proper subset of the hypotheses that are not further extended by connections made, during the forming, with at least one additional object of interest registered by the object detection system;

specifying at least one three-dimensional track of at least one ball in motion in three-dimensional space by applying a full three-dimensional physics model to data for the three-dimensional positions of the registered objects of interest used in the forming of at least one hypothesis that survives the eliminating; and

outputting for display the at least one three-dimensional track of the at least one ball in motion in three-dimensional space;

wherein the obtaining comprises receiving or generating the three-dimensional positions of the registered objects of interest in which a majority of the registered objects of interest are false positives including false detections by individual sensors and false combinations of detections from respective ones of the individual sensors including combined detections that are not from a same object.

2. The method of claim 1 , wherein forming the hypotheses of objects in motion in three-dimensional space using the filter comprises:

using a first model of motion in three-dimensional space when a number of registered detections for a given hypothesis is less than a threshold value, wherein the first model provides a very loose filter for what constitutes an object in Newtonian motion; and

using a second model of motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value, wherein the second model incorporates some knowledge of 3D physics.

3. The method of claim 2 , wherein the first model of motion in three-dimensional space is a linear model, and the second model of motion in three-dimensional space is a recurrent neural network model.

4. The method of claim 2 , wherein forming the hypotheses of objects in motion in three-dimensional space comprises:

predicting a next three-dimensional position for a registered detection for the given hypothesis using the first model or the second model, in accordance with the number of registered detections for the given hypothesis and the threshold value;

searching a spatial data structure containing all of the three-dimensional positions of objects of interest registered by the object detection system for a next time slice to find a set of three-dimensional positions within a predefined distance of the predicted next three-dimensional position;

using the predicted next three-dimensional position for the given hypothesis when the set of three-dimensional positions is a null set;

using a single three-dimensional position for the given hypothesis when the set of three-dimensional positions contains only the single three-dimensional position; and

when the set of three-dimensional positions contains two or more three-dimensional positions,

sorting the two or more three-dimensional positions based on proximity to the predicted next three-dimensional position to form a sorted set,

removing any less proximate three-dimensional positions in the sorted set beyond a predefined threshold number greater than two, and

branching the given hypothesis into two or more hypotheses in a group of hypotheses using the two or more three-dimensional positions remaining in the sorted set.

5. The method of claim 2 , wherein forming the hypotheses of objects in motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value comprises:

identifying a data point in the given hypothesis that has a local minimum vertical position;

checking respective vertical components of estimated three-dimensional velocity vectors before and after the data point in the given hypothesis; and

designating the data point as a ground impact when a first of the respective vertical components is negative and a second of the respective vertical components is positive.

6. The method of claim 5 , wherein the respective vertical components are from averages of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and the method comprises:

selecting at least one time window based on a noise level associated with the estimated three-dimensional velocity vectors, a minimum expected flight time on one or both sides of the data point, or both the noise level and the minimum expected flight time; and

computing the averages of the estimated three-dimensional velocity vectors before and after the data point that fall within the at least one time window.

7. The method of claim 5 , wherein forming the hypotheses of objects in motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value comprises dividing the given hypothesis into distinct segments comprising a flight segment, one or more bounce segments, and a roll segment, and the dividing comprises the identifying, the checking, the designating, and:

classifying a first segment of the given hypothesis before a first designated ground impact as the flight segment;

classifying a next segment of the given hypothesis after each next designated ground impact as one of the one or more bounce segments when an angle between a first estimated velocity vector before the next designated ground impact and a second estimated velocity vector after the next designated ground impact is greater than a threshold angle; and

classifying the next segment of the given hypothesis after a next designated ground impact as the roll segment when the angle between the first estimated velocity vector before the next designated ground impact and the second estimated velocity vector after the next designated ground impact is less than or equal to the threshold angle.

8. The method of claim 7 , wherein the given hypothesis is the at least one hypothesis that survives the eliminating, and specifying the at least one three-dimensional track of the at least one ball in motion in three-dimensional space comprises:

generating the data for the three-dimensional positions of the registered objects used in the forming of the at least one hypothesis by triangulating a more accurate 3D path using sensor observations registered by the object detection system in at least the flight segment; and

confirming the at least one ball in motion by applying the full three-dimensional physics model to the generated data for at least the flight segment.

9. The method of claim 1 , wherein the individual sensors comprise three or more sensors, and the obtaining comprises generating the three-dimensional positions by making respective combinations of the detections from respective pairs of the three or more sensors for a current time slice.

10. The method of claim 9 , wherein the three or more sensors comprise two cameras and a radar device, and the making comprises:

combining detections of a single object of interest by the two cameras using a stereo pairing of the two cameras to generate a first three-dimensional position of the single object of interest; and

combining detections of the single object of interest by the radar device and at least one of the two cameras to generate a second three-dimensional position of the single object of interest.

11. The method of claim 10 , wherein the three or more sensors comprise an additional camera, and the making comprises combining detections of the single object of interest by the additional camera and the two cameras using stereo pairings of the additional camera with each of the two cameras to generate a third three-dimensional position of the single object of interest and a fourth three-dimensional position of the single object of interest.

12. A system comprising:

two or more sensors and one or more first computers of an object detection system configured to register objects of interest and to allow more false positives for the registered objects of interest so as to minimize false negatives; and

one or more second computers configured to perform operations comprising

obtaining three-dimensional positions of objects of interest registered by the object detection system,

forming hypotheses of objects in motion in three-dimensional space using a filter applied to the three-dimensional positions of the registered objects of interest, wherein the filter allows connections between specific ones of the registered objects of interest when estimated three-dimensional velocity vectors for the specific objects of interest roughly correspond to an object in motion in three-dimensional space across time,

eliminating a proper subset of the hypotheses that are not further extended by connections made, during the forming, with at least one additional object of interest registered by the object detection system,

specifying at least one three-dimensional track of at least one ball in motion in three-dimensional space by applying a full three-dimensional physics model to data for the three-dimensional positions of the registered objects of interest used in the forming of at least one hypothesis that survives the eliminating, and

outputting for display the at least one three-dimensional track of the at least one ball in motion in three-dimensional space;

wherein the obtaining comprises receiving or generating the three-dimensional positions of the registered objects of interest in which a majority of the registered objects of interest are false positives including false detections by individual ones of the two or more sensors and false combinations of detections from respective ones of the two or more sensors including combined detections that are not from a same object.

13. The system of claim 12 , wherein the two or more sensors comprise cameras and at least one radar device.

14. The system of claim 12 , wherein the two or more sensors comprise a hybrid camera-radar sensor.

15. The system of claim 12 , wherein forming the hypotheses of objects in motion in three-dimensional space using the filter comprises:

using a first model of motion in three-dimensional space when a number of registered detections for a given hypothesis is less than a threshold value, wherein the first model provides a very loose filter for what constitutes an object in Newtonian motion; and

using a second model of motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value, wherein the second model incorporates some knowledge of 3D physics.

16. The system of claim 15 , wherein the first model of motion in three-dimensional space is a linear model, and the second model of motion in three-dimensional space is a recurrent neural network model.

17. The system of claim 15 , wherein forming the hypotheses of objects in motion in three-dimensional space comprises:

predicting a next three-dimensional position for a registered detection for the given hypothesis using the first model or the second model, in accordance with the number of registered detections for the given hypothesis and the threshold value;

searching a spatial data structure containing all of the three-dimensional positions of objects of interest registered by the object detection system for a next time slice to find a set of three-dimensional positions within a predefined distance of the predicted next three-dimensional position;

using the predicted next three-dimensional position for the given hypothesis when the set of three-dimensional positions is a null set;

using a single three-dimensional position for the given hypothesis when the set of three-dimensional positions contains only the single three-dimensional position; and

when the set of three-dimensional positions contains two or more three-dimensional positions,

sorting the two or more three-dimensional positions based on proximity to the predicted next three-dimensional position to form a sorted set,

removing any less proximate three-dimensional positions in the sorted set beyond a predefined threshold number greater than two, and

branching the given hypothesis into two or more hypotheses in a group of hypotheses using the two or more three-dimensional positions remaining in the sorted set.

18. The system of claim 15 , wherein forming the hypotheses of objects in motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value comprises:

identifying a data point in the given hypothesis that has a local minimum vertical position;

checking respective vertical components of estimated three-dimensional velocity vectors before and after the data point in the given hypothesis; and

designating the data point as a ground impact when a first of the respective vertical components is negative and a second of the respective vertical components is positive.

19. The system of claim 12 , wherein the two or more sensors comprise three or more sensors, and the obtaining comprises generating the three-dimensional positions by making respective combinations of the detections from respective pairs of the three or more sensors for a current time slice.

20. The system of claim 19 , wherein the three or more sensors comprise two cameras and a radar device, and the making comprises:

combining detections of a single object of interest by the two cameras using a stereo pairing of the two cameras to generate a first three-dimensional position of the single object of interest; and

combining detections of the single object of interest by the radar device and at least one of the two cameras to generate a second three-dimensional position of the single object of interest.

21. The system of claim 20 , wherein the three or more sensors comprise an additional camera, and the making comprises combining detections of the single object of interest by the additional camera and the two cameras using stereo pairings of the additional camera with each of the two cameras to generate a third three-dimensional position of the single object of interest and a fourth three-dimensional position of the single object of interest.

22. A non-transitory computer-readable medium encoding instructions that cause data processing apparatus associated with an object detection system comprising two or more sensors to perform operations comprising:

obtaining three-dimensional positions of objects of interest registered by the object detection system, which is configured to allow more false positives for the registered objects of interest so as to minimize false negatives;

forming hypotheses of objects in motion in three-dimensional space using a filter applied to the three-dimensional positions of the registered objects of interest, wherein the filter allows connections between specific ones of the registered objects of interest when estimated three-dimensional velocity vectors for the specific objects of interest roughly correspond to an object in motion in three-dimensional space across time;

eliminating a proper subset of the hypotheses that are not further extended by connections made, during the forming, with at least one additional object of interest registered by the object detection system;

specifying at least one three-dimensional track of at least one ball in motion in three-dimensional space by applying a full three-dimensional physics model to data for the three-dimensional positions of the registered objects of interest used in the forming of at least one hypothesis that survives the eliminating; and

outputting for display the at least one three-dimensional track of the at least one ball in motion in three-dimensional space;

wherein the obtaining comprises receiving or generating the three-dimensional positions of the registered objects of interest in which a majority of the registered objects of interest are false positives including false detections by individual sensors and false combinations of detections from respective ones of the individual sensors including combined detections that are not from a same object.

23. The non-transitory computer-readable medium of claim 22 , wherein forming the hypotheses of objects in motion in three-dimensional space using the filter comprises:

using a first model of motion in three-dimensional space when a number of registered detections for a given hypothesis is less than a threshold value, wherein the first model provides a very loose filter for what constitutes an object in Newtonian motion; and

using a second model of motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value, wherein the second model incorporates some knowledge of 3D physics.

24. The non-transitory computer-readable medium of claim 23 , wherein the first model of motion in three-dimensional space is a linear model, and the second model of motion in three-dimensional space is a recurrent neural network model.

25. The non-transitory computer-readable medium of claim 23 , wherein forming the hypotheses of objects in motion in three-dimensional space comprises:

predicting a next three-dimensional position for a registered detection for the given hypothesis using the first model or the second model, in accordance with the number of registered detections for the given hypothesis and the threshold value;

searching a spatial data structure containing all of the three-dimensional positions of objects of interest registered by the object detection system for a next time slice to find a set of three-dimensional positions within a predefined distance of the predicted next three-dimensional position;

using the predicted next three-dimensional position for the given hypothesis when the set of three-dimensional positions is a null set;

using a single three-dimensional position for the given hypothesis when the set of three-dimensional positions contains only the single three-dimensional position; and

when the set of three-dimensional positions contains two or more three-dimensional positions,

sorting the two or more three-dimensional positions based on proximity to the predicted next three-dimensional position to form a sorted set,

removing any less proximate three-dimensional positions in the sorted set beyond a predefined threshold number greater than two, and

branching the given hypothesis into two or more hypotheses in a group of hypotheses using the two or more three-dimensional positions remaining in the sorted set.

26. The non-transitory computer-readable medium of claim 23 , wherein forming the hypotheses of objects in motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value comprises:

identifying a data point in the given hypothesis that has a local minimum vertical position;

checking respective vertical components of estimated three-dimensional velocity vectors before and after the data point in the given hypothesis; and

designating the data point as a ground impact when a first of the respective vertical components is negative and a second of the respective vertical components is positive.

27. The non-transitory computer-readable medium of claim 26 , wherein the respective vertical components are from averages of the estimated three-dimensional velocity vectors before and after the data point in the given hypothesis, and the operations comprise:

selecting at least one time window based on a noise level associated with the estimated three-dimensional velocity vectors, a minimum expected flight time on one or both sides of the data point, or both the noise level and the minimum expected flight time; and

computing the averages of the estimated three-dimensional velocity vectors before and after the data point that fall within the at least one time window.

28. The non-transitory computer-readable medium of claim 26 , wherein forming the hypotheses of objects in motion in three-dimensional space when the number of registered detections for the given hypothesis is greater than or equal to the threshold value comprises dividing the given hypothesis into distinct segments comprising a flight segment, one or more bounce segments, and a roll segment, and the dividing comprises the identifying, the checking, the designating, and:

classifying a first segment of the given hypothesis before a first designated ground impact as the flight segment;

classifying a next segment of the given hypothesis after each next designated ground impact as one of the one or more bounce segments when an angle between a first estimated velocity vector before the next designated ground impact and a second estimated velocity vector after the next designated ground impact is greater than a threshold angle; and

classifying the next segment of the given hypothesis after a next designated ground impact as the roll segment when the angle between the first estimated velocity vector before the next designated ground impact and the second estimated velocity vector after the next designated ground impact is less than or equal to the threshold angle.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 6, 2025
From: FORSGREN, DANIEL; JANSSON, ANTON MIKAEL; NORHEIM, STEIN
To: TOPGOLF SWEDEN AB
Reel/Frame 071337/0938 →
Continuity (3)
Continuation 17516637 · Nov 1, 2021
Provisional Application 63109290 · Nov 3, 2020
Related Publication 20240273738A1 · Aug 15, 2024
References Cited (53)
US 5938545A · Cooper et al. · 1999 [cited by applicant]
US 8335345B2 · White et al. · 2012 [cited by applicant]
US 8670604B2 · Eggert et al. · 2014 [cited by applicant]
US 9036864B2 · Johnson et al. · 2015 [cited by applicant]
US 9283464B2 · Nipper et al. · 2016 [cited by applicant]
US 10444339B2 · Tuxen et al. · 2019 [cited by applicant]
US 10596416B2 · Forsgren · 2020 [cited by applicant]
US 10751569B2 · Guerci et al. · 2020 [cited by applicant]
US 10765944B2 · Dantas de Castro · 2020 [cited by applicant]
US 10803598B2 · Chaurasia et al. · 2020 [cited by applicant]
US 10898757B1 · Johansson et al. · 2021 [cited by applicant]
US 10987566B2 · Ferras · 2021 [cited by applicant]
US 10989791B2 · Tuxen et al. · 2021 [cited by applicant]
US 11537819B1 · Das · 2022 [cited by examiner]
US 11995846B2 · Forsgren et al. · 2024 [cited by applicant]
US 20080240497A1 · Porikli et al. · 2008 [cited by applicant]
US 20140163915A1 · Baek · 2014 [cited by examiner]
US 20160110913A1 · Kosoy et al. · 2016 [cited by applicant]
US 20160193501A1 · Nipper et al. · 2016 [cited by applicant]
US 20160306036A1 · Johnson · 2016 [cited by applicant]
US 20160320476A1 · Johnson · 2016 [cited by examiner]
US 20180101732A1 · Uchiyama · 2018 [cited by examiner]
US 20180197296A1 · Liu et al. · 2018 [cited by applicant]
US 20180357472A1 · Dreessen · 2018 [cited by applicant]
US 20190147219A1 · Thornbrue · 2019 [cited by examiner]
US 20200143094A1 · Haaland et al. · 2020 [cited by applicant]
US 20200164258A1 · Tuxen · 2020 [cited by examiner]
US 20200391077A1 · Forsgren · 2020 [cited by applicant]
US 20210069569A1 · Oh et al. · 2021 [cited by applicant]
US 20210220718A1 · Tuxen et al. · 2021 [cited by applicant]
US 20210264141A1 · Chojnacki · 2021 [cited by examiner]
US 20210275873A1 · Johansson et al. · 2021 [cited by applicant]
US 20220138969A1 · Forsgren · 2022 [cited by examiner]
JP 5180733 · 2013 [cited by applicant]
KR 20170092929A · 2017 [cited by applicant]
Ren, Jinchang, et al. “Real-time modeling of 3-D soccer ball trajectories from multiple fixed cameras.” IEEE Transactions on Circuits and Systems for Video Technology 18.3 (2008): 350-362. (Year: 2008). [cited by examiner]
Kittler, Josef, et al. “A memory architecture and contextual reasoning framework for cognitive vision.” Image Analysis: 14th Scandinavian Conference, SCIA 2005, Joensuu, Finland, Jun. 19-22, 2005. Proceedings 14. Spring… [cited by examiner]
Birbach, Oliver, and Udo Frese. “A multiple hypothesis approach for a ball tracking system.” International Conference on Computer Vision Systems. Berlin, Heidelberg: Springer Berlin Heidelberg, 2009. (Year: 2009). [cited by examiner]
Botha, Frik J., Corne E. van Daalen, and Johann Treurnicht. “Data fusion of radar and stereo vision for detection and tracking of moving objects.” 2016 Pattern Recognition Association of South Africa and Robotics and Me… [cited by examiner]
Ankur Handa, “Analysing High Frame-Rate Camera Tracking”, Imperial College London, Department of Computing, Sep. 2013, 224 pages. [cited by applicant]
Birbach et al., “A multiple hypothesis approach for a ball tracking system,” International Conference on Computer Vision Systems, Oct. 2009, 11 pages. [cited by applicant]
Birbach et al., “Estimation and prediction of multiple flying balls using probability hypothesis density filtering,” IEEE/RSJ International Conference on Intelligent Robots and Systems, Sep. 2011, 8 pages. [cited by applicant]
Botha et al., “Data fusion of radar and stereo vision for detection and tracking of moving objects,” 2016 Pattern Recognition Association of South Africa and Robotics and Mechatronics International Conference (PRASA-Rob… [cited by applicant]
Denny Britz, “WILDML Artificial Intelligence, Deep Learning, and NPL”, Recurrent Neural Networks Tutorial, Part 1—Introduction to RNNs—WildML, www.wildml.com/2015/09/recurrent-neural-networks-tutorial-part-1-intruductio… [cited by applicant]
International Search Report and Written Opinion in International Appln. No. PCT/EP2021/079919, dated Feb. 9, 2022, 16 pages. [cited by applicant]
Ishii et al., “3D Tracking of a Soccer Ball Using Two Synchronized Cameras”, University of Tsukuba, Graduate School of System and Information Engineering, Dec. 2007, 11 pages. [cited by applicant]
Jansson, “Predicting Trajectories of Golf Balls Using Recurrent Neural Networks,” Thesis for the degree of Master in Computer Science, KTH Royal Institute of Technology, School of Computer Science and Communication, Jun… [cited by applicant]
Joo et al., “A Multiple-Hypothesis Approach for Multiobject Visual Tracking,” in IEEE Transactions on Image Processing, Nov. 2007, 16(11):2849-2854. [cited by applicant]
Kamble et al., “Ball tracking in sports: a survey,” Artificial Intelligence Review, Oct. 2017, 52: 51 pages. [cited by applicant]
Peter Sodergren, “High accuracy stereo Structure from Motion through symmetry cost functions,” KTH Computer Science and Communication, Master's Thesis at CSC, Apr. 14, 2014, 81 pages. [cited by applicant]
Ren et al. “Real-time modeling of 3-D soccer ball trajectories from multiple fixed cameras,” IEEE Transactions on Circuits and Systems for Video Technology 18.3 (2008):350-362. [cited by applicant]
Office Action in Korean Appln. No. 20237015000, dated Mar. 17, 2025, 4 pages (with English translation). [cited by applicant]
Wikipedia, “Kalman filter”, https://en.wikipedia.org/wiki/Kalman_filter, Nov. 2, 2020, 37 pages. [cited by applicant]
Cited By (1)
US 12,582,870