IP Library › Granted Patent US 12,462,557
Granted Patent B2
US 12,462,557 · App. 17/853,445 · Granted Nov 4, 2025

Method, apparatus, and system for pole extraction from optical imagery

Inventors: Robert Ledner (Boulder, CO); Xiaoying Jin (Boulder, CO); Holly Russell (Lafayette, CO); Joseph Tankovich (Chicago, IL)
Assignee: HERE Global B.V.
G06V20/176G01C11/08G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,557
App. No.
17/853,445
Granted
Nov 4, 2025
Kind
B2
Abstract

An approach is provided for pole extraction from optical imagery. The approach involves, for instance, processing a plurality of images using a machine learning model to generate a plurality of redundant observations of a pole-like object and/or their semantic keypoints respectively depicted in the plurality of images. The approach also involves performing a photogrammetric triangulation of the plurality of redundant observations to determine three-dimensional coordinate data of the pole-like object and/or their semantic keypoints. The approach further involves providing the three-dimensional coordinate data of the pole-like object as an output.

Claims (36)

1 . A computer-implemented method comprising:

processing a plurality of images using a machine learning model to generate a plurality of redundant observations of a pole-like object respectively depicted in the plurality of images;

performing a photogrammetric triangulation of the plurality of redundant observations to determine three-dimensional coordinate data of the pole-like object, wherein the machine learning model generates the plurality of redundant observations based on detecting one or more semantic keypoints comprising a geometric representation of the pole-like object, and wherein the three-dimensional coordinate data include respective three-dimensional coordinates of the one or more semantic keypoints;

determining a location, a geometric attribute, or a combination thereof of the pole-like object based on the three-dimensional coordinate data; wherein the geometric attribute includes a length, an orientation, or a combination thereof; and

providing the three-dimensional coordinate data and the geometric attribute of the pole-like object as an output.

2 . The method of claim 1 , wherein the one or more semantic keypoints include a bottom point, a top point, a midpoint, or a combination thereof of a straight section of the pole-like object.

3 . The method of claim 1 , wherein the machine learning model is a Mask Region-Based Convolutional Neural Network (Mask R-CNN) that is trained using a plurality of reference images respectively labeled with a plurality of ground truth semantic keypoints of a plurality of pole-like objects.

4 . The method of claim 1 , wherein the plurality of pole-like objects in the plurality of reference images is respectively labeled with the plurality of ground truth semantic keypoints in a designated order.

5 . The method of claim 1 , further comprising:

back projecting the three-dimensional coordinate data into the plurality of images to perform a validation of the three-dimensional coordinated data,

wherein the validation is based on determining that the back-projected three-dimensional coordinate data falls within a threshold distance of the plurality of redundant observations.

6 . The method of claim 5 , wherein the validation is further based on at least one of (1) determining that the back-projected three-dimensional coordinate data corresponds to respective two-dimensional features in a minimum number of the plurality of images; (2) determining that a minimum intersection angle of confirmed rays of the photogrammetric triangulation associated with the three-dimensional coordinate data meets an angular threshold; or (3) determining that a maximum misclosure of the confirmed rays meets a distance threshold.

7 . The method of claim 1 , further comprising:

determining that the plurality of redundant observations correspond to a same object of the pole-like object based on a distance threshold from respective epipolar lines of the plurality of images.

8 . An apparatus comprising:

at least one processor; and

at least one memory including computer program code for one or more programs,

the at least one memory and the computer program code configured to, within the at least one processor, cause the apparatus to perform at least the following:

retrieve a machine learning model that is trained to detect one or more semantic keypoints associated with a pole-like object in a plurality of images;

perform a photogrammetric triangulation of the plurality of redundant observations to determine three-dimensional coordinate data of the pole-like object, the one or more semantic keypoints, or a combination thereof;

process the plurality of images using the machine learning model to generate a plurality of redundant observations of the pole-like object, wherein the plurality of redundant observations includes respective detections of the one or more semantic keypoints;

determine a location, a geometric attribute, or a combination thereof of the pole-like object based on the three-dimensional coordinate data; wherein the geometric attribute includes a length, an orientation, or a combination thereof; and

provide the plurality of redundant observations as an output.

9 . The apparatus of claim 8 , wherein the one or more semantic keypoints include a bottom point, a top point, a midpoint, or a combination thereof of a straight section of the pole-like object.

10 . The apparatus of claim 8 , wherein the machine learning model is trained using a plurality of reference images respectively labeled with the one or more semantic keypoints under one or more different contexts.

11 . The apparatus of claim 10 , wherein one or more semantic keypoints are labeled in a designated order.

12 . A non-transitory computer-readable storage medium carrying one or more sequences of one or more instructions which, when executed by one or more processors, cause an apparatus to perform:

processing a plurality of images using a machine learning model to generate a plurality of redundant observations of an object respectively depicted in the plurality of images;

performing a photogrammetric triangulation of the plurality of redundant observations to determine three-dimensional coordinate data of the object;

wherein the machine learning model generates the plurality of redundant observations based on detecting one or more semantic keypoints comprising a geometric representation of the pole-like object, and wherein the three-dimensional coordinate data include respective three-dimensional coordinates of the one or more semantic keypoints;

determining a location, a geometric attribute, or a combination thereof of the pole-like object based on the three-dimensional coordinate data; wherein the geometric attribute includes a length, an orientation, or a combination thereof; and

providing the three-dimensional coordinate data and the geometric attribute of the object as an output.

13 . The non-transitory computer-readable storage medium of claim 12 , wherein the machine learning model is a Mask Region-Based Convolutional Neural Network (Mask R-CNN) that is trained using a plurality of reference images respectively labeled with a plurality of ground truth semantic keypoints of a plurality of objects.

14 . The non-transitory computer-readable storage medium of claim 12 , wherein the apparatus is caused to further perform:

back projecting the three-dimensional coordinate data into the plurality of images to perform a validation of the three-dimensional coordinated data,

wherein the validation is based on determining the back-projected three-dimensional coordinate data falls within a threshold distance of the one or more redundant observations.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 19, 2022
From: LEDNER, ROBERT; JIN, XIAOYING; RUSSELL, HOLLY; TANKOVICH, JOSEPH
To: HERE GLOBAL B.V.
Reel/Frame 061139/0930 →
Continuity (2)
Provisional Application 63293360 · Dec 23, 2021
Related Publication 20230206625A1 · Jun 29, 2023
References Cited (23)
US 10846511B2 · Ozkucur et al. · 2020 [cited by applicant]
US 20180268256A1 · Di Febbo · 2018 [cited by examiner]
US 20190163990A1 · Mei · 2019 [cited by examiner]
US 20200134311A1 · Pojman · 2020 [cited by examiner]
US 20210020073A1 · Asmari et al. · 2021 [cited by applicant]
US 20210118165A1 · Strong · 2021 [cited by examiner]
US 20210166426A1 · McCormac · 2021 [cited by examiner]
US 20210201569A1 · Marschner · 2021 [cited by examiner]
US 20230206584A1 · Jin · 2023 [cited by examiner]
GB 2576322B · 2020 [cited by examiner]
Zhang et al, An enhanced multi-view vertical line locus matching algorithm of object space ground primitives based on positioning consistency for aerial and space images, ISPRS J. of Photogrammetry and Remote Sensing 13… [cited by examiner]
Mao et al, Deep Neural Networks for Road Sign Detection and Embedded Modeling Using Oblique Aerial Images, 2021, Remote Sensors, 13 (879): 1-24. (Year: 2021). [cited by examiner]
He et al., “Mask R-CNN”, Published in: 2017 IEEE International Conference on Computer Vision (ICCV), pp. 2961-2969. [cited by applicant]
Gomes et al., “Mapping Utility Poles in Aerial Orthoimages Using ATSS Deep Learning Method”, Letter, Published: Oct. 26, 2020, 14 pages. [cited by applicant]
Alam et al., “Automatic assessment and prediction of the resilience of utility poles using unmanned aerial vehicles and computer vision techniques”, International Journal of Disaster Risk Science, vol. 11, Feb. 19, 2020… [cited by applicant]
Zang et al., “Using deep learning to identify utility poles with crossarms and estimate their locations from google street view images”, Article, Sensors (Basel), Aug. 1, 2018;18(8):2484, 21 pages. [cited by applicant]
Yang et al., “A skeleton-based hierarchical method for detecting 3-d pole-like objects from mobile lidar point clouds”, Article, Published in: IEEE Geoscience and Remote Sensing Letters (vol. 16, Issue: 5, May 2019), pu… [cited by applicant]
Eigen et al., “Depth Map Prediction from a Single Image using a Multi-Scale Deep Network”, Article, NIPS'14: Proceedings of the 27th International Conference on Neural Information Processing Systems—vol. 2, Published:De… [cited by applicant]
Ranftl et al., “Vision Transformers for Dense Prediction”, International Conference on Computer Vision (ICCV), Mar. 24, 2021, 10 pages. [cited by applicant]
Godard et al., “Unsupervised Monocular Depth Estimation with Left-Right Consistency”, Published in: 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Sep. 13, 2016, 14 pages. [cited by applicant]
Godard et al., “Digging Into Self-Supervised Monocular Depth Estimation”, Published in: 2019 IEEE/CVF International Conference on Computer Vision (ICCV), Jun. 2018, 18 pages. [cited by applicant]
Office Action for related European Patent Application No. 22216109.3-1224, dated May 22, 2023, 9 pages. [cited by applicant]
He et al., “Mask R-CNN”, 2017 IEEE International Conference on Computer Vision (ICCV), Mar. 20, 2017, 9 pages. [cited by applicant]