IP Library Granted Patent US 12,307,710
Granted Patent B2
US 12,307,710 · App. 17/813,188 · Granted May 20, 2025

Machine-learned monocular depth estimation and semantic segmentation for 6-DOF absolute localization of a delivery drone

Inventor: Ali Shoeb (San Rafael, CA)
Assignee: Wing Aviation LLC
G06T7/74G01S19/485G05D1/101G06T7/50B64U50/19B64U2101/30B64U2101/60G06T2207/10032G06T2207/20081
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,307,710
App. No.
17/813,188
Granted
May 20, 2025
Kind
B2
Abstract

A method includes receiving a two-dimensional (2D) image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV. The method further includes applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, where the semantic image comprises one or more semantic labels. The method additionally includes retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels. The method also includes aligning the depth image of the environment with the reference depth data representative of the environment to determine a location of the UAV in the environment, where the aligning associates the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data.

Claims (43)

1. A method comprising:

receiving a two-dimensional (2D) image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV;

applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, wherein the machine learning model has been trained with a semantics branch to produce the semantic image and a depth branch to produce the depth image, and wherein the semantic image comprises one or more semantic labels;

retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; and

determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data; and

controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment.

2. The method of claim 1 , further comprising:

controlling the UAV to navigate in the environment using a Global Navigation Satellite System (GNSS) system;

detecting a disruption in service from the GNSS system, wherein the location of the UAV in the environment is determined responsive to detecting the disruption in service from the GNSS system; and

subsequent to detecting the disruption in service from the GNSS system, controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment.

3. The method of claim 1 , further comprising:

controlling the UAV to navigate in the environment using a GNSS system; and

using the determined location of the UAV in the environment to cross-check location data from the GNSS system.

4. The method of claim 1 , further comprising:

determining a GNSS location of the UAV in the environment using a GNSS system;

determining a refined location of the UAV in the environment based on the GNSS location of the UAV in the environment and the determined location of the UAV in the environment; and

controlling the UAV to navigate in the environment based on the refined location of the UAV in the environment.

5. The method of claim 1 , wherein the aligning the depth image of the environment with the reference depth data representative of the environment to determine the location of the UAV in the environment comprises using an iterative closest point (ICP) algorithm.

6. The method of claim 5 , wherein the ICP algorithm aligns points from the reference depth data with points from the depth image such that the reference semantic labels from the reference depth data correspond to the one or more semantic labels from the semantic image.

7. The method of claim 1 , wherein the semantic image of the environment and the depth image of the environment produced by the machine learning model have the same dimensions.

8. The method of claim 1 , wherein the camera on the UAV faces downward, and wherein the 2D image captured by the camera is representative of a terrain in the environment below the UAV.

9. The method of claim 1 , wherein the semantics branch and the depth branch of the machine learning model operate on a commonly generated feature set.

10. The method of claim 1 , wherein the machine learning model has been trained based on ground truth depth data, wherein the ground truth depth data is based on performance of a structure from motion (SfM) algorithm on images captured by one or more UAVs.

11. The method of claim 1 , wherein the machine learning model has been trained based on ground truth semantic data, wherein the ground truth semantic data is based on operator labeling of images captured by one or more UAVs.

12. The method of claim 1 , wherein the machine learning model has been trained using a scale invariant loss for training of the depth branch.

13. The method of claim 12 , further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to a median of a monocular depth map, wherein the monocular depth map is based on the reference depth data.

14. The method of claim 12 , further comprising applying a scale factor to the depth image, wherein the scale factor comprises a ratio of an altitude of the UAV above ground level relative to an above ground level estimate from a monocular depth map, wherein the monocular depth map is based on the reference depth data.

15. The method of claim 1 , wherein the one or more semantic labels are selected from a predetermined set of labels, wherein the predetermined set of labels comprises at least the following labels: building, road, vegetation, vehicle, driveway, lawn, and sidewalk.

16. The method of claim 1 , further comprising retrieving the reference depth data in advance of a flight of the UAV, wherein the reference depth data is selected based on a planned flight path of the UAV.

17. The method of claim 1 , further comprising applying a Kalman filter to the determined location of the UAV in the environment to control navigation of the UAV in the environment.

18. An unmanned aerial vehicle (UAV), comprising:

a camera; and

a control system configured to:

receive a two-dimensional (2D) image captured by the camera and representative of an environment of the UAV;

apply a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, wherein the machine learning model has been trained with a semantics branch to produce the semantic image and a depth branch to produce the depth image, and wherein the semantic image comprises one or more semantic labels;

retrieve reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; and

determine a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data.

19. A non-transitory computer readable medium comprising program instructions executable by one or more processors to perform operations, the operations comprising:

receiving a two-dimensional (2D) image captured by a camera on a unmanned aerial vehicle (UAV) and representative of an environment of the UAV;

applying a trained machine learning model to the 2D image to produce a semantic image of the environment and a depth image of the environment, wherein the machine learning model has been trained with a semantics branch to produce the semantic image and a depth branch to produce the depth image, and wherein the semantic image comprises one or more semantic labels;

retrieving reference depth data representative of the environment, wherein the reference depth data includes reference semantic labels; and

determining a location of the UAV in the environment, wherein determining the location of the UAV in the environment comprises aligning the depth image of the environment with the reference depth data representative of the environment, wherein aligning the depth image of the environment with the reference depth data representative of the environment is based on associating the one or more semantic labels from the semantic image with the reference semantic labels from the reference depth data; and

controlling the UAV to navigate in the environment based on the determined location of the UAV in the environment.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 19, 2022
From: SHOEB, ALI
To: WING AVIATION LLC
Reel/Frame 060552/0340 →
Continuity (1)
Related Publication 20240020876A1 · Jan 18, 2024
References Cited (26)
US 9359074B2 · Ganesh et al. · 2016 [cited by applicant]
US 9580173B1 · Burgess et al. · 2017 [cited by applicant]
US 10049589B1 · Boyd et al. · 2018 [cited by applicant]
US 10353388B2 · Schubert et al. · 2019 [cited by applicant]
US 10642284B1 · Barazovsky · 2020 [cited by examiner]
US 10679509B1 · Yarlagadda · 2020 [cited by applicant]
US 11195011B2 · Klaus · 2021 [cited by applicant]
US 20120195471A1 · Newcombe · 2012 [cited by examiner]
US 20130147911A1 · Karsch · 2013 [cited by examiner]
US 20180101782A1 · Gohl · 2018 [cited by examiner]
US 20180259960A1 · Cuban et al. · 2018 [cited by applicant]
US 20190286153A1 · Rankawat et al. · 2019 [cited by applicant]
US 20200241574A1 · Lin · 2020 [cited by examiner]
US 20210004976A1 · Guizilini · 2021 [cited by examiner]
US 20210041889A1 · Chen et al. · 2021 [cited by applicant]
US 20210150917A1 · Kubie · 2021 [cited by examiner]
US 20210174513A1 · Chidlovskii et al. · 2021 [cited by applicant]
US 20210224556A1 · Xu et al. · 2021 [cited by applicant]
US 20210240195A1 · Atherton · 2021 [cited by examiner]
US 20220026918A1 · Guizilini et al. · 2022 [cited by applicant]
US 20220111869A1 · Liu et al. · 2022 [cited by applicant]
US 20220343521A1 · Wofk · 2022 [cited by examiner]
US 20230406536A1 · Kojima · 2023 [cited by examiner]
CN 107444665 · 2020 [cited by applicant]
Guizilini et al., “Semantically-Guided Representation Learning for Self-Supervised Monocular Depth,” arXiv:2002.12319v1, ICLR 2020, 14 pages. [cited by applicant]
Parkison et al., “Semantic Iterative Closest Point through Expectation-Maximization,” Proceedings of the British Machine Vision Conference, 2018, 18 pages. [cited by applicant]