IP Library Granted Patent US 12,651,387
Granted Patent B2
US 12,651,387 · App. 18/054,585 · Granted Jun 9, 2026

Systems and methods for monocular based object detection

Inventors: Sean Foley (Atlanta, GA); James Hays (Decatur, GA)
Assignee: Ford Global Technologies, LLC
G06T11/60G01C21/3822G01C21/3859G05D1/0251G05D1/0274G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,651,387
App. No.
18/054,585
Granted
Jun 9, 2026
Kind
B2
Abstract

Disclosed herein are systems, methods, and computer program products for object detection. The methods comprise performing the following operations by a computing device: obtaining an image that comprises a plurality of layers superimposed on each other; identifying a center point of a robot on a map; selecting a portion of the map contained in a geometric shape overlaid on the map so as to have a center set to the center point of the robot; obtaining map information associated with the selected portion of the map; generating at least one additional layer using the map information; superimposing the at least one additional layer onto the image to generate a modified image; and performing an object detection algorithm to detect at least one object in the modified image.

Claims (51)

1 . A method for object detection, comprising:

obtaining, by a computing device, an image that comprises a plurality of layers superimposed on each other;

identifying, by the computing device, a center point of a robot on a map;

selecting, by the computing device, a portion of the map contained in a geometric shape overlaid on the map so as to have a center set to the center point of the robot;

obtaining, by the computing device, map information associated with the selected portion of the map;

generating, by the computing device, at least one additional layer using the map information;

superimposing, by the computing device, the at least one additional layer onto the image to generate a modified image; and

performing, by the computing device, an object detection algorithm to detect at least one object in the modified image.

2 . The method according to claim 1 , further comprising causing, by the computing device, control operations of the robot based on an output of the object detection algorithm.

3 . The method according to claim 1 , wherein the identifying the center point of the robot is based on pose information of the robot that comprises a location defined in 3D map coordinates, an angle of the robot relative to a reference point, and a pointing direction of the robot.

4 . The method according to claim 1 , wherein the selected portion of the map represents a geographic area encompassing the robot.

5 . The method according to claim 1 , further comprising performing machine learning operations to select a first combination of different types of map information when a first scenario exists for the robot and a second combination of different types of map information when a second scenario exists for the robot.

6 . The method according to claim 1 , wherein the map information comprises at least one of ground information, drivable geographical area information, ground depth information, map point distance-to-lane center information, lane direction information, and intersection information.

7 . The method according to claim 1 , wherein the at least one additional layer comprises a ground height layer, a ground depth layer, a drivable geographical area layer, a map point distance-to-lane center layer, a lane direction layer, or an intersection layer.

8 . The method according to claim 1 , wherein the generating the at least one additional layer comprises:

plotting the map information on a graph;

defining a 2D grid on the graph, the 2D grid comprising a plurality of cells with each said cell having four corners respectively associated with geometric points of the map;

using ground height values associated with the geometric points to define polygons in 3D space; and

projecting the polygons into a camera frame.

9 . The method according to claim 8 , further comprising setting a color value of each said projected polygon to (i) a color value associated with a select one of the geometric points used to define a corresponding one of the polygons or (ii) an average color value for the geometric points used to define a corresponding one of the polygons.

10 . The method according to claim 1 , wherein at least two different additional layers are generated using the map information and superimposed onto the image to generate the modified image.

11 . A system, comprising:

a processor;

a non-transitory computer-readable storage medium comprising programming instructions that are configured to cause the processor to implement a method for object detection, wherein the programming instructions comprise instructions to:

obtain an image that comprises a plurality of layers superimposed on each other;

identify a center point of a robot on a map;

select a portion of the map contained in a geometric shape overlaid on the map so as to have a center set to the center point of the robot;

obtain map information associated with the selected portion of the map;

generate at least one additional layer using the map information;

superimpose the at least one additional layer onto the image to generate a modified image; and

perform an object detection algorithm to detect at least one object in the modified image.

12 . The system according to claim 11 , wherein the programming instructions comprise instructions to control operations of the robot based on an output of the object detection algorithm.

13 . The system according to claim 11 , wherein the center point of the robot is identified based on pose information of the robot that comprises a location defined in 3D map coordinates, an angle of the robot relative to a reference point, and a pointing direction of the robot.

14 . The system according to claim 11 , wherein the selected portion of the map represents a geographic area encompassing the robot.

15 . The system according to claim 11 , wherein the programming instructions comprise instructions to perform machine learning operations to select a first combination of different types of map information when a first scenario exists for the robot and a second combination of different types of map information when a second scenario exists for the robot.

16 . The system according to claim 11 , wherein the map information comprises at least one of ground information, drivable geographical area information, ground depth information, map point distance-to-lane center information, lane direction information, and intersection information.

17 . The system according to claim 11 , wherein the at least one additional layer comprises a ground height layer, a ground depth layer, a drivable geographical area layer, a map point distance-to-lane center layer, a lane direction layer, or an intersection layer.

18 . The system according to claim 11 , wherein the at least one additional layer is generated by:

plotting the map information on a graph;

defining a 2D grid on the graph, the 2D grid comprising a plurality of cells with each said cell having four corners respectively associated with geometric points of the map;

using ground height values associated with the geometric points to define polygons in 3D space; and

projecting the polygons into a camera frame.

19 . The system according to claim 18 , wherein the programming instructions comprise instructions to set a color value of each said projected polygon to (i) a color value associated with a select one of the geometric points used to define a corresponding one of the polygons or (ii) an average color value for the geometric points used to define a corresponding one of the polygons.

20 . A non-transitory computer-readable medium that stores instructions that are configured to, when executed by at least one computing device, cause the at least one computing device to perform operations comprising:

obtaining an image that comprises a plurality of layers superimposed on each other;

identifying a center point of a robot on a map;

selecting a portion of the map contained in a geometric shape overlaid on the map so as to have a center set to the center point of the robot;

obtaining map information associated with the selected portion of the map;

generating at least one additional layer using the map information, wherein the additional layer includes a drivable geographical area layer;

superimposing the at least one additional layer onto the image to generate a modified image such that a least a portion of a non-drivable geographical area is blocked in the modified image; and

performing an object detection algorithm to detect at least one object in the modified image.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 9, 2023
From: ARGO AI, LLC
To: FORD GLOBAL TECHNOLOGIES, LLC
Reel/Frame 063025/0346 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: FOLEY, SEAN; HAYES, JAMES
To: ARGO AI, LLC
Reel/Frame 061733/0121 →
Continuity (2)
Continuation 17105199 · Nov 25, 2020
Related Publication 20230063845A1 · Mar 2, 2023
References Cited (17)
US 10410328B1 · Liu · 2019 [cited by examiner]
US 20020049532A1 · Nakamura · 2002 [cited by applicant]
US 20100217512A1 · Vu et al. · 2010 [cited by applicant]
US 20190220002A1 · Huang · 2019 [cited by examiner]
US 20200132490A1 · Yu · 2020 [cited by applicant]
US 20210256849A1 · Peranadam et al. · 2021 [cited by applicant]
IN 202027039221A · 2020 [cited by applicant]
International Preliminary Report on Patentability of PCT/US2021/057905 issued May 30, 2023, 5 pages. [cited by applicant]
Chabot, F. et al., “Deep MANTA: A Coarse-to-fine Many-Task Network for joint 2D and 3D Vehicle Analysis from Monocular Image”, (2017) arXiv:1703.07570v1 [cs.CV], available at https://arxiv.org/abs/1703.07570. [cited by applicant]
Chen, X. et al., “Monocular 3D Object Detection for Automous Driving”, International Conference on Computer Vision and Pattern Recognition (CVPR), 2016, available at https://www.cs.toronto.edu/˜urtasun/publications/chen… [cited by applicant]
Kim, Y. et al., “Deep Learning Based Vehicle Position and Orientation Estimation via Inverse Perspective Mapping Image”, 2019 IEEE Intelligent Vehicles Symposium (IV), available at https://ieeexplore.ieee.org/abstract/d… [cited by applicant]
Kundu, A. et al., “3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT, 2018, pp. 3559-3568, available at htt… [cited by applicant]
Mousavian, A. et al., “3D Bounding Box Estimation Using Deep Learning and Geometry”, (2017) 1612.00496v2 [cs.CV], available at https://arxiv.org/abs/1612.00496. [cited by applicant]
Roddick, T. et al., “Orthographic Feature Transform for Monocular 3D Object Detection”, 1811.08188v1 [cs.CV] (2018), available at https://arxiv.org/pdf/1811.08188.pdf. [cited by applicant]
Srivastava, S. et al., “Learning 2D to 3D Lifting for Object Detection in 3D for Autonomous Vehicles”, 1904.08494v2 [cs.CV] (2019), available at https://arxiv.org/abs/1904.08494. [cited by applicant]
Xu, B et al., “Multi-Level Fusion Based 3D Object Detection From Monocular Images”, 2018 IEEE/CVF Conference pn Computer Vision and Pattern Recognition, Salt Lake City, UT, 2018, pp. 2345-2353, available at https://open… [cited by applicant]
International Search Report and Written Opinion dated Dec. 27, 2021, issued in International Application No. PCT/2021/057905 (7 pages). [cited by applicant]