Multi-modal contextualization for object detection
Techniques for multi-modal contextualization of object detection for use with an agricultural vehicle are described herein. The techniques can provide additional context information to radar detection to assess the likelihood of an object detected by radar of being an object of interest. Examples of detection that might be detected by a radar sensor that, based on the additional context information, the agricultural vehicle can make a determination to ignore can include ground targets, ghost targets, side or overhead reflections from obstacles, detections from tall crops/weeds in a field.
1 . A system for autonomously-operating an agricultural vehicle, the system comprising:
one or more a radar detectors coupled to the agricultural vehicle to generate radar data including a first set of one or more objects in a navigation path of the agricultural vehicle;
a camera coupled to the agricultural vehicle to generate image data;
a 3D sensor coupled to the agricultural vehicle to generate 3D sensor data; and
a processor communicatively coupled to the radar detector, camera, and 3D sensor, the processor configured to:
generating a view including a second set of one or more objects in the navigation path of the autonomously-operated vehicle based on the image data, 3D sensor data, and a model, comprising:
inputting the image data into a machine learning (ML) model trained to identify drivable and non-drivable objects in the image data to generate a drivable area model image including the second set of one or more objects; and
projecting the drivable area model image into three dimensions based on the 3D sensor data to generate a 3D projected drivable area model image;
converting the 3D projected drivable area model image into a 3D point cloud including the second set of one or more objects;
comparing the first set of one or more objects and the second set of one or more objects based on the radar data and the 3D point cloud to generated augmented radar detection data;
determining whether a corresponding non-drivable target is present in the 3D point cloud;
in an event there is a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed non-drivable object in the augmented radar detection data; and
in an event there is not a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed drivable object in the augmented radar detection data; and
outputting the augmented radar detection data to control navigation of the agricultural vehicle.
2 . The system of claim 1 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on presence of the non-drivable target within a circle with a radius of a predetermined length.
3 . The system of claim 1 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on at least one of min-max-average-st dev values of drivable/non-drivable area projection in a circle of association, quantity of non-drivable area points in the circle of association; proximity of the respective object in the first set to the non-drivable area target; the respective object in a first set return signal power, probability of target existence, standard deviation of measured angle, speed and distance, history of target, and target ID; or tracking accuracy of targets through linear projections.
4 . The system of claim 1 , wherein the 3D sensor includes a stereo camera system and the 3D sensor data includes a disparity image.
5 . The system of claim 1 , wherein the 3D sensor includes a lidar sensor.
6 . The system of claim 1 , wherein the model includes a deep learning model trained to detect drivable and non-drivable areas in images.
7 . A method for operating an agricultural vehicle, the method comprising:
receiving radar data including a first set of one or more objects in a navigation path of the agricultural vehicle;
receiving image data;
receiving 3D sensor data;
generating a view including a second set of one or more objects in the navigation path of the agricultural vehicle based on the image data, 3D sensor data, and a model, comprising:
inputting the image data into a machine learning (ML) model trained to identify drivable and non-drivable objects in the image data to generate a drivable area model image including the second set of one or more objects; and
projecting the drivable area model image into three dimensions based on the 3D sensor data to generate a 3D projected drivable area model image;
converting the 3D projected drivable area model image into a 3D point cloud including the second set of one or more objects;
comparing the first set of one or more objects and the second set of one or more objects based on the radar data and the 3D point cloud to generated augmented radar detection data;
determining whether a corresponding non-drivable target is present in the 3D point cloud;
in an event there is a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed non-drivable object in the augmented radar detection data; and
in an event there is not a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed drivable object in the augmented radar detection data; and
outputting the augmented radar detection data to control navigation of the agricultural vehicle.
8 . The method of claim 7 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on presence of the non-drivable target within a circle with a radius of a predetermined length.
9 . The method of claim 7 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on at least one of min-max-average-st dev values of drivable/non-drivable area projection in a circle of association, quantity of non-drivable area points in the circle of association; proximity of the respective object in the first set to the non-drivable area target; the respective object in a first set return signal power, probability of target existence, standard deviation of measured angle, speed and distance, history of target, and target ID; or tracking accuracy of targets through linear projections.
10 . The method of claim 7 , wherein the 3D sensor data includes a disparity image captured by a stereo camera system.
11 . The method of claim 7 , wherein the 3D sensor data includes lidar data.
12 . The method of claim 7 , wherein the model includes a deep learning model trained to detect drivable and non-drivable areas in images.
13 . A machine-readable medium embodying instructions that, when executed by a machine, cause the machine to perform operations comprising:
receiving radar data including a first set of one or more objects in a navigation path of an agricultural vehicle;
receiving image data;
receiving 3D sensor data;
generating a view including a second set of one or more objects in the navigation path of the agricultural vehicle based on the image data, 3D sensor data, and a model, comprising:
inputting the image data into a machine learning (ML) model trained to identify drivable and non-drivable objects in the image data to generate a drivable area model image including the second set of one or more objects; and
projecting the drivable area model image into three dimensions based on the 3D sensor data to generate a 3D projected drivable area model image;
converting the 3D projected drivable area model image into a 3D point cloud including the second set of one or more objects;
comparing the first set of one or more objects and the second set of one or more objects based on the radar data and the 3D point cloud to generated augmented radar detection data;
determining whether a corresponding non-drivable target is present in the 3D point cloud;
in an event there is a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed non-drivable object in the augmented radar detection data; and
in an event there is not a corresponding non-drivable target present in the 3D point cloud, identifying the respective object in the first set as a confirmed drivable object in the augmented radar detection data; and
outputting the augmented radar detection data to control navigation of the agricultural vehicle.
14 . The machine-readable medium of claim 13 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on presence of the non-drivable target within a circle with a radius of a predetermined length.
15 . The machine-readable medium of claim 13 , wherein determining whether a corresponding non-drivable target is present in the 3D point cloud is based on at least one of min-max-average-st dev values of drivable/non-drivable area projection in a circle of association, quantity of non-drivable area points in the circle of association; proximity of the respective object in the first set to the non-drivable area target; the respective object in a first set return signal power, probability of target existence, standard deviation of measured angle, speed and distance, history of target, and target ID; or tracking accuracy of targets through linear projections.
16 . The machine-readable medium of claim 13 , wherein the 3D sensor data includes a disparity image captured by a stereo camera system.
17 . The machine-readable medium of claim 13 , wherein the 3D sensor data includes lidar data.
18 . The machine-readable medium of claim 13 , wherein the model includes a deep learning model trained to detect drivable and non-drivable areas in images.