IP Library › Granted Patent US 12,462,417
Granted Patent B2
US 12,462,417 · App. 18/080,482 · Granted Nov 4, 2025

Method and electronic device for 3D object detection using neural networks

Inventors: Danila Dmitrievich Rukhovich (Moscow, RU); Anna Borisovna Vorontsova (Moscow, RU); Anton Sergeevich Konushin (Moscow, RU)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G06T7/70G06T17/00G06V10/771G06V10/82G06T2207/20081G06T2207/30252G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,417
App. No.
18/080,482
Granted
Nov 4, 2025
Kind
B2
Abstract

A method of 3D object detection using an object detection neural network includes: receiving one or more monocular images; extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through a 2D feature extracting part of the object detection neural network, generating an averaged 3D voxel volume based on the 2D feature maps, extracting a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of a 3D feature extracting part of the object detection neural network, and performing 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through thane outdoor object detecting part of the object detection neural network.

Claims (65)

1 . A method of three-dimensional (3D) object detection using an object detection neural network comprising a two-dimensional (2D) feature extracting part, a 3D feature extracting part, and an outdoor object detecting part comprising parallel 2D convolutional layers for classification and location, which are pre-trained in end-to-end manner based on posed monocular images, the method comprising:

receiving one or more monocular images;

extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through the 2D feature extracting part of the object detection neural network,

generating an averaged 3D voxel volume based on the 2D feature maps,

extracting a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of the 3D feature extracting part of the object detection neural network, and

performing 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through the outdoor object detecting part of the object detection neural network.

2 . The method of claim 1 , wherein the method is performed by an electronic device of a vehicle comprising at least one camera configured to detect one or more 3D objects in a surrounding area of the vehicle.

3 . The method of claim 1 , further comprising, for each image of the one or more monocular images, aggregating features in the 2D feature maps corresponding to the respective image via a Feature Pyramid Network (FPN).

4 . The method of claim 1 , wherein the generating the averaged 3D voxel volume further comprises:

for each image of the one or more monocular images, generating a 3D voxel volume corresponding to the respective image using a pinhole camera model that determines a correspondence between 2D coordinates in the corresponding 2D feature maps and 3D coordinates in the 3D voxel volume;

for each of the 3D voxel volume, defining a binary mask corresponding to the respective 3D voxel volume, wherein the binary mask indicates, for each voxel in the 3D voxel volume, whether the respective voxel is inside a camera frustrum of the corresponding image;

for each of the 3D voxel volume, the corresponding binary mask, and the corresponding 2D feature maps, projecting features of the 2D feature maps for each of the voxel inside in the 3D voxel volume as defined by the binary mask;

producing an aggregated binary mask by aggregating the binary masks defined for the 3D voxel volumes of all of the one or more monocular images; and

generating the averaged 3D voxel volume by averaging the features projected in the 3D voxel volumes of all of the one or more monocular images for each of the voxel inside in the 3D voxel volume in the averaged 3D voxel volume as defined by the aggregated binary mask.

5 . The method of claim 1 , wherein the object detection neural network is trained on one or more outdoor training datasets by optimizing a total outdoor loss function based at least on a location loss, a focal loss for classification, and a cross-entropy loss for direction.

6 . A method of three-dimensional (3D) object detection using an object detection neural network comprising a two-dimensional (2D) feature extracting part, a 3D feature extracting part, and an indoor object detecting part comprising 3D convolutional layers for classification, centerness, and location, which are pre-trained in end-to-end manner based on posed monocular images, the method comprising:

receiving one or more monocular images,

extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through the 2D feature extracting part of the object detection neural network,

generating an averaged 3D voxel volume based on the 2D feature maps,

extracting refined 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through the 3D feature extracting part of the object detection neural network, and

performing 3D object detection by passing the refined 3D feature maps through the indoor object detecting part of the object detection neural network using dense voxel representation of intermediate features.

7 . The method of claim 6 , wherein the method is performed by an electronic device comprising at least one camera configured to detect one or more 3D objects in a surrounding area of the electronic device, and

wherein the electronic device is one of a smartphone, a tablet, smart glasses, or a mobile robot.

8 . The method of claim 6 , further comprising, for each image of the one or more monocular images, aggregating features in the 2D feature maps corresponding to the image via a Feature Pyramid Network (FPN).

9 . The method of claim 6 , wherein the generating the averaged 3D voxel volume further comprises:

for each image of the one or more monocular images, generating a 3D voxel volume corresponding to the respective image using a pinhole camera model that determines a correspondence between 2D coordinates in the corresponding 2D feature maps and 3D coordinates in the 3D voxel volume;

for each of the 3D voxel volume, defining a binary mask corresponding to the respective 3D voxel volume, wherein the binary mask indicates, for each voxel in the 3D voxel volume, whether the respective voxel is inside a camera frustrum of the corresponding image;

for each of the 3D voxel volume, the corresponding binary mask, and the corresponding 2D feature maps, projecting features of the 2D feature maps for each of the voxel inside in the 3D voxel volume as defined by the binary mask;

producing an aggregated binary mask by aggregating the binary masks defined for the 3D voxel volumes of all of the one or more monocular images; and

generating the averaged 3D voxel volume by averaging the features projected in the 3D voxel volumes of all of the one or more monocular images for each of the voxel inside in the 3D voxel volume in the averaged 3D voxel volume as defined by the aggregated binary mask.

10 . The method of claim 6 , wherein the object detection neural network is trained on one or more indoor training datasets by optimizing a total indoor loss function based on at least one of a focal loss for classification, a cross-entropy loss for centerness, or an Intersection-over-Union (IoU) loss for location.

11 . The method of claim 6 , wherein the object detection neural network further comprises a scene understanding part, which is pre-trained in end-to-end manner,

the method further comprising:

performing global average pooling of the extracted 2D feature maps to obtain a tensor representing the extracted 2D feature maps, and

estimating a 1 by passing the tensor through the scene understanding part configured to jointly estimate camera rotation and 3D layout of a scene, and

wherein the scene understanding part comprises two parallel branches: two fully connected layers that output 3D scene layout and two fully other connected layers that estimate camera rotation.

12 . The method of claim 11 , wherein the method is performed by an electronic device to understand the scene in a surrounding area of the electronic device,

wherein the electronic device is one of a smartphone, a tablet, smart glasses, or a mobile robot, and

wherein the electronic device includes at least one camera configured to capture the scene.

13 . The method of claim 11 , wherein the scene understanding part of the object detection neural network is trained on one or more indoor training datasets by optimizing a total scene understanding loss function based on at least one of a layout loss or a camera rotation estimation loss.

14 . An electronic device mounted on a vehicle comprising at least one camera, and configured to use an object detection neural network including a two-dimensional 2D feature extracting part, a 3D feature extracting part, and an outdoor object detecting part including parallel 2D convolutional layers for classification and location, which are pre-trained in end-to-end manner based on posed monocular images, the electronic device comprising:

a memory storing one or more instructions; and

a processor configured to execute the one or more instructions to:

receive one or more monocular images;

extract 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through the 2D feature extracting part of the object detection neural network,

generate an averaged 3D voxel volume based on the 2D feature maps,

extract a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of the 3D feature extracting part of the object detection neural network,

perform 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through the outdoor object detecting part of the object detection neural network; and

detect one or more 3D objects in a surrounding area of the vehicle based on 3D object detection.

15 . An electronic device comprising at least one camera, and configured to use an object detection neural network including a two-dimensional (2D) feature extracting part, a 3D feature extracting part, and an indoor object detecting part including 3D convolutional layers for classification, centerness, and location, which are pre-trained in end-to-end manner based on posed monocular images, the electronic device comprising:

a memory storing one or more instructions; and

a processor configured to execute the one or more instructions to:

receive one or more monocular images,

extract 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through the 2D feature extracting part of the object detection neural network,

generate an averaged 3D voxel volume based on the 2D feature maps,

extract refined 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through the 3D feature extracting part of the object detection neural network,

perform 3D object detection by passing the refined 3D feature maps through the indoor object detecting part of the object detection neural network using dense voxel representation of intermediate features; and

detect one or more 3D objects in a surrounding area of the electronic device based on the 3D object detection.

16 . The electronic device of claim 15 , wherein the electronic device is one of a smartphone, a tablet, smart glasses, or a mobile robot.

17 . A non-transitory computer readable medium storing computer executable instructions that when executed by a processor cause the processor to perform a method of three-dimensional (3D) object detection using an object detection neural network including a two-dimensional 2D feature extracting part, a 3D feature extracting part, and an outdoor object detecting part including parallel 2D convolutional layers for classification and location, which are pre-trained in end-to-end manner based on posed monocular images, the method comprising:

receiving one or more monocular images;

extracting 2D feature maps from each one of the one or more monocular images by passing the one or more monocular images through the 2D feature extracting part of the object detection neural network,

generating an averaged 3D voxel volume based on the 2D feature maps,

extracting a 2D representation of 3D feature maps from the averaged 3D voxel volume by passing the averaged 3D voxel volume through an encoder of the 3D feature extracting part of the object detection neural network, and

performing 3D object detection as 2D object detection in a Bird's Eye View (BEV) plane, the 2D object detection in the BEV plane being performed by passing the 2D representation of 3D feature maps through the outdoor object detecting part of the object detection neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 13, 2022
From: RUKHOVICH, DANILA DMITRIEVICH; VORONTSOVA, ANNA BORISOVNA; KONUSHIN, ANTON SERGEEVICH
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062073/0393 →
Priority Claims (2)
RU 2021114905 · May 26, 2021 · national
RU 2021128885 · Oct 4, 2021 · national
Continuity (2)
Continuation PCTKR2022007472 · May 26, 2022
Related Publication 20230121534A1 · Apr 20, 2023
References Cited (70)
US 9424461B1 · Yuan et al. · 2016 [cited by applicant]
US 10970518B1 · Zhou et al. · 2021 [cited by applicant]
US 11037051B2 · Kim et al. · 2021 [cited by applicant]
US 11100669B1 · Zhou · 2021 [cited by examiner]
US 20160196480A1 · Heifets et al. · 2016 [cited by applicant]
US 20190138786A1 · Trenholm et al. · 2019 [cited by applicant]
US 20200294257A1 · Yoo · 2020 [cited by applicant]
US 20210065440A1 · Sunkavalli et al. · 2021 [cited by applicant]
US 20210097717A1 · Wang et al. · 2021 [cited by applicant]
US 20210103776A1 · Jiang et al. · 2021 [cited by applicant]
US 20210146839A1 · Kim et al. · 2021 [cited by applicant]
US 20210146952A1 · Vora et al. · 2021 [cited by applicant]
US 20210390714A1 · Guizilini · 2021 [cited by examiner]
US 20230026857A1 · Di · 2023 [cited by examiner]
US 20230037958A1 · Liba · 2023 [cited by examiner]
US 20230222817A1 · Sun · 2023 [cited by examiner]
CN 110689008A · 2020 [cited by applicant]
CN 111932530A · 2020 [cited by applicant]
CN 112116700A · 2020 [cited by applicant]
CN 112287860A · 2021 [cited by applicant]
CN 112712062A · 2021 [cited by applicant]
RU 2693267C1 · 2019 [cited by applicant]
WO 2021042208A1 · 2021 [cited by applicant]
Kim, Jiwon, et al. “Monocular 3D object detection for an indoor robot environment.” 2020 29th IEEE International Conference on Robot and Human Interactive Communication (RO-MAN). IEEE, 2020. (Year: 2020). [cited by examiner]
Communication issued Jun. 8, 2022 by the Russian Patent Office in counterpart Russian Patent Application No. 2021128885. [cited by applicant]
Communication issued Jun. 9, 2022 by the Russian Patent Office in Russian Patent Application No. 2021128885. [cited by applicant]
International Search Report (PCT/ISA/210) issued Sep. 14, 2022 by the International Searching Authority in International Application No. PCT/KR2022/007472. [cited by applicant]
Written Opinion (PCT/ISA/237) issued Sep. 14, 2022 by the International Searching Authority in International Application No. PCT/KR2022/007472. [cited by applicant]
W. Bao, B. Xu, and Z. Chen, “MonoFENet: Monocular 3D Object Detection with Feature Enhancement Networks”, IEEE Transactions on Image Processing, Journal of Latex Xlass Files, vol. 14, No. 8, Aug. 2015, (13 pages total). [cited by applicant]
G. Brazil and X. Liu, “M3D-RPN: Monocular 3D Region Proposal Network for Object Detection”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9287-9296, 2019, (10 pages total). [cited by applicant]
H. Caesar et al., “nuScenes: A multimodal dataset for autonomous driving”, Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 11621-11631, 2020, (11 pages total). [cited by applicant]
F. Chabot et al., “Deep Manta: A Coarse-to-fine Many-Task Network for joint 2D and 3D vehicle analysis from monocular image”, In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 2040-20… [cited by applicant]
K. Chen et al., “MMDetection: Open MMLab Detection Toolbox and Benchmark”, arXiv:1906.07155v1 [cs.CV], Jun. 17, 2019, (13 total pages). [cited by applicant]
Chen et al., “Monocular 3D Object Detection for Autonomous Driving”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2147-2156, 2016, (10 pages total). [cited by applicant]
Chen et al., “MonoPair: Monocular 3D Object Detection Using Pairwise Spatial Relationships”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12093- 12102, 2020, (10 pages total). [cited by applicant]
Choi et al., “Understanding Indoor Scenes using 3D Geometric Phrases”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 33-40, 2013, (8 pages total). [cited by applicant]
Dai et al., “ScanNet: Richly-annotated 3D Reconstructions of Indoor Scenes”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, www.scan-net.org, pp. 5828-5839, 2017, (12 pages total). [cited by applicant]
Ding et al., “Learning Depth-Guided Convolutions for Monocular 3D Object Detection”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 1000-1001, 2020, (10 pages total). [cited by applicant]
Ding et al., “VoteNet: A Deep Learning Label Fusion Method for Multi-Atlas Segmentation”, International Conference on Medical Image Computing and Computer-Assisted Intervention, pp. 202-210, arXiv:1904.08963v2 [cs. CV],… [cited by applicant]
A. Geiger, P. Lenz, and R. Urtasun, “Are we ready for Autonomous Driving? The KITTI Vision Benchmark Suite”, In 2012 IEEE Conference on Computer Vision and Pattern Recognition, pp. 3354-3361, IEEE, 2012, (8 pages total). [cited by applicant]
K. He et al., “Deep Residual Learning for Image Recognition”, In Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, pp. 770-778, 2016, (9 pages total). [cited by applicant]
J. Hou et al., “3D-SIS: 3D Semantic Instance Segmentation of RGB-D Scans”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4421-4430, 2019, (10 pages total). [cited by applicant]
S. Huang et al., “Cooperative Holistic Scene Understanding: Unifying 3D Object, Layout, and Camera Pose Estimation”, arXiv:1810.13049, 2018, (12 pages total). [cited by applicant]
S. Huang et al., “Holistic 3D Scene Parsing and Reconstruction from a Single RGB Image”, In Proceedings of the European Conference on Computer Vision (ECCV), pp. 187-203, 2018, (17 pages total). [cited by applicant]
M. Jaritz et al., “Multi-view PointNet for 3D Scene Understanding”, In Proceedings of the IEEE/CVF International Conference on Computer Vision Workshops, 2019, (9 pages total). [cited by applicant]
E. Jorgensen et al., “Monocular 3D Object Detection and Box Fitting Trained End-to-End Using Intersection-over-Union Loss”, arXiv:1906.08070v2 [cs.CV], Jun. 20, 2019, (10 pages total). [cited by applicant]
H. Konigshof et al., “Realtime 3D Object Detection for Automated Driving Using Stereo Vision and Semantic Information”, In 2019 IEEE Intelligent Transportation Systems Conference (ITSC), pp. 1405-1410, IEEE, 2019, (6 pa… [cited by applicant]
J. Ku et al., “Monocular 3D Object Detection Leveraging Accurate Proposals and Shape Reconstruction”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 11867-11876, 2019, (10 page… [cited by applicant]
A. Kundu et al., “3D-RCNN: Instance-level 3D Object Reconstruction via Render-and-Compare”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3559-3568, 2018, (10 pages total). [cited by applicant]
Lang et al., “PointPillars: Fast Encoders for Object Detection from Point Clouds”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 12697-12705, 2019, (9 pages total). [cited by applicant]
B. Li et al., “GS3D: An Efficient 3D Object Detection Framework for Autonomous Driving”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 1019-1028, 2019, (10 pages total). [cited by applicant]
P. Li, X. Chen, and S. Shen, “Stereo R-CNN based 3D Object Detection for Autonomous Driving”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 7644-7652, 2019, (9 pages total). [cited by applicant]
P. Li et al., “RTM3D: Real-time Monocular 3D Detection from Object Keypoints for Autonomous Driving”, arXiv:2001.03343v1 [cs.CV] ,Jan. 10, 2020, (11 pages total). [cited by applicant]
Z. Liu et al., “SMOKE: Single-Stage Monocular 3D Object Detection via Keypoint Estimation”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition Workshops, pp. 996-997, 2020, (10 pages to… [cited by applicant]
X. Ma et al., “Accurate Monocular 3D Object Detection via Color-Embedded 3D Reconstruction for Autonomous Driving”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 6851-6860, 2019, (10 pa… [cited by applicant]
A. Mousavian et al., “3D Bounding Box Estimation Using Deep Learning and Geometry”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 7074-7082, 2017, (9 pages total). [cited by applicant]
Z. Murez et al., “Atlas: End-to-End 3D Scene Reconstruction from Posed Images”, arXiv:2003.10432v3 [cs.CV], Oct. 14, 2020, (18 pages total). [cited by applicant]
Y. Nie et al., “Total3DUnderstanding: Joint Layout, Object Pose and Mesh Reconstruction for Indoor Scenes from a Single Image”, In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 5… [cited by applicant]
C. Qi et al., “Frustum PointNets for 3D Object Detection from RGB-D Data”, In Proceedings of the IEEE on Computer Vision and Pattern Recognition, pp. 918-927, 2018, (10 pages total). [cited by applicant]
C. Qi et al., “ImVoteNet: Boosting 3D Object Detection in Point Clouds with Image Votes”, In Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, pp. 4404-4413, 2020, (10 pages total). [cited by applicant]
Z. Qin et al., “MonoGRNet: A Geometric Reasoning Network for Monocular 3D Object Localization”, In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 33, pp. 8851-8858, 2019, (8 total pages). [cited by applicant]
T. Roddick et al., “Orthographic Feature Transform for Monocular 3D Object Detection”, arXiv:1811.08188v1 [cs.CV], Nov. 20, 2018, (10 pages total). [cited by applicant]
A. Simonelli et al., “Disentangling Monocular 3D Object Detection”, arXiv:1905.12365v1 [cs.CV], May 29, 2019, (15 total pages). [cited by applicant]
V. A. Sindagi et al., “MVX-Net: Multimodal VoxelNet for 3D Object Detection”, In 2019 International Conference on Robotics and Automation (ICRA), pp. 7276-7282, IEEE, arXiv:1904.01649v1 [cs.CV], Apr. 2, 2019, (7 pages t… [cited by applicant]
S. Song et al., “SUN RGB-D: A RGB-D Scene Understanding Benchmark Suite”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 567-576, 2015, (10 pages total). [cited by applicant]
Z. Tian et al., “FCOS: Fully Convolutional One-Stage Object Detection”, In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9627-9636, 2019, (10 pages total). [cited by applicant]
Y. Yan et al., “SECOND: Sparsely Embedded Convolutional Detection”, Sensors, 18, 3337, doi.10.3390/s18103337, www.mdpi.com/journal/sensors, 2018, (17 pages total). [cited by applicant]
Z. Zhang et al., “H3DNet: 3D Object Detection Using Hybrid Geometric Primitives”, In European Conference on Computer Vision, pp. 311-329, arXiv:2006.05682v3 [cs.CV], Jul. 23, 2020, (30 pages total). [cited by applicant]
D. Zhou et al., “IoU Loss for 2D/3D Object Detection”, In 2019 International Conference on 3D Vision (3DV), pp. 85-94, IEEE, arXiv:1908.03851v1 [cs.CV], Aug. 11, 2019, (10 pages total). [cited by applicant]
Y. Zhou et al., “VoxelNet: End-to-End Learning for Point Cloud Based 3D Object Detection”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 4490-4499, 2018, (10 pages total). [cited by applicant]
Cited By (2)
US 12,626,394 US 12,694,574