IP Library Granted Patent US 12,469,170
Granted Patent B2
US 12,469,170 · App. 17/967,534 · Granted Nov 11, 2025

Apparatus and method for detecting a 3D object

Inventor: Hyun Kyu Lim (Seoul, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION
G06T7/75G06T7/10G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,469,170
App. No.
17/967,534
Granted
Nov 11, 2025
Kind
B2
Abstract

An apparatus and a method for detecting a three-dimensional (3D) object are provided. The apparatus includes a camera and a processor. The processor obtains a two-dimensional (2D) image using the camera, segments the 2D image to obtain free space information, extracts feature information associated with an object included in the 2D image from the 2D image, determines an attention score based on the free space information and the extracted feature information, using an attention mechanism, and detects 3D location information of the object from the image based on the attention score.

Claims (55)

1 . An apparatus for detecting a three-dimensional (3D) object, the apparatus comprising:

a camera;

storage; and

a processor configured to

obtain a two-dimensional (2D) image using the camera,

segment the 2D image to obtain feature information associated with a free space,

obtain feature information associated with an object included in the 2D image from the 2D image,

determine an attention score based on the feature information associated with the free space and the feature information associated with the object, using an attention mechanism, wherein the processor is configured to

compare the feature information associated with the free space with previously stored segmentation reference information to determine a segmentation loss for the feature information associated with the free space, and

extract a key from the feature information associated with the free space with regard to the segmentation loss, and

detect 3D location information of the object based on the attention score.

2 . The apparatus of claim 1 , wherein the processor is configured to:

segment the 2D image into a plurality of regions;

perform prediction of the free space with respect to the plurality of regions; and

obtain the feature information associated with the free space based on a result of performing the prediction.

3 . The apparatus of claim 1 , wherein the processor is configured to:

extract a query used in the attention mechanism from the feature information associated with the object; and

determine the attention score by multiplying the query by the key.

4 . The apparatus of claim 3 , wherein the processor is configured to:

determine a value used in the attention mechanism to be the same as the query;

apply the value as a weight to the attention score to determine an attention value; and

decode the attention value to detect the 3D location information.

5 . The apparatus of claim 4 , wherein the 3D location information includes 3D coordinate information of the object.

6 . The apparatus of claim 4 , wherein the processor is configured to compare the 3D location information with previously stored 3D location reference information to determine a loss of the 3D location information.

7 . The apparatus of claim 1 , wherein the processor includes a backbone network configured to extract the feature information associated with the object from the 2D image based on a hierarchical structure of a convolutional neural network (CNN).

8 . The apparatus of claim 1 , wherein the processor is configured to detect 2D and 3D dimensions of the object and a 3D orientation of the object from the 2D image, based on the feature information associated with the object.

9 . The apparatus of claim 1 , wherein the camera is a monocular camera.

10 . A method for detecting a three-dimensional (3D) object, the method comprising:

obtaining a 2D image using a camera;

segmenting the 2D image to obtain feature information associated with a free space;

obtaining feature information associated with an object included in the 2D image from the 2D image;

determining an attention score based on the feature information associated with the free space and the feature information associated with the object, using an attention mechanism; and

detecting 3D location information of the object based on the attention score,

wherein determining the attention score includes extracting a key used in the attention mechanism from the feature information associated with the free space, and

wherein obtaining the key includes

comparing the feature information associated with the free space with previously stored segmentation reference information to determine a segmentation loss for the feature information associated with the free space, and

extracting the key from the feature information associated with the free space with regard to the segmentation loss.

11 . The method of claim 10 , wherein obtaining the feature information associated with the free space includes:

segmenting the 2D image into a plurality of regions,

performing prediction of the free space with respect to the plurality of regions, and

obtaining the feature information associated with the free space based on a result of performing the prediction.

12 . The method of claim 10 , wherein determining of the attention score includes:

extracting a query used in the attention mechanism from the feature information associated with the object; and

determining the attention score by multiplying the query by the key.

13 . The method of claim 12 , further comprising:

determining a value used in the attention mechanism to be the same as the query; and

applying the value as a weight to the attention score to determine an attention value, and

wherein detecting of the 3D location information includes decoding the attention value to detect the 3D location information.

14 . The method of claim 13 , wherein the 3D location information includes 3D coordinate information of the object.

15 . The method of claim 13 , further comprising comparing the 3D location information with previously stored 3D location reference information to determine a loss of the 3D location information.

16 . The method of claim 10 , wherein obtaining the feature information associated with the object includes extracting the feature information associated with the object from the 2D image using a backbone network based on a hierarchical structure of a convolutional neural network (CNN).

17 . The method of claim 10 , further comprising:

detecting 2D and 3D dimensions of the object from the 2D image, based on the feature information associated with the object; and

detecting an orientation of the object from the 2D image, based on the feature information associated with the object.

18 . The method of claim 10 , wherein the 2D image is a monocular image obtained using a monocular camera.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2022
From: LIM, HYUN KYU
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION
Reel/Frame 061450/0144 →
Priority Claims (1)
KR 10-2022-0042186 · Apr 5, 2022 · national
Continuity (1)
Related Publication 20230316569A1 · Oct 5, 2023
References Cited (21)
US 10019637B2 · Chen et al. · 2018 [cited by applicant]
US 10176390B2 · Chen et al. · 2019 [cited by applicant]
US 20170140231A1 · Chen et al. · 2017 [cited by applicant]
US 20200143557A1 · Choi · 2020 [cited by examiner]
US 20200320327A1 · Fathi · 2020 [cited by examiner]
US 20210110202A1 · Vu · 2021 [cited by examiner]
US 20210209341A1 · Ye et al. · 2021 [cited by applicant]
US 20210241522A1 · Guler et al. · 2021 [cited by applicant]
US 20220020158A1 · Ning · 2022 [cited by examiner]
US 20220084238A1 · Tang · 2022 [cited by examiner]
CN 110349138B · 2021 [cited by examiner]
CN 109214349B · 2021 [cited by examiner]
JP 2017091549A · 2017 [cited by applicant]
KR 20190063153A · 2019 [cited by applicant]
WO 2020099338A1 · 2020 [cited by applicant]
Nguyen, Duy-Kien, et al. “BoxeR: Box-Attention for 2D and 3D Transformers.” arXiv preprint arXiv:2111.13087 (2021). (Year: 2021). [cited by examiner]
Heylen, Jonas, et al. “Monocinis: Camera independent monocular 3d object detection using instance segmentation.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2021. (Year: 2021). [cited by examiner]
Shen, Xiaoke, and Ioannis Stamos. “3d object detection and instance segmentation from 3d range and 2d color images.” Sensors 21.4 (2021): 1213. (Year: 2021). [cited by examiner]
Li, Buyu, et al. “Gs3d: An efficient 3d object detection framework for autonomous driving.” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition. 2019. (Year: 2019). [cited by examiner]
Brazil, Garrick, and Xiaoming Liu. “M3d-rpn: Monocular 3d region proposal network for object detection.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. [cited by applicant]
Dijk, Tom van, and Guido de Croon. “How do neural networks see depth in single images?.” Proceedings of the IEEE/CVF International Conference on Computer Vision. 2019. [cited by applicant]