IP Library › Granted Patent US 12,482,230
Granted Patent B2
US 12,482,230 · App. 18/217,904 · Granted Nov 25, 2025

Apparatus and method for detecting an object and a computer readable recording medium therefor

Inventors: Young Hyun Kim (Seoul, KR); Seo Won Lee (Seoul, KR); Jin Kyu Kim (Seoul, KR); Sang Pil Kim (Seoul, KR); Won Seok Roh (Seoul, KR)
Assignees: HYUNDAI MOTOR COMPANY; KIA CORPORATION; Korea University Research and Business Foundation
G06V10/764G06N3/094G06V10/82G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,482,230
App. No.
18/217,904
Granted
Nov 25, 2025
Kind
B2
Abstract

An apparatus and method for detecting an object and a computer readable recording medium therefor are disclosed. The apparatus includes: a first camera that obtains a first image of a first view angle range; a second camera that obtains a second image of a second view angle range having an overlapping region with the first view angle range; an object detection network that detects an object in the first and second images; a stereo matching device that corrects a deviation of a depth value between the first and second images in the overlapping region based on a feature value output by a backbone network of the object detection network; and a feature vector determination device that performs adversarial learning to reduce a deviation of a feature vector between objects in the overlapping region and a non-overlapping region while the object detection network detects the object.

Claims (48)

1 . An apparatus for detecting an object, the apparatus comprising:

a first camera configured to obtain a first image of a first view angle range;

a second camera configured to obtain a second image of a second view angle range having an overlapping region with the first view angle range;

an object detection network configured to detect an object in the first image and the second image;

a stereo matching device configured to correct a deviation of a depth value between the first image and the second image in the overlapping region based on a feature value output by a backbone network of the object detection network; and

a feature vector determination device configured to perform adversarial learning to reduce a deviation of a feature vector between an object in the overlapping region and an object in a non-overlapping region while the object detection network detects the object.

2 . The apparatus of claim 1 , wherein the stereo matching device is configured to correct the deviation of the depth value by reducing a deviation between a predicted disparity map and a ground truth disparity map in the overlapping region.

3 . The apparatus of claim 2 , wherein the stereo matching device is configured to:

generate a matched feature value by matching a first image feature value of the first image and a second image feature value of the second image output by the backbone network of the object detection network;

obtain multi-scale feature values for the matched feature value by using a multi-scale layer; and

generate a cost volume based on the multi-scale feature values.

4 . The apparatus of claim 3 , wherein the stereo matching device is configured to obtain a stereo focus loss corresponding to a deviation between the cost volume and the ground truth disparity map, and wherein the object detection network is configured to learn an object detection model to reduce the stereo focus loss.

5 . The apparatus of claim 4 , wherein the feature vector determination device is configured to output a regional classification loss proportional to a deviation of the feature vector based on the feature vector provided from the object detection network.

6 . The apparatus of claim 5 , wherein the object detection network includes:

a transformer configured to output prediction location information and score information of the object based on the feature value output from the backbone network; and

a class classifier configured to output a bounding box loss and a class loss based on the prediction location information, the score information, and the regional classification loss.

7 . The apparatus of claim 6 , wherein the object detection network is configured to perform learning to reduce the regional classification loss output from the feature vector determination device.

8 . The apparatus of claim 7 , wherein the object detection network is configured to

perform the learning to reduce a total loss obtained by subtracting the regional classification loss from a sum of the bounding box loss, the class loss, and the stereo focus loss.

9 . A method of detecting an object, the method comprising:

extracting, by overlapping region detection device, an overlapping region between a first image captured by a first camera and a second image captured by a second camera;

correcting, by a stereo matching device, a deviation of a depth value between the first image and the second image in the overlapping region while an object is detected in the first image and the second image by using an object detection network; and

performing, by a feature vector determination device, adversarial learning to reduce a feature vector deviation between objects in the overlapping region and an non-overlapping region.

10 . The method of claim 9 , wherein correcting the deviation of the depth value includes:

obtaining a predicted disparity map in the overlapping region; and

reducing a deviation between the predicted disparity map and a ground truth disparity map.

11 . The method of claim 10 , wherein obtaining the predicted disparity map includes:

generating a matched feature value by matching a first image feature value of the first image and a second image feature value of the second image output from a backbone network of the object detection network;

obtaining multi-scale feature values for the matched feature value by using a multi-scale layer; and

generating a cost volume based on the multi-scale feature values.

12 . The method of claim 11 , wherein correcting the depth value further includes:

obtaining a stereo focus loss corresponding to a deviation between the cost volume and the ground truth disparity map; and

learning an object detection model to reduce the stereo focus loss.

13 . The method of claim 12 , wherein performing the adversarial learning further includes outputting a regional classification loss proportional to a deviation of the feature vector based on the feature vector provided from the object detection network.

14 . The method of claim 13 , wherein performing the adversarial learning further includes providing, by a network, the regional classification loss to a class classifier that outputs a bounding box loss and a class loss.

15 . The method of claim 14 , wherein performing the adversarial learning further includes learning the object detection network to reduce the regional classification loss.

16 . The method of claim 9 , wherein extracting the overlapping region includes obtaining 4D information including the depth value based on 2D information extracted from the first image and the second image.

17 . A computer readable recording medium that stores computer readable instructions for performing operations, wherein the operations include:

extracting an overlapping region between a first image captured by a first camera and a second image captured by a second camera;

correcting a deviation of a depth value between the first image and the second image in the overlapping region; and

performing adversarial learning to reduce a deviation of a feature vector between an object in the overlapping region and an object in a non-overlapping region while an object is detected in the first image and the second image by using an object detection network.

18 . The computer readable recording medium of claim 17 , wherein correcting the deviation of the depth value includes:

obtaining, by a backbone network of the object detection network, a predicted disparity map based on a feature value; and

obtaining a stereo focus loss based on a deviation between the predicted disparity map and a ground truth disparity map.

19 . The computer readable recording medium of claim 18 , wherein correcting the depth value further includes learning an object detection model to reduce the stereo focus loss.

20 . The computer readable recording medium of claim 17 , wherein the performing of the adversarial learning includes:

outputting, by a feature vector determination device, a regional classification loss proportional to a deviation of the feature vector based on the feature vector provided from the object detection network; and

learning the object detection network to reduce the regional classification loss.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 3, 2023
From: KIM, YOUNG HYUN; LEE, SEO WON; KIM, JIN KYU; KIM, SANG PIL; ROH, WON SEOK
To: HYUNDAI MOTOR COMPANY; KIA CORPORATION; KOREA UNIVERSITY RESEARCH AND BUSINESS FOUNDATION
Reel/Frame 064141/0322 →
Priority Claims (1)
KR 10-2023-0003614 · Jan 10, 2023 · national
Continuity (1)
Related Publication 20240233326A1 · Jul 11, 2024
References Cited (31)
US 10511769B2 · Edpalm et al. · 2019 [cited by applicant]
US 20070291130A1 · Broggi · 2007 [cited by examiner]
US 20080199069A1 · Schick · 2008 [cited by examiner]
US 20110118608A1 · Lindner · 2011 [cited by examiner]
US 20110216208A1 · Matsuzawa · 2011 [cited by examiner]
US 20140211989A1 · Ding · 2014 [cited by examiner]
US 20150009149A1 · Gharib · 2015 [cited by examiner]
US 20150054958A1 · Kim · 2015 [cited by examiner]
US 20170019655A1 · Mueller · 2017 [cited by examiner]
US 20180338084A1 · Edpalm et al. · 2018 [cited by applicant]
US 20200029023A1 · Wippermann et al. · 2020 [cited by applicant]
US 20200334860A1 · Hsieh · 2020 [cited by examiner]
US 20200336655A1 · Hsieh · 2020 [cited by examiner]
US 20200336729A1 · Hsieh · 2020 [cited by examiner]
US 20210101791A1 · Ishizaki · 2021 [cited by examiner]
US 20210110180A1 · Wang et al. · 2021 [cited by applicant]
US 20210352259A1 · Jiang · 2021 [cited by examiner]
EP 3404913A1 · 2018 [cited by applicant]
KR 101699014B1 · 2017 [cited by applicant]
KR 102261323B1 · 2021 [cited by applicant]
Tsun-Hsuan Wang et al. , “3D LiDAR and Stereo Fusion using Stereo Matching Network with Conditional Cost Volume Normalization,” Jan. 28, 2020, 2019 IEEE/RSJ International Conference on Intelligent Robots and Systems (IR… [cited by examiner]
Hamid Laga et al.,“A Survey on Deep Learning Techniques for Stereo-Based Depth Estimation,” Mar. 4, 2022, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 4, Apr. 2022,pp. 1738-1759. [cited by examiner]
Weiqin Chuah et al.,“Deep Learning-Based Incorporation of Planar Constraints for Robust Stereo Depth Estimation in Autonomous Vehicle Applications,” Jul. 8, 2022, IEEE Transactions on Intelligent Transportation Systems,… [cited by examiner]
Jinshan Liu et al.,““Seeing is Not Always Believing”: Detecting Perception Error Attacks Against Autonomous Vehicles,” Sep. 1, 2021, IEEE Transactions on Dependable and Secure Computing, vol. 18, No. 5, Sep./Oct. 2021, … [cited by examiner]
Yilun Chen et al.,“DSGN: Deep Stereo Geometry Network for 3D Object Detection,” Jun. 2020, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2020, pp. 12536-12543. [cited by examiner]
Jia-Ren Chang et al.,“Pyramid Stereo Matching Network,” Jun. 2018, Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2018, pp. 5410-5412. [cited by examiner]
Yuxuan Liu et al.,“YOLOStereo3D: A Step Back to 2D for Efficient Stereo 3D Detection,” Oct. 18, 2021, 2021 IEEE International Conference on 1Robotics and Automation (ICRA 2021),May 31-Jun. 4, 2021, Xi'an, China,pp. 1301… [cited by examiner]
Yue Wang et al.,“DETR3D: 3D Object Detection from Multi-view Images via 3D-to-2D Queries,” Sep. 13, 2021, Proceedings of the 5th Conference on Robot Learning,pp. 1-5. [cited by examiner]
Tai Wang et al.,“Probabilistic and Geometric Depth: Detecting Objects in Perspective,” Jul. 2021, Proceedings of the 5th Conference on Robot Learning, PMLR 164, pp. 1-5. [cited by examiner]
Wonseok Roh et al., ORA3D: Overlap Region Aware Multi-view 3D Object Detection, Jul. 2, 2022, URL: [2207.00865] ORA3D: Overlap Region Aware Multi-view 3D Object Detection (arxiv.org); 14 pp. [cited by applicant]
Wonseok Roh et al., ORA3D_Overlap Region Aware Multi-view 3D Object Detection, The 33rd British Machine Vision Conference 2022; Nov. 21, 2022; 19 pp. [cited by applicant]