IP Library › Granted Patent US 12,333,830
Granted Patent B2
US 12,333,830 · App. 17/836,898 · Granted Jun 17, 2025

Target detection method, device, terminal device, and medium

Inventor: Yi Xu (Palo Alto, CA)
Assignee: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
G06V20/64G06T7/74G06V10/25G06V10/40G06V20/10G06T2207/10028G06T2207/30244G06V2201/07
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,830
App. No.
17/836,898
Filed
Jun 9, 2022
Granted
Jun 17, 2025
Kind
B2
Art Unit
2676
USPC
382/103
Abstract

The present disclosure provides a target detection method. The method includes: acquiring a scene image of a scene; acquiring a three-dimensional point cloud corresponding to the scene; identifying a supporting plane in the scene based on the three-dimensional point cloud; generating a plurality of region proposals based on the scene image and the supporting plane; and performing a target detection on the plurality of region proposals to determine a target object to be detected in the scene image. In addition, the present disclosure also provides a target detection device, a terminal device, and a medium.

Claims (67)

1. A target detection method, comprising:

acquiring a scene image of a scene;

acquiring a three-dimensional point cloud corresponding to the scene;

identifying a supporting plane in the scene based on the three-dimensional point cloud;

generating a plurality of region proposals for a target object to be detected based on the scene image and the supporting plane; and

performing a target detection on the plurality of region proposals to determine the target object to be detected in the scene image,

wherein generating the plurality of region proposals for the target object to be detected based on the scene image and the supporting plane, comprises:

generating a plurality of sampling boxes;

determining a moving range of the plurality of sampling boxes based on a position and an orientation of the supporting plane;

controlling the plurality of sampling boxes to move on the supporting plane by a preset step size in the moving range to generate a plurality of three-dimensional regions; and

projecting the plurality of three-dimensional regions to the scene image to form the plurality of region proposals;

wherein the plurality of sampling boxes and the supporting plane are three-dimensional virtual area.

2. The target detection method according to claim 1 , wherein acquiring the three-dimensional point cloud corresponding to the scene, comprises:

scanning the scene by a simultaneous localization and mapping (SLAM) system to generate the three-dimensional point cloud corresponding to the scene.

3. The target detection method according to claim 1 , wherein identifying the supporting plane in the scene based on the three-dimensional point cloud, comprises:

identifying a horizontal plane or a vertical plane in the scene based on the three-dimensional point cloud, and determining the horizontal plane or the vertical plane as the supporting plane.

4. The target detection method according to claim 1 , wherein generating the plurality of sampling boxes, comprises:

acquiring a target object of interest to the user;

acquiring parameters of the target object of interest; and

generating parameters of the plurality of sampling boxes based on the parameters of the target object of interest.

5. The target detection method according to claim 1 , wherein said identifying the supporting plane in the scene based on the three-dimensional point cloud comprises:

identifying the supporting plane in the scene based on the three-dimensional point cloud by a Random Sample Consensus (RANSAC) algorithm, the three-dimensional point cloud being an input of the RANSAC algorithm.

6. The target detection method according to claim 1 , wherein the supporting plane is a plane used to place actual objects or virtual objects in the scene.

7. A terminal device, comprising: a memory, a processor, and computer programs stored in the memory and executable by the processor, wherein the processor is configured to execute the computer programs to:

acquire a scene image of a scene;

acquire a three-dimensional point cloud corresponding to the scene;

identify a supporting plane in the scene based on the three-dimensional point cloud;

generate a plurality of region proposals for a target object to be detected based on the scene image and the supporting plane; and

perform a target detection on the plurality of region proposals to determine the target object to be detected in the scene image,

wherein the processor is further configured to execute the computer programs to:

generate a plurality of sampling boxes;

determine a moving range of the plurality of sampling boxes based on a position and an orientation of the supporting plane;

control the plurality of sampling boxes to move on the supporting plane by a preset step size in the moving range to generate a plurality of three-dimensional regions; and

project the plurality of three-dimensional regions to the scene image to form the plurality of region proposals;

wherein the plurality of sampling boxes and the supporting plane are three-dimensional virtual area.

8. The terminal device according to claim 7 , wherein the processor is further configured to execute the computer programs to:

scan the scene by a simultaneous localization and mapping (SLAM) system to generate the three-dimensional point cloud corresponding to the scene.

9. The terminal device according to claim 7 , wherein the processor is further configured to execute the computer programs to:

identify a horizontal plane or a vertical plane in the scene based on the three-dimensional point cloud, and determine the horizontal plane or the vertical plane as the supporting plane.

10. The terminal device according to claim 7 , wherein the processor is further configured to execute the computer programs to:

acquire a target object of interest to the user;

acquire parameters of the target object of interest; and

generate parameters of the plurality of sampling boxes based on the parameters of the target object of interest.

11. The terminal device according to claim 7 , wherein the processor is further configured to execute the computer programs to:

identify the supporting plane in the scene based on the three-dimensional point cloud by a Random Sample Consensus (RANSAC) algorithm, the three-dimensional point cloud being an input of the RANSAC algorithm.

12. The terminal device according to claim 7 , wherein the supporting plane is a plane used to place actual objects or virtual objects in the scene.

13. A non-transitory computer readable storage medium, storing computer programs therein, wherein the computer programs, when executed by a processor, cause the processor to:

acquire a scene image of a scene;

acquire a three-dimensional point cloud corresponding to the scene;

identify a supporting plane in the scene based on the three-dimensional point cloud;

generate a plurality of region proposals for a target object to be detected based on the scene image and the supporting plane; and

perform a target detection on the plurality of region proposals to determine the target object to be detected in the scene image,

wherein the computer programs, when executed by a processor, further cause the processor to:

generate a plurality of sampling boxes;

determine a moving range of the plurality of sampling boxes based on a position and an orientation of the supporting plane;

control the plurality of sampling boxes to move on the supporting plane by a preset step size in the moving range to generate a plurality of three-dimensional regions; and

project the plurality of three-dimensional regions to the scene image to form the plurality of region proposals;

wherein the plurality of sampling boxes and the supporting plane are three-dimensional virtual area.

14. The non-transitory computer readable storage medium according to claim 13 , wherein the computer programs, when executed by a processor, further cause the processor to:

scan the scene by a simultaneous localization and mapping (SLAM) system to generate the three-dimensional point cloud corresponding to the scene.

15. The non-transitory computer readable storage medium according to claim 13 , wherein the computer programs, when executed by a processor, further cause the processor to:

identify a horizontal plane or a vertical plane in the scene based on the three-dimensional point cloud, and determine the horizontal plane or the vertical plane as the supporting plane.

16. The non-transitory computer readable storage medium according to claim 13 , wherein the computer programs, when executed by a processor, further cause the processor to:

acquire a target object of interest to the user;

acquire parameters of the target object of interest; and

generate parameters of the plurality of sampling boxes based on the parameters of the target object of interest.

17. The target detection method according to claim 1 , wherein the plurality of region proposals are located on the supporting plane.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 9, 2022
From: XU, YI
To: GUANGDONG OPPO MOBILE TELECOMMUNICATIONS CORP., LTD.
Reel/Frame 060155/0154 →
Continuity (3)
Continuation PCTCN2020114034 · Sep 8, 2020
Provisional Application 62947419 · Dec 12, 2019
Related Publication 20220309761A1 · Sep 29, 2022
References Cited (52)
US 9807365B2 · Cansizoglu · 2017 [cited by examiner]
US 10430691B1 · Kim et al. · 2019 [cited by applicant]
US 10740920B1 · Ebrahimi Afrouzi · 2020 [cited by examiner]
US 20140192050A1 · Qiu et al. · 2014 [cited by applicant]
US 20170243352A1 · Kutliroff · 2017 [cited by examiner]
US 20180144458A1 · Xu et al. · 2018 [cited by applicant]
US 20180322623A1 · Memo · 2018 [cited by examiner]
US 20180330149A1 · Uhlenbrock et al. · 2018 [cited by applicant]
US 20190147245A1 · Qi et al. · 2019 [cited by applicant]
US 20190163958A1 · Li et al. · 2019 [cited by applicant]
US 20190287258A1 · Mano et al. · 2019 [cited by applicant]
US 20200048045A1 · Xiong et al. · 2020 [cited by applicant]
US 20200160487A1 · Kanzawa et al. · 2020 [cited by applicant]
US 20210090242A1 · Hever et al. · 2021 [cited by applicant]
US 20210192793A1 · Engelland-gay et al. · 2021 [cited by applicant]
US 20220020215A1 · Banerjee et al. · 2022 [cited by applicant]
US 20240427144A1 · Kern et al. · 2024 [cited by applicant]
CN 104143194A · 2014 [cited by applicant]
CN 103247041B · 2016 [cited by applicant]
CN 106529573A · 2017 [cited by applicant]
CN 107291093A · 2017 [cited by applicant]
CN 108133191A · 2018 [cited by applicant]
CN 108491818A · 2018 [cited by applicant]
CN 108710818A · 2018 [cited by applicant]
CN 109145677A · 2019 [cited by applicant]
CN 109345510A · 2019 [cited by applicant]
CN 109670517A · 2019 [cited by applicant]
CN 109816664A · 2019 [cited by applicant]
CN 110032962A · 2019 [cited by applicant]
CN 110110802A · 2019 [cited by applicant]
CN 110400304A · 2019 [cited by applicant]
WO 2019144300A1 · 2019 [cited by applicant]
Extended European Search Report dated Jan. 2, 2023 received in European Patent Application No. EP20898408.8. [cited by applicant]
Georgios Georgakis et al:“Multiview RGB-D Dataset for Object Instance Detection” ,2016 Fourth International Conference on 3DVision(3DV), Sep. 26, 2016 (Sep. 26, 2016) ,pp. 426-434, XP055548477. [cited by applicant]
International Search Report and Written Opinion dated Dec. 11, 2020 in International Application No. PCT/CN2020/114034. [cited by applicant]
Dry goods _ Introduction to target detection, reading this is enough (completed)—Programmer Sought; https://zhuanlan.zhihu.com/p/34142321. [cited by applicant]
A Survey of Object Detection Algorithms Based on Deep Learning; https://zhuanlan.zhihu.com/p/33981103. [cited by applicant]
Ross Girshick et al. “Rich feature hierarchies for accurate object detection and semantic segmentation”, 2014 IEEE Conference on Computer Vision and Pattern Recognition. [cited by applicant]
“Fast R-CNN”, Ross Girshick Microsoft Research, 2015. [cited by applicant]
Uijlings et al. “Selective Search for Object Recognition”. [cited by applicant]
Shaoqing Ren et al. “Faster R-CNN: Towards Real Time Object Detection with Region Proposal Networks”, 2016. [cited by applicant]
Joseph Redmon et al. “You Only Look Once: Unified, Real-Time Object Detection”, 2015 [cited by applicant]
Wei Liu et al. “SSD: Single Shot Multibox Detector”, 2015. [cited by applicant]
The First Office Action from corresponding Chinese Application No. 202080084583.4 dated Jun. 26, 2024. [cited by applicant]
Ren et al., “3D Object Detection with Latent Support Surfaces”, 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Abstract, Sections 1 and 4, Dec. 6, 2018. [cited by applicant]
Du, “Research and Implementation of Indoor Scene Real-time Segmentation Algorithm Based on the Depth and Color Data”, China Excellent Master's Degree Thesis Full-text Database Information Technology Series, Issue 05, Ma… [cited by applicant]
Chinese Second Office Action, Chinese Application No. 202080084583.4, mailed on Nov. 27, 2024 (14 pages). [cited by applicant]
Nguyen et al., “3D Point Cloud Segmentation: A survey”, IEEE, 2013. [cited by applicant]
Chinese Rejection decision for Chinese Application No. 202080084583.4, mailed Feb. 20, 2025 (15 pages). [cited by applicant]
Non-Final Office Action for U.S. Appl. No. 17/837,165, mailed Jan. 2, 2025 (62 pages). [cited by applicant]
European Examination Report for European Application No. 20898408.8 mailed Mar. 13, 2025 (5 pages). [cited by applicant]
International Search Report and Written Opinion, International Application No. PCT/CN2020/114063, mailed Dec. 8, 2020 (9 pages). [cited by applicant]