IP Library Granted Patent US 12,646,306
Granted Patent B2
US 12,646,306 · App. 18/475,988 · Granted Jun 2, 2026

Joint 3D detection and segmentation using bird's eye view and perspective view

Inventors: Zhe Huang (San Diego, CA); Lingting Ge (San Diego, CA); Yizhe Zhao (San Diego, CA)
Assignee: CreateAI, Inc.
G06V10/806G06V10/7715G06V20/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,306
App. No.
18/475,988
Granted
Jun 2, 2026
Kind
B2
Abstract

An image processing method includes performing, using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction; performing, a 3-dimensional (3D) feature extraction on the images; detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction.

Claims (38)

1 . A method of detecting objects in sensor data, comprising:

performing, using images obtained from one or more sensors onboard a vehicle, a 2-dimensional (2D) feature extraction;

performing, a 3-dimensional (3D) feature extraction on the images;

detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction;

refining 3D feature estimates using dual-space object queries that include joint proposals formed based on 2D features resulting from the 2D feature extraction and 3D features resulting from the 3D feature extraction;

wherein the refining comprises performing a multi-level refinement wherein, at each layer of the multi-level refinement, a self-attention layer that acts on both the 2D features and the 3D features, a first cross-attention layer that acts only on the 2D features, and a second cross-attention layer that acts only on the 3D features are used.

2 . The method of claim 1 , wherein the 2D feature extraction comprises a perspective view (PV) analysis of the images.

3 . The method of claim 1 , wherein the 3D feature extraction comprises a bird's eye view (BEV) analysis of the images.

4 . The method of any claim 1 , wherein the 3D feature extraction is performed by:

generating 3D features from the 2D features resulting from the 2D feature extraction.

5 . The method of claim 4 , wherein the generating the 3D features from the 2D features comprises applying a back-projection model to the 2D features.

6 . The method of claim 4 , wherein a shared pose is further used during the refining.

7 . The method of claim 1 , wherein the 2D feature extraction comprises a 3D object detection method.

8 . An apparatus for detecting objects in images from sensor data, the apparatus comprising at least one processor configured to:

perform, from the sensor data, a 2-dimensional (2D) feature extraction;

perform, from the sensor data, a 3-dimensional (3D) feature extraction;

detect the objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction;

refine 3D feature estimates using dual-space object queries that include joint proposals formed based on 2D features resulting from the 2D feature extraction and 3D features resulting from the 3D feature extraction;

wherein the refining comprises performing a multi-level refinement wherein, at each layer of the multi-level refinement, a self-attention layer that acts on both the 2D features and the 3D features, a first cross-attention layer that acts only on the 2D features, and a second cross-attention layer that acts only on the 3D features are used.

9 . The apparatus of claim 8 , wherein the 2D feature extraction comprises a perspective view (PV) analysis of the images and the 3D feature extraction comprises a bird's eye view (BEV) analysis of the images.

10 . The apparatus of claim 8 , wherein the at least one processor performs the 3D feature extraction by:

generating 3D features from the 2D features resulting from the 2D feature extraction.

11 . The apparatus of claim 10 , wherein the generating the 3D features from the 2D features comprises applying a back-projection model to the 2D features.

12 . The apparatus of claim 10 , wherein the refining is performed using a Conv3D or a Conv2D algorithm.

13 . The apparatus of claim 8 , wherein the 2D feature extraction comprises a 3D object detection method.

14 . The apparatus of claim 8 , wherein the 3D feature extraction comprises a dense segmentation and/or a detection method.

15 . A system for deployment on an autonomous vehicle, comprising:

one or more sensors configured to generate sensor data of an environment of the autonomous vehicle; and

at least one processor configured to detect objects in the sensor data by:

performing, using images from the sensor data obtained from the one or more sensors, a 2-dimensional (2D) feature extraction;

performing a 3-dimensional (3D) feature extraction on the sensor data;

detecting objects in the images by fusing detection results from the 2D feature extraction and the 3D feature extraction;

refining 3D feature estimates using dual-space object queries that include joint proposals formed based on 2D features resulting from the 2D feature extraction and 3D features resulting from the 3D feature extraction;

wherein the refining comprises performing a multi-level refinement wherein, at each layer of the multi-level refinement, a self-attention layer that acts on both the 2D features and the 3D features, a first cross-attention layer that acts only on the 2D features, and a second cross-attention layer that acts only on the 3D features are used.

16 . The system of claim 15 , wherein the one or more processor performs the 3D feature extraction by:

generating 3D features from the 2D features resulting from the 2D feature extraction.

17 . The system of claim 15 , wherein hybrid detection proposals are used for querying for the objects.

18 . The system of claim 15 , wherein the one or more sensors include a camera and a lidar.

Assignments (2)
CHANGE OF NAME Recorded Dec 3, 2025
From: TUSIMPLE, INC.
To: CREATEAI, INC.
Reel/Frame 073832/0553 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 27, 2023
From: HUANG, ZHE; GE, LINGTING; ZHAO, YIZHE
To: TUSIMPLE, INC.
Reel/Frame 065053/0481 →
Continuity (2)
Provisional Application 63518084 · Aug 7, 2023
Related Publication 20250054286A1 · Feb 13, 2025
References Cited (20)
US 11494937B2 · Urtasun · 2022 [cited by examiner]
US 20230260266A1 · Karasev · 2023 [cited by examiner]
US 20250029355A1 · Balachandran · 2025 [cited by examiner]
CN 113435232A · 2021 [cited by examiner]
CN 116168384A · 2023 [cited by examiner]
CN 116363615A · 2023 [cited by examiner]
Yang, Chenyu, Yuntao Chen, Hao Tian, Chenxin Tao, Xizhou Zhu, Zhaoxiang Zhang, Gao Huang et al. “BEVFormer v2: Adapting Modern Image Backbones to Bird's-Eye-View Recognition via Perspective Supervision.” arXiv preprint … [cited by examiner]
Li, Jing, Rui Li, Jiehao Li, Junzheng Wang, Qingbin Wu, and Xu Liu. “Dual-view 3d object recognition and detection via lidar point cloud and camera image.” Robotics and Autonomous Systems 150 (2022): 103999. (Year: 2022… [cited by examiner]
Ku, Jason, Melissa Mozifian, Jungwook Lee, Ali Harakeh, and Steven L. Waslander. “Joint 3d proposal generation and object detection from view aggregation.” In 2018 IEEE/RSJ international conference on intelligent robots… [cited by examiner]
Chen, Yun, Bin Yang, Ming Liang, and Raquel Urtasun. “Learning joint 2d-3d representations for depth completion.” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 10023-10032. 2019. (Year:… [cited by examiner]
Liu, Yingfei, Junjie Yan, Fan Jia, Shuailin Li, Aqi Gao, Tiancai Wang, and Xiangyu Zhang. “Petrv2: A unified framework for 3d perception from multi-camera images.” In Proceedings of the IEEE/CVF international conference… [cited by examiner]
Jonah Philon, et al. “Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d” In European Conference on Computer Vision, pp. 194-210. Springer, [2022]. [cited by applicant]
Yue Wang, et al. “Detr3d: 3d object detection from multi-view images via 3d-to-2d queries” In In Conference on Robot Learning, [2022], (12 pages). [cited by applicant]
Yinhao Li, et al. “BEVDepth: Acquisition of Reliable Depth for Multi-View 3D Object Detection” Proceedings of the AAAI Conference on Artificial Intelligence, 37(2): 1477-1485 [2023]. [cited by applicant]
Xuewu Lin, et al. “Sparse4D: Multi-view 3D Object Detection with Sparse Spatial-Temporal Fusion” arXiv preprint arXiv:2211.10581v2, [Feb. 10, 2023], (10 pages). [cited by applicant]
Adam W. Harley, et al. “Simple-BEV: What Really Matters for Multi-Sensor BEV Perception?” In IEEE International Conference on Robotics and Automation (ICRA) [Sep. 29, 2022], (7 pages). [cited by applicant]
Enze Xie, et al. “M2BEV: Multi-Camera Joint 3D Detection and Segmentation with Unified Birds-Eye View Representation” arXiv:2204.05088 [2022], (21 pages). [cited by applicant]
Zhiqi Li, et al. “BEVFormer: Learning Bird's-Eye-View Representation from Multi-Camera Images via Spatiotem poral Transformers” arXiv:2203.17270v2, [Jul. 13, 2022] (20 pages). [cited by applicant]
Junjie Huang, et al. “BEVDet4D: Exploit Temporal Cues in Multi-camera 3D Object Detection” arXiv:2203.17054v3 [Jun. 16, 2022], (11 pages). [cited by applicant]
Junjie Huang, et al. “BEVDet:High-performance Multi-camera 3D Object Detection in Bird-Eye-View” arXiv:2112.11790v3 [Jun. 16, 2022], (19 pages). [cited by applicant]