IP Library › Granted Patent US 12,646,292
Granted Patent B2
US 12,646,292 · App. 18/353,453 · Granted Jun 2, 2026

Surround scene perception using multiple sensors for autonomous systems and applications

Inventors: Minwoo Park (Saratoga, CA); Trung Pham (San Jose, CA); Junghyun Kwon (Santa Clara, CA); Sayed Mehdi Sajjadi Mohammadabadi (San Jose, CA); Bor-Jeng Chen (San Jose, CA); Xin Liu (Pleasanton, CA); Bala Siva Sashank Jujjavarapu (Sunnyvale, CA); Mehran Maghoumi (Santa Clara, CA)
Assignee: NVIDIA Corporation
G06V10/7715G06V10/82G06V20/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,292
App. No.
18/353,453
Granted
Jun 2, 2026
Kind
B2
Abstract

In various examples, feature values corresponding to a plurality of views are transformed into feature values of a shared orientation or perspective to generate a feature map—such as a Bird's-Eye-View (BEV), top-down, orthogonally projected, and/or other shared perspective feature map type. Feature values corresponding to a region of a view may be transformed into feature values using a neural network. The feature values may be assigned to bins of a grid and values assigned to at least one same bin may be combined to generate one or more feature values for the feature map. To assign the transformed features to the bins, one or more portions of a view may be projected into one or more bins using polynomial curves. Radial and/or angular bins may be used to represent the environment for the feature map.

Claims (72)

1 . A method comprising:

transforming one or more first feature values corresponding to a first perspective view of an environment into one or more first transformed feature values corresponding to a common perspective of the environment;

transforming one or more second feature values corresponding to a second perspective view of the environment into one or more second transformed feature values corresponding to the common perspective of the environment;

assigning the one or more first transformed feature values and the one or more second transformed feature values to one or more bins corresponding to one or more feature maps of the environment based at least on performing one or more perspective transformations using geometric relationships between positions in the first perspective view and the second perspective view and positions in the one or more feature maps; and

performing, using one or more machine learning models (MLMs) and based at least on the one or more feature maps, one or more operations for a machine.

2 . The method of claim 1 , further comprising:

computing a statistical combination of at least one of the one or more first transformed feature values and at least one of the one or more second transformed feature values; and

determining one or more combined feature values of the one or more feature maps using the statistical combination.

3 . The method of claim 1 , wherein the assigning is of at least one of the one or more first transformed feature values and at least one of the one or more second transformed feature values into a same bin of the one or more bins.

4 . The method of claim 1 , wherein the one or more perspective transformations includes fitting one or more curves to one or more portions of the one or more first transformed feature values and the one or more second transformed feature values.

5 . The method of claim 1 , wherein the one or more bins include one or more radial bins corresponding to the one or more feature maps.

6 . The method of claim 1 , wherein the one or more first feature values comprise one or more first two-dimensional (2D) image features of a first 2D image map obtained using a first image corresponding to the first perspective view and the one or more second feature values comprise one or more second 2D image features of a second 2D image map obtained using a second image corresponding to the second perspective view.

7 . The method of claim 1 , wherein the transforming the one or more first feature values into the one or more first transformed feature values includes applying the one or more first feature values to one or more second MLMs that encode global contextual information for one or more columns corresponding to the first perspective view and the one or more first feature values.

8 . The method of claim 1 , wherein the one or more MLMs are trained to detect, using the one or more feature maps, one or more of: objects, parking spaces, or freespace in the environment.

9 . A system comprising:

one or more processors to perform operations including:

converting feature values corresponding to a plurality of views of an environment into Bird's-Eye View (BEV) feature values corresponding to a BEV of the environment;

assigning the BEV feature values to one or more bins corresponding to one or more BEV feature maps based at least on performing one or more perspective transformations using geometric relationships between positions in the plurality of views and positions in the one or more BEV feature maps;

determining, using the BEV feature values, one or more BEV feature values for the one or more BEV feature maps based at least on the assigning; and

determining, using one or more machine learning models (MLMs) and based at least on the one or more BEV feature maps, one or more operations for a machine.

10 . The system of claim 9 , wherein the determining the one or more BEV feature values includes:

computing a statistical combination of a group of the BEV feature values based at least on the group of BEV feature values being geometrically projected into at least one same bin of the one or more bins; and

determining the one or more BEV feature values using the statistical combination.

11 . The system of claim 9 , wherein the one or more bins include one or more radial bins corresponding to the one or more BEV feature maps.

12 . The system of claim 9 , wherein the one or more perspective transformations includes fitting one or more polynomial curves to one or more portions of the BEV feature values.

13 . The system of claim 9 , wherein the feature values comprise one or more first two-dimensional (2D) image features of a first 2D image map generated using a first image corresponding to a first view of the plurality of views and one or more second 2D image features of a second 2D image map generated using a second image corresponding to a second view of the plurality of views.

14 . The system of claim 9 , wherein the converting includes applying a set of the feature values to one or more second MLMs that encode global contextual information for one or more columns corresponding to a view of the plurality of views and the set of the one or more BEV feature values.

15 . The system of claim 9 , wherein the system is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system for performing generative AI operations;

a system for performing operations using a large language model;

a system for performing conversational AI operations;

a system for generating synthetic data;

a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

16 . At least one processor comprising:

one or more circuits to perform one or more operations for a machine using one or more feature maps, the one or more feature maps being determined based at least on:

transforming feature values corresponding a plurality of perspective views of an environment into shared perspective feature values corresponding to a shared perspective of the environment, and

assigning the shared perspective feature values to one or more bins corresponding to the one or more feature maps based at least on one or more perspective transformations using geometric relationships between positions in the plurality of perspective views and positions in the one or more feature maps.

17 . The at least one processor of claim 16 , wherein the one or more feature maps are further determined based at least on:

computing a statistical combination of a group of the shared perspective feature values based at least on the group of the shared perspective feature values being geometrically projected into at least one same bin of one or more bins corresponding to the one or more feature maps; and

determining one or more shared perspective feature values for the one or more feature maps using the statistical combination.

18 . The at least one processor of claim 17 , wherein the group of the shared perspective feature values are projected into the at least one same bin based at least on fitting one or more curves to one or more portions of the shared perspective feature values.

19 . The at least one processor of claim 16 , wherein the one or more circuits are further to assign the shared perspective feature values to one or more radial bins corresponding to the one or more feature maps, wherein the one or more feature maps are determined based at least on the assigning.

20 . The at least one processor of claim 16 , wherein the at least one processor is comprised in at least one of:

a control system for an autonomous or semi-autonomous machine;

a perception system for an autonomous or semi-autonomous machine;

a system for performing simulation operations;

a system for performing digital twin operations;

a system for performing light transport simulation;

a system for performing collaborative content creation for 3D assets;

a system for performing generative AI operations;

a system for performing operations using a large language model;

a system for performing deep learning operations;

a system implemented using an edge device;

a system implemented using a robot;

a system for performing conversational AI operations;

a system for generating synthetic data;

a system for presenting at least one of virtual reality content, augmented reality content, or mixed reality content;

a system incorporating one or more virtual machines (VMs);

a system implemented at least partially in a data center; or

a system implemented at least partially using cloud computing resources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 9, 2023
From: PARK, MINWOO; PHAM, TRUNG; KWON, JUNGHYUN; SAJJADI MOHAMMADABADI, SAYED MEHDI; CHEN, BOR-JENG; LIU, XIN; JUJJAVARAPU, BALA SIVA SASHANK; MAGHOUMI, MEHRAN
To: NVIDIA CORPORATION
Reel/Frame 064541/0185 →
Continuity (2)
Provisional Application 63389828 · Jul 15, 2022
Related Publication 20240020953A1 · Jan 18, 2024
References Cited (53)
US 11328517B2 · Vaquero Gomez · 2022 [cited by examiner]
US 11397242B1 · Zhang · 2022 [cited by examiner]
US 11430218B2 · Li · 2022 [cited by examiner]
US 12315236B2 · Choi · 2025 [cited by examiner]
US 12412403B2 · Rezaei · 2025 [cited by examiner]
US 20140035775A1 · Zeng et al. · 2014 [cited by applicant]
US 20200160559A1 · Urtasun et al. · 2020 [cited by applicant]
US 20220035376A1 · Laddah et al. · 2022 [cited by applicant]
US 20220036650A1 · Gomez · 2022 [cited by examiner]
US 20230082097A1 · Choi · 2023 [cited by examiner]
US 20230260266A1 · Karasev · 2023 [cited by examiner]
US 20230267720A1 · Marvasti · 2023 [cited by examiner]
US 20240312177A1 · Redford · 2024 [cited by examiner]
WO 2024015632A1 · 2024 [cited by applicant]
Park, Minwoo; International Preliminary Report on Patentability for PCT Application No. PCT/US2023/027909, filed Jul. 17, 2023, mailed Jan. 30, 2025, 7 pgs. [cited by applicant]
Qi, Charles R., et al. “Frustum pointnets for 3d object detection from rgb-d data.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition. 2018. 15 Pages. [cited by applicant]
Godard, Clement et al: “Unsupervised Monocular Depth Estimation with Left-Right Consistency”, IEEE Computer Society Conference on Computer Vision and Pattern Recognition. Proceedings, IEEE Computer Society, US, Jul. 21,… [cited by applicant]
Reiher, L., et al., “A Sim2Real Deep Learning Approach for the Transformation of Images from Multiple Vehicle-Mounted Cameras to a Semantically Segmented Image in Bird's Eye View,” 23rd IEEE International Conference on … [cited by applicant]
Levi, et al.; “Stixelnet: A deep convolutional network for obstacle detection and road segmentation”, 26th British Machine Vision Conference (BMVC) 2015. [cited by applicant]
Park, Minwoo; International Search Report and Written Opinion for PCT Application No. PCT/US2023/027909; filed Jul. 17, 2023, mailed Sep. 28, 2023, 10 pgs. [cited by applicant]
Zhang, et al.; “RVDet: Feature-level Fusion of Radar and Camera for Object Detection”, 2021 IEEE Intelligent Trasportation Systems Conference (ITSC) Sep. 19-21, 2021, 7 pgs. [cited by applicant]
Wang, et al.: “MCF3D: Multi-Stage Complementary Fusion for Multi-Sensor 3D Object Detection,” IEEE Access, vol. 7, Jul. 24, 2019, 14 pgs. [cited by applicant]
Brazil, et al.; “M3D-RPN: Monocular 3D Region Proposal Network for Object Detection,” https://arxiv.org/abs/1907.06038; Aug. 11, 2019, 10 pgs. [cited by applicant]
Can, et al.; “Structured Bird's-eye-view Traffic Scene Understanding from Onboard Images,” 2021 IEEE/CVF International Conference on Computer Vision (ICCV), pp. 15641-15650, 2021. [cited by applicant]
Chen, et al.; “Monocular 3D Object Detection for Autonomous Driving,” in 2016 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), pp. 2147-2156, 2016. [cited by applicant]
Chitta, et al.; “Neat: Neural Attention Fields for End-to-End Autonomous Driving,” https://arxiv.org/abs/2109.04456; Sep. 9, 2021, 11 pgs. [cited by applicant]
Fu, et al.; “Deep Ordinal Regression Network for Monocular Depth Estimation,” In 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 2002-2011, 2018. [cited by applicant]
Hendy, et al.; “FISHING Net: Future Inference of Semantic Heatmaps in Grids,” https://arxiv.org/abs/2006.09917; Jun. 17, 2020, 9 pgs. [cited by applicant]
Liu, et al.; “SMOKE: Single-stage Monocular 3D Object Detection via Keypoint Estimation,” https://arxiv.org/abs/2002.10111; Feb. 24, 2020, 10 pgs. [cited by applicant]
Ma, et al.; “Accurate monocular 3D object Detection via color-embedded 3D reconstruction for autonomous driving,” In ICCV, 2019, 10 pgs. [cited by applicant]
Mani, et al.; “Mono Lay Out: Amodal Scene Layout from a Single Image,” 2020 IEEE Winter Conference on Applications of Computer Vision (WACV), 2020, 9 pgs. [cited by applicant]
Mousavian, et al.; “3D Bounding Box Estimation Using Deep Learning and Geometry, ”https://arxiv.org/abs/1612.00496; Apr. 10, 2017, 10 pgs. [cited by applicant]
Ng, et al.; “BEV-Seg: Bird's Eye View Semantic Segmentation Using Geometry and Semantic Point Cloud,” https://arxiv.org/abs/2006.11436; Jun. 23, 2020, 8 pgs. [cited by applicant]
Pan, et al.; “Cross-view Semantic Segmentation for Sensing Surroundings,” https://arxiv.org/abs/1906.03560, Jun. 18, 2020, 7 pgs. [cited by applicant]
Philion, et al., “Lift, Splat, Shoot: Encoding Images from Arbitrary Camera Rigs by Implicitly Unprojecting to 3D,” https://arxiv.org/abs/2008.05711, Aug. 13, 2020, 17 pgs. [cited by applicant]
Reading, et al.; “Categorical Depth Distribution Network for Monocular 3D Object Detection, ”https://arxiv.org/abs/2103.01100, Mar. 23, 2021, 11 pgs. [cited by applicant]
Roddick, et al.; “Predicting semantic map representations from images using pyramid occupancy networks,” https://arxiv.org/abs/2003.13402, Mar. 30, 2020, 11 pgs. [cited by applicant]
Roddick, et al.; “Orthographic feature transform for monocular 3d object detection,” https://arxiv.org/abs/1811.08188, Nov. 20, 2018, 10 pgs. [cited by applicant]
Saha, et al.; “Enabling spatio-temporal aggregation in birds-eye-view vehicle estimation,” in 2021 IEEE International Conference on Robotics and Automation (ICRA), 2021, 7 pgs. [cited by applicant]
Saha, et al.; “Translating images into maps,” https://arxiv.org/abs/2110.00966, Mar. 20, 2022, 7 pgs. [cited by applicant]
Sanberg, et al.; “Free-Space detection with Self-supervised and online trained fully convolutional networks,” https://arxiv.org/abs/1604.02316, Jan. 5, 2017, 8 pgs. [cited by applicant]
Scheck, et al.; “Where to drive: Free space detection with one fisheye camera,” https://arxiv.org/abs/2011.05822, Nov. 11, 2020, 10 pgs. [cited by applicant]
Schulter, et al.; “Learning to look around objects for top-view representations of outdoor scenes,” https://arxiv.org/abs/1803.10870; Mar. 28, 2018, 28 pgs. [cited by applicant]
Shi, et al.; “PointrRCNN: 3D object proposal generation and detection from point cloud,” https://arxiv.org/abs/1812.04244, May 16, 2019, 10 pgs. [cited by applicant]
Sun, et al.; “What makes for end-to-end object detection?” https://arxiv.org/abs/2012.05780, Jul. 12, 2021, 15 pgs. [cited by applicant]
Wang, et al.; “Fully Convolutional One-Stage Monocular 3D Object Detection,” https://arxiv.org/abs/2104.10956, Sep. 24, 2021, 11 pgs. [cited by applicant]
Wang, et al.; “Pseudo-Lidar from Visual Depth Estimation: Bridging the Gap in 3D Object Detection for Autonomous Driving,” https://arxiv.org/abs/1812.07179, Feb. 22, 2020, 16 pgs. [cited by applicant]
Wang, et al.; “DETR3D: 3D Object Detection from Multi-View Images via 3D-to-2D Queries,” https://arxiv.org/abs/2110.06922, Oct. 13, 2021, 12 pgs. [cited by applicant]
Wang, et al.; “DETR3D: 3D Object Detection from Multi-View Images via 3D-to-2D Queries,” In 5th Annual Conference on Robot Learning, 2021, 12 pgs. [cited by applicant]
Yang, et al.; “PIXOR: Real-time 3D Object Detection from Point Clouds,” https://arxiv.org/abs/1902.06326, Mar. 2, 2019, 10 pgs. [cited by applicant]
Yang, et al.; “Projecting Your View Attentively: Monocular Road Scene Layout Estimation Via Cross-View Transformation,” In 2021 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2021, 10 pgs. [cited by applicant]
Zhou, et al.; “Objects As Points,” https://arxiv.org/abs/1904.07850, Apr. 25, 2019, 12 pgs. [cited by applicant]
Cao, et al.; “Surround-view Free Space Boundary Detection with Polar Representation,” In Proceedings of the British Machine Vision Conference (BMVC). BMVA Press, Nov. 2021, 12 pgs. [cited by applicant]