IP Library › Granted Patent US 12,620,170
Granted Patent B2
US 12,620,170 · App. 18/515,016 · Granted May 5, 2026

Sparse voxel transformer for camera-based 3D semantic scene completion

Inventors: Yiming Li (Jersey City, NJ); Zhiding Yu (Santa Clara, CA); Christopher B. Choy (Los Angeles, CA); Chaowei Xiao (Tempe, AZ); Jose Manuel Alvarez Lopez (Mountain View, CA); Sanja Fidler (Toronto, CA); Animashree Anandkumar (Pasadena, CA)
Assignee: NVIDIA Corporation
G06T17/00B60W50/14G06T3/40G06V10/44G06V10/771G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,620,170
App. No.
18/515,016
Granted
May 5, 2026
Kind
B2
Abstract

An artificial intelligence framework is described that incorporates a number of neural networks and a number of transformers for converting a two-dimensional image into three-dimensional semantic information. Neural networks convert one or more images into a set of image feature maps, depth information associated with the one or more images, and query proposals based on the depth information. A first transformer implements a cross-attention mechanism to process the set of image feature maps in accordance with the query proposals. The output of the first transformer is combined with a mask token to generate initial voxel features of the scene. A second transformer implements a self-attention mechanism to convert the initial voxel features into refined voxel features, which are up-sampled and processed by a lightweight neural network to generate the three-dimensional semantic information, which may be used by, e.g., an autonomous vehicle for various advanced driver assistance system (ADAS) functions.

Claims (160)

1 . A computer-implemented method, comprising:

receiving one or more input images, wherein each image of the one or more input images is a two-dimensional (2D) image of a scene; and

processing, via a plurality of models implemented by one or more processors, the one or more input images to generate three-dimensional (3D) semantic information for the scene, the processing comprising:

extracting a set of image feature maps from the one or more input images by at least one feature extraction network,

generating a depth map by processing the one or more input images by a depth estimation network,

generating 3D point cloud data based on the depth map,

generating a first binary voxel grid occupancy map at a first resolution based on the 3D point cloud,

converting the first binary voxel grid occupancy map at the first resolution to a second binary voxel grid occupancy map at a second resolution by a depth correction network, and

generating the three-dimensional semantic information based on the set of image feature maps by using at least one transformer.

2 . The method of claim 1 , wherein the processing, via the plurality of models implemented by the one or more processors, the one or more input images to generate the three-dimensional semantic information for the scene further comprises:

generating, via a query proposal network, a set of query proposals by processing at least one of the depth map or the second binary voxel grid occupancy map.

3 . The method of claim 1 , wherein the at least one feature extraction network is a convolutional neural network (CNN).

4 . The method of claim 2 , wherein the generating the three-dimensional semantic information based on the set of image feature maps by using the at least one transformer comprises:

processing, via a first transformer, the set of image feature maps using a deformable cross-attention (DCA) mechanism in accordance with the set of query proposals to generate an updated set of query proposals.

5 . The method of claim 4 , wherein the generating the three-dimensional semantic information based on the set of image feature maps by using the at least one transformer comprises:

generating initial voxel features by combining the updated set of query proposals with a mask token; and

processing, via a second transformer, the initial voxel features using a deformable self-attention (DSA) mechanism to generate refined voxel features.

6 . The method of claim 5 , wherein the generating the three-dimensional semantic information based on the set of image feature maps by using the at least one transformer comprises:

up-sampling the refined voxel features; and

processing the up-sampled refined voxel features via a neural network comprising one or more fully connected layers to generate the three-dimensional semantic information.

7 . The method of claim 1 , wherein the plurality of models are trained in accordance with a loss criteria as defined by:

ℒ

=

-

Σ

k

=

1

K

⁢

Σ

c

=

c

0

c

m

c

y

ˆ

k

,

c

⁢

log

⁢

(

e

y

k

,

c

Σ

c

⁢

e

y

k

,

c

)

,

where k is a voxel index, K is a total number of the voxels, c indexes a plurality of semantic classes, y k,c is a predicted logits for the k-th voxel belonging to class c, ŷ k,c is a k-th element of Ŷ t ; and c is a weight for each class according to an inverse of a class frequency.

8 . The method of claim 1 , further comprising:

capturing, via an image sensor, the one or more input images.

9 . The method of claim 7 , wherein the image sensor is integrated in an autonomous vehicle, the method further comprising:

performing at least one advanced driver assistance systems (ADAS) function based on the three-dimensional semantic information, wherein the at least one ADAS function includes one or more of the following:

emergency braking;

pedestrian detection;

collision avoidance;

route planning;

lane departure warning; or

object avoidance.

10 . The method of claim 1 , wherein the at least one feature extraction network includes a convolutional neural network (CNN) configured to process the one or more images to generate the set of image feature maps, and wherein the at least one transformer includes a first transformer configured to implement a deformable cross-attention mechanism and a second transformer configured to implement a deformable self-attention mechanism.

11 . A system, comprising:

a memory storing one or more input images, wherein each image of the one or more input images is a two-dimensional (2D) image of a scene; and

one or more processors, connected to the memory, to:

process, via a plurality of models, the one or more input images to generate three-dimensional (3D) semantic information for the scene, by:

extracting a set of image feature maps from the one or more input images by at least one feature extraction network,

generating a depth map by processing the one or more input images by a depth estimation network,

generating 3D point cloud data based on the depth map,

generating a first binary voxel grid occupancy map at a first resolution based on the 3D point cloud,

converting the first binary voxel grid occupancy map at the first resolution to a second binary voxel grid occupancy map at a second resolution by a depth correction network, and

generating the three-dimensional semantic information based on the set of image feature maps by at least one transformer.

12 . The system of claim 11 , wherein the processing, via the plurality of models, the one or more input images to generate the 3D semantic information comprises:

generating, via a query proposal network, a set of query proposals by processing at least one of the depth map or the second binary voxel grid occupancy map.

13 . The system of claim 12 , wherein the processing, via the plurality of models, the one or more input images to generate the 3D semantic information comprises:

processing, via a first transformer, the set of image feature maps using a deformable cross-attention (DCA) mechanism in accordance with the set of query proposals to generate an updated set of query proposals;

generating initial voxel features by combining the updated set of query proposals with a mask token;

processing, via a second transformer, the initial voxel features using a deformable self-attention (DSA) mechanism to generate refined voxel features;

up-sampling the refined voxel features; and

processing the up-sampled refined voxel features via a neural network comprising one or more fully connected layers to generate the three-dimensional semantic information.

14 . The system of claim 11 , wherein the plurality of models are trained in accordance with a loss criteria as defined by:

ℒ

=

-

Σ

k

=

1

K

⁢

Σ

c

=

c

0

c

m

c

y

ˆ

k

,

c

⁢

log

⁢

(

e

y

k

,

c

Σ

c

⁢

e

y

k

,

c

)

,

where k is a voxel index, K is a total number of the voxels, c indexes a plurality of semantic classes, y k,c is a predicted logits for the k-th voxel belonging to class c, ŷ k,c is a k-th element of Ŷt; and c is a weight for each class according to an inverse of a class frequency.

15 . The system of claim 11 , further comprising:

an image sensor, wherein the one or more input images are captured by the image sensor.

16 . The system of claim 15 , wherein the system comprises an autonomous vehicle, and wherein the autonomous vehicle performs at least one advanced driver assistance systems (ADAS) function based on the three-dimensional semantic information, wherein the at least one ADAS functions includes one or more of the following:

emergency braking;

pedestrian detection;

collision avoidance;

route planning;

lane departure warning; or

object avoidance.

17 . A non-transitory computer-readable media storing computer instructions that, responsive to being executed by one or more processors, cause a device to perform the steps of:

receiving one or more input images, wherein each image of the one or more input images is a two-dimensional (2D) image of a scene; and

processing, via a plurality of models implemented by one or more processors, the one or more input images to generate three-dimensional (3D) semantic information for the scene, the processing comprising:

extracting a set of image feature maps from the one or more input images by at least one feature extraction network,

generating a depth map by processing the one or more input images by a depth estimation network,

generating 3D point cloud data based on the depth map,

generating a first binary voxel grid occupancy map at a first resolution based on the 3D point cloud,

converting the first binary voxel grid occupancy map at the first resolution to a second binary voxel grid occupancy map at a second resolution by a depth correction network, and

generating the three-dimensional semantic information based on the set of image feature maps by at least one transformer.

18 . The non-transitory computer-readable media of claim 17 , wherein the processing, via the plurality of models implemented by the one or more processors, the one or more input images to generate the 3D semantic information comprises:

generating, via a query proposal network, a set of query proposals by processing at least one of the depth map or the second binary voxel grid occupancy map;

processing, via a first transformer, the set of image features using a deformable cross-attention (DCA) mechanism in accordance with the set of query proposals to generate an updated set of query proposals;

generating initial voxel features by combining the updated set of query proposals with a mask token;

processing, via a second transformer, the initial voxel features using a deformable self-attention (DSA) mechanism to generate refined voxel features;

up-sampling the refined voxel features; and

processing the up-sampled refined voxel features via a neural network comprising one or more fully connected layers to generate the three-dimensional semantic information.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 26, 2024
From: LI, YIMING; YU, ZHIDING; CHOY, CHRISTOPHER B.; XIAO, CHAOWEI; ALVAREZ LOPEZ, JOSE MANUEL; FIDLER, SANJA; ANANDKUMAR, ANIMASHREE
To: NVIDIA CORPORATION
Reel/Frame 066562/0920 →
Continuity (2)
Provisional Application 63426497 · Nov 18, 2022
Related Publication 20240087222A1 · Mar 14, 2024
References Cited (25)
US 20230326215A1 · Yu · 2023 [cited by examiner]
Song, S., et al., “Semantic scene completion from a single depth image,” In Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1746-1754, 2017. [cited by applicant]
He, K., et al., “Masked autoencoders are scalable vision learners,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 16000-16009, 2022. [cited by applicant]
Cao, A.Q., et al., “Monoscene: Monocular 3D semantic scene completion,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 3991-4001, 2022. [cited by applicant]
Behley, J., et al., “SemanticKITTI: A dataset for semantic scene understanding of LiDAR sequences,” In Proceedings of the IEEE/CVF International Conference on Computer Vision, pp. 9297-9307, 2019. [cited by applicant]
Roldao, L., et al., “LMSCNet: Lightweight multiscale 3D semantic completion,” In 2020 International Conference on 3D Vision (3DV), pp. 111-119, 2020. [cited by applicant]
Cheng, R., et al., “S3CNet: A sparse semantic scene completion network for LiDAR point clouds,” In Conference on Robotic Learning, pp. 2148-2161, PMLR, 2021. [cited by applicant]
Yan, X., et al., “Sparse single sweep LiDAR point cloud segmentation via learning contextual shape priors from scene completion,” In Proceedings of the AAAI Conference on Artificial Intelligence, vol. 35, pp. 3101-3109,… [cited by applicant]
Tulsiani, S., et al., “Multi-view supervision for single-view reconstruction via differentiable ray consistency,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 2626-2634, 2017. [cited by applicant]
Ummenhofer, B., et al., “DeMON: Depth and motion network for learning monocular stereo,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 5038-5047, 2017. [cited by applicant]
Xie, E., et al., “SegFormer:: Simple and efficient design for semantic segmentation with transformers,” In Advances in Neural Information Processing Systems, vol. 34, pp. 12077-12090, 2021. [cited by applicant]
Long, J., et al., “Fully convolutional networks for semantic segmentation,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, pp. 3431-3440, 2015. [cited by applicant]
Noh, H., et al., “Learning deconvolution network for semantic segmentation,” In Proceedings of the IEEE International Conference on Computer Vision (ICCV), Dec. 2015. [cited by applicant]
He, K., et al., Deep residual learning for image recognition, “In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition,” pp. 770-778, 2016. [cited by applicant]
Carion, N., et al., “End-to-end object detection with transformers,” In European Conference on Computer Vision, pp. 213-229, Springer, 2020. [cited by applicant]
Wang, Y., et al., “DETR3D: 3D object detection from multi-view images via 3D-to-2D queries,” In Conference on Robot Learning, pp. 180-191, PMLR, 2022. [cited by applicant]
Xie, E., et al., “M2BEV: Multi-camera joint 3D detection and segmentation with unified birds-eye view representation,” arXiv preprint arXiv:2204.05088, 2022. [cited by applicant]
Li, Z., et al., “BEVFormer: Learning birds-eye-view representation from multi-camera images viai spatiotemporal transformers,” In European Conference on Computer Vision, 2022. [cited by applicant]
Zhu, X., et al., “Demformable DETR: Deformable transformers for end-to-end object detection,” In International Conference on Learning Representations, 2020. [cited by applicant]
Vaswani, A.,, et al., “Attention is all you need,” In Advances in neural information processing systems, vol. 30, 2017. [cited by applicant]
Ren, S., et al., “Faster R-CNN: Towards real-time object detection with region proposal networks,” In Advances in Neural Information Processing Systems, vol. 28, 2015. [cited by applicant]
Bhat, S.F., et al., “AdaBins: Depth estimation using adaptive bins,” In Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, pp. 4009-4018, 2021. [cited by applicant]
You, Y., et al., “Pseudo-LiDAR++: Accurate depth for 3D object detection in autonomous driving,” In International Conference on Learning Representations, 2019. [cited by applicant]
Yuan, W., et al., “Neural window fully-connected CRFs for monocular depth estimation,” In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, 2022. [cited by applicant]
Shamsafar, F, et al., “MobileStereoNet: Towards lightweight deep networks for stereo matching,” In Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, pp. 2417-2426, 2022. [cited by applicant]