IP Library Patent Application 18808906
Patent Application
App. No. 18/808,906

SYSTEMS AND METHODS FOR END TO END SCENE RECONSTRUCTION FROM MULTIVIEW IMAGES

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/808,906
Abstract

Systems and methods of generating a three-dimensional (3D) reconstruction of a scene or environment surrounding a user of a spatial computing system, such as a virtual reality, augmented reality or mixed reality system, using only multiview images comprising, and without the need for depth sensors or depth data from sensors. Features are extracted from a sequence of frames of RGB images and back-projected using known camera intrinsics and extrinsics into a 3D voxel volume wherein each pixel of the voxel volume is mapped to a ray in the voxel volume. The back-projected features are fused into the 3D voxel volume. The 3D voxel volume is passed through a 3D convolutional neural network to refine the and regress truncated signed distance function values at each voxel of the 3D voxel volume.

Claims (40)

1 . A method of generating a three-dimensional (3D) reconstruction of a scene from multiview images, the method comprising:

back-projecting features from each frame of a sequence of frames into a 3D voxel volume wherein each pixel of the 3D voxel volume is mapped to a ray in the 3D voxel volume;

passing the 3D voxel volume through a 3D convolutional neural network (3D CNN) having an encoder-decoder to refine the features in the 3D voxel volume and regress output truncated signed distance function (TSDF) values at each voxel of the 3D voxel volume; and

after passing the 3D voxel volume through all layers of the 3D CNN, passing the refined features in the 3D voxel volume and TSDF values at each voxel of the 3D voxel volume through a batch normalization (batchnorm) function and a rectified linear unit (reLU) function,

wherein the 3D reconstruction is generated without the use of depth data from depth sensors.

2 . The method of claim 1 , further comprising:

fusing/accumulating features from each frame into the 3D voxel volume.

3 . The method of claim 2 , wherein the sequence of frames of images is fused into a single 3D feature volume using a running average.

4 . The method of claim 3 , wherein the running average is a simple running average.

5 . The method of claim 3 , wherein the running average is a weighted running average.

6 . The method of claim 1 , wherein additive skip connections are included from an encoder to a decoder of the 3D CNN, and the method further comprises:

using the additive skip connections to skip one or more features in the 3D voxel volume from the encoder to the decoder of the 3D CNN.

7 . The method of claim 6 , wherein one or more null voxels of the 3D voxel volume do not have features back-projected into them corresponding to voxels which were not observed during the sequence of frames of images, and the method further comprises:

not using the additive skip connections from the encoder for the null voxels;

passing the null voxels through the batchnorm function and the reLU function to match the magnitude of the voxels undergoing the skip connections.

8 . The method of claim 1 , wherein the 3D CNN has a plurality of layers each having a set of 3×3×3 residual blocks, and the 3D CNN implements downsampling with 3×3×3 stride 2 convolution and upsampling using trilinear interpolation followed by a 1×1×1 convolution.

9 . The method of claim 1 , wherein the 3D CNN further comprises an additional head for predicting semantic segmentation, and the method further comprises:

the 3D CNN predicting semantic segmentation of the features in the 3D voxel volume.

10 . The method of claim 1 , further comprising training the 2D CNN using short frame sequences covering portions of scenes.

11 . The method of claim 10 , wherein the short frame sequences include ten or fewer frame sequences.

12 . The method of claim 11 , further comprising:

fine tuning the training of the 2D CNN using larger frame sequences having more frame sequences than the short frame sequences.

13 . The method of claim 12 , wherein the larger frame sequences include 100 or more frame sequences.

14 . A cross reality system, comprising:

a head-mounted display device having a display system;

a computing system in operable communication with the head-mounted display device;

a plurality of camera sensors in operable communication with the computing system;

wherein the computing system is configured to generate a three-dimensional (3D) reconstruction of a scene from a sequence of frames of images captured by the camera sensors by a process comprising:

back-projecting features from each frame of a sequence of frames into a 3D voxel volume wherein each pixel of the 3D voxel volume is mapped to a ray in the 3D voxel volume;

passing the 3D voxel volume through a 3D convolutional neural network (3D CNN) having an encoder-decoder to refine the features in the 3D voxel volume and regress output truncated signed distance function (TSDF) values at each voxel of the 3D voxel volume; and

after passing the 3D voxel volume through all layers of the 3D convolutional encoder-decoder, passing the refined features in the 3D voxel volume and TSDF values at each voxel of the 3D voxel volume through a batch normalization (batchnorm) function and a rectified linear unit (reLU) function,

wherein the 3D reconstruction is generated without the use of depth data from depth sensors.

15 . The system of claim 14 , wherein the process further comprises:

fusing/accumulating features from each frame into the 3D voxel volume.

16 . The system of claim 15 , wherein the sequence of frames of images is fused into a single 3D feature volume using one of a running average, a simple running average, and a weighted running average.

17 . The system of claim 14 , wherein additive skip connections are included from an encoder to a decoder of the 3D CNN, and the process further comprises:

using the additive skip connections to skip one or more features in the 3D voxel volume from the encoder to the decoder of the 3D CNN.

18 . The system of claim 17 , wherein one or more null voxels of the 3D voxel volume do not have features back-projected into them corresponding to voxels which were not observed during the sequence of frames of images, and the process for generating a three-dimensional (3D) reconstruction of the scene from the sequence of frames of images further comprises:

not using the additive skip connections from the encoder for the null voxels;

passing the null voxels through the batchnorm function and the reLU function to match a magnitude of the voxels undergoing the skip connections.

Assignments (2)
SECURITY INTEREST Recorded Oct 28, 2025
From: MAGIC LEAP, INC.; MENTOR ACQUISITION ONE, LLC; MOLECULAR IMPRINTS, INC.
To: CITIBANK, N.A., AS COLLATERAL AGENT
Reel/Frame 073387/0487 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2024
From: MUREZ, ZACHARY PAUL
To: MAGIC LEAP, INC.
Reel/Frame 068583/0483 →