IP Library › Granted Patent US 12,737,906
Granted Patent B2
US 12,737,906 · App. 18/144,099 · Granted Sep 15, 2026

3D reconstruction without 3D convolutions

Inventors: Mohamed Sayed (London, GB); John Gibson (Sunnyvale, CA); James Watson (London, GB); Victor Adrian Prisacariu (Oxford, GB); Michael David Firman (London, GB); Clément Godard (San Francisco, CA)
Assignee: Niantic Spatial, Inc.
G06T7/55G06T15/08G06T17/00G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,737,906
App. No.
18/144,099
Granted
Sep 15, 2026
Kind
B2
Abstract

A depth estimation module may receive a reference image and a set of source images of an environment. The depth module may receive image features of the reference image and the set of source images. The depth module may generate a 4D feature volume that includes the image features and metadata associated with the reference image and set of source images. The image features and the metadata may be arranged in the feature volume based on relative pose distances between the reference image and the set of source images. The depth module may reduce the 4D feature volume to generate a 3D cost volume. The depth module may apply a depth estimation model to the 3D cost volume and data based on the reference image to generate a two dimensional (2D) depth map for the reference image.

Claims (228)

1 . A method comprising:

receiving a reference image of an environment and a set of one or more source images of the environment;

receiving image features of the reference image and the set of source images, the image features representing visual information of the reference image and the set of source images;

generating a four dimensional (4D) feature volume that includes the image features and metadata associated with the reference image and set of source images, the image features and the metadata arranged in the 4D feature volume based on relative pose distances between the reference image and the set of source images;

reducing the 4D feature volume to generate a three dimensional (3D) cost volume; and

applying a depth estimation model to the 3D cost volume and data based on the reference image to generate a two dimensional (2D) depth map for the reference image,

wherein the 4D feature volume is a 4D tensor of dimension C×D×H×W, where C, D, H, and W are constants greater than zero, wherein for each spatial location (k, i, j), the 4D feature volume includes a C dimensional vector that includes (1) image features of the reference image

f

k

,

i

,

j

0

,

(2) image features of the set of source images

〈

f

〉

k

,

i

,

j

n

for n∈[1, N], where indicates that the image features of the set of source images are perspective-warped into a reference frame of the reference image, and (3) the metadata.

2 . The method of claim 1 , wherein the image features and metadata are arranged in the 4D feature volume according to ascending or descending order of relative pose distance.

3 . The method of claim 1 , wherein the metadata in the 4D feature volume includes at least one of:

a ray direction of the reference image

r

k

,

i

,

j

0

;

a ray direction of one of the source images

r

k

,

i

,

j

n

;

a reference plane depth

𝓏

k

,

i

,

j

0

;

a source plane depth

𝓏

k

,

i

,

j

n

;

a relative ray angle θ 0,n ;

a relative pose distance p 0,n ; or

a depth validity mask

m

k

,

i

,

j

n

.

4 . The method of claim 1 , wherein the depth estimation model includes a 2D convolutional neural network including an encoder-decoder architecture augmented with the cost volume.

5 . The method of claim 1 , wherein reducing the 4D feature volume includes reducing volumetric cells of the 4D feature volume in parallel into a feature map.

6 . The method of claim 1 , further comprising generating a 3D representation of the environment based on the 2D depth map of the reference image.

7 . The method of claim 6 , wherein at least one of:

the 3D representation is generated without performing a 3D convolution or

generating the 3D representation includes fusing the 2D depth map of the reference image with another 2D depth map.

8 . The method of claim 1 , wherein the image features of the reference image are generated by a first feature extractor model and the data based on the reference image includes second image features of the reference image generated by a second feature extractor model different from the first feature extractor model.

9 . A method comprising:

receiving a reference image of an environment and a set of one or more source images of the environment;

receiving image features of the reference image and the set of source images, the image features representing visual information of the reference image and the set of source images;

generating a four dimensional (4D) feature volume that includes the image features and metadata associated with the reference image and set of source images, the image features and the metadata arranged in the 4D feature volume based on relative pose distances between the reference image and the set of source images;

reducing the 4D feature volume to generate a three dimensional (3D) cost volume; and

applying a depth estimation model to the 3D cost volume and data based on the reference image to generate a two dimensional (2D) depth map for the reference image,

wherein a relative pose distance for the reference image and one of the source images of the set of source images p 0,n is given by

p

o

,

n

=

t

0

,

n

+

2

3

⁢

tr

⁡

(

-

R

0

,

n

)

,

where is an identity matrix, t 0,n is a relative position of source camera n to reference camera, R 0,n is a relative rotation transformation between reference camera and source camera n, and tr () is a trace function.

10 . A non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receiving a reference image of an environment and a set of one or more source images of the environment;

receiving image features of the reference image and the set of source images, the image geatures representing visual information of the reference image and the set of source images;

generating a four dimensional (4D) feature volume that includes the image features and metadata associated with the reference image and the set of source images, the image features and the metadata arranged in the 4D feature volume based on relative pose distances between the reference image and the set of source images;

reducing the 4D feature volume to generate a three dimensional (3D) cost volume; and

applying a depth estimation model to the 3D cost volume and data based on the reference image to generate a two dimensional (2D) depth map for the reference image,

wherein the 4D feature volume is a 4D tensor of dimension C×D×H×W, where C, D, H, and W are constants greater than zero, wherein for each spatial location (k, i, j), the 4D feature volume includes a C dimensional vector that includes (1) image features of the reference image

f

k

,

i

,

j

0

,

(2) image features of the set of source images

〈

f

〉

k

,

i

,

j

n

for n∈[1, N], indicates that the image features of the set of source images are perspective-warped into a reference frame of the reference image, and (3) the metadata.

11 . The non-transitory computer-readable medium of claim 10 , wherein the image features and metadata are arranged in the 4D feature volume according to ascending or descending order of relative pose distance.

12 . The non-transitory computer-readable medium of claim 10 , wherein the metadata in the 4D feature volume includes at least one of:

a ray direction of the reference image

r

k

,

i

,

j

0

;

a ray direction of one of the source images

r

k

,

i

,

j

n

;

a reference plane depth

𝓏

k

,

i

,

j

0

;

a source plane depth

𝓏

k

,

i

,

j

n

;

a relative ray angle θ 0,n ;

a relative pose distance p 0,n ; or

a depth validity mask

m

k

,

i

,

j

n

.

13 . The non-transitory computer-readable medium of claim 10 , wherein the depth estimation model includes a 2D convolutional neural network including an encoder-decoder architecture augmented with the cost volume.

14 . The non-transitory computer-readable medium of claim 10 , wherein reducing the 4D feature volume includes reducing volumetric cells of the 4D feature volume in parallel into a feature map.

15 . The non-transitory computer-readable medium of claim 10 , further comprising generating a 3D representation of the environment based on the 2D depth map of the reference image.

16 . The non-transitory computer-readable medium of claim 15 , wherein the 3D representation is generated without performing a 3D convolution.

17 . The non-transitory computer-readable medium of claim 15 , wherein generating the 3D representation includes fusing the 2D depth map of the reference image with another 2D depth map.

18 . A non-transitory computer-readable medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receiving a reference image of an environment and a set of one or more source images of the environment;

receiving image features of the reference image and the set of source images, the image features representing visual information of the reference image and the set of source images;

generating a four dimensional (4D) feature volume that includes the image features and metadata associated with the reference image and the set of source images, the image features and the metadata arranged in the 4D feature volume based on relative pose distances between the reference image and the set of source images;

reducing the 4D feature volume to generate a three dimensional (3D) cost volume; and

applying a depth estimation model to the 3D cost volume and data based on the reference image to generate a two dimensional (2D) depth map for the reference image,

wherein a relative pose distance for the reference image and one of the source images of the set of source images p 0,n is given by:

p

o

,

n

=

t

0

,

n

+

2

3

⁢

tr

⁡

(

-

R

0

,

n

)

,

where is an identity matrix, t 0,n is a relative position of source camera n to reference camera, R 0,n is a relative rotation transformation between reference camera and source camera n, and tr () is a trace function.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 16, 2025
From: NIANTIC, INC.
To: NIANTIC SPATIAL, INC.
Reel/Frame 071555/0833 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 22, 2024
From: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
To: NIANTIC, INC.
Reel/Frame 066197/0211 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 14, 2023
From: SAYED, MOHAMED; WATSON, JAMES; PRISACARIU, VICTOR ADRIAN; FIRMAN, MICHAEL DAVID; GODARD, CLEMENT; GIBSON, JOHN
To: NIANTIC INTERNATIONAL TECHNOLOGY LIMITED
Reel/Frame 065871/0712 →
Continuity (2)
Provisional Application 63339090 · May 6, 2022
Related Publication 20230360241A1 · Nov 9, 2023
References Cited (89)
US 10546424B2 · Pang · 2020 [cited by examiner]
US 10600233B2 · Lakshman · 2020 [cited by examiner]
US 11062471B1 · Zhong · 2021 [cited by examiner]
US 20160148433A1 · Petrovskaya · 2016 [cited by examiner]
Sun, Penghui, Suping Wu, and Kui Lin. “Attention-guided multi-view stereo network for depth estimation.” 2020 IEEE 22nd International Conference on High Performance Computing and Communications; IEEE 18th International … [cited by examiner]
Bhat, S.F. et al., “AdaBins: Depth Estimation Using Adaptive Bins,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Nov. 28, 2020, pp. 4009-4018. [cited by applicant]
Bozic, A. et al., “Transformerfusion: Monocular rgb scene reconstruction using transformers,” Advances in Neural Information Processing Systems, 34, Dec. 6, 2021, pp. 1403-1414. [cited by applicant]
Casser, V. et al., “Depth prediction without the sensors: Leveraging structure for unsupervised learning from monocular videos,” Proceedings of the AAAI conference on artificial intelligence, vol. 33. No. 01. Jul. 17, 2… [cited by applicant]
Chang, J.R. et al., “Pyramid stereo matching network,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2018, pp. 5410-5418. [cited by applicant]
Chen, Y. et al., “Self-supervised learning with geometric constraints in monocular video: Connecting flow, depth, and camera,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Jul. 12, 2019, pp. … [cited by applicant]
Cheng, X. et al., “Learning depth with convolutional spatial propagation network,” IEEE transactions on pattern analysis and machine intelligence, 42(10) arXiv:1810.02695v3, Oct. 4, 2019, pp. 2361-2379. [cited by applicant]
Choe, J. et al., “Volumefusion: Deep depth fusion for 3d scene reconstruction,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Aug. 19, 2021, pp. 16086-16095. [cited by applicant]
Collins, R.T., “A space-sweep approach to true multi-image matching,” Proceedings CVPR IEEE computer society conference on computer vision and pattern recognition, Dec. 1995 pp. 358-363. [cited by applicant]
Curless, B. et al., “A volumetric method for building complex models from range images,” Proceedings of the 23rd annual conference on Computer graphics and interactive techniques, Aug. 1, 1996, pp. 303-312. [cited by applicant]
Dai, A. et al., “Scannet: Richly-annotated 3d reconstructions of indoor scenes,” Proceedings of the IEEE conference on computer vision and pattern recognition, Feb. 2017, pp. 5828-5839. [cited by applicant]
Drory, A. et al., “Semi-global matching: a principled derivation in terms of message passing,” Pattern Recognition: 36th German Conference, GCPR 2014, Proceedings 36, Springer International Publishing, Oct. 15, 2014, pp… [cited by applicant]
Duzceker, A. et al., “Deepvideomvs: Multi-view stereo on video with recurrent spatio-temporal fusion,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Dec. 2020, pp. 15324-15333. [cited by applicant]
Eigen, D. et al., “Depth map prediction from a single image using a multi-scale deep network,” Advances in neural information processing systems, (27) Dec. 2014, pp. 1-9. [cited by applicant]
Facil, J.M. et al., “CAM-Convs: Camera-aware multi-scale convolutions for single-view depth,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 11826-11835. [cited by applicant]
Fischer, P. et al., “FlowNet: Learning Optical Flow with Convolutional Networks,” arXiv:1504.06852v2, May 4, 2015, pp. 1-13. [cited by applicant]
Furukawa, Y. et al., “Multi-view stereo: A tutorial,” Foundations and Trends® in Computer Graphics and Vision, 9(1-2) Jun. 23, 2015, pp. 1-148. [cited by applicant]
Github, “huggingface/pytorch-image-models,” May 11, 2023, 38 pages, [Online] [Retrieved on Jul. 20, 2023] Retrieved from the Internet URL: <https://github.com/huggingface/pytorch-image-models>. [cited by applicant]
Github, “Lightning-AI / lightning,” Apr. 1, 2019, 14 pages, [Online] [Retrieved on Jul. 20, 2023] Retrieved from the Internet URL: <https://github.com/Lightning-AI/lightning>. [cited by applicant]
Glocker, B. et al., “Real-time RGB-D camera relocalization,” 2013 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), Oct. 1, 2013, pp. 173-179. [cited by applicant]
Godard, C. et al., “Unsupervised monocular depth estimation with left-right consistency,” Proceedings of the IEEE conference on computer vision and pattern recognition, Nov. 9, 2017, pp. 270-279. [cited by applicant]
Godard, C. et al., “Digging into self-supervised monocular depth estimation,” Proceedings of the IEEE/CVF international conference on computer vision, Jun. 2018, pp. 3828-3838. [cited by applicant]
Gu, X. et al., “Cascade Cost Volume for High-Resolution Multi-View Stereo and Stereo Matching,” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 2020, pp. 2495-2504. [cited by applicant]
He, K. et al., “Deep residual learning for image recognition,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2016, pp. 770-778. [cited by applicant]
Hirschmuller, H., “Stereo processing by semiglobal matching and mutual information,” IEEE Transactions on pattern analysis and machine intelligence, 30(2) Dec. 18, 2007, pp. 328-341. [cited by applicant]
Hou, Y. et al., “Multi-view stereo by temporal nonparametric fusion,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 2651-2660. [cited by applicant]
Huang, P.H. et al., “Deepmvs: Learning multi-view stereopsis,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition, Jun. 2018, pp. 2821-2830. [cited by applicant]
Im, S. et al., “Dpsnet: End-to-end deep plane sweep stereo,” arXiv preprint arXiv:1905.00538, May 2, 2019, pp. 1-12. [cited by applicant]
Kähler, O. et al., “Hierarchical voxel block hashing for efficient integration of depth images,” IEEE Robotics and Automation Letters, 1(1) Dec. 2015, pp. 192-197. [cited by applicant]
Kang, S.B. et al., “Handling occlusions in dense multi-view stereo,” Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, CVPR 2001, vol. 1. Dec. 8, 2001, pp. 1-35. [cited by applicant]
Kar, A. et al., “Learning a multi-view stereo machine,” Advances in neural information processing systems (30) Dec. 2017, pp. 1-12. [cited by applicant]
Kazhdan, M. et al., “Poisson surface reconstruction,” Proceedings of the fourth Eurographics symposium on Geometry processing, vol. 7, Jun. 2006, pp. 1-10. [cited by applicant]
Kendall, A. et al., “End-to-end learning of geometry and context for deep stereo regression,” Proceedings of the IEEE international conference on computer vision, Oct. 2017, pp. 66-75. [cited by applicant]
Kuznietsov, Y. et al., “Comoda: Continuous monocular depth adaptation using past experiences,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Jan. 2021, pp. 2907-2917. [cited by applicant]
Lee, D.T. et al., “Two algorithms for constructing a Delaunay triangulation,” International Journal of Computer & Information Sciences, vol. 9, No. 3, Jun. 1980, pp. 219-242. [cited by applicant]
Li, Z. et al., “Megadepth: Learning single-view depth prediction from internet photos,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2018, pp. 2041-2050. [cited by applicant]
Liang, Z. et al., “Learning for disparity estimation through feature constancy,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2018, pp. 2811-2820. [cited by applicant]
Lin, T.Y. et al., “Feature Pyramid Networks for Object Detection,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jul. 2017, pp. 2117-2125. [cited by applicant]
Long, X. et al., “Occlusion-Aware Depth Estimation with Adaptive Normal Constraints,” arXiv:2004.00845v4, Jul. 12, 2021, pp. 1-17. [cited by applicant]
Lorensen, W.E. et al., “Marching cubes: A high resolution 3D surface construction algorithm,” Seminal graphics: pioneering efforts that shaped the field, 1, 1998, pp. 347-353. [cited by applicant]
Loshchilov, I. et al., “Decoupled Weight Decay Regularization,” arXiv:1711.05101v3, Jan. 4, 2019, pp. 1-19. [cited by applicant]
Lou, X. et al., “Consistent video depth estimation,” ACM Transactions on Graphics (ToG), vol. 39, No. 4, Article 71, Jul. 8, 2020, pp. 1-13. [cited by applicant]
Marcel, S. et al., “Torchvision the machine-vision package of torch,” Proceedings of the 18th ACM international conference on Multimedia, Oct. 25, 2010, pp. 1485-1488. [cited by applicant]
Mayer, N. et al., “A large dataset to train convolutional networks for disparity, optical flow, and scene flow estimation,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2016, pp. 4… [cited by applicant]
Mccraith, R. et al., “Monocular depth estimation with self-supervised instance adaptation,” arXiv preprint arXiv:2004.05821, Apr. 13, 2020, pp. 1-7. [cited by applicant]
Mures, Z. et al., “Atlas: End-to-End 3D Scene Reconstruction from Posed Images,” arXiv:2003.10432v3, Oct. 14, 2020, pp. 1-18. [cited by applicant]
Newcombe, R.A. et al., “Kinectfusion: Real-time dense surface mapping and tracking,” 2011 10th IEEE international symposium on mixed and augmented reality, Oct. 26, 2011, pp. 127-136. Ieee. [cited by applicant]
Newcombe, R.A. et al., “DTAM: Dense Tracking and Mapping in Real-Time,” 2011 international conference on computer vision, IEEE, Nov. 6, 2011, pp. 2320-2327. [cited by applicant]
Niebner, M. et al., “Real-time 3D reconstruction at scale using voxel hashing,” ACM Transactions on Graphics (ToG), 32(6), Nov. 1, 2013, pp. 1-11. [cited by applicant]
Paszke, A. et al., “Pytorch: An imperative style, high-performance deep learning library,” Advances in neural information processing systems, (32) Dec. 2019, pp. 1-12. [cited by applicant]
Patil, V. et al., “Don't Forget The Past: Recurrent Depth Estimation from Monocular Video,” IEEE Robotics And Automation Letters, arXiv:2001.02613v2, Jul. 28, 2020, pp. 1-8. [cited by applicant]
Prisacariu, V.A. et al., “Infinitam v3: A framework for large-scale 3D reconstruction with loop closure,” arXiv:1708.00783v1, Aug. 2, 2017, pp. 1-19. [cited by applicant]
Ranftl, R. et al., “Towards Robust Monocular Depth Estimation: Mixing Datasets for Zero-Shot Cross-Dataset Transfer,” IEEE Transactions on pattern analysis and machine intelligence, vol. 44, No. 3, Mar. 2022, pp. 1623-1… [cited by applicant]
Rich, A. et al., “3DVNet: Multi-View Depth Prediction and Volumetric Refinement,” International Conference on 3D Vision, arXiv:2112.00202v1, Dec. 1, 2021, pp. 700-709. [cited by applicant]
Runz, M. et al., “Maskfusion: Real-time recognition, tracking and reconstruction of multiple moving objects,” 2018 IEEE International Symposium on Mixed and Augmented Reality (ISMAR), arXiv:1804.09194v2, Oct. 22, 2018, … [cited by applicant]
Sayed, M. et al., “SimpleRecon: 3D Reconstruction Without 3D Convolutions Supplementary Material,” arXiv:2208.14743v1, Aug. 31, 2022, pp. 1-19. [cited by applicant]
Scharstein, D. et al., “A taxonomy and evaluation of dense two-frame stereo correspondence algorithms,” Proceedings of the IEEE Workshop on Stereo and Multi-Baseline Vision, Dec. 2001, pp. 1-35. [cited by applicant]
Schonberger, J.L. et al., “Structure-from-motion revisited,” Proceedings of the IEEE conference on computer vision and pattern recognition, Jun. 2016, pp. 4104-4113. [cited by applicant]
Schonberger, J.L. et al., “Pixelwise view selection for unstructured multi-view stereo,” In: European Conference on Computer Vision (ECCV) 14th European Conference Proceedings, Part III 14, Oct. 2016, pp. 501-518. [cited by applicant]
Scona, R. et al., “Staticfusion: Background reconstruction for dense rgb-d slam in dynamic environments,” 2018 IEEE international conference on robotics and automation (ICRA), IEEE, May 21, 2018, pp. 3849-3856. [cited by applicant]
Shu, C. et al., “Feature-metric Loss for Self-supervised Learning of Depth and Egomotion,” European Conference on Computer Vision, arXiv:2007.10603v1, Jul. 21, 2020, pp. 572-588. [cited by applicant]
Sinha, A. et al., “DELTAS: Depth Estimation by Learning Triangulation And densification of Sparse points,” Computer Vision—ECCV, Part XXI 16, arXiv:2003.08933v2, Aug. 25, 2020, pp. 104-121. [cited by applicant]
Sitzmann, V. et al., “Deepvoxels: Learning persistent 3d feature embeddings,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2019, pp. 2437-2446. [cited by applicant]
Stier, N. et al., “Vortx: Volumetric 3d reconstruction with transformers for voxelwise view selection and fusion,” 2021 International Conference on 3D Vision (3DV), IEEE. arXiv:2112.00236v1, Dec. 1, 2021, pp. 320-330. [cited by applicant]
Sun, J. et al., “NeuralRecon: Real-time coherent 3D reconstruction from monocular video,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 15598-15607. [cited by applicant]
Tan, M. et al., “Mnasnet: Platform-aware neural architecture search for mobile, ” Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jun. 2019, pp. 2820-2828. [cited by applicant]
Tan, M. et al., “Efficientnetv2: Smaller models and faster training,” International conference on machine learning, PMLR, Jul. 1, 2021, pp. 10096-10106. [cited by applicant]
Tananaev, D. et al., “Temporally consistent depth estimation in videos with recurrent architectures,” Proceedings of the European Conference on Computer Vision (ECCV) Workshops, Sep. 2018, pp. 1-14. [cited by applicant]
Vaswani, A. et al., “Attention Is All You Need,” Advances in neural information processing systems, (30) Dec. 2017, pp. 1-11. [cited by applicant]
Wang, K. et al., “Mvdepthnet: Real-time multiview depth estimation neural network,” 2018 International conference on 3d vision (3DV), IEEE, arXiv:1807.08563v1, Jul. 23, 2018, pp. 248-257. [cited by applicant]
Watson, J. et al., “The Temporal Opportunist: Self-Supervised Multi-Frame Monocular Depth,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 1164-1174. [cited by applicant]
Watson, J. et al., “Self-Supervised Monocular Depth Hints,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 2162-2171. [cited by applicant]
Whelan, T. et al., “Kintinuous: Spatially extended KinectFusion,” In: RSS Workshop on RGB-D: Advanced Reasoning with Depth Camera, Jul. 19, 2012, pp. 1-10. [cited by applicant]
Whelan, T. et al., “ElasticFusion: Dense SLAM without a pose graph,” Robotics: Science and Systems, Jul. 13, 2015, pp. 1-9. [cited by applicant]
Wimbauer, F. et al., “MonoRec: Semi-supervised dense reconstruction in dynamic environments from a single moving camera,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp.… [cited by applicant]
Yang, X. et al., “Mobile3DRecon: real-time monocular 3D reconstruction on a mobile phone,” IEEE Transactions on Visualization and Computer Graphics, vol. 26, No. 12, Sep. 21, 2020, pp. 3446-3456. [cited by applicant]
Yao, Y. et al., “Mvsnet: Depth inference for unstructured multi-view stereo,” Proceedings of the European conference on computer vision (ECCV), Sep. 2018, pp. 767-783. [cited by applicant]
Yee, K. et al., “Fast deep stereo with 2D convolutional processing of cost signatures,” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision, Mar. 2020, pp. 183-191. [cited by applicant]
Yin, W. et al., “Enforcing geometric constraints of virtual normal for depth prediction,” Proceedings of the IEEE/CVF International Conference on Computer Vision, Oct. 2019, pp. 5684-5693. [cited by applicant]
Yin, W. et al., “Learning to recover 3d scene shape from a single image,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 2021, pp. 204-213. [cited by applicant]
Zbontar, J. et al., “Stereo matching by training a convolutional neural network to compare image patches,” J. Mach. Learn. Res., 17(1), Jan. 1, 2016, pp. 2287-2318. [cited by applicant]
Zhang, F. et al., “Ga-net: Guided aggregation net for end-to-end stereo matching,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Oct. 2019 pp. 185-194. [cited by applicant]
Zhang, F. et al., “Domain-invariant stereo matching networks,” Computer Vision—ECCV 2020: 16th European Conference, Proceedings, Part II 16, arXiv:1911.13287v1, Nov. 29, 2019, pp. 420-439. [cited by applicant]
Zhao, Y. et al., “Camera pose matters: Improving depth prediction by mitigating pose distribution bias,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 19, 2021, pp. 15759-15768. [cited by applicant]
Zhou, Z. et al., “UNet++: A Nested U-Net Architecture for Medical Image Segmentation,” Deep Learn Med Image Anal Multimodal Learn Clin Decis Support, Sep. 20, 2018, pp. 3-11. [cited by applicant]