IP Library › Granted Patent US 12,354,237
Granted Patent B2
US 12,354,237 · App. 17/547,731 · Granted Jul 8, 2025

Feature-free photogrammetric 3D imaging with cameras under unconstrained motion

Inventors: Kevin Zhou (Durham, NC); Colin Cooke (Durham, NC); Jaehee Park (Durham, NC); Ruobing Qian (Durham, NC); Roarke Horstmeyer (Durham, NC); Joseph Izatt (Durham, NC); Sina Farsiu (Durham, NC)
Assignees: DUKE UNIVERSITY; RAMONA OPTICS INC.
G06T5/50G06T5/70G06T5/80G06T2207/20084G06T2207/20221
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,354,237
App. No.
17/547,731
Granted
Jul 8, 2025
Kind
B2
Abstract

A method of mesoscopic photogrammetry can be carried out using a set of images captured from a camera on a mobile computing device. Upon receiving the set of images, the method generates a composite image, which can include applying homographic rectification to warp all images of the set of images onto a common plane; applying a rectification model to undo perspective distortion in each image of the set of images; and applying an undistortion model for adjusting for camera imperfections of a camera that captured each image of the set of images. A height map is generated co-registered with the composite image, for example, by using an untrained CNN whose weights/parameters are optimized in order to optimize the height map. The height map and the composite image can be output for display.

Claims (102)

1. A method of mesoscopic photogrammetry, comprising:

receiving a set of images;

generating a composite image, the generating of the composite image comprising:

applying homographic rectification to warp all images of the set of images onto a common plane;

applying a perspective distortion rectification model to undo perspective distortion in each image of the set of images; and

applying an undistortion model for adjusting for camera imperfections of a camera that captured each image of the set of images;

generating a height map co-registered with the composite image, the generating of the height map comprising:

applying a raw image form of each image of the set of images as input to a neural network that outputs the height map, wherein at least a plurality of images of the set of images are utilized to generate the height map; and

optimizing parameters of the neural network to optimize the height map, wherein the parameters of the neural network to optimize the height map comprise a 6D camera pose and distortion and the height map; and

outputting the height map and the composite image for display.

2. The method of claim 1 , wherein the neural network is an untrained convolutional neural network.

3. The method of claim 1 , wherein the perspective distortion rectification model is an orthorectification model.

4. The method of claim 1 , wherein the perspective distortion rectification model is an arbitrary reference rectification model.

5. The method of claim 1 , wherein the undistortion model comprises a piecewise linear, nonparametric model.

6. The method of claim 5 , wherein the piecewise linear, nonparametric model includes a radially dependent relative magnification factor that is discretized into n r points, {{tilde over (M)}} t=0 n r -1 , paced by δ r with intermediate points linearly interpolated as:

M

~

⁡

(

r

)

=

(

1

+

⌊

r

δ

r

⌋

-

r

δ

r

)

⁢

M

~

⌊

r

δ

τ

⌋

+

(

r

δ

r

-

⌊

r

δ

r

⌋

)

⁢

M

~

⌊

r

δ

r

⌋

+

1

where └˜┘ is a flooring operation, 0≤r<(n r −1)δ r is a radial distance from a distortion center;

wherein for a given point in each image, r im , the applying of the undistortion model is given by r im ←{tilde over (M)}(|r im |)r im .

7. A computing device, comprising:

a processor, memory, and instructions stored on the memory that when executed by the processor, direct the computing device to:

receive a set of images;

generate a composite image by:

applying homographic rectification to warp all images of the set of images onto a common plane;

applying a rectification model to undo perspective distortion in each image of the set of images; and

applying an undistortion model for adjusting for camera imperfections of a camera that captured each image of the set of images;

generate a height map co-registered with the composite image by:

applying a raw image form of each image of the set of images as input to a neural network that outputs the height map, wherein at least a plurality of images of the set of images are utilized to generate the height map; and

optimizing parameters of the neural network to optimize the height map, wherein the parameters of the neural network to optimize the height map comprise a 6D camera pose and distortion and the height map; and

output the height map and the composite image for display.

8. The computing device of claim 7 , further comprising:

a camera, wherein the camera captures the set of images.

9. The computing device of claim 7 , further comprising:

a display, wherein the instructions to output the height map and the composite image for display direct the computing device to display the height map or the composite image at the display.

10. The computing device of claim 7 , further comprising:

a network interface, wherein the computing device receives the set of images via the network interface.

11. The computing device of claim 10 , further comprising:

instructions for implementing the neural network stored on the memory.

12. The computing device of claim 7 , wherein the rectification model is an orthorectification model.

13. The computing device of claim 7 , wherein the rectification model is an arbitrary reference rectification model.

14. The computing device of claim 7 , wherein the undistortion model comprises a piecewise linear, nonparametric model.

15. A computer-readable storage medium having instructions stored thereon that when executed by a computing device, direct the computing device to:

receive a set of images;

generate a composite image by:

applying homographic rectification to warp all images of the set of images onto a common plane;

applying a rectification model to undo perspective distortion in each image of the set of images; and

applying an undistortion model for adjusting for camera imperfections of a camera that captured each image of the set of images;

generate a height map co-registered with the composite image by:

applying a raw image form of each image of the set of images as input to a neural network that outputs the height map, wherein at least a plurality of images of the set of images are utilized to generate the height map; and

optimizing parameters of the neural network to optimize the height map, wherein the parameters of the neural network to optimize the height map comprise a 6D camera pose and distortion and the height map; and

output the height map and the composite image for display.

16. The computer-readable storage medium of claim 15 , wherein the applying of the homographic rectification, the applying of the rectification model, and the applying of the undistortion model are performed simultaneously.

17. The computer-readable storage medium of claim 15 , wherein the undistortion model comprises a piecewise linear, nonparametric model.

18. The method of claim 1 , wherein the set of images are captured by a camera of a mobile computing device with unstabilized, freehand motion and without precalibration of camera distortion.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2023
From: ZHOU, KEVIN; COOKE, COLIN; QIAN, ROUBING; HORSTMEYER, ROARKE; IZATT, JOSEPH; FARSIU, SINA
To: DUKE UNIVERSITY
Reel/Frame 065648/0132 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 22, 2023
From: PARK, JAEHEE
To: RAMONA OPTICS INC.
Reel/Frame 065648/0336 →
Continuity (2)
Provisional Application 63123698 · Dec 10, 2020
Related Publication 20220188996A1 · Jun 16, 2022
References Cited (72)
US 20020057279A1 · Jouppi · 2002 [cited by examiner]
US 20180139378A1 · Moriuchi · 2018 [cited by examiner]
US 20190026400A1 · Fuscoe · 2019 [cited by examiner]
US 20210004933A1 · Wong · 2021 [cited by examiner]
Zuurmond (“Accurate Camera Position Determination by Means of Moiré Pattern Analysis”, Stellenbosch University 2015) (Year: 2015). [cited by examiner]
MVDepthNet: Real-time Multiview Depth Estimation Neural Network, 3D Vision, 2018 (Year: 2018). [cited by examiner]
Zhou, Kevin C. et al., “Mesoscopic Photogrammetry With an Unstabilized Phone Camera,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2021, pp. 7535-7545. [cited by applicant]
Abadi, Martin et al., “TensorFlow: A System for Large-Scale Machine Learning,” 12th USENIX Symposium on Operating Systems Design and Implementation (OSDI 16), Nov. 2016, pp. 265-283, USENIX Association, Savannah, GA. [cited by applicant]
Aguerrebere, Cecilia et al., “A Practical Guide to Multi-Image Alignment,” 2018 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP), Apr. 2018 (accessible Feb. 9, 2018), pp. 1927-1931, IEEE… [cited by applicant]
Alismail, Hatem et al., “Photometric Bundle Adjustment for Vision-Based SLAM,” Computer Vision—ACCV 2016, Mar. 12, 2017 (accessible Aug. 5, 2016), pp. 324-341, vol. 10114, , Springer International Publishing, Cham. [cited by applicant]
Goodfellow, Ian et al., “Deep learning, vol. 1”, Oct. 3, 2015, 706 pages, MIT press Cambridge. [cited by applicant]
Bianco, Simone et al., “Evaluating the Performance of Structure from Motion Pipelines,” Journal of Imaging, Aug. 1, 2018, p. 98, vol. 4, issue 8. [cited by applicant]
Brown, Duane C. “Decentering Distortion of Lenses,” PE, 1966, pp. 444-462, vol. 32, issue 3. [cited by applicant]
Camposeco, Federico et al., “Non-parametric structure-based calibration of radially symmetric cameras,” Proceedings of the IEEE International Conference on Computer Vision, Dec. 2015, pp. 2192-2200. [cited by applicant]
Chen, Tianqi et al., “Training Deep Nets with Sublinear Memory Cost”, Apr. 22, 2016, arXiv, pp. 1-12. [cited by applicant]
Delaunoy, Amael et al., “Photometric Bundle Adjustment for Dense Multi-View 3D Modeling,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2014, 8 pages. [cited by applicant]
Engel, Jakob et al., “Direct Sparse Odometry,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Mar. 1, 2018, pp. 611-625, vol. 40, issue 3. [cited by applicant]
Engel, Jakob et al., “LSD-SLAM: Large-Scale Direct Monocular SLAM,” Computer Vision—ECCV 2014, Sep. 2014, pp. 834-849, vol. 8690, Springer International Publishing, Cham. [cited by applicant]
Fitzgibbon, A. W. “Simultaneous linear estimation of multiple view geometry and lens distortion,” Proceedings of the 2001 IEEE Computer Society Conference on Computer Vision and Pattern Recognition. CVPR 2001, Dec. 2001… [cited by applicant]
Fuentes-Pacheco, Jorge et al., “Visual simultaneous localization and mapping: a survey,” Artificial Intelligence Review, Nov. 2015 (accessible Nov. 13, 2012), pp. 55-81, vol. 43, issue 1. [cited by applicant]
Fuhrmann, Simon et al., “MVE—A Multi—View Reconstruction Environment,” Eurographics Workshop on Graphics and Cultural Heritage, Oct. 2014, 8 pages, The Eurographics Association. [cited by applicant]
Furukawa, Yasutaka et al., “Multi-View Stereo: A Tutorial,” Foundations and Trends® in Computer Graphics and Vision, Jun. 23, 2015, pp. 1-148, vol. 9, issue 1-2. [cited by applicant]
Ghosh, Pallabi et al., “Deep depth prior for multi-view stereo,” ArXiv Preprint ArXiv:2001.07791, 2020 (accessible Dec. 1, 2020), 12 pages, vol. 2. [cited by applicant]
Godard, Clement et al., “Digging Into Self-Supervised Monocular Depth Estimation,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 3828-3838. [cited by applicant]
Gordon, Ariel et al., “Depth From Videos in the Wild: Unsupervised Monocular Depth Learning From Unknown Cameras,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 8977-8986. [cited by applicant]
Griewank, Andreas et al., “Algorithm 799: revolve: an implementation of checkpointing for the reverse or adjoint mode of computational differentiation,” ACM Transactions on Mathematical Software, Mar. 2000, pp. 19-45, v… [cited by applicant]
Hanna, K. J. “Direct multi-resolution estimation of ego-motion and structure from motion,” Proceedings of the IEEE Workshop on Visual Motion, 1991, pp. 156-162, IEEE Comput. Soc. Press, Princeton, NJ, USA. [cited by applicant]
Hartley, R. et al. “Parameter-Free Radial Distortion Correction with Center of Distortion Estimation,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Aug. 2007 (accessible Jan. 18, 2007), pp. 1309-1321,… [cited by applicant]
Heckel, Reinhard et al. “Deep Decoder: Concise Image Representations from Untrained Non-convolutional Networks”, Mar. 12, 2019 (accessible Oct. 2, 2018), arXiv, 17 pages. [cited by applicant]
Heel, Joachim “Direct estimation of structure and motion from multiple frames”, Mar. 1, 1990, Massachusetts Inst of Tech Cambridge Artificial Intelligence Lab, 70 pages. [cited by applicant]
Horn, Berthold K. P. et al. “Direct methods for recovering motion,” International Journal of Computer Vision, Jun. 1988, pp. 51-76, vol. 2, issue 1. [cited by applicant]
Hou, Yuxin et al., “Multi-View Stereo by Temporal Nonparametric Fusion,” Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), Oct. 2019, pp. 2651-2660. [cited by applicant]
Huang, Po-Han et al., “DeepMVS: Learning Multi-View Stereopsis,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 2821-2830. [cited by applicant]
Im, Sunghoon et al., “DPSNet: End-to-end Deep Plane Sweep Stereo”, May 2, 2019, arXiv, 12 pages. [cited by applicant]
Kingma, Diederik P. et al. “Adam: A Method for Stochastic Optimization”, Dec. 22, 2014, arXiv, 15 pages. [cited by applicant]
Kumar, R., et al. “Direct recovery of shape from multiple views: a parallax based approach,” Proceedings of 12th International Conference on Pattern Recognition, Oct. 1994, pp. 685-688, vol. 1, , IEEE Comput. Soc. Press… [cited by applicant]
Le, Tung D. et al., “TFLMS: Large Model Support in TensorFlow by Graph Rewriting”, Jul. 5, 2018, arXiv, 10 pages. [cited by applicant]
Li, Ruihao et al., “UnDeepVO: Monocular Visual Odometry Through Unsupervised Deep Learning,” 2018 IEEE International Conference on Robotics and Automation (ICRA), May 2018, pp. 7286-7291, IEEE, Brisbane, QLD. [cited by applicant]
Lowe, David G. “Distinctive Image Features from Scale-Invariant Keypoints,” International Journal of Computer Vision, Nov. 2004, pp. 91-110, vol. 60, issue 2. [cited by applicant]
Luhmann, Thomas “Close range photogrammetry for industrial applications,” ISPRS Journal of Photogrammetry and Remote Sensing, Nov. 2010 (accessible Jul. 15, 2010), pp. 558-569, vol. 65, issue 6. [cited by applicant]
Mahjourian, Reza et al., “Unsupervised Learning of Depth and Ego-Motion From Monocular Video Using 3D Geometric Constraints,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 20… [cited by applicant]
Matthies, Larry et al., “Kalman filter-based algorithms for estimating depth from image sequences,” International Journal of Computer Vision, Sep. 1989, pp. 209-238, vol. 3, issue 3. [cited by applicant]
Matthies, Larry H. et al., “Incremental estimation of dense depth maps from image sequences.,” CVPR, Jun. 1988, pp. 366-374. [cited by applicant]
Moulon, Pierre et al., “OpenMVG: Open Multiple View Geometry,” Reproducible Research in Pattern Recognition. RRPR 2016, Cacun, Mexico, Dec. 4, 2016, pp. 60-74, vol. 10214, Springer International Publishing, Cham. [cited by applicant]
Murez, Zak et al., “Atlas: End-to-End 3D Scene Reconstruction from Posed Images,” Computer Vision—ECCV 2020, Aug. 2020, pp. 414-431, vol. 12352, Springer International Publishing, Cham. [cited by applicant]
Newcombe, Richard A. et al., “DTAM: Dense tracking and mapping in real-time,” 2011 International Conference on Computer Vision, Nov. 2011, pp. 2320-2327, IEEE, Barcelona, Spain. [cited by applicant]
Oliensis, J. “Direct multi-frame structure from motion for hand-held cameras,” Proceedings 15th International Conference on Pattern Recognition. ICPR-2000, Sep. 2000, pp. 889-895, vol. 1, IEEE Comput. Soc, Barcelona, Sp… [cited by applicant]
Peggs, G. N. et al., “Recent developments in large-scale dimensional metrology,” Proceedings of the Institution of Mechanical Engineers, Part B: Journal of Engineering Manufacture, Jun. 1, 2009, pp. 571-595, vol. 223, i… [cited by applicant]
Ranjan, Anurag et al., “Competitive Collaboration: Joint Unsupervised Learning of Depth, Camera Motion, Optical Flow and Motion Segmentation,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recogn… [cited by applicant]
Robinson, D. et al., “Optimal Registration Of Aliased Images Using Variable Projection With Applications To Super-Resolution,” The Computer Journal, Feb. 19, 2008 (accessible Apr. 12, 2007), pp. 31-42, vol. 52, issue 1. [cited by applicant]
Rupnik, Ewelina et al., “MicMac—a free, open-source solution for photogrammetry,” Open Geospatial Data, Software and Standards, Dec. 2017, 9 pages, vol. 2, issue 1, article 14. [cited by applicant]
Sawhney, “3D geometry from planar parallax,” Proceedings of IEEE Conference on Computer Vision and Pattern Recognition CVPR-94, Jun. 21, 1994, pp. 929-934, IEEE Comput. Soc. Press, Seattle, WA, USA. [cited by applicant]
Schonberger, Johannes L. et al., “Structure-From-Motion Revisited,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2016, pp. 4104-4113. [cited by applicant]
Schops, Thomas et al., “Why Having 10,000 Parameters in Your Camera Model Is Better Than Twelve,” Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2020, pp. 2535-2544. [cited by applicant]
Stein, G. P. et al., “Model-based brightness constraints: on direct estimation of structure and motion,” IEEE Transactions on Pattern Analysis and Machine Intelligence, Sep. 2000, pp. 992-1015, vol. 22, issue 9. [cited by applicant]
Tang, Chengzhou et al., “BA-Net: Dense Bundle Adjustment Network”, Jun. 13, 2018, arXiv, 18 pages. [cited by applicant]
Teed, Zachary et al., “DeepV2D: Video to Depth with Differentiable Structure from Motion”, Dec. 11, 2018, arXiv, 20 pages. [cited by applicant]
Ullman, S. “The interpretation of structure from motion,” Proceedings of the Royal Society of London. Series B. Biological Sciences, Jan. 15, 1979, pp. 405-426, vol. 203, issue 1153. [cited by applicant]
Ulyanov, Dmitry et al., “Deep Image Prior,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 9446-9454. [cited by applicant]
Vijayanarasimhan, Sudheendra et al., “SfM-Net: Learning of Structure and Motion from Video”, Apr. 25, 2017, arXiv, 9 pages. [cited by applicant]
Wang, Chaoyang et al., “Learning Depth From Monocular Videos Using Direct Methods,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018, pp. 2022-2030. [cited by applicant]
Wang, Kaixuan et al., “MVDepthNet: Real-Time Multiview Depth Estimation Neural Network,” 2018 International Conference on 3D Vision (3DV), Sep. 2018 (accessible Jul. 23, 2018), pp. 248-257, IEEE, Verona. [cited by applicant]
Wei, Xingkui et al., “DeepSFM: Structure from Motion via Deep Bundle Adjustment,” Computer Vision—ECCV 2020, Aug. 2020 (accessible Dec. 20, 2019), pp. 230-247, vol. 12346, Springer International Publishing, Cham. [cited by applicant]
Wu, Changchang, “Towards Linear-Time Incremental Structure from Motion,” 2013 International Conference on 3D Vision, Jun. 2013, pp. 127-134, IEEE, Seattle, WA, USA. [cited by applicant]
Wu, Changchang, “Critical Configurations For Radial Distortion Self-Calibration,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2014, pp. 25-32. [cited by applicant]
Yao, Yao et al., “MVSNet: Depth Inference for Unstructured Multi-view Stereo,” Proceedings of the European Conference on Computer Vision (ECCV), Sep. 2018 (accessible Apr. 7, 2018), pp. 767-783. [cited by applicant]
Yin, Zhichao et al., “GeoNet: Unsupervised Learning of Dense Depth, Optical Flow and Camera Pose,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jun. 2018 (accessible Mar. 6, 2018… [cited by applicant]
Zhan, Huangying et al., “Unsupervised Learning of Monocular Depth Estimation and Visual Odometry With Deep Feature Reconstruction,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), J… [cited by applicant]
Zhou, Huizhong et al., “DeepTAM: Deep Tracking and Mapping,” Proceedings of the European Conference on Computer Vision (ECCV), Sep. 2018 (accessible Aug. 6, 2018), pp. 822-838. [cited by applicant]
Zhou, Kevin C. et al., “Diffraction tomography with a deep image prior,” Optics Express, Apr. 27, 2020 (accessible Dec. 10, 2019), pp. 12872-12896, vol. 28, issue 9. [cited by applicant]
Zhou, Kevin C. et al., “Optical coherence refraction tomography,” Nature Photonics, Nov. 2019 (accessible Aug. 19, 2019), pp. 794-802, vol. 13, issue 11. [cited by applicant]
Zhou, Tinghui et al., “Unsupervised Learning of Depth and Ego-Motion From Video,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Jul. 2017, pp. 1851-1858. [cited by applicant]