IP Library Granted Patent US 12,236,521
Granted Patent B1
US 12,236,521 · App. 17/494,742 · Granted Feb 25, 2025

Techniques for determining a three-dimensional textured representation of a surface of an object from a set of images with varying formats

Inventor: Henry Harlyn Baker (Los Altos, CA)
G06T15/205G06T7/13G06T7/55G06T2207/10024G06T2207/10028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,521
App. No.
17/494,742
Granted
Feb 25, 2025
Kind
B1
Abstract

Systems and methods of the present disclosure can facilitate determining a three-dimensional surface representation of an object. In some embodiments, the system includes a computer, a calibration module, which is configured to determine a camera geometry of a set of cameras, and an imaging module, which is configured to capture spatial images using the cameras. The computer is configured to determine epipolar lines in the spatial images, transform the spatial images with a collineation transformation, determine second derivative spatial images with a second derivative filter, construct epipolar plane edge images based on zero crossings of second derivative epipolar planes image based on the epipolar lines, select edges and compute depth estimates, sequence the edges based on contours in a spatial edge image, filter the depth estimates, and create a three-dimensional surface representation based on the filtered depth estimates and the original spatial images.

Claims (156)

1. A method of computing a depth estimate of a point on an object, comprising:

receiving a first plurality of images of the object from a first source;

receiving a second plurality of images of the object from a second source; and

with at least one processor,

creating a first epipolar plane image based on the first plurality of images,

creating a second epipolar plane image based on the second plurality of images, and

computing the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein computing the depth estimate further comprises selecting a first edge from the first epipolar plane image, corresponding to the point on the object, selecting a second edge from the second epipolar plane image, corresponding to the first edge, and computing the depth estimate based on the first edge and the second edge;

wherein the method further comprises, with the at least one processor:

identifying a first contour, in a first reference image, the first contour comprising a first plurality of contour edges;

identifying a second contour, in a second reference image, the second contour comprising a second plurality of contour edges;

mapping the first edge to at least one contour edge of the first plurality of contour edges; and

mapping the second edge to at least one contour edge of the second plurality of contour edges; and

wherein the computing of the depth estimate is based on the first contour and the second contour.

2. The method of claim 1 , wherein:

the first source comprises a first camera module having two or more cameras arranged along the first line, within a threshold; and

the second source comprises a second camera module having two or more cameras arranged along the second line, within a threshold.

3. The method of claim 1 , wherein:

the method further comprises receiving a third plurality of images from a third source and, with the at least one processor, creating a third epipolar plane image based on the third plurality of images; and

the computing of the depth estimate is also based on the third epipolar plane image.

4. A method of computing a depth estimate of a point on an object, comprising:

receiving a first plurality of images of the object from a first source;

receiving a second plurality of images of the object from a second source; and

with at least one processor,

creating a first epipolar plane image based on the first plurality of images,

creating a second epipolar plane image based on the second plurality of images, and

computing the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein the first source comprises a first camera module having two or more cameras arranged along a first line, within a threshold;

wherein the second source comprises a second camera module having two or more cameras arranged along a second line, within a threshold; and

wherein:

the method further comprises receiving a third plurality of images from a third source and, with the at least one processor, creating a third epipolar plane image based on the third plurality of images;

the computing of the depth estimate is also based on the third epipolar plane image;

the third source comprises a third camera module having two or more cameras arranged along a third line, within a threshold;

the first line does not intersect the second; and

the third line intersects the first line and the second line.

5. The method of claim 4 , wherein computing the depth estimate further comprises:

selecting a first edge from the first epipolar plane image, corresponding to the point on the object;

selecting a second edge from the second epipolar plane image, corresponding to the first edge; and

computing the depth estimate based on the first edge and the second edge.

6. The method of claim 5 , wherein:

the method further comprises, with the at least one processor,

identifying a first contour, in a first reference image, the first contour comprising a first plurality of contour edges,

identifying a second contour, in a second reference image, the second contour comprising a second plurality of contour edges,

mapping the first edge to at least one contour edge of the first plurality of contour edges,

mapping the second edge to at least one contour edge of the second plurality of contour edges; and

the computing of the depth estimate is based on the first contour and the second contour.

7. A method of computing a depth estimate of a point on an object, comprising:

receiving a first plurality of images of the object from a first source;

receiving a second plurality of images of the object from a second source; and

with at least one processor,

creating a first epipolar plane image based on the first plurality of images,

creating a second epipolar plane image based on the second plurality of images, and

computing the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein the first source comprises a first camera module having two or more cameras arranged along a first line, within a threshold;

wherein the second source comprises a second camera module having two or more cameras arranged along a second line, within a threshold; and

wherein:

the first line intersects the second line;

the specific image is produced by a specific camera which is positioned on the first line, within a threshold, and on the second line, within a threshold; and

the method further comprises using the specific image as a reference image.

8. The method of claim 7 , further comprising, with the at least one processor, creating a three-dimensional representation of the object based on the depth estimate.

9. The method of claim 7 , wherein:

the first source comprises a first plurality of cameras;

the second source comprises a second plurality of cameras;

the cameras of the first plurality of cameras are colinear within a threshold; and

the cameras of the second plurality of cameras are collinear within a threshold.

10. A system for computing a depth estimate of a point on an object, comprising:

a first source;

a second source; and

at least one processor configured to

receive a first plurality of images of the object, from the first source,

receive a second plurality of images of the object, from the second source,

create a first epipolar plane image based on the first plurality of images,

create a second epipolar plane image based on the second plurality of images, and

compute the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein the at least one processor is further configured to select a first edge from the first epipolar plane image, corresponding to the point on the object, select a second edge from the second epipolar plane image, corresponding to the first edge, and compute the depth estimate based on the first edge and the second edge; and

wherein the at least one processor is further configured to:

identify a first contour, in a first reference image, the first contour comprising a first plurality of contour edges;

identify a second contour, in a second reference image, the second contour comprising a second plurality of contour edges;

map the first edge to at least one contour edge of the first plurality of contour edges;

map the second edge to at least one contour edge of the second plurality of contour edges; and

compute the depth estimate using the first contour and the second contour.

11. The system of claim 10 , wherein:

the first source comprises a first camera module comprising two or more cameras arranged along the first line, within a threshold; and

the second source comprises a second camera module having two or more cameras arranged along the second line, within a threshold.

12. The system of claim 10 , wherein:

the system further comprises a third source; and

the at least one processor is further configured to

receive a third plurality of images from the third source and create a third epipolar plane image based on the third plurality of images, and

compute the depth estimate also based on the third epipolar plane image.

13. A system for computing a depth estimate of a point on an object, comprising:

a first source;

a second source; and

at least one processor configured to

receive a first plurality of images of the object, from the first source,

receive a second plurality of images of the object, from the second source,

create a first epipolar plane image based on the first plurality of images,

create a second epipolar plane image based on the second plurality of images, and

compute the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein the first source comprises a first camera module comprising two or more cameras arranged along the first line, within a threshold;

wherein the second source comprises a second camera module having two or more cameras arranged along the second line, within a threshold; and

wherein:

the system further comprises a third source;

the at least one processor is further configured to

receive a third plurality of images from the third source,

create a third epipolar plane image based on the third plurality of images, and

compute the depth estimate based on the third epipolar plane image;

the third source comprises a third camera module having two or more cameras arranged along a third line, within a threshold;

the first line does not intersect the second line; and

the third line intersects each of the first line and the second line.

14. The system of claim 13 , wherein the at least one processor is further configured to:

select a first edge from the first epipolar plane image, corresponding to the point on the object;

select a second edge from the second epipolar plane image, corresponding to the first edge; and

compute the depth estimate based on the first edge and the second edge.

15. The system of claim 14 , wherein the at least one processor is further configured to:

identify a first contour, in a first reference image, the first contour comprising a first plurality of contour edges;

identify a second contour, in a second reference image, the second contour comprising a second plurality of contour edges;

map the first edge to at least one contour edge of the first plurality of contour edges;

map the second edge to at least one contour edge of the second plurality of contour edges; and

compute the depth estimate using the first contour and the second contour.

16. A system for computing a depth estimate of a point on an object, comprising:

a first source;

a second source; and

at least one processor configured to

receive a first plurality of images of the object, from the first source,

receive a second plurality of images of the object, from the second source,

create a first epipolar plane image based on the first plurality of images,

create a second epipolar plane image based on the second plurality of images, and

compute the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein the first source comprises a first camera module comprising two or more cameras arranged along the first line, within a threshold;

wherein the second source comprises a second camera module having two or more cameras arranged along the second line, within a threshold; and

wherein:

the first line intersects the second line;

the specific image is produced by a specific camera, the specific camera positioned along the first line, within a threshold, and along the second line, within a threshold; and

the at least one processor is further configured to use the specific image as a reference image.

17. The system of claim 16 , wherein:

the first source comprises a first plurality of cameras;

the second source comprises a second plurality of cameras;

the cameras of the first plurality of cameras are collinear within a threshold; and

the cameras of the second plurality of cameras are collinear within a threshold.

18. The system of claim 16 , wherein the at least one processor is configured to create a three-dimensional representation of the object based on the depth estimate.

19. An apparatus for computing a depth estimate of a point on an object, the apparatus comprising instructions stored on at least one computer-readable non-transient storage medium, said instructions when executed to cause at least one processor to:

receive a first plurality of images from a first source;

receive a second plurality of images from a second source;

create a first epipolar plane image based on the first plurality of images;

create a second epipolar plane image based on the second plurality of images; and

compute the depth estimate based on the first epipolar plane image and the second epipolar plane image;

wherein said instructions, when executed, are to cause the at least one processor to:

identify a first contour, in a first reference image, the first contour comprising a first plurality of contour edges;

identify a second contour, in a second reference image, the second contour comprising a second plurality of contour edges;

map the first edge to at least one contour edge of the first plurality of contour edges;

map the second edge to at least one contour edge of the second plurality of contour edges; and

compute the depth estimate using the first contour and the second contour.

20. The apparatus of claim 19 , wherein said instructions, when executed, are to cause the at least one processor to:

select a first edge from the first epipolar plane image, corresponding to the point on the object;

select a second edge from the second epipolar plane image, corresponding to the first edge; and

compute the depth estimate based on the first edge and the second edge.

Continuity (3)
Continuation 16538029 · Aug 12, 2019
Division 15802777 · Nov 3, 2017
Provisional Application 62418718 · Nov 7, 2016
References Cited (95)
US 5359362A · Lewis et al. · 1994 [cited by applicant]
US 5613048A · Chen et al. · 1997 [cited by applicant]
US 5710875A · Harashima et al. · 1998 [cited by applicant]
US 5764236A · Tanaka et al. · 1998 [cited by applicant]
US 5852672A · Lu · 1998 [cited by applicant]
US 5937105A · Katayama et al. · 1999 [cited by applicant]
US 5963664A · Kumar et al. · 1999 [cited by applicant]
US 6072496A · Guenter et al. · 2000 [cited by applicant]
US 6188777B1 · Darrell et al. · 2001 [cited by applicant]
US 6191808B1 · Katayama et al. · 2001 [cited by applicant]
US 6215898B1 · Woodfill et al. · 2001 [cited by applicant]
US 6608622B1 · Katayama et al. · 2003 [cited by applicant]
US 6661918B1 · Gordon et al. · 2003 [cited by applicant]
US 6990228B1 · Wiles et al. · 2006 [cited by applicant]
US 7003134B1 · Covell et al. · 2006 [cited by applicant]
US 7142726B2 · Ziegler et al. · 2006 [cited by applicant]
US 7317830B1 · Gordon et al. · 2008 [cited by applicant]
US 7664315B2 · Woodfill et al. · 2010 [cited by applicant]
US 7970177B2 · Hilaire et al. · 2011 [cited by applicant]
US 8743214B2 · Grossmann et al. · 2014 [cited by applicant]
US 8872897B2 · Grossmann et al. · 2014 [cited by applicant]
US 8988317B1 · Liang et al. · 2015 [cited by applicant]
US 9113043B1 · Kim et al. · 2015 [cited by applicant]
US 9240048B2 · Tao · 2016 [cited by examiner]
US 9300946B2 · Do et al. · 2016 [cited by applicant]
US 9361660B2 · Tanaka · 2016 [cited by applicant]
US 9462164B2 · Venkataraman · 2016 [cited by examiner]
US 9674504B1 · Salvagnini · 2017 [cited by examiner]
US 9674506B2 · Tian et al. · 2017 [cited by applicant]
US 9786062B2 · Sorkine-Hornung et al. · 2017 [cited by applicant]
US 10008027B1 · Baker · 2018 [cited by examiner]
US 10119808B2 · Venkataraman · 2018 [cited by examiner]
US 10122994B2 · Yücer · 2018 [cited by examiner]
US 10430994B1 · Baker · 2019 [cited by examiner]
US 10832429B2 · Blasco Claret · 2020 [cited by examiner]
US 10887581B2 · Yücer · 2021 [cited by examiner]
US 11170561B1 · Baker · 2021 [cited by examiner]
US 20030072483A1 · Chen · 2003 [cited by applicant]
US 20120019530A1 · Baker · 2012 [cited by applicant]
US 20120045100A1 · Ishigami · 2012 [cited by applicant]
US 20130044181A1 · Baker · 2013 [cited by applicant]
US 20140152647A1 · Tao et al. · 2014 [cited by applicant]
US 20140327674A1 · Sorkine-Hornung · 2014 [cited by applicant]
US 20140328535A1 · Sorkine-Hornung · 2014 [cited by applicant]
US 20150279043A1 · Bakhtiari et al. · 2015 [cited by applicant]
US 20150304634A1 · Karvounis · 2015 [cited by applicant]
US 20150381966A1 · Tian et al. · 2015 [cited by applicant]
US 20160210776A1 · Wanner · 2016 [cited by examiner]
US 20170094243A1 · Venkataraman · 2017 [cited by examiner]
US 20180139436A1 · Yücer · 2018 [cited by examiner]
US 20190236796A1 · Blasco Claret · 2019 [cited by examiner]
CN 205365796U · 2016 [cited by applicant]
WO 2018072817A1 · 2018 [cited by applicant]
WO 2018072858A1 · 2018 [cited by applicant]
Baker, H. Harlyn and Robert C. Bolles, “Generalizing Epipolar-Plane Image Analysis on the Spatiotemporal Surface,” International Journal of Computer Vision, 3, pp. 33-49 (1989). [cited by applicant]
Berent, Jesse, and Pier Luigi Dragotti, “Plenoptic Manifolds—Exploiting structure and coherence in multiview Images,” IEEE Signal Processing Magazine, Nov. 2007, 11 pages. [cited by applicant]
Berent, Jesse, and Pier Luigi Dragotti, “Segmentation of epipolar-plane image vols. with occlusion and disocclusion competition,” Multimedia Signal Processing, 2006 IEEE 8th Workshop on, IEEE 2006, 4 pages. [cited by applicant]
Bishop, Tom E., and Paolo Favaro, “Full-resolution depth map estimation from an aliased plenoptic light field.” Computer Vision—ACCV 2010, Springer Berlin Heidelberg, 2011, pp. 186-200. [cited by applicant]
Bolles, Robert C., H. Harlyn Baker, and David H. Marimont, “Epipolar-Plane Image Analysis: an Approach to Determining Structure from Motion,” International Journal of Computer Vision, I, 7-55 (1987). [cited by applicant]
Criminisi, Antonio, Jamie Shotton, Andrew Blake, and Philip HS Torr, “Gaze manipulation for one-to-one teleconferencing,” Computer Vision, 2003 Proceedings. Ninth IEEE International Conference on, pp. 191-198, IEEE, 200… [cited by applicant]
Criminisi, Antonio, et al., Extracting layers and analyzing their specular properties using epipolar-plane-image analysis., Computer vision and image understanding 97.1 (2005): 51-85. [cited by applicant]
Dansereau, Don, and Len Bruton, “Gradient-based depth estimation from 4d light fields,” Circuits and Systems, 2004, ISCAS'04, Proceedings of the 2004 International Symposium on. vol. 3, IEEE 2004, 4 pages. [cited by applicant]
Diebold, M., O. Blum, M. Gutsche, S. Wanner, C. Garbe, H. Baker, and B. Jahne, “Light-field camera design for high-accuracy depth estimation,” SPIE Optical Metrology, pp. 952803-952803. International Society for Optics … [cited by applicant]
Ishibashi, Takashi, et al., “3D space representation using epipolar plane depth image,” Picture Coding Symposium (PCS), IEEE 2010, 4 pages. [cited by applicant]
Kang, Sing Bing, Richard Szeliski, and Jinxiang Chai, “Handling occlusions in dense multi-view stereo.” Computer Vision and Pattern Recognition, 2001, CVPR 2001, Proceedings of the 2001 IEEE Computer Society Conference … [cited by applicant]
Katayama, Akihiro, Koichiro Tanaka, Takahiro Oshino, and Hideyuki Tamura, “Viewpoint-dependent stereoscopic display using interpolation of multiviewpoint images,” IS&T/SPIE Symposium on Electronic Imaging: Science & Tec… [cited by applicant]
Kawasaki, Hiroshi, et al., “Enhanced navigation system with real images and real-time information,” ITSWC'01 (2001), 11 pages. [cited by applicant]
Lv, Huijin, et al., “Light field depth estimation exploiting linear structure in EPI,” Multimedia & Expo Workshops (ICMEW), 2015 IEEE International Conference on, IEEE 2015. [cited by applicant]
Madanayake, Arjuna, et al., “VLSI architecture for 4-D depth filtering,” Signal, Image and Video Processing 9.4 (2013): 809-818. [cited by applicant]
Matousek, Martin, Tomcs Werner, and Vaclav Hlavac, “Accurate correspondences from epipolar plane images.” Proc. Computer Vision Winter Workshop 2001, 9 pages. [cited by applicant]
Mellor, J. P., Seth Teller, and Tomas Lozano-Perez, “Dense depth maps from epipolar images” (1996), 13 pages. [cited by applicant]
Seitz, Steven M., and Jiwon Kim, “The space of all stereo images,” International Journal of Computer Vision 48.1 (2002): 21-38. [cited by applicant]
Tao, Michael W., Sunil Hadap, Jitendra Malik, and Ravi Ramamoorthi, “Depth from Combining Defocus and Correspondence Using Light-Field Cameras,” Computer Vision (ICCV), 2013 IEEE International Conference on, pp. 673-680… [cited by applicant]
Wanner, Sven, and Bastian Goldluecke, “Spatial and angular variational super-resolution of 4D light fields.” In Computer Vision-ECCV 2012, pp. 608-621, Springer Berlin Heidelberg, 2012. [cited by applicant]
Wanner, Sven, and Bastian Goldluecke, “Globally consistent depth labeling of 4D light fields,” Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, IEEE 2012, 8 pages. [cited by applicant]
Wanner, Sven, and Bastian Goldluecke, “Variational light field analysis for disparity estimation and super-resolution,” Pattern Analysis and Machine Intelligence, IEEE Transactions on, 36.3 (2014): 606-619. [cited by applicant]
Wanner, Sven, Janis Fehr, and Bernd Jahne, “Generating EPI representations of 4D light fields with a single lens focused plenoptic camera,” Advances in Visual Computing. Springer Berlin Heidelberg 2011, 90-101. [cited by applicant]
Wanner, Sven, Christoph Straehle, and Bastian Goldluecke, “Globally consistent multi-label assignment on the ray space of 4d light fields,” Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, IEEE 2… [cited by applicant]
Wanner, Sven, and Bastian Goldluecke, “Reconstructing reflective and transparent surfaces from epipolar plane images,” Pattern Recognition, Springer Berlin Heidelberg, 2013, pp. 1-10. [cited by applicant]
Zheng, Jiang Yu, “Acquiring 3-D models from sequences of contours,” Pattern Analysis and Machine Intelligence, IEEE Transactions on 16.2 (1994): 163-178. [cited by applicant]
Zhu, Zhigang, Guangyou Xu, and Xueyin Lin, “Efficient Fourier-based approach for detecting orientations and pcclusions in epipolar plane images for 3D scene modeling,” International journal of computer vision 61.3 (2005… [cited by applicant]
Chen, Jie, and Lap-Pui Chau, “A fast adaptive guided filtering algorithm for light field depth interpolation,” 2014 IEEE International Symposium on Circuits and Systems (ISCAS), IEEE 2014, 10 pages. [cited by applicant]
Dragotti, Pier Luigi, and Mike Brookes. “Efficient segmentation and representation of multi-view images.” SEAS-DTC workshop, SEAS-DTC workshop, Edinburgh 2007, 7 pages. [cited by applicant]
Saksen, Aaron. “Dynamically Reparameterized Light Fields”, masters' thesis, Massachusetts Institute of Technology, Nov. 28, 2000, 79 pages. [cited by applicant]
Kim, Changil, et al. “Scene reconstruction from high spatio-angular resolution light fields.” ACM Trans. Graph. 32.4 (2013): 73-1, 10 pages. [cited by applicant]
Tanimoto, Masayuki. “Overview of free viewpoint television”, in Signal Processing: Image Communication, vol. 21, iss. 6, pp. 454-461, Jul. 2006. [cited by applicant]
Tosic, Ivana, and Kathrin Berkner. “Light field scale-depth space transform for dense depth estimation.” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops. 2014, pp. 435-442. [cited by applicant]
Wilburn, Bennett, Neel Joshi, Vaibhav Vaish, Marc Levoy, and Mark Horowitz. “High-speed videography using a dense camera array”, in Proc. 2004 IEEE Computer Society Conference on Computer Vision and Pattern Recognition,… [cited by applicant]
Wilburn, Bennett. “High Performance Imaging Using Arrays of Inexpensive Cameras”, PhD thesis, Stanford University, Dec. 2004, 128 pages. [cited by applicant]
Zhang, Cha, and Tsuhan Chen. “A Self-Reconfigurable Camera Array”, in H. W. Jensen and A. Keller, eds., Eurographics Symposium on Rendering, 2004. [cited by applicant]
Ekmekcioglu, Erhan, Vladan Velisavljević, and Stewart T. Worrall, “Efficient edge, motion and depth-range adaptive processing for enhancement of multi-view depth map sequences,” 2009 16th IEEE International Conference o… [cited by applicant]
Bebis, George. CS491E/791E: Computer Vision, class notes. Spring 2004. Retrieved from: http://www.cse.unr.edu/˜bebis/CS791E/Notes/EpipolarGeonetry.pdf, 16 pages. [cited by applicant]
Johnston, Douglas V. CS229: Machine Learning, class project, “Learning Depth in Light Field Images” 2005. Retrieved from: http://cs229.stanford.edu/proj2005/Johnston-LearningDepthInLightfieldImages.pdf, 4 pages. [cited by applicant]
Levin, Anat, et al. “Image and depth from a conventional camera with a coded aperture.” ACM transactions on graphics (TOG) 26.3 (Jul. 2007): 70, 10 pages. [cited by applicant]
Ng, Ren. “Digital light field photography”, PhD thesis, Stanford University, Jul. 2006, 203 pages. [cited by applicant]