IP Library › Granted Patent US 12,243,250
Granted Patent B1
US 12,243,250 · App. 18/519,154 · Granted Mar 4, 2025

Image capture apparatus for synthesizing a gaze-aligned view

Inventor: Henry Harlyn Baker (Los Altos, CA)
G06T7/50G06T7/13G06T7/593G06T7/70G06T2207/10012G06T2207/10024G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,250
App. No.
18/519,154
Filed
Nov 27, 2023
Granted
Mar 4, 2025
Kind
B1
Art Unit
2677
USPC
382/154
Abstract

Systems and methods of the present disclosure can facilitate determining a three-dimensional surface representation of an object. In some embodiments, the system includes a computer, a calibration module, which is configured to determine a camera geometry of a set of cameras, and an imaging module, which is configured to capture spatial images using the cameras. The computer is configured to determine epipolar lines in the spatial images, transform the spatial images with a collineation transformation, determine second derivative spatial images with a second derivative filter, construct epipolar plane edge images based on zero crossings of second derivative epipolar planes image based on the epipolar lines, select edges and compute depth estimates, sequence the edges based on contours in a spatial edge image, filter the depth estimates, and create a three-dimensional surface representation based on the filtered depth estimates and the original spatial images.

Claims (54)

1. An apparatus to facilitate synthesis of a gaze-aligned view of a face of a human from the perspective of a viewer, the apparatus comprising:

three or more cameras positionally arranged in common plane, each of the three or more cameras in a positional relationship so as to capture a respective perspective view of the face, a first subset of the three or more cameras forming a first linear camera array and a second subset of the three or more cameras forming a second linear camera array, the first linear camera array and the second linear camera array corresponding to independent dimensions of the common plane;

a frame or chassis that mechanically mounts the three or more cameras in the positional relationship;

wherein eyes of the face define a horizontal axis, which intersects the eyes, and a vertical axis, which bisects the eyes and which passes through the horizontal axis;

wherein the three or more cameras are positioned such that the respective views of the face correspond to perspectives taken of the face from each side of the vertical axis and perspective views of the face taken from each side of the horizontal axis; and

at least one electrical interface to couple electrical outputs of the three or more cameras, corresponding to images of the face, with at least one processor, the at least one processor to perform the synthesis of the gaze-aligned view of the face using the images taken of the face.

2. The apparatus of claim 1 wherein the frame or chassis is structured such that neither it nor any camera of the three or more cameras blocks a ray which extends parallel to the gaze-aligned view of the face and passes through an intersection of the horizontal axis and the vertical axis and wherein the viewer is an individual viewing an electronic display device.

3. The apparatus of claim 1 wherein: the apparatus comprises an electronic display device; the frame or chassis is mechanically coupled to the electronic display device; the at least one processor is further configured to display an image on the electronic display device; and the frame or chassis mounts each camera of the three or more cameras outside of a region imaged by the electronic display device.

4. The apparatus of claim 3 , embodied as a live video call device, wherein the apparatus is configured to transmit, to a recipient device, an image of the gaze-aligned view, and wherein the apparatus is configured to display, on the electronic display device, an image received from the recipient device.

5. The apparatus of claim 1 wherein the apparatus further comprises the at least one processor.

6. The apparatus of claim 5 wherein the at least one processor is configured to:

receive digital values corresponding to a first image from a first camera of the three or more cameras, a second image from a second one of the three or more cameras, and a third image from a third one of the three or more cameras;

create a first epipolar-plane image from the digital values;

process the first epipolar-plane image, by identifying a plurality of edges in the first epipolar-plane image and by fitting a line to the plurality of edges;

synthesize a three-dimensional model of the face from the fitted line and from image data from at least one of the three or more cameras; and

create the gaze-aligned view of the face using the three-dimensional model.

7. The apparatus of claim 6 wherein one of the first camera, the second camera and the third camera is a reference camera, wherein a corresponding one of the first image, the second image and the third image is a reference image, and wherein the at least one processor is configured to synthesize the gaze-aligned view using the fitted line and the reference image.

8. The apparatus of claim 7 wherein the reference camera is a camera which is in each of the first linear camera array and the second linear camera array.

9. The apparatus of claim 1 wherein the apparatus further comprises electronic storage, wherein the at least one processor is to store an image of the gaze-aligned view of the face, in digital electronic form, in the electronic storage, and wherein the apparatus further comprises circuitry to transmit the digital, electronic form of the gaze-aligned view of the face to remote electronic circuitry.

10. The apparatus of claim 1 wherein the three or more cameras are mechanically mounted, by the frame or chassis, such that the positional relationship corresponds to form a polygon.

11. An apparatus to synthesize a gaze-aligned view of a face of a human from the perspective of a viewer, the apparatus comprising:

a camera assembly, the camera assembly having

three or more cameras positionally arranged in common plane, each of the three or more cameras in a positional relationship so as to capture a respective perspective view of the face, a first subset of the three or more cameras forming a first linear camera array and a second subset of the three or more cameras forming a second linear camera array, the first linear camera array and the second linear camera array corresponding to independent dimensions of the common plane,

a frame or chassis that mechanically mounts the three or more cameras in the positional relationship, and

at least one electrical interface;

wherein

eyes of the face define a horizontal axis, which intersects the eyes, and a vertical axis, which bisects the eyes and which passes through the horizontal axis, and

the three or more cameras are positioned such that the respective views of the face correspond to perspectives taken of the face from each side of the vertical axis and perspective views of the face taken from each side of the horizontal axis;

electronic storage; and

at least one processor configured to receive electrical outputs of the three or more cameras, corresponding to images of the face, from the electrical interface, to synthesize the gaze-aligned view of the face using the images taken of the face, and to store an image of the gaze-aligned view of the face in the electronic storage.

12. The apparatus of claim 11 wherein the frame or chassis is structured such that neither it nor any camera of the three or more cameras blocks a ray which extends parallel to the gaze-aligned view of the face and passes through an intersection of the horizontal axis and the vertical axis, and wherein the viewer is an individual viewing the image.

13. The apparatus of claim 11 wherein: the apparatus comprises an electronic display device; the frame or chassis is mechanically coupled to the electronic display device; the at least one processor is further configured to display an image on the electronic display device; and the frame or chassis mounts each camera of the three or more cameras outside of a region imaged by the electronic display device.

14. The apparatus of claim 13 , embodied as a live video call device, wherein the apparatus is configured to transmit, to a recipient device, an image of the gaze-aligned view, and wherein the apparatus is configured to display, on the electronic display device, an image received from the recipient device.

15. The apparatus of claim 14 wherein the three or more cameras are mechanically mounted, by the frame or chassis, such that the positional relationship corresponds to form a polygon, the electronic display device positioned within the polygon.

16. The apparatus of claim 11 wherein the at least one processor is configured to:

receive digital values corresponding to a first image from a first camera of the three or more cameras, a second image from a second one of the three or more cameras, and a third image from a third one of the three or more cameras;

create a first epipolar-plane image from the digital values;

process the first epipolar-plane image, by identifying a plurality of edges in the first epipolar-plane image and by fitting a line to the plurality of edges;

synthesize a three-dimensional model of the face from the fitted line and from image data from at least one of the three or more cameras; and

create the gaze-aligned view of the face using the three-dimensional model.

17. The apparatus of claim 16 wherein one of the first camera, the second camera and the third camera is a reference camera, wherein a corresponding one of the first image, the second image and the third image is a reference image, and wherein the at least one processor is configured to synthesize the gaze-aligned view using the fitted line and the reference image.

18. The apparatus of claim 17 wherein the reference camera is a camera which is in each of the first linear camera array and the second linear camera array.

19. A video conferencing apparatus to synthesize a gaze-aligned view of a face of a first human from the perspective of a viewer, the video conferencing apparatus comprising:

an electronic display device to display video received from a recipient device associated with the viewer;

a camera assembly, the camera assembly having

three or more cameras positionally arranged in common plane, each of the three or more cameras in a positional relationship so as to capture a respective perspective view of the face, a first subset of the three or more cameras forming a first linear camera array and a second subset of the three or more cameras forming a second linear camera array, the first linear camera array and the second linear camera array corresponding to independent dimensions of the common plane,

a frame or chassis that mechanically mounts the three or more cameras in the positional relationship, and

at least one electrical interface;

wherein

eyes of the face define a horizontal axis, which intersects the eyes, and a vertical axis, which bisects the eyes,

the frame is structured such that neither it nor any camera of the three or more cameras blocks a ray which extends parallel to the gaze-aligned view of the face, passes through an intersection of the horizontal axis and the vertical axis, and intersects the electronic display device, and

the three or more cameras are positioned such that the respective views of the face correspond to perspectives taken of the face from each side of the vertical axis and perspective views of the face taken from each side of the horizontal axis;

electronic storage; and

at least one processor configured to receive electrical outputs of the three or more cameras, corresponding to images of the face, from the electrical interface, to synthesize the gaze-aligned view of the face using the images taken of the face, to store an image of the gaze-aligned view of the face in the electronic storage, and to transmit the image from the electronic storage to the recipient device.

Continuity (5)
Continuation 17744487 · May 13, 2022
Continuation 16735638 · Jan 6, 2020
Continuation 15986736 · May 22, 2018
Continuation 14887462 · Oct 20, 2015
Provisional Application 62066321 · Oct 20, 2014
References Cited (109)
US 4654872A · Hisano · 1987 [cited by applicant]
US 5359362A · Lewis et al. · 1994 [cited by applicant]
US 5613048A · Chen et al. · 1997 [cited by applicant]
US 5703961A · Rogina · 1997 [cited by examiner]
US 5710875A · Harashima et al. · 1998 [cited by applicant]
US 5764236A · Tanaka et al. · 1998 [cited by applicant]
US 5852672A · Lu · 1998 [cited by applicant]
US 5937105A · Katayama et al. · 1999 [cited by applicant]
US 5963664A · Kumar · 1999 [cited by applicant]
US 6191808B1 · Katayama et al. · 2001 [cited by applicant]
US 6608622B1 · Katayama · 2003 [cited by applicant]
US 6990228B1 · Wiles · 2006 [cited by applicant]
US 7015951B1 · Yoshigahara · 2006 [cited by examiner]
US 7142726B2 · Ziegler · 2006 [cited by applicant]
US 7224357B2 · Chen · 2007 [cited by applicant]
US 8044989B2 · Mareachen · 2011 [cited by applicant]
US 8565557B2 · Au · 2013 [cited by applicant]
US 8885882B1 · Yin · 2014 [cited by examiner]
US 9113043B1 · Kim · 2015 [cited by applicant]
US 9674506B2 · Tian · 2017 [cited by applicant]
US 9684953B2 · Kuster · 2017 [cited by examiner]
US 9786062B2 · Sorkine-Hornung · 2017 [cited by applicant]
US 10008027B1 · Baker · 2018 [cited by examiner]
US 10430994B1 · Baker · 2019 [cited by examiner]
US 10565722B1 · Baker · 2020 [cited by examiner]
US 11170561B1 · Baker · 2021 [cited by examiner]
US 11367197B1 · Baker · 2022 [cited by examiner]
US 11869205B1 · Baker · 2024 [cited by examiner]
US 20020015522A1 · Rogina · 2002 [cited by applicant]
US 20020071026A1 · Agraharam · 2002 [cited by applicant]
US 20030072483A1 · Chen · 2003 [cited by examiner]
US 20050078866A1 · Criminisi · 2005 [cited by examiner]
US 20050129325A1 · Wu · 2005 [cited by examiner]
US 20080136895A1 · Mareachen · 2008 [cited by applicant]
US 20090169057A1 · Wu · 2009 [cited by applicant]
US 20090304232A1 · Tsukizawa · 2009 [cited by examiner]
US 20100329543A1 · Li · 2010 [cited by applicant]
US 20110211068A1 · Yokota · 2011 [cited by applicant]
US 20110235897A1 · Watanabe · 2011 [cited by examiner]
US 20120019530A1 · Baker · 2012 [cited by applicant]
US 20120045100A1 · Ishigami · 2012 [cited by applicant]
US 20120263448A1 · Winter · 2012 [cited by examiner]
US 20130038696A1 · Ding · 2013 [cited by examiner]
US 20130088489A1 · Schmeitz · 2013 [cited by examiner]
US 20130208098A1 · Alcolado · 2013 [cited by applicant]
US 20140016857A1 · Richards · 2014 [cited by applicant]
US 20140132736A1 · Chang · 2014 [cited by examiner]
US 20140152647A1 · Tao · 2014 [cited by applicant]
US 20140168380A1 · Heidemann · 2014 [cited by applicant]
US 20140192164A1 · Tenn · 2014 [cited by examiner]
US 20140328535A1 · Sorkine-Hornung · 2014 [cited by applicant]
US 20150244987A1 · Delegue · 2015 [cited by applicant]
US 20150279043A1 · Bakhtiari · 2015 [cited by applicant]
US 20150304634A1 · Karvounis · 2015 [cited by applicant]
US 20150339527A1 · Plummer · 2015 [cited by examiner]
US 20150341557A1 · Chapdelaine-Couture · 2015 [cited by examiner]
US 20150381966A1 · Tian · 2015 [cited by applicant]
US 20160210776A1 · Wanner · 2016 [cited by examiner]
US 20170083087A1 · Plummer · 2017 [cited by examiner]
US 20170094243A1 · Venkataraman · 2017 [cited by applicant]
CN 106599810A1 · 2017 [cited by applicant]
EP 0637815A2 · 1995 [cited by applicant]
EP 1063614A2 · 2000 [cited by examiner]
EP 2479998A1 · 2012 [cited by applicant]
JP 2001109879A · 2001 [cited by applicant]
JP 2001183133A · 2001 [cited by applicant]
JP 2012002683A · 2012 [cited by applicant]
WO 9958927A1 · 1999 [cited by applicant]
WO WO2013127418A1 · 2013 [cited by examiner]
WO WO2015194084A1 · 2015 [cited by examiner]
Gaze Estimation From Eye Appearance: A Head Pose-Free Method via Eye Image Synthesis, Feng Lu et al., IEEE, 2015, pp. 3680-3693 (Year: 2015). [cited by examiner]
Gaze-corrected View Generation Using Stereo Camera System for Immersive Videoconferencing, Sang-Beom Lee et al., IEEE, 2011, pp. 1033-1040 (Year: 2011). [cited by examiner]
Gaze Manipulation for One-to-one Teleconferencing, A. Criminisi et al., ICCV, 2003, pp. 1-8 (Year: 2003). [cited by examiner]
Communicating Eye-gaze Across a Distance: Comparing an Eye-gaze enabled Immersive Collaborative Virtual Environment, Aligned Video Conferencing, and Being Together, David Roberts et al., IEEE, 2009, pp. 135-142 (Year: 2… [cited by examiner]
GAZE-2: Conveying Eye Contact in Group Video Conferencing Using Eye-Controlled Camera Direction, Roel Vertegaa et al., CHI, 2003, pp. 521-528 (Year: 2003). [cited by examiner]
Baker, H. Harlyn and Robert C. Bolles, “Generalizing Epipolar-Plane Image Analysis on the Spatiotemporal Surface”, International Journal of Computer Vision, 3, 33-49 (1989). [cited by applicant]
Berent, Jesse, and Pier Luigi Dragotti. “Plenoptic Manifolds—Exploiting structure and coherence in multiview images.” IEEE Signal Processing Magazine, Nov. 2007. [cited by applicant]
Berent, Jesse, and Pier Luigi Dragotti. “Segmentation of epipolar-plane image voumes with occlusion and disocclusion competition.” Multimedia Signal Processing, 2006 IEEE 8th Workshop on. IEEE, 2006. [cited by applicant]
Bishop, Tom E., and Paolo Favaro. “Full-resolution depth map estimation from an aliased plenoptic light field.” Computer Vision—ACCV 2010. Springer Berlin Heidelberg, 2011. 186-200. [cited by applicant]
Bolles, Robert C., H. Harlyn Baker, and David H. Marimont, “Epipolar-Plane Image Analysis: An Approach to Determining Structure from Motion”, International Journal of Computer Vision, I, 7-55 (1987). [cited by applicant]
Criminisi, Antonio, Jamie Shotton, Andrew Blake, and Philip HS Torr. “Gaze manipulation for one-to-one teleconferencing.” In Computer Vision, 2003. Proceedings. Ninth IEEE International Conference on, pp. 191-198. IEEE,… [cited by applicant]
Criminisi, Antonio, et al. “Extracting layers and analyzing their specular properties using epipolar-plane-image analysis.” Computer vision and image understanding 97.1 (2005): 51-85. [cited by applicant]
Dansereau, Don, and Len Bruton. “Gradient-based depth estimation from 4d light fields.” Circuits and Systems, 2004. ISCAS'04. Proceedings of the 2004 International Symposium on. vol. 3. IEEE, 2004. [cited by applicant]
Diebold, M., O. Blum, M. Gutsche, S. Wanner, C. Garbe, H. Baker, and B. Jahne. “Light-field camera design for high-accuracy depth estimation.” In SPIE Optical Metrology, pp. 952803-952803. International Society for Opti… [cited by applicant]
Ishibashi, Takashi, et al. “3D space representation using epipolar plane depth image.” Picture Coding Symposium (PCS), 2010. IEEE, 2010. [cited by applicant]
Kang, Sing Bing, Richard Szeliski, and Jinxiang Chai. “Handling occlusions in dense multi-view stereo.” Computer Vision and Pattern Recognition, 2001. CVPR 2001. Proceedings of the 2001 IEEE Computer Society Conference … [cited by applicant]
Katayama, Akihiro, Koichiro Tanaka, Takahiro Oshino, and Hideyuki Tamura. “Viewpoint-dependent stereoscopic display using interpolation of multiviewpoint images.” In IS&T/SPIE's Symposium on Electronic Imaging: Science … [cited by applicant]
Kawasaki, Hiroshi, et al. “Enhanced navigation system with real images and real-time information.” ITSWC'01 (2001). [cited by applicant]
LV, Huijin, et al. “Light field depth estimation exploiting linear structure in EPI.” Multimedia & Expo Workshops (ICMEW), 2015 IEEE International Conference on. IEEE, 2015. [cited by applicant]
Madanayake, Arjuna, et al. “VLSI architecture for 4-D depth filtering.” Signal, Image and Video Processing 9.4 (2013): 809-818. [cited by applicant]
Matousek, Martin, Tomcs Werner, and Vaclav Hlavac. “Accurate correspondences from epipolar plane images.” Proc. Computer Vision Winter Workshop. 2001. [cited by applicant]
Mellor, J. P., Seth Teller, and Tomas Lozano-Perez. “Dense depth maps from epipolar images.” (1996). [cited by applicant]
Seitz, Steven M., and Jiwon Kim. “The space of all stereo images.” International Journal of Computer Vision 48.1 (2002): 21-38. [cited by applicant]
Tao, Michael W., Sunil Hadap, Jitendra Malik, and Ravi Ramamoorthi. “Depth from Combining Defocus and Correspondence Using Light-Field Cameras.” In Computer Vision (ICCV), 2013 IEEE International Conference on, pp. 673-… [cited by applicant]
Wanner, Sven, and Bastian Goldluecke. “Spatial and angular Variational super-resolution of 4D light fields.” In Computer Vision—ECCV 2012, pp. 608-621. Springer Berlin Heidelberg, 2012. [cited by applicant]
Wanner, Sven, and Bastian Goldluecke. “Globally consistent depth labeling of 4D light fields.” Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on. IEEE, 2012. [cited by applicant]
Wanner, Sven, and Bastian Goldluecke. “Variational light field analysis for disparity estimation and super-resolution.” Pattern Analysis and Machine Intelligence, IEEE Transactions on 36.3 (2014): 606-619. [cited by applicant]
Wanner, Sven, Janis Fehr, and Bernd Jahne. “Generating EPI representations of 4D light fields with a single lens focused plenoptic camera.” Advances in Visual Computing. Springer Berlin Heidelberg, 2011. 90-101. [cited by applicant]
Wanner, Sven, Christoph Straehle, and Bastian Goldluecke. “Globally consistent multi-label assignment on the ray space of 4d light fields.” Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on. IEEE, … [cited by applicant]
Wanner, Sven, and Bastian Goldluecke. “Reconstructing reflective and transparent surfaces from epipolar plane images.” Pattern Recognition. Springer Berlin Heidelberg, 2013. 1-10. [cited by applicant]
Zheng, Jiang Yu. “Acquiring 3-D models from sequences of contours.” Pattern Analysis and Machine Intelligence, IEEE Transactions on 16.2 (1994): 163-178. [cited by applicant]
Zhu, Zhigang, Guangyou Xu, and Xueyin Lin. “Efficient Fourier-based approach for detecting orientations and occlusions in epipolar plane images for 3D scene modeling.” International journal of computer vision 61.3 (2005… [cited by applicant]
T. Wang, A. A. Efros and R. Ramamoorthi, “Depth Estimation with Occlusion Modeling Using Light-Field Cameras,” in IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 38, No. 11, pp. 2170-2181, Nov. 1, 2… [cited by applicant]
Teleconferencing system using virtual camera, Shibuichi et al., SPIE 2000, pp. 350-361 (2000). [cited by applicant]
Virtual space teleconferencing using sea of cameras, Henry Fuchs et al., ResearchGate, 1994, pp. 1-8 (1994). [cited by applicant]
Multi-view stereoscopic 3D image communication system for web based reality teleconferencing system, Park et al., SPIE 2004, pp. 263-272 (2004). [cited by applicant]
Xu et al., “A novel ray-space based view generation algorithm via radon transform,” Springer, 2013, pp. 1-15 (2013). [cited by applicant]
Matusik et al., “Image-based visual hulls,” ACM 2000, pp. 369-374 (2000). [cited by applicant]
U.S. Appl. No. 17/494,742, Office Action mailed Sep. 9, 2024, 13 pages. [cited by applicant]