IP Library Granted Patent US 12,382,009
Granted Patent B2
US 12,382,009 · App. 18/131,182 · Granted Aug 5, 2025

System and method for corrected video-see-through for head mounted displays

Inventor: Dae Hyun Lee (Etobicoke, CA)
Assignee: INTERAPTIX INC.
H04N13/344H04N13/117H04N13/239
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,382,009
App. No.
18/131,182
Granted
Aug 5, 2025
Kind
B2
Abstract

A head mounted display system with video-see-through (VST) is taught. The system and method process video images captured by at least two forward facing video cameras mounted to the headset to produce generated images whose viewpoints correspond to the viewpoint of the user if the user was not wearing the display system. By generating VST images which have viewpoints corresponding to the user's viewpoint, errors in sizing, distances and positions of objects in the VST images are prevented.

Claims (30)

1. A head mounted display system comprising:

a display capable of being worn by a user in front of their eyes and displaying images to the user;

at least eight video cameras at fixed locations relative to the display and having respective fields of view that are narrower than an expected field of view of the user and that cover a combined field of view angle of at least 180 degrees horizontally and vertically, the at least eight video cameras operable to capture video images from the respective fields of view and configured to detect distances between the display and objects located in the respective fields of view;

a computational device operable to:

obtain, via at least one of the at least eight video cameras, the distances between the display and the objects located in the respective fields of view;

receive user-selected virtual locations of the pupils of the user, the user-selected virtual locations being different to actual locations of the pupils of the user;

determine, based on the respective fields of view of the at least eight video cameras and the distances, a transformation to transform captured video images from each of the at least eight video cameras to the user-selected virtual locations of the pupils of the user; and

apply the transformation to the captured video images to generate images for providing video-see-through on the display, the generated images corresponding to a virtual viewpoint from the user-selected virtual locations of the pupils of the user to provide the user with a different real-world view from an actual real-world view that would be experienced by the user if not wearing the display, including displaying the objects located in the respective fields of view of the at least eight video cameras at depths corresponding to distances of the objects from the user-selected virtual locations of the pupils of the user, the depths computed based on the obtained distances and the respective fields of view of the at least eight video cameras relative to the user-selected virtual locations of the pupils of the user.

2. The head mounted display system according to claim 1 wherein the computational device is operable to compute the distances between the display and the objects located in the respective fields of view of the at least eight video cameras from the captured video images.

3. The head mounted display system according to claim 2 , wherein the computational device is operable to generate an image for each pupil of the user, each generated image corresponding to the virtual viewpoint of the respective pupil of the user and each generated image is displayed to the respective pupil of the user providing the user with a stereoscopic image.

4. The head mounted display system according to claim 1 wherein the computational device is operable to generate an image for each pupil of the user, each generated image corresponding to the virtual viewpoint of the respective pupil of the user and each generated image is displayed to the respective pupil of the user providing the user with a stereoscopic image.

5. The head mounted display system according to claim 1 wherein the computational device is mounted to the display.

6. The head mounted display system according to claim 1 , wherein the computational device is further operable to obtain an inter-pupil distance between the pupils of the user and determine the transformation further based on the inter-pupil distance.

7. The head mounted display system according to claim 1 wherein the computational device is connected to the display by a wire tether.

8. The head mounted display system according to claim 1 wherein the computational device is wirelessly connected to the display.

9. The head mounted display system of claim 1 , wherein the computational device is operable to determine the fixed respective fields of view of the at least eight video cameras relative to the pupils of the user based on an inter-pupil distance and an eye-to-display distance.

10. A method of operating a head mounted display worn by a user in front of their eyes, the head mounted display having at least eight video cameras operable to capture video images, the at least eight video cameras having respective fields of view that are narrower than an expected field of view of the user and that cover a combined field of view angle of at least 180 degrees horizontally and vertically and configured to detect distances between the display and objects, the method comprising the steps of:

determining respective fields of view of the at least eight video cameras relative to each pupil of the eyes of the user;

obtaining video images captured by the at least eight video cameras;

computing distances between the display and objects located in the respective fields of view of the at least eight video cameras based on the video images captured by the at least eight video cameras;

receiving user-selected virtual locations of the pupils of the user, the user-selected virtual locations being different to actual locations of the pupils of the user;

determining, based on the respective fields of view of each of the at least eight video cameras relative to the pupil of each eye of the user, and the computed distances, a transformation to transform the captured video images from each of the at least eight video cameras to the user-selected virtual locations of the pupils of the user;

applying the transformation to the captured video images to render generated images corresponding to a virtual viewpoint from the user-selected virtual locations of the pupils of the user to provide the user with a different real-world view from an actual real-world view that would be experienced by the user if not wearing the display wherein the generated images display the objects at depths corresponding to distances of the objects from the user-selected virtual locations of the pupils of the user, the depths computed based on the computed distances and the respective fields of view of the at least eight video cameras relative to the user-selected virtual locations of the pupils of the user; and

providing video-see-through to the user by displaying the generated images to the user on the head mounted display.

11. The method of claim 10 , further comprising processing the captured video images to render a respective generated image for each pupil of the user, each respective generated image corresponding to the virtual viewpoint of the respective pupil of the user.

12. The method of claim 10 , further comprising obtaining an inter-pupil distance between the pupils of the user and determining the transformation further based on the inter-pupil distance.

13. The head mounted display system of claim 1 wherein the generated images appear enlarged or magnified with respect to the actual real-world view.

14. The head mounted display system of claim 1 wherein the virtual viewpoint defines the pupils of the eyes of the user as being located to one side of, above, or below the user.

15. The method of claim 10 wherein the generated images appear enlarged or magnified with respect to the actual real-world view.

16. The method of claim 10 wherein the virtual viewpoint defines the pupils of the eyes of the user as being located to one side of, above, or below the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2023
From: LEE, DAE HYUN
To: INTERAPTIX INC.
Reel/Frame 064232/0761 →
Continuity (3)
Continuation 16946860 · Jul 9, 2020
Provisional Application 62871783 · Jul 9, 2019
Related Publication 20230239457A1 · Jul 27, 2023
References Cited (37)
US 10437065B2 · Zhu et al. · 2019 [cited by applicant]
US 20150206329A1 · Devries · 2015 [cited by examiner]
US 20170118458A1 · Gronholm et al. · 2017 [cited by applicant]
US 20180033209A1 · Akeley · 2018 [cited by examiner]
Ballan, Luca, et al. “Unstructured video-based rendering.” ACM Transactions on Graphics 29.4 (2010): 1-11. [cited by applicant]
Barron, Jonathan T. et al., “The fast bilateral solver.” European conference on computer vision. Cham: Springer International Publishing, 2016. [cited by applicant]
Bleyer, Michael et al., “Patchmatch stereo-stereo matching with slanted support windows.” Bmvc. vol. 11. 2011. [cited by applicant]
Chaurasia, Gaurav, et al. “Depth synthesis and local warps for plausible image-based navigation.” ACM Transactions on Graphics (TOG) 32.3 (2013): 1-12. [cited by applicant]
Chaurasia, Gaurav et al. “Silhouette-Aware Warping for Image-Based Rendering.” Computer Graphics Forum. vol. 30. No. 4. Oxford, UK: Blackwell Publishing Ltd, 2011. [cited by applicant]
Chen, Shenchang Eric et al. “View interpolation for image synthesis.” Proceedings of the 20th annual conference on Computer graphics and interactive techniques. 1993. [cited by applicant]
Chen, Shenchang Eric. “Quicktime VR: An image-based approach to virtual environment navigation.” Proceedings of the 22nd annual conference on Computer graphics and interactive techniques. 1995. [cited by applicant]
Davis, James et al. “Spacetime stereo: A unifying framework for depth from triangulation.” 2003 IEEE Computer Society Conference on Computer Vision and Pattern Recognition, 2003. Proceedings.. vol. 2. IEEE, 2003. [cited by applicant]
Di Martino, Matias et al. “An Analysis and Implementation of Multigrid Poisson Solvers With Verified Linear Complexity.” Image Processing on Line 8 (2018): 192-218. [cited by applicant]
Fanello, Sean Ryan, et al. “Low compute and fully parallel computer vision with hashmatch.” Proceedings of the IEEE International Conference on Computer Vision. 2017. [cited by applicant]
Goesele, Michael, et al. “Ambient point clouds for view interpolation.” ACM Transactions on Graphics (TOG) 29.4 (2010): 1-6. [cited by applicant]
Hedman, Peter, et al. “Casual 3D photography.” ACM Transactions on Graphics (TOG) 36.6 (2017): 1-15. [cited by applicant]
Hedman, Peter et al. “Instant 3d photography.” ACM Transactions on Graphics (TOG) 37.4 (2018): 1-12. [cited by applicant]
Hirschmuller, Heiko. “Stereo processing by semiglobal matching and mutual information.” IEEE Transactions on pattern analysis and machine intelligence 30.2 (2008): 328-341. [cited by applicant]
Hirschmuller, Heiko, Maximilian Buder, and Ines Ernst. “Memory efficient semi-global matching.” ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences 1 (2012): 371-376. [cited by applicant]
Holynski, Aleksander et al. “Fast depth densification for occlusion-aware augmented reality.” ACM Transactions on Graphics (ToG) 37.6 (2018): 1-11. [cited by applicant]
Hornung, Alexander et al. “Interactive pixel-accurate free viewpoint rendering from images with silhouette aware sampling.” Computer Graphics Forum. vol. 28. No. 8. Oxford, UK: Blackwell Publishing Ltd, 2009. [cited by applicant]
Jakubowski, M. et al. “Block-based motion estimation algorithms—a survey.” Opto-Electronics Review 21 (2013): 86-102. [cited by applicant]
Kanade, Takeo, et al. “A Stereo Machine for Video-Rate Dense Depth Mapping and Its New Applications.” Proceedings of the 1996 Conference on Computer Vision and Pattern Recognition (CVPR'96). 1996. [cited by applicant]
Levin, Anat et al. “Colorization using optimization.” ACM SIGGRAPH 2004 Papers. 2004. 689-694. [cited by applicant]
Lipski, C., et al. “Virtual Video Camera: Image-Based Viewpoint Navigation Through Space and Time.” Computer Graphics Forum. vol. 8. No. 29. 2010. [cited by applicant]
Mazumdar, Amrita, et al. “A hardware-friendly bilateral solver for real-time virtual reality video.” Proceedings of High Performance Graphics. 2017. 1-10. [cited by applicant]
McMillan, Leonard et al. “Plenoptic modeling: An image-based rendering system.” Proceedings of the 22nd annual conference on Computer graphics and interactive techniques. 1995. [cited by applicant]
Nover Harris et al. “ESPReSSo: efficient slanted PatchMatch for real-time spacetime stereo.” 2018 International Conference on 3D Vision (3DV). IEEE, 2018. [cited by applicant]
Ostermann, Jorn, et al. “Video coding with H. 264/AVC: tools, performance, and complexity.” IEEE Circuits and Systems magazine 4.1 (2004): 7-28. [cited by applicant]
Perez, Patrick et al. “Poisson image editing.” ACM SIGGRAPH 2003 Papers. 2003. 313-318. [cited by applicant]
Richardt, Christian, et al. “Real-time spatiotemporal stereo matching using the dual-cross-bilateral grid.” Computer Vision-ECCV 2010: 11th European Conference on Computer Vision, Heraklion, Crete, Greece, Sep. 5-11, 20… [cited by applicant]
Scharstein, Daniel et al. “A taxonomy and evaluation of dense two-frame stereo correspondence algorithms.” International journal of computer vision 47 (2002): 7-42. [cited by applicant]
Chan, S. C et al. “Image-based rendering and synthesis.” IEEE Signal Processing Magazine 24.6 (2007): 22-33. [cited by applicant]
Stich, Timo, et al. “View and time interpolation in image space.” Computer Graphics Forum. vol. 27. No. 7. Oxford, UK: Blackwell Publishing Ltd, 2008. [cited by applicant]
Szeliski, Richard. “Locally adapted hierarchical basis preconditioning.” ACM SIGGRAPH 2006 Papers. 2006. 1135-1143. [cited by applicant]
Valentin, Julien, et al. “Depth from motion for smartphone AR.” ACM Transactions on Graphics (ToG) 37.6 (2018): 1-19. [cited by applicant]
Szeliski, Richard, “Computer Vision: Algorithms and Applications”, 2nd Ed., 2022, Retrieved from the Internet on Jul. 5, 2023 from URL: http://szeliski.org/Book/. [cited by applicant]