IP Library Granted Patent US 12,437,486
Granted Patent B2
US 12,437,486 · App. 18/327,605 · Granted Oct 7, 2025

Generation and rendering of extended-view geometries in video see-through (VST) augmented reality (AR) systems

Inventor: Yingen Xiong (Mountain View, CA)
Assignee: Samsung Electronics Co., Ltd.
G06T19/006G06T3/40G06T5/20G06V10/44
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,437,486
App. No.
18/327,605
Granted
Oct 7, 2025
Kind
B2
Abstract

A method includes obtaining multiple see-through image frames of an environment around an augmented reality (AR) device using multiple imaging sensors of the AR device. The method also includes generating a depth map based on the see-through image frames and generating a three-dimensional (3D) representation of the environment based on the depth map. The method further includes projecting the 3D representation onto a curved surface, mapping points of the projected 3D representation to multiple virtual view images, and presenting the virtual view images on one or more displays of the AR device. Generating the depth map may include generating an initial depth map using a trained machine learning model and modifying the initial depth map to provide both spatial consistency and temporal consistency in order to generate a refined depth map. The curved surface may include a portion of a cylindrical, spherical, or conical surface.

Claims (95)

1. A method comprising:

obtaining multiple see-through image frames of an environment around an augmented reality (AR) device using multiple imaging sensors of the AR device;

generating a depth map based on the see-through image frames;

generating a three-dimensional (3D) representation of the environment based on the depth map;

projecting the 3D representation onto a curved surface;

mapping points of the projected 3D representation to multiple virtual view images; and

presenting the virtual view images on one or more displays of the AR device.

2. The method of claim 1 , wherein generating the depth map comprises:

generating an initial depth map based on the see-through image frames;

identifying first image features associated with a first of the see-through image frames captured using a first of the imaging sensors;

identifying second image features associated with a second of the see-through image frames captured using a second of the imaging sensors;

generating predicted second image features associated with the second see-through image frame based on the first image features and relative positions of the first and second imaging sensors;

generating a confidence map associated with the initial depth map based on differences between the second image features and the predicted second image features; and

refining the initial depth map based on the confidence map to generate a refined depth map.

3. The method of claim 2 , wherein generating the initial depth map comprises providing the see-through image frames as inputs to a trained machine learning model, the trained machine learning model configured to generate the initial depth map using the inputs.

4. The method of claim 1 , wherein generating the depth map comprises:

identifying first features associated with a sequence of see-through image frames;

identifying second features associated with a sequence of initial depth maps, each initial depth map associated with one of the see-through image frames; and

performing image-guided depth filtering and refinement using the first and second features to generate a refined depth map, the image-guided depth filtering and refinement comprising weight-based depth filtering using weights that are based on the first and second features associated with a current one of the see-through image frames and one or more neighboring ones of the see-through image frames.

5. The method of claim 1 , wherein generating the depth map comprises:

generating an initial depth map using a trained machine learning model; and

modifying the initial depth map to provide both spatial consistency and temporal consistency in order to generate a refined depth map, the spatial consistency associated with depth consistencies between different ones of the see-through image frames captured at a common time, the temporal consistency associated with depth consistencies between different ones of the see-through image frames captured at different times.

6. The method of claim 1 , wherein:

a resolution of the depth map is lower than a resolution of the see-through image frames; and

projecting the 3D representation to the curved surface comprises:

increasing the resolution of the depth map to generate a higher-resolution depth map;

verifying depths contained in the higher-resolution depth map; and

projecting the 3D representation of the environment to the curved surface using verified depths contained in the higher-resolution depth map.

7. The method of claim 1 , wherein the curved surface comprises a portion of a cylindrical, spherical, or conical surface.

8. An augmented reality (AR) device comprising:

imaging sensors configured to capture multiple see-through image frames of an environment around the AR device;

one or more displays; and

at least one processing device configured to:

generate a depth map based on the see-through image frames;

generate a three-dimensional (3D) representation of the environment based on the depth map;

project the 3D representation onto a curved surface;

map points of the projected 3D representation to multiple virtual view images; and

initiate presentation of the virtual view images on the one or more displays.

9. The AR device of claim 8 , wherein, to generate the depth map, the at least one processing device is configured to:

generate an initial depth map based on the see-through image frames;

identify first image features associated with a first of the see-through image frames captured using a first of the imaging sensors;

identify second image features associated with a second of the see-through image frames captured using a second of the imaging sensors;

generate predicted second image features associated with the second see-through image frame based on the first image features and relative positions of the first and second imaging sensors;

generate a confidence map associated with the initial depth map based on differences between the second image features and the predicted second image features; and

refine the initial depth map based on the confidence map to generate a refined depth map.

10. The AR device of claim 9 , wherein, to generate the initial depth map, the at least one processing device is configured to provide the see-through image frames as inputs to a trained machine learning model, the trained machine learning model configured to generate the initial depth map using the inputs.

11. The AR device of claim 8 , wherein, to generate the depth map, the at least one processing device is configured to:

identify first features associated with a sequence of see-through image frames;

identify second features associated with a sequence of initial depth maps, each initial depth map associated with one of the see-through image frames; and

perform image-guided depth filtering and refinement using the first and second features to generate a refined depth map; and

wherein, to perform the image-guided depth filtering and refinement, the at least one processing device is configured to perform weight-based depth filtering using weights that are based on the first and second features associated with a current one of the see-through image frames and one or more neighboring ones of the see-through image frames.

12. The AR device of claim 8 , wherein, to generate the depth map, the at least one processing device is configured to:

generate an initial depth map using a trained machine learning model; and

modify the initial depth map to provide both spatial consistency and temporal consistency in order to generate a refined depth map, the spatial consistency associated with depth consistencies between different ones of the see-through image frames captured at a common time, the temporal consistency associated with depth consistencies between different ones of the see-through image frames captured at different times.

13. The AR device of claim 8 , wherein:

a resolution of the depth map is lower than a resolution of the see-through image frames; and

to project the 3D representation to the curved surface, the at least one processing device is configured to:

increase the resolution of the depth map to generate a higher-resolution depth map;

verify depths contained in the higher-resolution depth map; and

project the 3D representation of the environment to the curved surface using verified depths contained in the higher-resolution depth map.

14. The AR device of claim 8 , wherein the curved surface comprises a portion of a cylindrical, spherical, or conical surface.

15. A non-transitory machine readable medium containing instructions that when executed cause at least one processor of an augmented reality (AR) device to:

obtain multiple see-through image frames of an environment around the AR device using multiple imaging sensors of the AR device;

generate a depth map based on the see-through image frames;

generate a three-dimensional (3D) representation of the environment based on the depth map;

project the 3D representation onto a curved surface;

map points of the projected 3D representation to multiple virtual view images; and

initiate presentation of the virtual view images on one or more displays of the AR device.

16. The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to generate the depth map comprise:

instructions that when executed cause the at least one processor to:

generate an initial depth map based on the see-through image frames;

identify first image features associated with a first of the see-through image frames captured using a first of the imaging sensors;

identify second image features associated with a second of the see-through image frames captured using a second of the imaging sensors;

generate predicted second image features associated with the second see-through image frame based on the first image features and relative positions of the first and second imaging sensors;

generate a confidence map associated with the initial depth map based on differences between the second image features and the predicted second image features; and

refine the initial depth map based on the confidence map to generate a refined depth map.

17. The non-transitory machine readable medium of claim 16 , wherein the instructions that when executed cause the at least one processor to generate the initial depth map comprise:

instructions that when executed cause the at least one processor to provide the see-through image frames as inputs to a trained machine learning model, the trained machine learning model configured to generate the initial depth map using the inputs.

18. The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to generate the depth map comprise:

instructions that when executed cause the at least one processor to:

identify first features associated with a sequence of see-through image frames;

identify second features associated with a sequence of initial depth maps, each initial depth map associated with one of the see-through image frames; and

perform image-guided depth filtering and refinement using the first and second features to generate a refined depth map; and

wherein the instructions that when executed cause the at least one processor to perform the image-guided depth filtering and refinement comprise:

instructions that when executed cause the at least one processor to perform weight-based depth filtering using weights that are based on the first and second features associated with a current one of the see-through image frames and one or more neighboring ones of the see-through image frames.

19. The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to generate the depth map comprise:

instructions that when executed cause the at least one processor to:

generate an initial depth map using a trained machine learning model; and

modify the initial depth map to provide both spatial consistency and temporal consistency in order to generate a refined depth map, the spatial consistency associated with depth consistencies between different ones of the see-through image frames captured at a common time, the temporal consistency associated with depth consistencies between different ones of the see-through image frames captured at different times.

20. The non-transitory machine readable medium of claim 15 , wherein:

a resolution of the depth map is lower than a resolution of the see-through image frames; and

the instructions that when executed cause the at least one processor to project the 3D representation to the curved surface comprise instructions that when executed cause the at least one processor to:

increase the resolution of the depth map to generate a higher-resolution depth map;

verify depths contained in the higher-resolution depth map; and

project the 3D representation of the environment to the curved surface using verified depths contained in the higher-resolution depth map.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2023
From: XIONG, YINGEN, DR.
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 063831/0854 →
Continuity (2)
Provisional Application 63442015 · Jan 30, 2023
Related Publication 20240257475A1 · Aug 1, 2024
References Cited (23)
US 11037359B1 · Bleyer et al. · 2021 [cited by applicant]
US 11158108B2 · Bleyer et al. · 2021 [cited by applicant]
US 11288543B1 · Tovchigrechko · 2022 [cited by applicant]
US 11693242B2 · Fortin-Deschenes et al. · 2023 [cited by applicant]
US 11703323B2 · Hall et al. · 2023 [cited by applicant]
US 20130176405A1 · Jeong et al. · 2013 [cited by applicant]
US 20150249815A1 · Sandrew · 2015 [cited by examiner]
US 20180278916A1 · Kim · 2018 [cited by examiner]
US 20190332942A1 · Wang et al. · 2019 [cited by applicant]
US 20210125414A1 · Berkebile · 2021 [cited by examiner]
US 20210150747A1 · Liu et al. · 2021 [cited by applicant]
US 20210400249A1 · Edmonds et al. · 2021 [cited by applicant]
US 20220060677A1 · Park et al. · 2022 [cited by applicant]
US 20220148207A1 · Varekamp et al. · 2022 [cited by applicant]
US 20220254121A1 · Xu et al. · 2022 [cited by applicant]
CN 113179396A · 2021 [cited by applicant]
KR 1020120119774A · 2012 [cited by applicant]
KR 1020210058683A · 2021 [cited by applicant]
WO 2018110822A1 · 2018 [cited by applicant]
Takeo Kanade, Atsushi Yoshida, Kazuo Oda, Hiroshi Kano and Masaya Tanaka, A Stereo Machine for Video-rate Dense Depth Mapping and Its New Applications, 1996, In Proceedings CVPR IEEE Computer Society Conference on Compu… [cited by examiner]
Masayuki Kanbara, Takashi Okuma, Haruo Takemura and Naokazu Yokoya, Real-time Composition of Stereo Images for Video See-through Augmented Reality, 1999, In Proceedings IEEE International Conference on Multimedia Comput… [cited by examiner]
Masayuki Kanbara, Takashi Okuma, Haruo Takemura and Naokazu Yokoya, A Stereoscopic Video See-through Augmented Reality System Based on Real-time Vision-based Registration, 2000, In Proceedings IEEE Virtual Reality 2000,… [cited by examiner]
International Search Report and Written Opinion of the International Searching Authority dated Feb. 19, 2024 in connection with International Patent Application No. PCT/KR2023/018564, 9 pages. [cited by applicant]