IP Library › Granted Patent US 12,598,286
Granted Patent B2
US 12,598,286 · App. 18/526,726 · Granted Apr 7, 2026

Depth-varying reprojection passthrough in video see-through (VST) extended reality (XR)

Inventors: Yingen Xiong (Mountain View, CA); Christopher A. Peri (Mountain View, CA)
Assignee: Samsung Electronics Co., Ltd.
H04N13/344G06T19/006H04N13/128H04N13/239
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,598,286
App. No.
18/526,726
Granted
Apr 7, 2026
Kind
B2
Abstract

A method includes obtaining images of a scene captured using a stereo pair of imaging sensors of an XR device and depth data associated with the images, where the scene includes multiple objects. The method also includes obtaining volume-based 3D models of the objects. The method further includes, for one or more first objects, performing depth-based reprojection of the one or more 3D models of the one or more first objects to left and right virtual views based on one or more depths of the one or more first objects. The method also includes, for one or more second objects, performing constant-depth reprojection of the one or more 3D models of the one or more second objects to the left and right virtual views based on a specified depth. In addition, the method includes rendering the left and right virtual views for presentation by the XR device.

Claims (89)

1 . A method comprising:

obtaining (i) images of a scene captured using a stereo pair of imaging sensors of an extended reality (XR) device and (ii) depth data associated with the images, the scene including one or more first objects and one or more second objects;

obtaining volume-based three-dimensional (3D) models of the first and second objects included in the scene, the volume-based 3D models including: (i) one or more 3D models of the one or more first objects and (ii) one or more 3D models of the one or more second objects;

for the one or more first objects, performing depth-based reprojection of the one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects;

for the one or more second objects, performing constant-depth reprojection of the one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; and

rendering the left virtual view and the right virtual view for presentation by the XR device.

2 . The method of claim 1 , wherein obtaining the volume-based 3D models of the first and second objects in the scene comprises, for each of the first and second objects:

determining whether a corresponding volume-based 3D model for the object exists in a library of previously-generated volume-based 3D models; and

one of:

responsive to determining that the corresponding volume-based 3D model for the object does exist in the library, obtaining the corresponding volume-based 3D model from the library; or

responsive to determining that the corresponding volume-based 3D model for the object does not exist in the library, generating a new volume-based 3D model representing the object.

3 . The method of claim 2 , wherein generating the new volume-based 3D model representing the object comprises:

generating a dense depth map based on the images of the scene, the depth data, and position data associated with the XR device;

performing volume reconstruction based on the dense depth map; and

generating the new volume-based 3D model representing the object based on the volume reconstruction and at least one color texture captured in the images.

4 . The method of claim 2 , wherein:

the library of previously-generated volume-based 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects; and

each of the 3D models in the library comprises a 3D mesh representing a geometry of an associated scene or object, a color texture map representing at least one color of the associated scene or object, and one or more parameters identifying at least one of a pattern, dimensions, a pose, or a transformation of the associated scene or object.

5 . The method of claim 1 , wherein:

the one or more first objects are positioned within a threshold distance from the XR device;

the one or more second objects are positioned beyond the threshold distance from the XR device; and

the method further comprises identifying the threshold distance using a trained machine learning model.

6 . The method of claim 1 , wherein:

the method further comprises identifying a focus region of a user;

the one or more first objects are positioned within the focus region of the user; and

the one or more second objects are positioned outside the focus region of the user.

7 . The method of claim 1 , wherein:

the one or more first objects comprise one or more foreground objects in the scene; and

the one or more second objects comprise one or more background objects in the scene.

8 . An extended reality (XR) device comprising:

at least one display;

imaging sensors configured to capture images of a scene, the scene including one or more first objects and one or more second objects; and

at least one processing device configured to:

obtain (i) the images of the scene and (ii) depth data associated with the images;

obtain volume-based three-dimensional (3D) models of the first and second objects included in the scene, the volume-based 3D models including: (i) one or more 3D models of the one or more first objects and (ii) one or more 3D models of the one or more second objects;

for the one or more first objects, perform depth-based reprojection of the one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects;

for the one or more second objects, perform constant-depth reprojection of the one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; and

render the left virtual view and the right virtual view for presentation by the at least one display.

9 . The XR device of claim 8 , wherein, to obtain the volume-based 3D models of the first and second objects in the scene, the at least one processing device is configured, for each of the first and second objects, to:

determine whether a corresponding volume-based 3D model for the object exists in a library of previously-generated volume-based 3D models; and

one of:

responsive to determining that the corresponding volume-based 3D model for the object does exist in the library, obtain the corresponding volume-based 3D model from the library; or

responsive to determining that the corresponding volume-based 3D model for the object does not exist in the library, generate a new volume-based 3D model representing the object.

10 . The XR device of claim 9 , wherein, to generate the new volume-based 3D model representing the object, the at least one processing device is configured to:

generate a dense depth map based on the images of the scene, the depth data, and position data associated with the XR device;

perform volume reconstruction based on the dense depth map; and

generate the new volume-based 3D model representing the object based on the volume reconstruction and at least one color texture captured in the images.

11 . The XR device of claim 9 , wherein:

the library of previously-generated volume-based 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects; and

each of the 3D models in the library comprises a 3D mesh representing a geometry of an associated scene or object, a color texture map representing at least one color of the associated scene or object, and one or more parameters identifying at least one of a pattern, dimensions, a pose, or a transformation of the associated scene or object.

12 . The XR device of claim 8 , wherein:

the one or more first objects are positioned within a threshold distance from the XR device;

the one or more second objects are positioned beyond the threshold distance from the XR device; and

the at least one processing device is further configured to identify the threshold distance using a trained machine learning model.

13 . The XR device of claim 8 , wherein:

the at least one processing device is further configured to identify a focus region of a user;

the one or more first objects are positioned within the focus region of the user; and

the one or more second objects are positioned outside the focus region of the user.

14 . The XR device of claim 8 , wherein:

the one or more first objects comprise one or more foreground objects in the scene; and

the one or more second objects comprise one or more background objects in the scene.

15 . A non-transitory machine readable medium containing instructions that when executed cause at least one processor to:

obtain (i) images of a scene captured using a stereo pair of imaging sensors of an extended reality (XR) device and (ii) depth data associated with the images, the scene including one or more first objects and one or more second objects;

obtain volume-based three-dimensional (3D) models of the first and second objects included in the scene, the volume-based 3D models including: (i) one or more 3D models of the one or more first objects and (ii) one or more 3D models of the one or more second objects;

for the one or more first objects, perform depth-based reprojection of the one or more 3D models of the one or more first objects to a left virtual view and a right virtual view based on one or more depths of the one or more first objects;

for the one or more second objects, perform constant-depth reprojection of the one or more 3D models of the one or more second objects to the left virtual view and the right virtual view based on a specified depth; and

render the left virtual view and the right virtual view for presentation by the XR device.

16 . The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to obtain the volume-based 3D models of the first and second objects in the scene comprise:

instructions that when executed cause the at least one processor, for each of the first and second objects, to:

determine whether a corresponding volume-based 3D model for the object exists in a library of previously-generated volume-based 3D models; and

one of:

responsive to determining that the corresponding volume-based 3D model for the object does exist in the library, obtain the corresponding volume-based 3D model from the library; or

responsive to determining that the corresponding volume-based 3D model for the object does not exist in the library, generate a new volume-based 3D model representing the object.

17 . The non-transitory machine readable medium of claim 16 , wherein the instructions that when executed cause the at least one processor to generate the new volume-based 3D model representing the object comprise:

instructions that when executed cause the at least one processor to:

generate a dense depth map based on the images of the scene, the depth data, and position data associated with the XR device;

perform volume reconstruction based on the dense depth map; and

generate the new volume-based 3D model representing the object based on the volume reconstruction and at least one color texture captured in the images.

18 . The non-transitory machine readable medium of claim 16 , wherein:

the library of previously-generated volume-based 3D models is configured to store 3D models of scenes, real-world objects, and virtual objects; and

each of the 3D models in the library comprises a 3D mesh representing a geometry of an associated scene or object, a color texture map representing at least one color of the associated scene or object, and one or more parameters identifying at least one of a pattern, dimensions, a pose, or a transformation of the associated scene or object.

19 . The non-transitory machine readable medium of claim 15 , wherein:

the one or more first objects are positioned within a threshold distance from the XR device;

the one or more second objects are positioned beyond the threshold distance from the XR device; and

the instructions when executed further cause the at least one processor to identify the threshold distance using a trained machine learning model.

20 . The non-transitory machine readable medium of claim 15 , wherein:

the instructions when executed further cause the at least one processor to identify a focus region of a user;

the one or more first objects are positioned within the focus region of the user; and

the one or more second objects are positioned outside the focus region of the user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 1, 2023
From: XIONG, YINGEN, DR.; PERI, CHRISTOPHER A., DR.
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 065737/0022 →
Continuity (3)
Provisional Application 63525916 · Jul 10, 2023
Provisional Application 63436282 · Dec 30, 2022
Related Publication 20240223742A1 · Jul 4, 2024
References Cited (20)
US 9800864B2 · Aronsson et al. · 2017 [cited by applicant]
US 10948726B2 · Zhu et al. · 2021 [cited by applicant]
US 11200745B2 · Johnson et al. · 2021 [cited by applicant]
US 11310487B1 · Clemens et al. · 2022 [cited by applicant]
US 11386529B2 · Taylor · 2022 [cited by applicant]
US 20130176405A1 · Jeong et al. · 2013 [cited by applicant]
US 20140198189A1 · Aronsson et al. · 2014 [cited by applicant]
US 20150325040A1 · Stirbu · 2015 [cited by applicant]
US 20180033209A1 · Akeley · 2018 [cited by examiner]
US 20190101758A1 · Zhu · 2019 [cited by examiner]
US 20210174472A1 · Taylor · 2021 [cited by applicant]
US 20210174570A1 · Bleyer et al. · 2021 [cited by applicant]
US 20210233305A1 · Garcia et al. · 2021 [cited by applicant]
US 20210233314A1 · Johnson et al. · 2021 [cited by applicant]
US 20210400249A1 · Edmonds et al. · 2021 [cited by applicant]
US 20220327784A1 · Strandborg et al. · 2022 [cited by applicant]
WO 2018213131A1 · 2018 [cited by applicant]
WO 2022212109A1 · 2022 [cited by applicant]
Supplementary European Search Report dated Oct. 13, 2025 in connection with European Patent Application No. 23912935.6, 9 pages. [cited by applicant]
Muhlhausen et al., “Temporal Consistent Motion Parallax for Omnidirectional Stereo Panorama Video,” VRST 20, Nov. 2020, 9 pages. [cited by applicant]