IP Library Granted Patent US 12,530,076
Granted Patent B2
US 12,530,076 · App. 18/670,128 · Granted Jan 20, 2026

Dynamically-adaptive planar transformations for video see-through (VST) extended reality (XR)

Inventor: Yingen Xiong (Mountain View, CA)
Assignee: Samsung Electronics Co., Ltd.
G06F3/012G06F3/013
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,530,076
App. No.
18/670,128
Granted
Jan 20, 2026
Kind
B2
Abstract

A method includes obtaining multiple image frames captured using one or more imaging sensors of a video see-through (VST) extended reality (XR) device while a user's head is at a first head pose and depth data associated with the image frames. The method also includes predicting a second head pose of the user's head when rendered images will be displayed. The method further includes projecting at least one of the image frames onto one or more first planes to generate at least one projected image frame. The method also includes transforming the at least one projected image frame from the one or more first planes to one or more second planes corresponding to the second head pose to generate at least one transformed image frame. The method further includes rendering the at least one transformed image frame for presentation on one or more displays of the VST XR device.

Claims (50)

1 . A method comprising:

obtaining (i) multiple image frames captured using one or more imaging sensors of a video see-through (VST) extended reality (XR) device while a user's head is at a first head pose and (ii) depth data associated with the image frames;

predicting a second head pose of the user's head when rendered images will be displayed;

determining whether the user is focusing on no plane within a scene captured in the image frames or on one or more planes within the scene; and

in response to the user focusing on the one or more planes within the scene:

projecting at least one of the image frames onto one or more first planes to generate at least one projected image frame, the one or more first planes associated with the one or more planes on which the user is focusing;

transforming the at least one projected image frame from the one or more first planes to one or more second planes corresponding to the second head pose to generate at least one transformed image frame; and

rendering the at least one transformed image frame for presentation on one or more displays of the VST XR device.

2 . The method of claim 1 , wherein the one or more planes on which the user is focusing are identified based on a gaze of the user.

3 . The method of claim 2 , further comprising:

when the user is not focusing on any plane, performing a time warp to transform the images frames based on the second head pose.

4 . The method of claim 1 , wherein the at least one projected image frame is transformed from the one or more first planes to the one or more second planes based on one or more normal vectors of the one or more first planes, one or more depths of the one or more first planes, and a difference between the first head pose and the second head pose.

5 . The method of claim 1 , wherein the one or more first planes comprise two or more planes associated with one or more objects in the scene captured in the image frames.

6 . The method of claim 1 , wherein the one or more planes within the scene are associated with one or more objects in the scene.

7 . The method of claim 1 , wherein the depth data has a resolution that is equal to or less than half a resolution of each of the image frames.

8 . A video see-through (VST) extended reality (XR) device comprising:

one or more displays; and

at least one processing device configured to:

obtain (i) multiple image frames captured using one or more imaging sensors while a user's head is at a first head pose and (ii) depth data associated with the image frames;

predict a second head pose of the user's head when rendered images will be displayed;

determine whether the user is focusing on one or more planes within a scene captured in the image frames based on a gaze of the user;

generate at least one transformed image frame, wherein, to generate the at least one transformed image frame, the at least one processing device is configured to:

in response to determining that the user is focusing on the one or more planes, project at least one of the image frames onto one or more first planes to generate at least one projected image frame and transform the at least one projected image frame from the one or more first planes to one or more second planes corresponding to the second head pose, the one or more first planes associated with the one or more planes on which the user is focusing; or

in response to determining that the user is not focusing on any plane, perform a time warp to transform the images frames based on the second head pose; and

render the at least one transformed image frame for presentation on the one or more displays.

9 . The VST XR device of claim 8 , wherein:

the one or more planes within the scene comprise multiple planes associated with multiple objects in the scene at different depths within the scene; and

the at least one transformed image frame provides a clear view of the multiple objects.

10 . The VST XR device of claim 9 , wherein;

the multiple objects include multiple planar surfaces;

the at least one processing device is further configured to identify the planar surfaces in the image frames; and

the one or more first planes are associated with one or more of the planar surfaces.

11 . The VST XR device of claim 8 , wherein the at least one processing device is configured to transform the at least one projected image frame from the one or more first planes to the one or more second planes based on one or more normal vectors of the one or more first planes, one or more depths of the one or more first planes, and a difference between the first head pose and the second head pose.

12 . The VST XR device of claim 8 , wherein the one or more first planes comprise two or more planes associated with one or more objects in the scene captured in the image frames.

13 . The VST XR device of claim 8 , wherein, to perform the time warp, the at least one processing device is configured to apply a translation, a rotation, or both to the images frames.

14 . The VST XR device of claim 8 , wherein the depth data has a resolution that is equal to or less than half a resolution of each of the image frames.

15 . A non-transitory machine readable medium containing instructions that when executed cause at least one processor of a video see-through (VST) extended reality (XR) device to:

obtain (i) multiple image frames captured using one or more imaging sensors of the VST XR device while a user's head is at a first head pose and (ii) depth data associated with the image frames;

predict a second head pose of the user's head when rendered images will be displayed;

determine whether the user is focusing on no plane within a scene captured in the image frames or on one or more planes within the scene; and

in response to the user focusing on the one or more planes within the scene:

project at least one of the image frames onto one or more first planes to generate at least one projected image frame, the one or more first planes associated with the one or more planes on which the user is focusing;

transform the at least one projected image frame from the one or more first planes to one or more second planes corresponding to the second head pose to generate at least one transformed image frame; and

render the at least one transformed image frame for presentation on one or more displays of the VST XR device.

16 . The non-transitory machine readable medium of claim 15 , wherein the non-transitory machine readable medium further contains instructions that when executed cause the at least one processor to identify the one or more planes on which the user is focusing based on a gaze of the user.

17 . The non-transitory machine readable medium of claim 16 , wherein the non-transitory machine readable medium further contains instructions that when executed cause the at least one processor, when the user is not focusing on any plane, to perform a time warp to transform the images frames based on the second head pose.

18 . The non-transitory machine readable medium of claim 15 , wherein the instructions that when executed cause the at least one processor to transform the at least one projected image frame from the one or more first planes to the one or more second planes comprise:

instructions that when executed cause the at least one processor to transform the at least one projected image frame from the one or more first planes to the one or more second planes based on one or more normal vectors of the one or more first planes, one or more depths of the one or more first planes, and a difference between the first head pose and the second head pose.

19 . The non-transitory machine readable medium of claim 15 , wherein the one or more first planes comprise two or more planes associated with one or more objects in the scene captured in the image frames.

20 . The non-transitory machine readable medium of claim 15 , wherein the one or more first planes represent the one or more planes on which the user is focusing.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 21, 2024
From: XIONG, YINGEN, DR.
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 067481/0392 →
Continuity (2)
Provisional Application 63535454 · Aug 30, 2023
Related Publication 20250076969A1 · Mar 6, 2025
References Cited (21)
US 9514571B2 · Williams et al. · 2016 [cited by applicant]
US 9721395B2 · Chan et al. · 2017 [cited by applicant]
US 10313661B2 · Kass · 2019 [cited by applicant]
US 10529063B2 · Rodriguez et al. · 2020 [cited by applicant]
US 10955685B2 · Osmanis et al. · 2021 [cited by applicant]
US 11017712B2 · Kuwahara · 2021 [cited by examiner]
US 11043018B2 · Zhang · 2021 [cited by examiner]
US 11704877B2 · Peri et al. · 2023 [cited by applicant]
US 12008723B2 · Miller et al. · 2024 [cited by applicant]
US 20140375547A1 · Katz · 2014 [cited by examiner]
US 20150310665A1 · Michail · 2015 [cited by examiner]
US 20160364904A1 · Parker et al. · 2016 [cited by applicant]
US 20170243324A1 · Mierle et al. · 2017 [cited by applicant]
US 20180253868A1 · Bratt et al. · 2018 [cited by applicant]
US 20200368616A1 · Delamont · 2020 [cited by applicant]
US 20220326527A1 · Taylor · 2022 [cited by examiner]
US 20220397955A1 · Lahr et al. · 2022 [cited by applicant]
US 20230082705A1 · Trepte · 2023 [cited by applicant]
KR 1020180121994A · 2018 [cited by applicant]
KR 1020230122177 · 2023 [cited by applicant]
International Search Report and Written Opinion of the International Searching Authority dated Sep. 26, 2024 in connection with International Patent Application No. PCT/KR2024/008878, 13 pages. [cited by applicant]