IP Library Granted Patent US 12,462,408
Granted Patent B2
US 12,462,408 · App. 18/449,352 · Granted Nov 4, 2025

Reducing texture lookups in extended-reality applications

Inventors: Mikko Strandborg (Helsinki, FI); Antti Hirvonen (Helsinki, FI)
Assignee: Varjo Technologies Oy
G06T7/50G06T7/11G06T2207/10028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,408
App. No.
18/449,352
Granted
Nov 4, 2025
Kind
B2
Abstract

A display apparatus has a display area of an output image divided into a plurality of regions. For a given region of the output image, a corresponding first region in a video-see-through (VST) image and a corresponding second region in at least one virtual-reality (VR) image are determined. From a depth map corresponding to the VST image, a minimum optical depth and a maximum optical depth in the first region of the VST image are determined. From a depth map corresponding to the at least one VR image, a minimum optical depth and a maximum optical depth in the second region of the at least one VR image are determined. When the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, pixel data for pixels in the given region of the output image from the second region of the at least one VR image is fetched. Otherwise, when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image, the pixel data for the pixels in the given region of the output image from the first region of the VST image is fetched.

Claims (57)

1 . A method implemented by a display apparatus, the method comprising:

dividing a display area of an output image into a plurality of regions;

for a given region of the output image, determining a corresponding first region in a video-see-through (VST) image and a corresponding second region in at least one virtual-reality (VR) image;

determining, from a depth map corresponding to the VST image, a minimum optical depth and a maximum optical depth in the first region of the VST image;

determining, from a depth map corresponding to the at least one VR image, a minimum optical depth and a maximum optical depth in the second region of the at least one VR image;

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, fetching pixel data for pixels in the given region of the output image from the second region of the at least one VR image; and

when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image, fetching the pixel data for the pixels in the given region of the output image from the first region of the VST image.

2 . The method of claim 1 , further comprising when the minimum optical depth in the first region of the VST image is not greater than the maximum optical depth in the second region of the at least one VR image, and the minimum optical depth in the second region of the at least one VR image is not greater than the maximum optical depth in the first region of the VST image,

fetching, from the depth map corresponding to the VST image, a first optical depth for a given pixel in the given region of the output image;

fetching, from the depth map corresponding to the at least one VR image, a second optical depth for the given pixel;

when the first optical depth is greater than or equal to the second optical depth, fetching pixel data for the given pixel from the second region of the at least one VR image; and

when the first optical depth is smaller than the second optical depth, fetching pixel data for the given pixel from the first region of the VST image.

3 . The method of claim 1 , further comprising:

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, detecting whether a transparency of at least a part of the second region of the at least one VR image is more than a predefined threshold; and

when it is detected that the transparency of at least a part of the second region of the at least one VR image is more than the predefined threshold, fetching pixel data for at least a subset of the pixels in the given region of the output image from the first region of the VST image, wherein said subset of the pixels in the given region of the output image correspond to said part of the second region of the at least one VR image whose transparency is more than the predefined threshold.

4 . The method of claim 1 , wherein the plurality of regions of the display area are in a form of tiles.

5 . The method of claim 1 , wherein the display area is divided into the plurality of regions iteratively.

6 . A method implemented by at least one server that is communicably coupled to at least one display apparatus, the method comprising:

receiving, from the at least one display apparatus, depth data corresponding to a video-see-through (VST) image captured at the at least one display apparatus;

dividing a display area into a plurality of regions;

for a given region of the display area, determining a corresponding first region in the VST image and a corresponding second region in at least one virtual-reality (VR) image;

determining, from the depth data corresponding to the VST image, a minimum optical depth and a maximum optical depth in the first region of the VST image;

determining, from at least one depth map corresponding to the at least one VR image, a minimum optical depth and a maximum optical depth in the second region of the at least one VR image;

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, classifying the given region of the display area as a VR region that is to be rendered from the at least one VR image only;

when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image, classifying the given region of the display area as a VST region that is to be rendered from the VST image only;

when the minimum optical depth in the first region of the VST image is not greater than the maximum optical depth in the second region of the at least one VR image, and the minimum optical depth in the second region of the at least one VR image is not greater than the maximum optical depth in the first region of the VST image, classifying the given region of the display area as a mixed region that is to be rendered from the at least one VR image and the VST image; and

sending, to the at least one display apparatus, information indicative of respective classifications of the plurality of regions.

7 . The method of claim 6 , further comprising skipping encoding and transport of the second region of the at least one VR image to the at least one display apparatus, when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image.

8 . A display apparatus comprising a processor configured to:

divide a display area of an output image into a plurality of regions;

for a given region of the output image, determine a corresponding first region in a video-see-through (VST) image and a corresponding second region in at least one virtual-reality (VR) image;

determine, from a depth map corresponding to the VST image, a minimum optical depth and a maximum optical depth in the first region of the VST image;

determine, from a depth map corresponding to the at least one VR image, a minimum optical depth and a maximum optical depth in the second region of the at least one VR image;

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, fetch pixel data for pixels in the given region of the output image from the second region of the at least one VR image; and

when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image, fetch the pixel data for the pixels in the given region of the output image from the first region of the VST image.

9 . The display apparatus of claim 8 , wherein the processor is configured to: when the minimum optical depth in the first region of the VST image is not greater than the maximum optical depth in the second region of the at least one VR image, and the minimum optical depth in the second region of the at least one VR image is not greater than the maximum optical depth in the first region of the VST image,

fetch, from the depth map corresponding to the VST image, a first optical depth for a given pixel in the given region of the output image;

fetch, from the depth map corresponding to the at least one VR image, a second optical depth for the given pixel;

when the first optical depth is greater than or equal to the second optical depth, fetch pixel data for the given pixel from the second region of the at least one VR image; and

when the first optical depth is smaller than the second optical depth, fetch pixel data for the given pixel from the first region of the VST image.

10 . The display apparatus of claim 8 , wherein the processor is configured to:

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, detect whether a transparency of at least a part of the second region of the at least one VR image is more than a predefined threshold; and

when it is detected that the transparency of at least a part of the second region of the at least one VR image is more than the predefined threshold, fetch pixel data for at least a subset of the pixels in the given region of the output image from the first region of the VST image, wherein said subset of the pixels in the given region of the output image correspond to said part of the second region of the at least one VR image whose transparency is more than the predefined threshold.

11 . The display apparatus of claim 8 , wherein the plurality of regions of the display area are in a form of tiles.

12 . The display apparatus of claim 8 , wherein the display area is divided into the plurality of regions iteratively.

13 . A system comprising at least one server that is communicably coupled to at least one display apparatus, the at least one server being configured to:

receive, from the at least one display apparatus, depth data corresponding to a video-see-through (VST) image captured at the at least one display apparatus;

divide a display area into a plurality of regions;

for a given region of the display area, determine a corresponding first region in the VST image and a corresponding second region in at least one virtual-reality (VR) image;

determine, from the depth data corresponding to the VST image, a minimum optical depth and a maximum optical depth in the first region of the VST image;

determine, from at least one depth map corresponding to the at least one VR image, a minimum optical depth and a maximum optical depth in the second region of the at least one VR image;

when the minimum optical depth in the first region of the VST image is greater than the maximum optical depth in the second region of the at least one VR image, classify the given region of the display area as a VR region that is to be rendered from the at least one VR image only;

when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image, classify the given region of the display area as a VST region that is to be rendered from the VST image only;

when the minimum optical depth in the first region of the VST image is not greater than the maximum optical depth in the second region of the at least one VR image, and the minimum optical depth in the second region of the at least one VR image is not greater than the maximum optical depth in the first region of the VST image, classify the given region of the display area as a mixed region that is to be rendered from the at least one VR image and the VST image; and

send, to the at least one display apparatus, information indicative of respective classifications of the plurality of regions.

14 . The system of claim 13 , wherein the at least one server is configured to skip encoding and transport of the second region of the at least one VR image to the at least one display apparatus, when the minimum optical depth in the second region of the at least one VR image is greater than the maximum optical depth in the first region of the VST image.

15 . The system of claim 13 , wherein the display area is divided into the plurality of regions iteratively.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 17, 2023
From: STRANDBORG, MIKKO; HIRVONEN, ANTTI
To: VARJO TECHNOLOGIES OY
Reel/Frame 064617/0163 →
Continuity (1)
Related Publication 20250061593A1 · Feb 20, 2025
References Cited (15)
US 9542010B2 · Roberts · 2017 [cited by examiner]
US 11652976B1 · Ollila · 2023 [cited by examiner]
US 11966045B2 · Ollila · 2024 [cited by examiner]
US 20180160048A1 · Rainisto · 2018 [cited by examiner]
US 20180335925A1 · Hsiao · 2018 [cited by examiner]
US 20210096368A1 · Carlsson · 2021 [cited by examiner]
US 20220327784A1 · Strandborg · 2022 [cited by examiner]
US 20230419595A1 · Xiong · 2023 [cited by examiner]
US 20240319377A1 · Feyerabend · 2024 [cited by examiner]
US 20250076969A1 · Xiong · 2025 [cited by examiner]
US 20250148701A1 · Xiong · 2025 [cited by examiner]
US 20250157123A1 · Räntilä · 2025 [cited by examiner]
WO WO2016099556A1 · 2016 [cited by examiner]
WO WO2024233094A1 · 2024 [cited by examiner]
Kundu, J. et al. “VRT-Net: Real-Time Scene Parsing via Variable Resolution Transform” Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision (WACV), 2020, pp. 2049-2056 (Year: 2020). [cited by examiner]