IP Library › Granted Patent US 12,659,443
Granted Patent B2
US 12,659,443 · App. 18/214,265 · Granted Jun 16, 2026

Viewpoint synthesis with enhanced 3D perception

Inventors: Seppo Valli (Espoo, FI); Pekka Siltanen (Helsinki, FI)
Assignee: Adeia Guides Inc.
H04N13/117G06T5/77G06T7/194G06T7/20G06T7/40G06T7/50G06T15/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,659,443
App. No.
18/214,265
Granted
Jun 16, 2026
Kind
B2
Abstract

Systems and methods relate to segmenting object(s) of the first image data from a first viewpoint; generating respective segment texture data and segment depth data for each segmented object; arranging the segmented object(s) into a configuration based on the second viewpoint; compiling depth data corresponding to each pixel location in the configuration based on the second viewpoint; forming a synthesized depth map based at least in part on a closest depth value for each overlapping depth value at each pixel location of the compiled depth data; forming a synthesized texture map based on correspondence between segment depth data and respective segment texture data; and generating multiple focal planes (MFPs) based on the synthesized texture map and the synthesized depth map to enable generating for display second image data based on the MFPs.

Claims (68)

1 . A method comprising:

accessing first image data comprising texture data and depth data from a first viewpoint;

segmenting one or more objects of the first image data from the first viewpoint, wherein at least one segmented object of the one or more segmented objects is an occluding object;

generating respective segment texture data and segment depth data for each segmented object of the one or more segmented objects;

identifying a second viewpoint;

arranging at least one of the one or more segmented objects into a configuration by:

updating texture data corresponding to the occluding object to move the occluding object to a position based on the second viewpoint relative to a previous position of the occluding object in the texture data;

identifying at least one hole in the texture data caused by the moving of the occluding object; and

filling the at least one hole in the texture data based at least in part on at least one pixel prediction method;

compiling, based on the segment depth data for each segmented object in the configuration, updated depth data corresponding to each pixel location in the configuration based on the second viewpoint;

forming a synthesized depth map for the updated depth data with the at least one hole caused by the moving of the occluding object based at least in part on a closest depth value for each overlapping depth value at each pixel location of the compiled depth data;

forming a synthesized texture map based on correspondence between the updated depth data and the updated texture data; and

generating multiple focal planes (MFPs) based on the synthesized texture map and the synthesized depth map to enable generating for display second image data based on the MFPs.

2 . The method of claim 1 , wherein accessing the first image data comprises:

receiving coded data and decoding the coded data to the texture data and the depth data.

3 . The method of claim 1 , wherein segmenting the one or more objects of the first image data from the first viewpoint is based on color, depth, or both color and depth.

4 . The method of claim 1 , further comprising:

generating a background image by removing at least one object from the first image data with corresponding depth data closer to the first viewpoint than other depth data of the first image data and combining areas of the first image data with corresponding depth data farther away from the first viewpoint, and by using a background prediction method based on collecting areas of still background behind moving objects.

5 . The method of claim 4 , wherein the at least one pixel prediction method comprises:

using spatial interpolation or extrapolation to predict missing texture data of pixels.

6 . The method of claim 1 , further comprising:

tracking a position of a user, wherein the identifying the second viewpoint is based on the tracking.

7 . The method of claim 1 , further comprising:

identifying a third viewpoint; and

in response to determining that the third viewpoint is less than a threshold distance from the second viewpoint:

generating modified MFPs by shifting and scaling the MFPs based on the third viewpoint; and

generating for display modified image data based on the modified MFPs.

8 . The method of claim 1 , wherein the arranging comprises:

moving the at least one segmented object based on the second viewpoint, wherein the moving is based on 3D coordinates of the at least one segmented object and a scaling factor to define a position and a size of the at least one segmented object as viewed from the second viewpoint.

9 . The method of claim 1 , wherein the arranging comprises:

moving the at least one segmented object in a fixed pose.

10 . The method of claim 1 , wherein the arranging comprises:

moving the at least one segmented object to have a same orientation relative to the second viewpoint as a previous orientation relative to the first viewpoint.

11 . A system comprising:

input/output circuitry configured to:

access first image data comprising texture data and depth data from a first viewpoint; and

control circuitry configured to:

segment one or more objects of the first image data from the first viewpoint, wherein at least one segmented object of the one or more segmented objects is an occluding object;

generate respective segment texture data and segment depth data for each segmented object of the one or more segmented objects;

identify a second viewpoint;

arrange at least one of the one or more segmented objects into a configuration by:

updating texture data corresponding to the occluding object to move the occluding object to a position based on the second viewpoint relative to a previous position of the occluding object; and

identifying at least one hole in the texture data based at least in part on at least one pixel prediction method;

compile, based on the segment depth data for each segmented object in the configuration, updated depth data corresponding to each pixel location in the configuration based on the second viewpoint;

form a synthesized depth map for the updated depth data with the least one hole caused by the moving of the occluding object based at least in part on a closest depth value for each overlapping depth value at each pixel location of the compiled depth data; and

form a synthesized texture map based on correspondence between the updated depth data and the updated texture data to enable generating multiple focal planes (MFPs) based on the synthesized texture map and the synthesized depth map to enable generating for display second image data based on the MFPs.

12 . The system of claim 11 , wherein:

the input/output circuitry is configured to access the first image data by receiving coded data, and

the control circuitry is further configured to decode the coded data to the texture data and the depth data.

13 . The system of claim 11 , wherein the control circuitry is further configured to:

segment the one or more objects of the first image data from the first viewpoint is based on color, depth, or both color and depth.

14 . The system of claim 11 , wherein the control circuitry is further configured to:

generate a background image by removing at least one object from the first image data with corresponding depth data closer to the first viewpoint than other depth data of the first image data and combining areas of the first image data with corresponding depth data farther away from the first viewpoint, and by using a background prediction method based on collecting areas of still background behind moving objects.

15 . The system of claim 14 , wherein the at least one pixel prediction method comprises:

using spatial interpolation or extrapolation to predict missing texture data of pixels.

16 . The system of claim 11 , wherein the control circuitry is further configured to:

track a position of a user, wherein the control circuitry is configured to identify the second viewpoint based on the tracking.

17 . The system of claim 11 , wherein the control circuitry is further configured to:

identify a third viewpoint; and

in response to determining that the third viewpoint is less than a threshold distance from the second viewpoint:

generate modified MFPs by shifting and scaling the MFPs based on the third viewpoint; and

generate for display modified image data based on the modified MFPs.

18 . The system of claim 11 , wherein the control circuitry is configured to arrange the at least one of the one or more segmented objects into the configuration based on the second viewpoint by:

moving the at least one segmented object based on the second viewpoint, wherein the moving is based on 3D coordinates of the at least one segmented object and a scaling factor to define a position and a size of the at least one segmented object as viewed from the second viewpoint.

19 . The system of claim 11 , wherein the control circuitry is configured to arrange the at least one of the one or more segmented objects into the configuration based on the second viewpoint by:

moving the at least one segmented object in a fixed pose.

20 . The system of claim 11 , wherein the control circuitry is configured to arrange the at least one of the one or more segmented objects into the configuration based on the second viewpoint by:

moving the at least one segmented object to have a same orientation relative to the second viewpoint as a previous orientation relative to the first viewpoint.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 6, 2023
From: VALLI, SEPPO; SILTANEN, PEKKA
To: ADEIA GUIDES INC.
Reel/Frame 064168/0873 →
Continuity (1)
Related Publication 20240430394A1 · Dec 26, 2024
References Cited (17)
US 10330936B2 · Fix et al. · 2019 [cited by applicant]
US 11184599B2 · Harviainen et al. · 2021 [cited by applicant]
US 11543655B1 · Weber · 2023 [cited by examiner]
US 12267477B2 · Valli et al. · 2025 [cited by applicant]
US 20170188002A1 · Chan · 2017 [cited by examiner]
US 20180270464A1 · Harviainen · 2018 [cited by examiner]
US 20210133994A1 · Valli · 2021 [cited by examiner]
US 20240080431A1 · Van Geest · 2024 [cited by examiner]
WO 2019183211A1 · 2019 [cited by applicant]
Zitnick, C.L., Kang, S.B., Uyttendaele, M., Winder, S. and Szeliski, R., 2004. High-quality video view interpolation using a layered representation. ACM transactions on graphics (TOG), 23(3), pp. 600-608. (Year: 2004). [cited by examiner]
Lee, C. and Ho, Y.S., Oct. 2009, View synthesis using depth map for 3D video. In Proceedings of 2009 APSIPA Annual Summit and conference, Sapporo, Japan (pp. 350-357). (Year: 2009). [cited by examiner]
Akeley, Kurt , et al., “A Stereo Display Prototype with Multiple Focal Distances”, ACM Trans. Graph. 23, 3, 804-813, 2004. [cited by applicant]
Schmeing, Michael , et al., “Faithful Spatio-Temporal Disocclusion Filling using Local Optimization”, 21st International Conference on Pattern Recognition (ICPR 2012), 2012. [cited by applicant]
Matsuda, N., et.al., “Focal Surface Displays,” ACM Transactions on Graphics, 36(4), 86:1-86:14, 2017. [cited by applicant]
Youtube, “El 2020 Plenary: Quality Screen Time: Leveraging Computational Displays for Spatial Computing,” https://www.youtube.com/watch?v=LQwMAI9bGNY, 2020. [cited by applicant]
Shade, J., et al., “Layered Depth Images,” https://dl.acm.org/doi/pdf/10.1145/280814.280882, 1998. [cited by applicant]
U.S. Appl. No. 18/214,272, filed Jun. 26, 2023, Seppo Valli. [cited by applicant]