IP Library › Granted Patent US 12,389,034
Granted Patent B2
US 12,389,034 · App. 18/029,593 · Granted Aug 12, 2025

Method and apparatus for signaling depth of multi-plane images-based volumetric video

Inventors: Bertrand Chupeau (Rennes, FR); Franck Thudor (Rennes, FR); Renaud Dore (Rennes, FR)
Assignee: InterDigital CE Patent Holdings, SAS
H04N19/597G06V10/54G06V10/762H04N13/178H04N19/124H04N19/59
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,389,034
App. No.
18/029,593
Granted
Aug 12, 2025
Kind
B2
Abstract

Methods, apparatus and data stream are described to encode, transmit and decode an atlas-based representation of a 3D scene based on a multiplane image (MPI) representation in which a depth component is encoded in each layer. Layers of the MPI are clustered on a transparency basis to generate texture, transparency and depth patch pictures. Patch pictures are packed in at least one atlas image. Metadata associating each patch to a layer and each layer to a depth and a depth quantization law are encoded in the data stream with the at least one atlas. At the decoding side, the MPI with a depth component is retrieved from the data stream and is used to render a viewport image from a viewpoint in the neighborhood of the center of the MPI.

Claims (47)

1. A method comprising:

obtaining a multiplane image representative of a three-dimensional (3D) scene, wherein layers of the multiplane image have a constant depth value and comprise a texture component, a transparency component, and a depth component, the depth component being determined according to a quantization law relative to the constant depth value of the layer;

generating patch pictures by clustering layers of the multiplane image on a transparency basis;

packing the patch pictures in at least one atlas image;

generating first metadata comprising, for each layer of the multiplane image, the constant depth value of the layer and one or more parameters representative of the quantization law relative to the constant depth value of the layer;

generating second metadata associating a patch picture with a layer of the multiplane image; and

encoding the at least one atlas image, the first metadata, and the second metadata in a data stream.

2. The method of claim 1 , wherein the texture component of the patch pictures is stored in a texture atlas image, wherein the transparency component of the patch pictures is stored in a transparency atlas image, and wherein the depth component of the patch pictures is stored in a depth atlas image.

3. The method of claim 2 , wherein the depth atlas image is downscaled.

4. The method of claim 1 , wherein the patch pictures comprise a texture component, a transparency component, and a depth component.

5. A device comprising circuitry, comprising a processor and a memory, configured for:

obtaining a multiplane image representative of a three-dimensional (3D) scene, wherein layers of the multiplane image have a constant depth value and comprise a texture component, a transparency component, and a depth component, the depth component being determined according to a quantization law relative to the constant depth value of the layer;

generating patch pictures by clustering layers of the multiplane image on a transparency basis;

packing the patch pictures in at least one atlas image;

generating first metadata comprising, for each layer of the multiplane image, the constant depth value of the layer and one or more parameters representative of the quantization law relative to the constant depth value of the layer;

generating second metadata associating a patch picture with a layer of the multiplane image; and

encoding the at least one atlas image, the first metadata, and the second metadata in a data stream.

6. The device of claim 5 , wherein the texture component of the patch pictures is stored in a texture atlas image, wherein the transparency component of the patch pictures is stored in a transparency atlas image, and wherein the depth component of the patch pictures is stored in a depth atlas image.

7. The device of claim 6 , wherein the depth atlas image is downscaled.

8. The device of claim 5 , wherein the patch pictures comprise a texture component, a transparency component, and a depth component.

9. A method comprising:

retrieving, from a data stream, at least one atlas image packing patch pictures comprising a texture component, a transparency component, and a depth component;

retrieving, from the data stream, first metadata comprising, for each layer of a multiplane image representative of a three-dimensional (3D) scene, a constant depth value and one or more parameters representative of a quantization law relative to the constant depth value of the layer;

retrieving, from the data stream, second metadata associating each patch picture with a layer of the multiplane image;

building the multiplane image according to the first metadata and the second metadata; and

rendering a viewport image of the 3D scene with the multiplane image, wherein the depth component of each patch picture is inverse quantized according to the quantization law relative to the constant depth value of the layer associated with the patch picture in the second metadata.

10. The method of claim 9 , wherein the texture component of the patch pictures is retrieved from a texture atlas image, wherein the transparency component of the patch pictures is retrieved from a transparency atlas image, and wherein the depth component of the patch pictures is retrieved from a depth atlas image.

11. The method of claim 10 , wherein the depth atlas image is upscaled.

12. A device comprising circuitry, comprising a processor and a memory, configured for:

retrieving, from a data stream, at least one atlas image packing patch pictures comprising a texture component, a transparency component, and a depth component;

retrieving, from the data stream, first metadata comprising, for each layer of a multiplane image representative of a three-dimensional (3D) scene, a constant depth value and one or more parameters representative of a quantization law relative to the constant depth value of the layer;

retrieving, from the data stream, second metadata associating each patch picture with a layer of the multiplane image;

building the multiplane image according to the first metadata and the second metadata; and

rendering a viewport image of the 3D scene with the multiplane image, wherein the depth component of each patch picture is inverse quantized according to the quantization law relative to the constant depth value of the layer associated with the patch picture in the second metadata.

13. The device of claim 12 , wherein the texture component of the patch pictures is retrieved from a texture atlas image, wherein the transparency component of the patch pictures is retrieved from a transparency atlas image, and wherein the depth component of the patch pictures is retrieved from a depth atlas image.

14. The device of claim 13 , wherein the depth atlas image is upscaled.

15. A non-transitory computer readable medium having stored thereon instructions for causing one or more processors to perform a method comprising:

obtaining a multiplane image representative of a three-dimensional (3D) scene wherein layers of the multiplane image have a constant depth value and comprise a texture component, a transparency component, and a depth component, the depth component being determined according to a quantization law relative to the constant depth value of the layer;

generating patch pictures by clustering layers of the multiplane image on a transparency basis;

packing the patch pictures in at least one atlas image;

generating first metadata comprising, for each layer of the multiplane image, the constant depth value of the layer and one or more parameters representative of the quantization law relative to the constant depth value of the layer;

generating second metadata associating a patch picture with a layer of the multiplane image; and

encoding the at least one atlas image, the first metadata, and the second metadata in a data stream.

16. The non-transitory computer readable medium of claim 15 , wherein the texture component of the patch pictures is stored in a texture atlas image, wherein the transparency component of the patch pictures is stored in a transparency atlas image, and wherein the depth component of the patch pictures is stored in a depth atlas image.

17. The non-transitory computer readable medium of claim 16 , wherein a size of the depth atlas image is smaller than a size of the other atlases.

18. The non-transitory computer readable medium of claim 16 , wherein the patch pictures comprise a texture component, a transparency component, and a depth component.

19. The non-transitory computer readable medium of claim 16 , wherein the depth atlas image is downscaled.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 30, 2023
From: CHUPEAU, BERTRAND; THUDOR, FRANCK; DORE, RENAUD
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 063180/0261 →
Priority Claims (1)
EP 20306134 · Sep 30, 2020 · regional
Continuity (1)
Related Publication 20230362409A1 · Nov 9, 2023
References Cited (46)
US 20200219290A1 · Tourapis · 2020 [cited by examiner]
US 20210005006A1 · Oh · 2021 [cited by examiner]
US 20210005016A1 · Oh · 2021 [cited by examiner]
US 20210006833A1 · Tourapis · 2021 [cited by examiner]
US 20210209807A1 · Oh · 2021 [cited by examiner]
US 20210211724A1 · Kim · 2021 [cited by examiner]
US 20210217200A1 · Oh · 2021 [cited by examiner]
US 20210217203A1 · Kim · 2021 [cited by examiner]
US 20210295567A1 · Lee · 2021 [cited by examiner]
US 20210314626A1 · Kammachi Sreedhar · 2021 [cited by examiner]
US 20210320961A1 · Lee · 2021 [cited by examiner]
US 20210320962A1 · Oh · 2021 [cited by examiner]
US 20210398323A1 · Lee · 2021 [cited by examiner]
US 20220141487A1 · Oh · 2022 [cited by examiner]
US 20220159261A1 · Oh · 2022 [cited by examiner]
US 20220172433A1 · Ricard · 2022 [cited by examiner]
US 20220217314A1 · Oh · 2022 [cited by examiner]
US 20220256131A1 · Oh · 2022 [cited by examiner]
US 20220377302A1 · Fleureau · 2022 [cited by examiner]
US 20230050860A1 · Ilola · 2023 [cited by examiner]
US 20230103016A1 · Oh · 2023 [cited by examiner]
CA 3143885A1 · 2020 [cited by examiner]
CN 113728625A · 2021 [cited by examiner]
EP 3709273A1 · 2020 [cited by applicant]
WO WO2020141248A1 · 2020 [cited by examiner]
WO WO2020146547A1 · 2020 [cited by examiner]
WO WO2021001193A1 · 2021 [cited by examiner]
WO WO2021071257A1 · 2021 [cited by examiner]
WO WO2021122983A1 · 2021 [cited by examiner]
WO WO2021136876A1 · 2021 [cited by examiner]
WO WO2021176133A1 · 2021 [cited by examiner]
WO WO2021201386A1 · 2021 [cited by examiner]
WO WO2021204700A1 · 2021 [cited by examiner]
WO WO2021206282A1 · 2021 [cited by examiner]
WO WO2021210860A1 · 2021 [cited by examiner]
WO WO2021224056A1 · 2021 [cited by applicant]
“2D Video Coding of Volumetric Video Data”—Schwarz et al., 2018 Picture Coding Symposium (PCS); Date of Conference: Jun. 24-27, 2018. (Year: 2018). [cited by examiner]
Anonymous, “Information technology—Coded Representation of Immersive Media—Part 12: MPEG Immersive Video”, International Organization for Standardization (ISO), Coding of Moving Pictures and Audio, ISO/IEC JTC 1/SC 29/W… [cited by applicant]
Fleureau et al., “MIV CE1-Related—Activation of Transparency Attribute and MPI Profile in MIV”, International Organization for Standardization (ISO), Coding of Moving Pictures and Audio, ISO/IEC JTC 1/SC 29/WG 4, Docume… [cited by applicant]
Zhou et al., “Stereo Magnification: Learning View Synthesis using Multiplane Images”, Association for Computing Machinery, ACM Transactions on Graphics, vol. 37, Issue 4, Article 65, Aug. 2018, 12 pages. [cited by applicant]
Anonymous, “ISO/IEC FDIS 23090-5, Information technology—Coded Representation of Immersive Media—Part 5: Visual Volumetric Video-based Coding (V3C) and Video-based Point Cloud Compression (V-PCC)”, International Organiz… [cited by applicant]
Anonymous, “High Efficiency Video Coding”, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Recommendati… [cited by applicant]
Anonymous, “Information Technology—Coding of Audio-Visual Objects—Part 10: Advanced Video Coding”, International Standard, ISO/IEC 14496-10, Second Edition, Oct. 1, 2004, 280 pages. [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems—infrastructure of audiovisual services—Coding of moving video: High Efficiency Video Coding”, International Telecommunication Union, Recommendation ITU-T H.265, O… [cited by applicant]
Anonymous, “Terminal Equipment and Protocols for Telematic Services”, Information Technology—Digital Compression and Coding of Continuous-Tone Still images—Requirements and Guidelines, International Telecommunication Un… [cited by applicant]
ITU_T, “Advanced video coding for generic audiovisual services”, ITU-T H.264, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, I… [cited by applicant]