IP Library › Granted Patent US 12,541,882
Granted Patent B2
US 12,541,882 · App. 17/785,542 · Granted Feb 3, 2026

Methods and apparatuses for encoding, decoding and rendering 6DOF content from 3DOF+ composed elements

Inventors: Charles Salmon-Legagneur (Rennes, FR); Charline Taibi (Chartres de Bretagne, FR); Jean Le Roux (Rennes, FR); Serge Travert (Dinan, FR)
Assignee: InterDigital VC Holdings, Inc.
G06T9/001G06T2219/028
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,882
App. No.
17/785,542
Granted
Feb 3, 2026
Kind
B2
Abstract

A volumetric content is encoded as a set of clusters by an encoder and transmitted to a decoder which retrieves the volumetric content. Clusters common to different viewpoints are obtained and mutualized. Clusters are projected onto 2D images and encoded as independent video streams. Reduction in visual artefacts and reduction of data for storage and streaming are achieved.

Claims (38)

1 . A method for encoding a 3D scene, the method comprising:

clustering points of the 3D scene into a plurality of clusters according to a depth range of the points of the 3D scene, wherein the plurality of clusters comprises at least a background cluster and a foreground cluster;

obtaining a first set of 2D images by projecting clusters visible from a first set of points of view according to first projection parameters, wherein the first set of points of view is encompassed in a first viewing box defined in the 3D scene and comprises at least two points of view;

obtaining a second set of 2D images by projecting clusters visible from a second set of points of view according to second projection parameters, wherein the second set of points of view is encompassed in a second viewing box defined in the 3D scene different from the first viewing box;

identifying one or more common clusters visible from both the first set of points of view and the second set of points of view; and

encoding the first set of 2D images and the first projection parameters in a first data stream and each 2D image of the second set of 2D images and the second projection parameters in a set of distinct data streams, wherein 2D common images corresponding to the one or more common clusters are encoded only once.

2 . The method of claim 1 , wherein the clustering is further based on semantics associated with points of the 3D scene, on color of the points of the 3D scene, or on motion of points of the 3D scene.

3 . The method of claim 1 , further comprising:

encoding metadata comprising:

a list of viewing boxes defined in the 3D scene and a list of common clusters for the 3D scene; and

for each viewing box, a list of sets of clusters representative of the 3D scene for the respective viewing box and, for each set of clusters associated with the viewing box, identifiers of the common clusters, and the list of clusters other than the common clusters.

4 . A non-transitory computer readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform the method of claim 1 .

5 . A method for decoding a 3D scene, the method comprising:

decoding a first set of first 2D images from a first data stream and a second 2D image from each data stream of a set of distinct data streams, the first 2D images being representative of a projection according to first projection parameters of at least one cluster of points of the 3D scene visible from a first set of points of view encompassed in a first viewing box defined in the 3D scene, the second 2D images being representative of a projection according to second projection parameters of at least one cluster of points of the 3D scene visible from a second set of points of view encompassed in a second viewing box defined in the 3D scene different from the first viewing box, wherein 2D common images corresponding to common clusters to the first and second viewing boxes are decoded only once, and the points of the 3D scene being clustered into a plurality of clusters according to a depth range of the points of the 3D scene, wherein the plurality of clusters comprises at least a background cluster and a foreground cluster; and

un-projecting pixels of the first 2D images according to the first projection parameters and to the first set of points of view and un-projecting pixels of the second 2D images according to the second projection parameters and to the second set of points of view.

6 . The method of claim 5 , further comprising

obtaining metadata comprising:

a list of viewing boxes defined in the 3D scene and a list of common clusters for the 3D scene; and

for each viewing box, a list of sets of clusters representative of the 3D scene for the respective viewing box and, for each set of clusters associated with the viewing box, identifiers of the common clusters, and the list of clusters other than the common clusters; and

decoding 2D images from data streams comprising clusters of 3D points visible from a current point of view.

7 . A non-transitory computer readable medium storing instructions which, when executed by one or more processors, cause the one or more processors to perform the method of claim 5 .

8 . A device for encoding a 3D scene comprising a memory associated with a processor configured for:

clustering points of the 3D scene into a plurality of clusters according to a depth range of the points of the 3D scene, wherein the plurality of clusters comprises at least a background cluster and a foreground cluster;

obtaining a first set of 2D images by projecting clusters visible from a first set of points of view according to first projection parameters, wherein the first set of points of view is encompassed in a first viewing box defined in the 3D scene and comprises at least two points of view;

obtaining a second set of 2D images by projecting clusters visible from a second set of points of view according to second projection parameters, wherein the second set of points of view is encompassed in a second viewing box defined in the 3D scene different from the first viewing box;

identifying one or more common clusters visible from both the first set of points of view and the second set of points of view; and

encoding the first set of 2D images and the first projection parameters in a first data stream and each 2D image of the second set of 2D images and the second projection parameters in a set of distinct data streams, wherein 2D common images corresponding to the one or more common clusters are encoded only once.

9 . The device of claim 8 , wherein the clustering is further based on semantics associated with points of the 3D scene, on color of the points of the 3D scene, or on motion of the points of the 3D scene.

10 . The device of claim 8 , wherein the processor is further configured for encoding metadata comprising:

a list of viewing boxes defined in the 3D scene and a list of common clusters for the 3D scene; and

for each viewing box, a list of sets of clusters representative of the 3D scene for the respective viewing box and, for each set of clusters associated with the viewing box, identifiers of the common clusters, and the list of clusters other than the common clusters.

11 . A device for decoding a 3D scene comprising a memory associated with a processor configured for:

decoding a first set of first 2D images from a first data stream and a second 2D image from each data stream of a set of distinct data streams, the first 2D images being representative of a projection according to first projection parameters of at least one cluster of points of the 3D scene visible from a first set of points of view encompassed in a first viewing box defined in the 3D scene, the second 2D images being representative of a projection according to second projection parameters of at least one cluster of points of the 3D scene visible from a second set of points of view encompassed in a second viewing box defined in the 3D scene different from the first viewing box, wherein 2D common images corresponding to common clusters to the first and second viewing boxes are decoded only once, and the points of the 3D scene being clustered into a plurality of clusters according to a depth range of the points of the 3D scene, wherein the plurality of clusters comprises at least a background cluster and a foreground cluster; and

un-projecting pixels of the first 2D images according to the first projection parameters and to the first set of points of view and un-projecting pixels of the second 2D images according to the second projection parameters and to the second set of points of view.

12 . The device of claim 11 , further the processor is further configured for obtaining metadata comprising:

a list of viewing boxes defined in the 3D scene and a list of common clusters for the 3D scene; and

for each viewing box, a list of sets of clusters representative of the 3D scene for the respective viewing box and, for each set of clusters associated with the viewing box, identifiers of the common clusters, and the list of clusters other than the common clusters; and

decoding 2D images from data streams comprising clusters of 3D points visible from a current point.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2022
From: SALMON-LEGAGNEUR, CHARLES; TAIBI, CHARLINE; LE ROUX, JEAN; TRAVENT, SERGE
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 060215/0140 →
Priority Claims (1)
EP 19306692 · Dec 19, 2019 · regional
Continuity (1)
Related Publication 20230032599A1 · Feb 2, 2023
References Cited (27)
US 10044944B2 · Adsumilli et al. · 2018 [cited by applicant]
US 20120195370A1 · Guerrero · 2012 [cited by applicant]
US 20130336582A1 · Dai · 2013 [cited by applicant]
US 20180060700A1 · Bleyer · 2018 [cited by examiner]
US 20180063505A1 · Lee et al. · 2018 [cited by applicant]
US 20180075634A1 · Kandler · 2018 [cited by applicant]
US 20180262684A1 · Lowry et al. · 2018 [cited by applicant]
US 20180309927A1 · Tanner et al. · 2018 [cited by applicant]
US 20180316902A1 · Tanaka · 2018 [cited by applicant]
US 20190313081A1 · Oh · 2019 [cited by applicant]
CN 103503454A · 2014 [cited by applicant]
CN 108353156A · 2018 [cited by applicant]
CN 109644262A · 2019 [cited by applicant]
EP 3429210A1 · 2019 [cited by examiner]
EP 3457688A1 · 2019 [cited by applicant]
EP 3547703A1 · 2019 [cited by examiner]
EP 3562159A1 · 2019 [cited by applicant]
JP 2013257843A · 2013 [cited by applicant]
WO 2017082079A1 · 2017 [cited by applicant]
WO 2019195547A1 · 2019 [cited by applicant]
Fleureau et al., “Description of Technicolor Intel response to MPEG-I 3DoF+ Call for Proposal”, International Organisation for Standardisation, ISO/IEC JTC1/SC29/WG11, Coding of Moving Pictures and Audio, Document: MPEG… [cited by applicant]
Rhee et al., “MR360: Mixed Reality Rendering for 360° Panoramic Videos”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Visualization and Computer Graphics, vol. 23, Issue: 4, Apr. 2017, … [cited by applicant]
Anonymous, “High Efficiency Video Coding”, ITU-T Telecommunication Standardization Sector of Itu, Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Recommendati… [cited by applicant]
Anonymous, “Information Technology—Coding of Audio-Visual Objects—Part 10: Advanced Video Coding”, International Standard, ISO/IEC 14496-10, Second Edition, Oct. 1, 2004, 280 pages. [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems—infrastructure of audiovisual services—Coding of moving video: High Efficiency Video Coding”, International Telecommunication Union, Recommendation ITU-T H.265, O… [cited by applicant]
Anonymous, “Terminal Equipment and Protocols for Telematic Services”, Information Technology—Digital Compression and Coding of Continuous-Tone Still images—Requirements and Guidelines, International Telecommunication Un… [cited by applicant]
ITU_T, “Advanced video coding for generic audiovisual services”, ITU-T H.264, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, I… [cited by applicant]