IP Library › Granted Patent US 12,238,272
Granted Patent B2
US 12,238,272 · App. 17/739,484 · Granted Feb 25, 2025

Adaptive mode selection for point cloud compression

Inventors: Alexandre Zaghetto (San Jose, CA); Ali Tabatabai (Cupertino, CA); Danillo Graziosi (Flagstaff, AZ)
Assignees: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
H04N19/103H04N19/119H04N19/147H04N19/176H04N19/42H04N19/597
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,238,272
App. No.
17/739,484
Granted
Feb 25, 2025
Kind
B2
Abstract

An electronic device and method for adaptive mode selection for point cloud compression, is provided. The electronic device receives a 3D point cloud geometry and partitions the 3D point cloud geometry into a set of 3D blocks. For a 3D block of the set of 3D blocks, mode decision information is determined. The mode decision information includes class information of the 3D point cloud geometry, operational conditions associated with an encoding stage of the 3D point cloud geometry, or mode-related information associated with one or more 3D blocks of the set of 3D blocks. Based on the mode decision information, one or more modes are selected for the 3D block from a plurality of modes. Each mode corresponds to a function that is used to encode the 3D block. The 3D block is encoded based on the one or more modes.

Claims (89)

1. An electronic device, comprising:

circuitry configured to:

receive a three-dimensional (3D) point cloud geometry;

partition the 3D point cloud geometry into a set of 3D blocks;

determine, for a 3D block of the set of 3D blocks, mode decision information that comprises at least one of:

class information associated with the 3D point cloud geometry,

one or more operational conditions associated with an encoding stage of the 3D point cloud geometry, or

mode-related information associated with one or more 3D blocks of the set of 3D blocks;

select one or more modes for the 3D block of the 3D point cloud geometry from a plurality of modes, based on the mode decision information,

wherein each mode of the plurality of modes corresponds to an encoding function for the 3D block of the point cloud geometry,

the plurality of modes further corresponds to a plurality of density levels,

each density level of the plurality of density levels corresponds to a median of a distribution of local density values associated with 3D points in a 3D block of a calibration point cloud of a plurality of calibration point clouds, and

each of the local density values is a number of neighborhood points within a spherical volume around each 3D point in the 3D block of the calibration point cloud; and

encode the 3D block of the 3D point cloud geometry based on the selected one or more modes.

2. The electronic device according to claim 1 , wherein the class information associated with the 3D point cloud geometry includes at least one of a geometry bit-depth, a density level of the plurality of density levels, or a point distribution associated with the 3D point cloud geometry.

3. The electronic device according to claim 1 , wherein the one or more operational conditions associated with the encoding stage of the 3D point cloud geometry includes a target rate-distortion cost associated with 3D point cloud geometry.

4. The electronic device according to claim 1 , wherein the circuitry is further configured to:

load a table that maps the plurality of modes with different classes associated with the plurality of calibration point clouds and operational conditions associated with an encoding stage of the plurality of calibration point clouds, wherein each calibration point cloud of the plurality of calibration point clouds is different from the received 3D point cloud geometry; and

search the table based on the class information and the one or more operational conditions to select the one or more modes.

5. The electronic device according to claim 4 , wherein the circuitry is further configured to:

partition the calibration point cloud of the plurality of calibration point clouds into a plurality of 3D blocks;

encode the plurality of 3D blocks based on each of the plurality of modes to generate a plurality of encoded 3D blocks;

determine a rate-distortion cost associated with each of the generated plurality of encoded 3D blocks;

determine statistical information that indicates, for each mode of the plurality of modes, a fraction of the plurality of encoded 3D blocks for which the rate-distortion cost is a minimum for the plurality of modes;

determine, from the generated plurality of encoded 3D blocks, a subset of encoded 3D blocks for which the fraction of the plurality of encoded 3D blocks is above a threshold, based on the determined statistical information;

determine, from the plurality of modes, a subset of modes that is used in the generation of the subset of encoded 3D blocks; and

generate the table based on the determined subset of modes, the different classes, and the operational conditions, wherein the table is generated prior to the encode of the 3D block of the 3D point cloud geometry.

6. The electronic device according to claim 1 , wherein the circuitry is further configured to:

encode the 3D block of the 3D point cloud geometry based on each of the selected one or more modes to determine one or more encoded 3D blocks;

determine rate-distortion costs associated with the determined one or more encoded 3D blocks;

determine a mode of the selected one or more modes as an optimal mode for the encoding stage, based on a determination that a rate-distortion cost associated with the mode corresponds to a minimum of the determined rate-distortion costs; and

encode the 3D block of the 3D point cloud geometry based on the determined mode to generate an encoded 3D block.

7. The electronic device according to claim 1 , wherein the encoding function corresponds to a Deep Neural Network (DNN) model that is trained to encode the 3D block of the 3D point cloud geometry to generate an encoded 3D block.

8. The electronic device according to claim 7 , wherein each mode of the plurality of modes corresponds to an alpha parameter of a focal loss function used in a training stage of the DNN model, and

the focal loss function is configured to penalize a removal of non-empty voxels from the 3D block of the 3D point cloud geometry.

9. The electronic device according to claim 1 , wherein the circuitry is further configured to:

determine subsets of the set of 3D blocks, based on a scan of the set of 3D blocks in a defined scan order;

encode each 3D block of a first subset of the determined subsets, based on the plurality of modes to generate a plurality of encoded 3D blocks;

determine a rate-distortion cost associated with each encoded 3D block of the plurality of encoded 3D blocks; and

determine mode usage statistics associated with the first subset based on the determined rate-distortion cost associated with each encoded 3D block of the plurality of encoded 3D blocks,

wherein the mode-related information includes the determined mode usage statistics associated with the first subset.

10. The electronic device according to claim 9 , wherein the circuitry is further configured to select the one or more modes for a second subset that includes the 3D block of the 3D point cloud geometry, wherein

the second subset is included in the determined subsets, and

the second subset succeeds the first subset in accordance with the defined scan order.

11. The electronic device according to claim 1 , wherein the circuitry is further configured to determine, from the set of 3D blocks, a subset of 3D blocks that is in a neighborhood of the 3D block of the 3D point cloud geometry, based on a spatial arrangement of the set of 3D blocks in the 3D point cloud geometry, and

wherein the selection of the one or more modes is based on a usage of the one or more modes to encode each 3D block of the subset of 3D blocks into a respective encoded 3D block.

12. The electronic device according to claim 1 , wherein the circuitry is further configured to determine point cloud metrics including the class information associated with the 3D block of the 3D point cloud geometry, wherein

the one or more modes is selected further based on an application of a classifier model on the point cloud metrics, and

the classifier model is a machine learning model that is trained on a task of mode prediction.

13. The electronic device according to claim 1 , wherein the circuitry is further configured to determine point cloud metrics including the class information associated with the 3D block of the 3D point cloud geometry and a subset of 3D blocks in a neighborhood of the 3D block of the 3D point cloud geometry,

wherein the one or more modes is selected further based on an application of a classifier model on the point cloud metrics, and

the classifier model is a machine learning model that is trained on a task of mode prediction.

14. The electronic device according to claim 1 , wherein the circuitry is further configured to apply a convolutional neural network on the 3D block of the 3D point cloud geometry to generate a mode prediction for the 3D block of the 3D point cloud geometry,

wherein the mode prediction is included in the mode decision information and the one or more modes are selected based on the mode prediction.

15. The electronic device according to claim 1 , wherein the circuitry is further configured to apply a convolutional neural network on the 3D block of the 3D point cloud geometry and a subset of 3D blocks in a neighborhood of the 3D block of the 3D point cloud geometry, to generate a mode prediction for the 3D block of the 3D point cloud geometry,

wherein the mode prediction is included in the mode decision information and the one or more modes are selected based on the mode prediction.

16. A method, comprising:

in an electronic device:

receiving a three-dimensional (3D) point cloud geometry;

partitioning the 3D point cloud geometry into a set of 3D blocks;

determining, for a 3D block of the set of 3D blocks, mode decision information that comprises at least one of:

class information associated with the 3D point cloud geometry,

one or more operational conditions associated with an encoding stage of the 3D point cloud geometry, or

mode-related information associated with one or more 3D blocks of the set of 3D blocks;

selecting one or more modes for the 3D block of the point cloud geometry from a plurality of modes, based on the mode decision information,

wherein each mode of the plurality of modes corresponds to an encoding function for the 3D block of the 3D point cloud geometry,

the plurality of modes further corresponds to a plurality of density levels,

each density level of the plurality of density levels corresponds to a median of a distribution of local density values associated with 3D points in a 3D block of a calibration point cloud of a plurality of calibration point clouds, and

each of the local density values is a number of neighborhood points within a spherical volume around each 3D point in the 3D block of the calibration point cloud; and

encoding the 3D block of the 3D point cloud geometry based on the selected one or more modes.

17. The method according to claim 16 , further comprising:

loading a table that maps the plurality of modes with different classes associated with the plurality of calibration point clouds and operational conditions associated with an encoding stage of the plurality of calibration point clouds, wherein each calibration point cloud of the plurality of calibration point clouds is different from the received 3D point cloud geometry; and

searching the table based on the class information and the one or more operational conditions to select the one or more modes.

18. The method according to claim 16 , wherein the encoding function corresponds to a Deep Neural Network (DNN) model that is trained to encode the 3D block of the 3D point cloud geometry to generate an encoded 3D block.

19. The method according to claim 18 , wherein each mode of the plurality of modes corresponds to an alpha parameter of a focal loss function used in a training stage of the DNN, and

the focal loss function is configured to penalize a removal of non-empty voxels from the 3D block of the 3D point cloud geometry.

20. A non-transitory computer-readable medium having stored thereon, computer-executable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprising:

receiving a three-dimensional (3D) point cloud geometry;

partitioning the 3D point cloud geometry into a set of 3D blocks;

determining, for a 3D block of the set of 3D blocks, mode decision information that comprises at least one of:

class information associated with the 3D point cloud geometry,

one or more operational conditions associated with an encoding stage of the 3D point cloud geometry, or

mode-related information associated with one or more 3D blocks of the set of 3D blocks;

selecting one or more modes for the 3D block of the point cloud geometry from a plurality of modes, based on the mode decision information,

wherein each mode of the plurality of modes corresponds to an encoding function for the 3D block of the 3D point cloud geometry,

the plurality of modes further corresponds to a plurality of density levels,

each density level of the plurality of density levels corresponds to a median of a distribution of local density values associated with 3D points in a 3D block of a calibration point cloud of a plurality of calibration point clouds, and

each of the local density values is a number of neighborhood points within a spherical volume around each 3D point in the 3D block of the calibration point cloud; and

encoding the 3D block of the point cloud geometry based on the selected one or more modes.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2022
From: ZAGHETTO, ALEXANDRE; TABATABAI, ALI; GRAZIOSI, DANILLO
To: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
Reel/Frame 060795/0125 →
Continuity (2)
Provisional Application 63262135 · Oct 5, 2021
Related Publication 20230104977A1 · Apr 6, 2023
References Cited (18)
US 20170347120A1 · Chou et al. · 2017 [cited by applicant]
US 20200219290A1 · Tourapis et al. · 2020 [cited by applicant]
US 20210105458A1 · Sugio · 2021 [cited by examiner]
US 20220377327A1 · Park · 2022 [cited by examiner]
US 20230164353A1 · Lee · 2023 [cited by examiner]
CN 112601082A · 2021 [cited by applicant]
A. F. R. Guarda, N. M. M. Rodrigues and F. Pereira, “Adaptive Deep Learning-Based Point Cloud Geometry Coding,” in IEEE Journal of Selected Topics in Signal Processing, vol. 15, No. 2, pp. 415-430, Feb. 2021, doi: 10.11… [cited by examiner]
3DG, “G-PCC codec description”, 128. MPEG Meeting; Oct. 7-Oct. 11, 2019; Geneva; (Motion Picture Expert Group or ISO/ IEC JTC1 /SC29/WG11), No. n18891 Dec. 13, 2019, XP030225589. [cited by applicant]
Tekalp A Murat et al: “Editorial: Introduction to the Issue on Deep Learning for Image/Video Restoration and Compression”, IEEE Journal of Selected Topics in Signal Processing, IEEE, US, vol. 15, No. 2, Feb. 22, 2021 (F… [cited by applicant]
Tabatabai (Sony) A et al: “[EE13.54 Related] Discussion points on the AI tools evaluation procedures for point cloud compression and analysis”, 135. MPEG Meeting; Jul. 12-Jul. 16, 2021; Online; (Motion Picture Expert Gr… [cited by applicant]
Shao, et al., “Hybrid Point Cloud Attribute Compression Using Slice-based Layered Structure and Block-based Intra Prediction”, Proceedings of the 26th ACM international conference on Multimedia, Oct. 2018, pp. 1199-1207. [cited by applicant]
Guarda, et al., “Adaptive Deep Learning-Based Point Cloud Geometry Coding,” IEEE, Journal of Selected Topics in Signal Processing, vol. 15, No. 2, Feb. 2021, pp. 415-430. [cited by applicant]
Wiegand, et al., “Overview of the H.264/AVC video coding standard”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 13, No. 7, Jul. 2003, pp. 560-576. [cited by applicant]
Sullivan, et al., “Overview of the High Efficiency Video Coding (HEVC) Standard”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 22, No. 12, Dec. 2012, pp. 1649-1668. [cited by applicant]
Bross, et al., “Overview of the Versatile Video Coding (VVC) Standard and Its Applications”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 31, No. 10, Oct. 2021, pp. 3736-3764. [cited by applicant]
Pan, et al., “Fast Mode Decision Algorithm for Intra prediction in H.264/AVC Video Coding”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, No. 7, Jul. 2005, pp. 813-822. [cited by applicant]
Wu, et al., “Fast intermode decision in H.264/AVC video coding”, IEEE Transactions on Circuits and Systems for Video Technology, vol. 15, No. 7, Jul. 2005, pp. 953-958. [cited by applicant]
Zhao, et al., “Fast Mode Decision Algorithm for Intra Prediction in HEVC”, Visual Communications and Image Processing (VCIP), 2011, 04 pages. [cited by applicant]