IP Library Granted Patent US 12,646,222
Granted Patent B2
US 12,646,222 · App. 18/706,208 · Granted Jun 2, 2026

State summarization for binary voxel grid coding

Inventors: Maurice Quach (Gif-sur-Yvette, FR); Jiahao Pang (New York, NY); Muhammad Asad Lodhi (New York, NY); Dong Tian (New York, NY); Giuseppe Valenzise (Gif-sur-Yvette, FR); Frederic Dufaux (Gif-sur-Yvette, FR)
Assignee: InterDigital Patent Holdings, Inc.
G06T9/001G06T9/002
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,222
App. No.
18/706,208
Granted
Jun 2, 2026
Kind
B2
Abstract

In one implementation, we improve the binary voxel-based octree coding method, via a proposed state summarization module for context modeling. Given a current voxel to be encoded or decoded, instead of directly estimating its occupancy probability based on the associated binary occupancy context, a proposed state summarization module is applied to convert the original binary context to a summarized representation. Under the summarized representation, the estimation of the occupancy probability becomes more affordable and effective. In particular, density-based state summarization, pattern-based, learning-based state summarization, and learning-based state summarization methods are provided.

Claims (46)

1 . An apparatus for encoding point cloud data, comprising:

at least one processor configured for

determining a first set of a first number of states associated with an occupancy state for each of a plurality of encoded voxels neighboring a current voxel, wherein the current voxel and the plurality of neighboring encoded voxels are in a point cloud represented by the point cloud data;

determining a density of occupied voxels in the plurality of neighboring encoded voxels;

processing the first set of states to obtain a second set of a second number of states, based on the density of occupied voxels in the plurality of neighboring encoded voxels;

predicting, based on the second set of states, a probability for an occupancy state for the current voxel; and

encoding the occupancy state for the current voxel, based on said predicted probability for the occupancy state for the current voxel.

2 . The apparatus of claim 1 , wherein processing the first set of states comprises converting the first set of states to a summarized state space having the second number of states.

3 . The apparatus of claim 1 , wherein determining the density of occupied voxels comprises:

classifying the voxels neighboring the current voxel into one or more classes of neighboring voxels based on a distance from each neighboring voxel to the current voxel; and

determining a number of occupied neighboring voxels included in each of the one or more classes of neighboring voxels.

4 . The apparatus of claim 2 , wherein converting the first set of states to the summarized state space comprises determining that one or more of a set of patterns exists in a voxel neighborhood comprising the plurality of neighboring voxels.

5 . The apparatus of claim 1 , wherein processing the first set of states comprises applying a neural network including a plurality of point-based MLP layers.

6 . The apparatus of claim 1 , wherein processing the first set of states comprises applying a first neural network including a plurality of convolutional layers and a second neural network including a plurality of point-based MLP layers, wherein outputs from said first and second neural networks are concatenated.

7 . A method for encoding point cloud data, comprising:

determining a first set of a first number of states associated with an occupancy state for each of a plurality of encoded voxels neighboring a current voxel, wherein the current voxel and the plurality of encoded neighboring voxels are in a point cloud represented by point cloud data;

determining a density of occupied voxels in the plurality of neighboring encoded voxels;

processing the first set of states to obtain a second set of a second number of states, based on the density of occupied voxels in the plurality of neighboring encoded voxels;

predicting, based on the second set of states, a probability for an occupancy state for the current voxel; and

encoding the occupancy state for the current voxel, based on said predicted probability for the occupancy state for the current voxel.

8 . The method of claim 7 , wherein determining the density of occupied voxels comprises:

classifying the voxels neighboring the current voxel into one or more classes of neighboring voxels based on a distance from each neighboring voxel to the current voxel; and

determining a number of occupied neighboring voxels included in each of the one or more classes of neighboring voxels.

9 . The method of claim 7 , wherein processing the first set of states comprises converting the first set of states to a summarized state space having the second number of states.

10 . The method of claim 9 , wherein converting the first set of states to the summarized state space comprises determining that one or more of a set of patterns exists in a voxel neighborhood comprising the plurality of neighboring voxels.

11 . The method of claim 7 , wherein processing the first set of states comprises applying a neural network including a plurality of point-based MLP layers.

12 . The method of claim 7 , wherein processing the first set of states comprises applying a first neural network including a plurality of convolutional layers and a second neural network including a plurality of point-based MLP layers, wherein outputs from said first and second neural networks are concatenated.

13 . An apparatus for decoding point cloud data, comprising:

at least one processor configured for

determining a first set of a first number of states associated with an occupancy state for each of a plurality of decoded voxels neighboring a current voxel, wherein the current voxel and the plurality of neighboring decoded voxels are in a point cloud represented by the point cloud data;

determining a density of occupied voxels in the plurality of neighboring encoded voxels;

processing the first set of states to obtain a second set of a second number of states, based on the density of occupied voxels in the plurality of neighboring encoded voxels;

predicting, based on the second set of states, a probability for an occupancy state for the current voxel; and

decoding the occupancy state for the current voxel, based on said predicted probability for the occupancy state for the current voxel.

14 . The apparatus of claim 13 , wherein processing the first set of states comprises converting the first set of states to a summarized state space having the second number of states.

15 . The apparatus of claim 14 , wherein converting the first set of states to the summarized state space comprises determining that one or more of a set of patterns exists in a voxel neighborhood comprising the plurality of neighboring voxels.

16 . The apparatus of claim 13 , wherein processing the first set of states comprises applying a neural network including a plurality of point-based MLP layers.

17 . A method for decoding point cloud data, comprising:

determining a first set of a first number of states associated with an occupancy state for each of a plurality of decoded voxels neighboring a current voxel, wherein the current voxel and the plurality of decoded neighboring voxels are in a point cloud represented by the point cloud data;

determining a density of occupied voxels in the plurality of neighboring encoded voxels;

processing the first set of states to obtain a second set of a second number of states, based on the density of occupied voxels in the plurality of neighboring encoded voxels;

predicting, based on the second set of states, a probability for an occupancy state for the current voxel; and

decoding the occupancy state for the current voxel, based on said predicted probability for the occupancy state for the current voxel.

18 . The method of claim 17 , wherein processing the first set of states comprises converting the first set of states to a summarized state space having the second number of states.

19 . The method of claim 18 , wherein converting the first set of states to the summarized state space comprises determining that one or more of a set of patterns exists in a voxel neighborhood comprising the plurality of neighboring voxels.

20 . The method of claim 17 , wherein processing the first set of states comprises applying a neural network including a plurality of point-based MLP layers.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 4, 2024
From: QUACH, MAURICE; PANG, JIAHAO; LODHI, MUHAMMAD ASAD; TIAN, DONG; VALENZISE, GIUSEPPE; DUFAUX, FREDERIC
To: INTERDIGITAL PATENT HOLDINGS INC.
Reel/Frame 067315/0076 →
Continuity (3)
Provisional Application 63332375 · Apr 19, 2022
Provisional Application 63275511 · Nov 4, 2021
Related Publication 20250014228A1 · Jan 9, 2025
References Cited (17)
US 20210192797A1 · Lasserre · 2021 [cited by examiner]
US 20210383575A1 · Zhang · 2021 [cited by examiner]
US 20230162402A1 · Zhang · 2023 [cited by examiner]
US 20240078715A1 · Lodhi · 2024 [cited by examiner]
US 20240406427A1 · Lodhi · 2024 [cited by examiner]
EP 3553745A1 · 2019 [cited by applicant]
WO WO2022150680A1 · 2022 [cited by applicant]
WO WO2023059727A1 · 2023 [cited by applicant]
Lasserre, Sébastien, David Flynn, and Shouxing Qu. “Using neighbouring nodes for the compression of octrees representing the geometry of point clouds.” Proceedings of the 10th ACM Multimedia Systems Conference. 2019. (Y… [cited by examiner]
“G-PCC codec description v12”, WG 7, MPEG 3D Graphics Coding, ISO/IEC JTC 1/SC 29/WG 7, N0151, Sep. 24, 2021. [cited by applicant]
Lodhi et al., “Point Cloud Geometry Compression Using Learned Octree Entropy Coding”, ISO/IEC JTC 1/SC 29/WG 7 m58167, Oct. 2021. [cited by applicant]
Lasserre et al., “On Improving The Coding of Neighbour-Based Occupancy of Octree in GCC”, m58303 presentation slides, 136, MPEG Meeting, Nov. 10, 2021. [cited by applicant]
Qi et al., “Pointnet++: Deep Hierarchical Feature Learning on Point Sets in a Metric Space”, Cornell University Library, arXiv Computer Science, Computer Vision and Pattern Recognition, Jun. 7, 2017, 14 pages. [cited by applicant]
Que et al., “VoxelContext-Net: An Octree based Framework for Point Cloud Compression”, Cornell University Library, Computer Science, Computer Vision and Pattern Recognition, Document: arXiv:2105.02158v1, May 5, 2021, 10… [cited by applicant]
Huang, et al., “OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, 2020, pp. 1313-1323. [cited by applicant]
Kaya et al., “Neural Network Modeling of Probabilities for Coding the Octree Representation of Point Clouds”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 11, 2021, 6… [cited by applicant]
Lasserre et al., “Using neighbouring nodes for the compression of octrees representing thegeometry of point clouds”, In Proceedings of the 10th ACM Multimedia Systems Conference (pp. 145-153), Jun. 30, 2019. [cited by applicant]