IP Library Granted Patent US 12,488,511
Granted Patent B2
US 12,488,511 · App. 18/271,738 · Granted Dec 2, 2025

Apparatus and method for point cloud processing

Inventors: Muhammad Asad Lodhi (Highland Park, NJ); Jiahao Pang (Plainsboro, NJ); Dong Tian (Boxborough, MA)
Assignee: InterDigital Patent Holdings, Inc.
G06T9/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,488,511
App. No.
18/271,738
Granted
Dec 2, 2025
Kind
B2
Abstract

A method, apparatus or system for processing point cloud information can involve a learned deep entropy model over octrees for lossless compression/decompression of 3D point cloud data, wherein self-supervised compression/decompression involves an adaptive entropy coder operating on a tree-structured conditional entropy model and utilizing information from the local neighborhood as well as the global topology from the tree structure.

Claims (39)

1 . An apparatus comprising a processor configured to receive a bitstream including a compressed point cloud represented as a tree structure; and to decode the bitstream by, starting from a root node of the tree structure as a current node:

constructing a context for the current node comprising at least one of a level, an octant, a location, or a parent of the current node;

predict an occupancy symbol distribution that indicates child nodes occupancy probabilities for the current node by using a learning-based entropy model based on the context for the current node and feature information from one or more available neighboring nodes and from one or more ancestor nodes of the current node; and

decode an occupancy symbol for the current node using an adaptive entropy decoder based on the occupancy symbol distribution;

expanding the tree structure based on the decoded occupancy symbol; and

output the expanded tree structure representative of a reconstruction of the compressed point cloud.

2 . The apparatus of claim 1 , wherein the tree structure is an octree, a kd-tree, a quad tree-binary tree, or a prediction tree.

3 . The apparatus of claim 1 , wherein the processor is configured to predict occupancy symbol distributions for each node of the tree structure in parallel.

4 . The apparatus of claim 1 , wherein the learning-based entropy model uses one or more deep features of one or more sibling nodes of the one or more ancestor nodes through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

5 . The apparatus of claim 1 , wherein the learning-based entropy model uses one or more deep features of all nodes at one or more ancestor levels through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

6 . A method comprising:

receiving an encoded bitstream including a compressed point cloud represented as a tree structure; and

decoding the encoded bitstream by, starting from a root node of the tree structure as a current node:

constructing a context for the current node comprising at least one of a level, an octant, a location, or a parent of the current node;

predicting an occupancy symbol distribution that indicates child nodes occupancy probabilities for the current node by using a learning-based entropy model and based on the context for the current node and feature information from one or more available neighboring nodes and from one or more ancestor nodes of the current node;

decoding an occupancy symbol for the current node using an adaptive entropy decoder based on the occupancy symbol distribution;

expanding the tree structure based on the decoded occupancy symbol; and

outputting the expanded tree structure representative of a reconstruction of the compressed point cloud.

7 . The method of claim 6 , wherein predicting the occupancy symbol distributions for each node of the tree structure is performed in parallel.

8 . The method of claim 6 , wherein the learning-based model uses one or more deep features of all nodes at a parent level k associated with the current node for all layers deeper than k+1.

9 . The method of claim 6 , wherein the tree structure is an octree, a kd-tree, a quad tree-binary tree, or a prediction tree.

10 . The method of claim 6 , wherein the learning-based entropy model uses one or more deep features of one or more sibling nodes of the one or more ancestor nodes through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

11 . The method of claim 6 , wherein the learning-based entropy model uses one or more deep features of all nodes at one or more ancestor levels through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

12 . An apparatus for compressing a point cloud represented as a tree structure; the apparatus comprising a processor configured to construct a context for each node of the tree structure comprising at least one of a level, an octant, a location, or a parent of the current node, by, starting from a root node of the tree structure as a current node:

predicting an occupancy symbol distribution that indicates child nodes occupancy probabilities for the current node by using a learning based entropy model based on the context for the current node and feature information from one or more neighboring nodes and from one or more ancestor nodes of the current node;

encoding an occupancy symbol for the current node into one of a plurality of bitstreams representative of the current node by using an adaptive entropy encoder based on the predicted occupancy symbol distribution; and

combine the plurality of bitstreams.

13 . The apparatus of claim 12 , wherein the tree structure is an octree, a kd-tree, a quad tree-binary tree, or a prediction tree.

14 . The apparatus of claim 12 , wherein the processor is configured to predict occupancy symbol distributions for each node of the tree structure in parallel.

15 . The apparatus of claim 12 , wherein the learning-based entropy model uses one or more deep features of one or more sibling nodes of the one or more ancestor nodes through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

16 . A method for compressing a point cloud represented as a tree structure, the method comprising:

constructing a context for each node of the tree structure comprising at least one of a level, an octant, a location, or a parent of the current node, by, starting from a root node of the tree structure as a current node:

predicting an occupancy symbol distribution that indicates child nodes occupancy probabilities for the current node by using a learning based entropy model based on the context for the current node and feature information from one or more neighboring nodes and from one or more ancestor nodes of the current node;

encoding an occupancy symbol for the current node into one of a plurality of bitstreams representative of the current node by using an adaptive entropy encoder based on the predicted occupancy symbol distribution; and

combining the plurality of bitstreams.

17 . The method of claim 16 , wherein predicting the occupancy symbol distributions for each node of the tree structure is performed in parallel.

18 . The method of claim 16 , wherein the learning-based model uses one or more deep features of all nodes at a parent level k associated with the current node for all layers deeper than k+1.

19 . The method of claim 16 , wherein the tree structure is an octree, a kd-tree, a quad tree-binary tree, or a prediction tree.

20 . The method of claim 16 , wherein the learning-based entropy model uses one or more deep features of one or more sibling nodes of the one or more ancestor nodes through the same learning-based model used for siblings of the current node, to predict the occupancy symbol distributions.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 11, 2023
From: LODHI, MUHAMMAD ASAD; PANG, JIAHAO; TIAN, DONG
To: INTERDIGITAL PATENT HOLDINGS, INC.
Reel/Frame 064213/0560 →
Continuity (2)
Provisional Application 63135775 · Jan 11, 2021
Related Publication 20240078715A1 · Mar 7, 2024
References Cited (33)
US 20140210652A1 · Bartnik · 2014 [cited by examiner]
US 20180174275A1 · Bourdev · 2018 [cited by examiner]
US 20180367818A1 · Liu · 2018 [cited by examiner]
US 20190068995A1 · Chono · 2019 [cited by examiner]
US 20190182508A1 · Zhao · 2019 [cited by examiner]
US 20190306538A1 · Karczewicz · 2019 [cited by examiner]
US 20190378271A1 · Takeshima · 2019 [cited by examiner]
US 20200221139A1 · Vosoughi et al. · 2020 [cited by applicant]
EP 3514966A1 · 2019 [cited by applicant]
WO WO2021002730A1 · 2021 [cited by applicant]
Chen et al., “BSP-Net: Generating Compact Meshes via Binary Space Partitioning”, Institute of Electrical and Electronics Engineers (IEEE), 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Seat… [cited by applicant]
Quach et al., “Improved Deep Point Cloud Geometry Compression”, Institute of Electrical and Electronics Engineers (IEEE), 2020 IEEE 22nd International Workshop on Multimedia Signal Processing (MMSP), Tampere, Finland, S… [cited by applicant]
Tulsiani et al., “Learning Shape Abstractions by Assembling Volumetric Primitives”, Institute of Electrical and Electronics Engineering (IEEE), 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Nov… [cited by applicant]
Meagher, Donald, “Geometric Modeling using Octree Encoding”, Computer Graphics and Image Processing, vol. 19, Issue No. 2, Jun. 1982, 19 pages. [cited by applicant]
Chen et al., “Fast Resampling of Three-Dimensional Point Clouds via Graphs”, Institute of Electrical and Electronics Engineering (IEEE), IEEE Transactions on Signal Processing, vol. 66, Issue No. 3, Feb. 1, 2018, 16 pag… [cited by applicant]
Zou et al., “3D-PRNN: Generating Shape Primitives with Recurrent Neural Networks”, Institute of Electrical and Electronics Engineering (IEEE), 2017 IEEE International Conference on Computer Vision (ICCV), Oct. 22, 2017,… [cited by applicant]
Li et al., “Supervised Fitting of Geometric Primitives to 3D Point Clouds”, Institute of Electrical and Electronics Engineers (IEEE), 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beac… [cited by applicant]
Huang et al., “DeepPrimitive: Image decomposition by layered primitive detection”, Computational Visual Media, vol. 4, No. 4, Dec. 2018, 13 pages. [cited by applicant]
Mammou, et al., “G-PCC codec description v1”, International Organization for Standardization, ISO/IEC JTC1/ SC29/WG11, Coding of Moving Pictures and Audio, Document: N18015, Oct. 2018, Macau, China, 31 pages. [cited by applicant]
Li et al., “Grass: Generative recursive autoencoders for shape structures”, ACM Transactions on Graphics (TOG), vol. 36, Issue 4, Article No. 52, Jul. 20, 2017, 14 pages. [cited by applicant]
Que et al., “VoxelContext-Net: An Octree based Framework for Point Cloud Compression”, Cornell University Library, Computer Science, Computer Vision and Pattern Recognition, Document: arXiv:2105.02158v1, May 5, 2021, 10… [cited by applicant]
Dovrat et al., “Learning to Sample”, Institute of Electrical and Electronics Engineering (IEEE), 2019 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), Long Beach, CA, USA, Jun. 15, 2019, 10 pages. [cited by applicant]
Anonymous, “Draco 3D Data Compression”, Draco Library Home Page; URL: https://google.github.io/draco/, 2017, 2 pages. [cited by applicant]
Yang et al., “FoldingNet: Point Cloud Auto-Encoder via Deep Grid Deformation”, Institute of Electrical and Electronics Engineering (IEEE), 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake C… [cited by applicant]
Huang et al., “OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression”, Institute of Electrical and Electronics Engineering (IEEE), 2020 IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), … [cited by applicant]
Eldar et al., “The Farthest Point Strategy for Progressive Image Sampling”, Institute of Electrical and Electronics Engineering (IEEE), IEEE Transactions on Image Processing, vol. 6, No. 9, Sep. 1997, 11 pages. [cited by applicant]
Biswas et al., “MUSCLE: Multi Sweep Compression of LiDAR using Deep Entropy Models”, Cornell University Library, Electrical Engineering and Systems Science, Image and Video Processing, Document: arXiv:2011.07590v1, Nov.… [cited by applicant]
Li et al., “Primitive Fitting using Deep Boundary Aware Geometric Segmentation”, Cornell University Library, ARXIV, Computer Science, Computer Vision and Pattern Recognition, Document: arXiv:1810.01604v1, Oct. 3, 2018, … [cited by applicant]
Qi et al., “PointNet: Deep Learning on Point Sets for 3D Classification and Segmentation”, Institute of Electrical and Electronics Engineering (IEEE), 2017 IEEE Conference on Computer Vision and Pattern Recognition (CVP… [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems—infrastructure of audiovisual services—Coding of moving video: High Efficiency Video Coding”, International Telecommunication Union, Recommendation ITU-T H.265, O… [cited by applicant]
ISO/IEC, “Information Technology—Generic Coding of Moving Pictures and Associated Audio”, International Organization for Standardization, ISO/IEC JTC1/SC29/WG11, N0702rev, Recommendation H.262, ISO/IEC 13818-2, Mar. 25,… [cited by applicant]
ISO/IEC, “Information technology—Generic coding of moving pictures and associated audio information: Systems”, International Standard ISO/IEC 13818-1:201x—Recommendation ITU-T H.222.0, Oct. 2014, 296 pages. [cited by applicant]
ITU-T, “Infrastructure of audiovisual services—Coding of moving video—Versatile Video Coding”, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Mult… [cited by applicant]