IP Library › Granted Patent US 12,348,774
Granted Patent B2
US 12,348,774 · App. 18/471,753 · Granted Jul 1, 2025

Multiscale inter-prediction for dynamic point cloud compression

Inventors: Pranav Aniruddha Kadam (Sunnyvale, CA); Alexandre Zaghetto (San Jose, CA); Danillo Bracco Graziosi (Flagstaff, AZ); Ali Tabatabai (Cupertino, CA)
Assignees: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
H04N19/597H04N19/105H04N19/124H04N19/30H04N19/503H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,348,774
App. No.
18/471,753
Granted
Jul 1, 2025
Kind
B2
Abstract

An electronic device and method for multiscale inter-prediction for dynamic point cloud compression is provided. The electronic device receives a set of reference point cloud frames and a current point cloud frame. The electronic device generates reference frame data comprising a feature set for each reference point cloud frame and a first set of features for the current point cloud frame. The electronic device predicts a second set of features for the current point cloud frame, using a first neural network predictor, based on the reference frame data. The electronic device computes a set of residual features based on the first set of features and the second set of features. The electronic device generates a set of quantized residual features based on the set of residual features and a bitstream of encoded point cloud data for the current 3D point cloud frame based on the set of quantized residual features.

Claims (61)

1. A first electronic device, comprising:

a memory configured to store a first neural network predictor; and

circuitry configured to:

receive a three-dimensional (3D) point cloud sequence that includes a set of reference 3D point cloud frames and a current 3D point cloud frame that is to be encoded;

generate reference frame data comprising a feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames;

generate current frame data associated with 3D points of the current 3D point cloud frame,

wherein the current frame data comprises a first set of features associated with an occupancy of the 3D points in the current 3D point cloud frame;

predict a second set of features associated with the 3D points of the current 3D point cloud frame based on application of the first neural network predictor on the reference frame data;

compute a set of residual features based on the first set of features and the second set of features;

generate a set of quantized residual features based on application of a quantization scheme on the set of residual features; and

generate a bitstream of encoded point cloud data for the current 3D point cloud frame based on application of an encoding scheme on the set of quantized residual features.

2. The first electronic device according to claim 1 , wherein the second set of features are predicted further based on coordinate information associated with the 3D points of the current 3D point cloud frame.

3. The first electronic device according to claim 2 , wherein the circuitry is further configured to encode the coordinate information based on an application of an octree-based encoder on the coordinate information, and

the bitstream of encoded point cloud data includes the encoded coordinate information.

4. The first electronic device according to claim 1 , wherein

the memory is further configured to store a first Point Cloud Compression (PCC) encoder,

the reference frame data is generated based on an application of the first PCC encoder on the set of reference 3D point cloud frames, and

the current frame data is generated based on application of the first PCC encoder on the current 3D point cloud frame.

5. The first electronic device according to claim 1 , wherein the feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames comprises:

reference features associated with an occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames, and

reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame.

6. The first electronic device according to claim 1 , wherein the set of reference 3D point cloud frames includes at least one reference 3D point cloud frame that precedes the current 3D point cloud frame or at least one reference 3D point cloud frame that succeeds the current 3D point cloud frame.

7. The first electronic device according to claim 1 , wherein the circuitry is further configured to down sample the feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames by at least one scaling factor.

8. A second electronic device, comprising:

a memory configured to store a second neural network predictor; and

circuitry configured to:

receive a three-dimensional (3D) point cloud sequence that includes a set of reference 3D point cloud frames;

generate reference frame data comprising a feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames;

receive a bitstream of encoded point cloud data associated with a current 3D point cloud frame that is to be decoded;

predict a third set of features associated with 3D points of the current 3D point cloud frame based on an application of the second neural network predictor on the reference frame data;

generate a fourth set of features associated with the 3D points of the current 3D point cloud frame based on the received bitstream of encoded point cloud data and the predicted third set of features; and

reconstruct the current 3D point cloud frame based on application of a decoding scheme on the generated fourth set of features.

9. The second electronic device according to claim 8 , wherein the received bitstream of encoded point cloud data further comprises encoded coordinate information associated with the 3D points of the current 3D point cloud frame.

10. The second electronic device according to claim 9 , wherein the circuitry is further configured to generate coordinate information associated with the 3D points of the current 3D point cloud frame based on application of an octree-based decoder on the encoded coordinate information.

11. The second electronic device according to claim 10 , wherein the third set of features are predicted further based on the generated coordinate information.

12. The second electronic device according to claim 8 , wherein the memory is further configured to store a Point Cloud Compression (PCC) encoder,

wherein the reference frame data is generated based on an application of the PCC encoder on the set of reference 3D point cloud frames.

13. The second electronic device according to claim 8 , wherein the feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames comprises:

reference features associated with an occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames, and

reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame.

14. A method, comprising:

in a first electronic device:

receiving a three-dimensional (3D) point cloud sequence that includes a set of reference 3D point cloud frames and a current 3D point cloud frame that is to be encoded;

generating reference frame data comprising a feature set associated with 3D points of each reference 3D point cloud frame of the set of 3D point cloud reference frames;

generating current frame data associated with 3D points of the current 3D point cloud frame,

wherein the current frame data comprises a first set of features associated with an occupancy of the 3D points in the current 3D point cloud frame;

predicting a second set of features associated with the 3D points of the current 3D point cloud frame based on application of a first neural network predictor on the reference frame data;

computing a set of residual features based on the first set of features and the second set of features;

generating a set of quantized residual features based on application of a quantization scheme on the set of residual features; and

generating a bitstream of encoded point cloud data for the current 3D point cloud frame based on an application of an encoding scheme on the set of quantized residual features.

15. The method according to claim 14 , wherein the second set of features are predicted further based on coordinate information associated with the 3D points of the current 3D point cloud frame.

16. The method according to claim 15 , further comprising encoding the coordinate information based on application of an octree-based encoder on the coordinate information,

wherein the bitstream of encoded point cloud data includes the encoded coordinate information.

17. The method according to claim 14 , wherein,

the reference frame data is generated based on an application of a first Point Cloud Compression (PCC) encoder on the set of reference 3D point cloud frames, and

the current frame data is generated based on application of the first PCC encoder on the current 3D point cloud frame.

18. The method according to claim 14 , wherein the feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames comprises:

reference features associated with an occupancy of the 3D points in a corresponding reference 3D point cloud frame of the set of reference 3D point cloud frames, and

reference coordinate information associated with the 3D points of the corresponding reference 3D point cloud frame.

19. The method according to claim 14 , wherein the set of reference 3D point cloud frames includes at least one reference 3D point cloud frame that precedes the current 3D point cloud frame or at least one reference 3D point cloud frame that succeeds the current 3D point cloud frame.

20. The method according to claim 14 , further comprising down sampling the feature set associated with 3D points of each reference 3D point cloud frame of the set of reference 3D point cloud frames by at least one scaling factor.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2023
From: KADAM, PRANAV ANIRUDDHA; ZAGHETTO, ALEXANDRE; GRAZIOSI, DANILLO BRACCO; TABATABAI, ALI
To: SONY GROUP CORPORATION; SONY CORPORATION OF AMERICA
Reel/Frame 064985/0473 →
Continuity (3)
Provisional Application 63380089 · Oct 19, 2022
Related Publication 20240137563A1 · Apr 25, 2024
Related Publication 20240236369A9 · Jul 11, 2024
References Cited (20)
US 10650588B2 · Hazeghi et al. · 2020 [cited by applicant]
US 20200258262A1 · Lasserre · 2020 [cited by examiner]
US 20230105257A1 · Qi · 2023 [cited by examiner]
CN 112489072A · 2021 [cited by applicant]
WO 2022141418A1 · 2022 [cited by applicant]
Cui, et al., “Cooperative Perception Technology of Autonomous Driving in the Internet of Vehicles Environment: A Review”, MDPI, Sensors, Jul. 25, 2022, 31 pages. [cited by applicant]
“MPEG 3DG, V-PCC Test Model”, Version 8, ISO/IEC JTC1/SC29/WG11, N18884, 2019. [cited by applicant]
“MPEG 3DG, G-PCC codec description” Version 4, ISO/IEC JTC1/SC29/ WG11, N18673, 2019. [cited by applicant]
Wang, et al., “Multiscale Point Cloud Geometry Compression”, 2021 Data Compression Conference (DCC), IEEE, Nov. 7, 2020, 10 pages. [cited by applicant]
Wang, et al., “Sparse Tensor-based Multiscale Representation for Point Cloud Geometry Compression”, Computer Vision and Pattern Recognition, Oct. 21, 2022, 17 pages. [cited by applicant]
Pang, et al., “GRASP-Net: Geometric Residual Analysis and Synthesis for Point Cloud Compression”, Image and Video Processing, Sep. 9, 2022, 09 pages. [cited by applicant]
Huang, et al., “OctSqueeze: Octree-Structured Entropy Model for LiDAR Compression”, Proceedings of the IEEE/CVF conference on computer vision and pattern recognition, Jan. 8, 2021, 20 pages. [cited by applicant]
Que, et al., “VoxelContext-Net: An Octree based Framework for Point Cloud Compression”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, May 5, 2021, 10 pages. [cited by applicant]
Fan, et al., “D-DPCC: Deep Dynamic Point Cloud Compression via 3D Motion Prediction”, Computer Vision and Pattern Recognition, May 2, 2022, 7 pages. [cited by applicant]
Wang, et al., “SparsePCGCv3: Dynamic SparsePCGC with Inter Frame Prediction”, MPEG m60354, Jul. 2022. [cited by applicant]
Akhtar, et al., “Dynamic Point Cloud Geometry Compression using Sparse Convolutions”, MPEG-137, Document: m59617, Apr. 2022. [cited by applicant]
Choy, et al., “4D Spatio-Temporal ConvNets: Minkowski Convolutional Neural Networks”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Jun. 13, 2019, 21 pages. [cited by applicant]
Vaswani, et al., “Attention Is All You Need”, 31st Conference on Neural Information Processing Systems, (NIPS 2017), Aug. 2, 2023, 15 pages. [cited by applicant]
Anique Akhtar (UMKC) et al: “[AI-3DGC] Dynamic Point Cloud Geometry Compression using Sparse Convolutions”, 138. MPEG Meeting; Apr. 25-Apr. 29, 2022; Online; (Motion Picture Expert Group or ISO/IEC JTC1/SC29/WG11), No. … [cited by applicant]
Wang Jianqiang et al: “Multiscale Point Cloud Geometry Compression”, 2021 Data Compression Conference (DCC), IEEE, Mar. 23, 2021, pp. 73-82. [cited by applicant]