IP Library Granted Patent US 12,183,044
Granted Patent B2
US 12,183,044 · App. 17/816,102 · Granted Dec 31, 2024

Three-dimensional content processing methods and apparatus

Inventors: Yaxian Bai (Guangdong, CN); Cheng Huang (Guangdong, CN)
Assignee: ZTE Corporation
G06T9/001H04N19/1883H04N19/597H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,183,044
App. No.
17/816,102
Granted
Dec 31, 2024
Kind
B2
Abstract

Methods, systems, and apparatus for processing of three-dimensional visual content are described. One example method of processing three-dimensional content includes parsing a level of detail (LoD) information of a bitstream containing three-dimensional (3D) content that is represented as one geometry sub-bitstream and one or more attribute sub-bitstreams; and generating, based on the LoD information, decoded information by decoding at least a portion of the geometry sub-bitstream and the one or more attribute sub-bitstreams corresponding to a desired level of detail; and reconstructing, using the decoded information, a three-dimensional scene corresponding at least to the desired level of detail. The bitstream conforms to a format organized according to multiple levels of details of the 3D content.

Claims (78)

1. A method of processing three-dimensional content, comprising:

parsing a level of detail (LoD) information of a bitstream containing three-dimensional (3D) content that is represented as one geometry sub-bitstream and one or more attribute sub-bitstreams, wherein the parsing the LoD information comprises identifying a first syntax structure in the bitstream that includes multiple levels of details;

generating, based on the LoD information, decoded information by decoding at least a portion of the geometry sub-bitstream and the one or more attribute sub-bitstreams corresponding to a desired level of detail;

reconstructing, using the decoded information, a three-dimensional scene corresponding at least to the desired level of detail,

wherein the bitstream conforms to a format organized according to multiple levels of details of the 3D content; and

using a sample entry type field in the bitstream for determining whether the bitstream supports a spatial scalability functionality and for the identifying the first syntax structure, wherein the first syntax structure of the bitstream with multiple levels of details comprises:

a first structure in which a complete set of levels of the bitstream is carried in one track with sub-sample structure;

a second structure that includes an extractor with each level of the bitstream in the one track; and

a third structure that includes one or more levels of the bitstream in the one track with redundant data from lower levels.

2. The method of claim 1 , wherein the parsing the LoD information comprises:

determining whether the bitstream comprises spatial scalability sub-bitstreams;

identifying the LoD information using a second syntax structure, a sub-sample structure, a sample entry or a descriptor; or

locating content corresponding to the desired level of detail or the complete set of levels.

3. The method of claim 2 , wherein the sub-sample structure comprises a codec_specific_parameters field extension representing the LoD information.

4. The method of claim 2 , further comprising:

identifying a value of a LoD of the bitstream using an LoD value in the sample entry.

5. The method of claim 1 , further comprising:

identifying one or more tracks containing sub-streams corresponding to complete levels of detail using a first track group type; and

decoding data in the one or more tracks corresponding to complete levels of detail.

6. The method of claim 1 , further comprising:

decoding a portion of the bitstream corresponding to the desired levels of detail and one or more lower levels in a single track.

7. The method of claim 2 , further comprising:

using a LoD descriptor to determine whether an Adaptation Set supports the spatial scalability functionality; and

identifying a LoD in the Adaptation Set using an LoD value in the LoD descriptor.

8. The method of claim 1 , further comprising:

decoding a portion of the bitstream corresponding to the desired LoD and one or more lower levels from a single Adaptation Set; or

identifying and decoding a portion of the bitstream corresponding to the desired LoD in one Adaptation Set and data with lower levels in other Adaptation Sets.

9. The method of claim 1 , further comprising:

identifying one or more Adaptation Sets containing data corresponding to all levels of detail using a complete track id; and

decoding complete data in one or more Adaptation Sets corresponding to complete levels of detail.

10. The method of claim 1 ,

wherein a portion of the bitstream corresponding to the desired LoD includes data corresponding to the desired LoD with a single-track encapsulation, and

wherein the single-track encapsulation comprises the one geometry sub-bitstream and the one or more attribute bitstreams that are encapsulated in a same track.

11. The method of claim 1 ,

wherein a portion of the bitstream corresponding to the desired LoD includes data corresponding to the desired LoD with a multiple-track encapsulation, and

wherein the multiple-track encapsulation comprises the one geometry sub-bitstream and the one or more attribute bitstreams that are encapsulated in separate tracks.

12. The method of claim 1 , wherein the reconstructing the three-dimensional scene comprises:

reconstructing a spatial position and one or more attribute values of each point in the 3D content; or

reconstructing a spatial position and attribute values of each point in the 3D content and rendering 3D scenes according to a viewing position and a viewport of a user.

13. An apparatus, comprising a processor configured to implement a method that causes the apparatus to:

parse a level of detail (LoD) information of a bitstream containing three-dimensional (3D) content that is represented as one geometry sub-bitstream and one or more attribute sub-bitstreams, wherein the parsing the LoD information comprises identify a first syntax structure in the bitstream that includes multiple levels of details;

generate, based on the LoD information, decoded information by decoding at least a portion of the geometry sub-bitstream and the one or more attribute sub-bitstreams corresponding to a desired level of detail;

reconstruct, using the decoded information, a three-dimensional scene corresponding at least to the desired level of detail,

wherein the bitstream conforms to a format organized according to multiple levels of details of the 3D content; and

use a sample entry type field in the bitstream for a determination of whether the bitstream supports a spatial scalability functionality and for the identify the first syntax structure, wherein the first syntax structure of the bitstream with multiple levels of details comprises:

a first structure in which a complete set of levels of the bitstream is carried in one track with sub-sample structure;

a second structure that includes an extractor with each level of the bitstream in the one track; and

a third structure that includes one or more levels of the bitstream in the one track with redundant data from lower levels.

14. The apparatus of claim 13 , wherein the LoD information is parsed by the processor configured to:

determine whether the bitstream comprises spatial scalability sub-bitstreams;

identify the LoD information using a second syntax structure, a sub-sample structure, a sample entry or a descriptor; or

locate content corresponding to the desired level of detail or the complete set of levels.

15. The apparatus of claim 14 , wherein the sub-sample structure comprises a codec_specific_parameters field extension representing the LoD information.

16. The apparatus of claim 14 , wherein the processor is further configured to:

identify a value of a LoD of the bitstream using an LoD value in the sample entry.

17. The apparatus of claim 14 , wherein the processor is further configured to:

use a LoD descriptor to determine whether an Adaptation Set supports the spatial scalability functionality; and

identify a LoD in the Adaptation Set using an LoD value in the LoD descriptor.

18. The apparatus of claim 13 , wherein the processor is further configured to:

identify one or more tracks containing sub-streams corresponding to complete levels of detail using a first track group type; and

decode data in the one or more tracks corresponding to complete levels of detail.

19. The apparatus of claim 13 , wherein the processor is further configured to:

decode a portion of the bitstream corresponding to the desired levels of detail and one or more lower levels in a single track.

20. The apparatus of claim 13 , wherein the processor is further configured to:

decode a portion of the bitstream corresponding to the desired LoD and one or more lower levels from a single Adaptation Set; or

identify and decoding a portion of the bitstream corresponding to the desired LoD in one Adaptation Set and data with lower levels in other Adaptation Sets.

21. The apparatus of claim 13 , wherein the processor is further configured to:

identify one or more Adaptation Sets containing data corresponding to all levels of detail using a complete track id; and

decode complete data in one or more Adaptation Sets corresponding to complete levels of detail.

22. The apparatus of claim 13 ,

wherein a portion of the bitstream corresponding to the desired LoD includes data corresponding to the desired LoD with a single-track encapsulation, and

wherein the single-track encapsulation comprises the one geometry sub-bitstream and the one or more attribute bitstreams that are encapsulated in a same track.

23. The apparatus of claim 13 ,

wherein a portion of the bitstream corresponding to the desired LoD includes data corresponding to the desired LoD with a multiple-track encapsulation, and

wherein the multiple-track encapsulation comprises the one geometry sub-bitstream and the one or more attribute bitstreams that are encapsulated in separate tracks.

24. The apparatus of claim 13 , wherein the three-dimensional scene is reconstructed by the processor configured to:

reconstruct a spatial position and one or more attribute values of each point in the 3D content; or

reconstruct a spatial position and attribute values of each point in the 3D content and render 3D scenes according to a viewing position and a viewport of a user.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 18, 2023
From: BAI, YAIXIAN; HUANG, CHENG
To: ZTE CORPORATION
Reel/Frame 065273/0104 →
Continuity (2)
Continuation PCTCN2020098010 · Jun 24, 2020
Related Publication 20220366611A1 · Nov 17, 2022