IP Library › Granted Patent US 12,101,508
Granted Patent B2
US 12,101,508 · App. 17/814,762 · Granted Sep 24, 2024

Volumetric media process methods and apparatus

Inventors: Cheng Huang (Guangdong, CN); Yaxian Bai (Guangdong, CN)
Assignee: ZTE Corporation
H04N19/597G06T15/08G06T15/20H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,101,508
App. No.
17/814,762
Granted
Sep 24, 2024
Kind
B2
Abstract

Methods and apparatus for processing of volumetric visual data are described. One example method includes decoding, by a decoder, a bitstream containing volumetric visual information for a 3-dimensional scene that is represented as one or more atlas sub-bitstreams and one or more encoded video sub-bitstreams, reconstructing, using a result of decoding the one or more atlas sub-bitstreams and a result of decoding the one or more encoded video sub-bitstreams, the 3-dimensional scene, and rendering a target view of the 3-dimensional scene based on a desired viewing position and/or a desired viewing orientation.

Claims (64)

1. A method of volumetric visual data processing, comprising:

decoding, by a decoder, a bitstream containing volumetric visual information for a 3-dimensional scene that is represented as one or more atlas sub-bitstreams and one or more encoded video sub-bitstreams;

reconstructing, using a result of decoding the one or more atlas sub-bitstreams and a result of decoding the one or more encoded video sub-bitstreams, the 3-dimensional scene; and

rendering a target view of the 3-dimensional scene based on a desired viewing position and/or a desired viewing orientation, and

wherein the decoding of the bitstream includes:

decapsulating a group of volumetric visual tracks corresponding to an atlas group based on a first syntax element of a volumetric visual parameter track identified according to a first sample entry type, the atlas group including all atlases generated from a same view group from which one or more views of a volumetric visual data have been selected for the rendering of the target view; and

decoding the atlas group corresponding to the same view group, and

wherein each of volumetric visual tracks in the group of volumetric visual tracks is associated with a second syntax element with a track group type equal to a certain value indicating that a corresponding volumetric visual track belongs to the group of volumetric visual tracks that correspond to the atlas group.

2. The method according to claim 1 , wherein

the atlas group corresponds to the same view group from which the one or more views of the volumetric visual data have been selected for the rendering of the target view.

3. The method according to claim 2 , wherein the decapsulating of the group of volumetric visual tracks corresponding to the atlas group is performed before the decoding the atlas group, and

wherein the group of volumetric visual tracks and the volumetric visual parameter track carry all atlas data for the atlas group.

4. The method according to claim 3 , further comprising:

identifying the group of volumetric visual tracks according to a specific track group type and a specific track group identity, wherein each of volumetric visual tracks in the group of volumetric visual tracks contains a specific track reference to the volumetric visual parameter track.

5. The method according to claim 2 , further comprising:

selecting, by the decoder, the one or more views of the volumetric visual data for the target view based on one or more view group information, wherein each view group information describes one or more views, wherein each view group information further includes camera parameters for the one or more views.

6. The method according to claim 1 , wherein

wherein the group of volumetric visual tracks and the volumetric visual parameter track carry all atlas data for the atlas group.

7. The method according to claim 2 , further comprising:

selecting, by the decoder, the one or more views of the volumetric visual data for rendering of the target view based on view information for the one or more views, wherein each view information describes camera parameters of a corresponding view.

8. The method according to claim 3 , further comprising:

identifying the volumetric visual parameter track according to the first sample entry type, wherein the volumetric visual parameter track specifies constant parameter sets and common atlas data for all referenced volumetric visual tracks with a specific track reference.

9. The method of claim 3 , further comprising:

identifying a timed metadata track according to a second sample entry type that indicates the one or more views of the volumetric visual data selected for the rendering of the target view are dynamic.

10. The method according to claim 1 , wherein the one or more encoded video sub-bitstreams include at least one of:

one or more video-coded elementary streams for geometry data, zero or one video-coded elementary stream for occupancy map data, or

zero or more video-coded elementary streams for attribute data, wherein the geometry data, the occupancy map data and the attribute data are descriptive of the 3-dimensional scene.

11. A method of volumetric visual data processing, comprising:

generating, by an encoder, a bitstream containing volumetric visual information for a 3-dimensional scene by representing the 3-dimensional scene using one or more atlas sub-bitstreams and one or more encoded video sub-bitstreams, and

including, in the bitstream, information enabling rendering of a target view of the 3-dimensional scene based on a desired viewing position and/or a desired viewing orientation,

wherein the generating comprises:

encoding, by the encoder, an atlas group corresponding to a view group from which one or more views of a volumetric visual data are selectable has been selected for the rendering of the target view, the atlas group including all atlases generated from the view group;

encapsulating a group of volumetric visual tracks corresponding to the atlas group based on a first syntax element of a volumetric visual parameter track identified according to a first sample entry type, and

wherein each of volumetric visual tracks in the group of volumetric visual tracks is associated with a second syntax element with a track group type equal to a certain value indicating that a corresponding volumetric visual track belongs to the group of volumetric visual tracks that correspond to the atlas group.

12. The method according to claim 11 ,

wherein the group of volumetric visual tracks and the volumetric visual parameter track carry all atlas data for the atlas group.

13. The method according to claim 12 , further comprising:

including, in the bitstream, information identifying the group of volumetric visual tracks according to a specific track group type and a specific track group identity, wherein each of volumetric visual tracks in the group of volumetric visual tracks contains a specific track reference to the volumetric visual parameter track.

14. The method according to claim 11 ,

wherein the one or more views of the volumetric visual data for the target view are encoded based on one or more view group information, wherein each view group information describes one or more views, wherein each view group information further includes camera parameters for the one or more views.

15. The method according to claim 11 , further comprising:

including information that identifies the one or more views of the volumetric visual data for rendering of the target view based on view information for the one or more views, wherein the view information describes camera parameters of a corresponding view.

16. The method according to claim 12 , further comprising:

including, in the bitstream, information for identifying the volumetric visual parameter track according to the first sample entry type, wherein the volumetric visual parameter track specifies constant parameter sets and common atlas data for all referenced volumetric visual tracks with a specific track reference.

17. The method of claim 12 , further comprising:

including, in the bitstream, information for identifying a timed metadata track according to a second sample entry type that indicates the one or more views of the volumetric visual data selected for the rendering of the target view are dynamic.

18. The method according to claim 11 , wherein the one or more encoded video sub-bitstreams include at least one of:

one or more video-coded elementary streams for geometry data, zero or one video-coded elementary stream for occupancy map data, or

zero or more video-coded elementary streams for attribute data, wherein the geometry data, the occupancy map data and the attribute data are descriptive of the 3-dimensional scene.

19. A communication apparatus comprising a processor configured to:

decode a bitstream containing volumetric visual information for a 3-dimensional scene that is represented as one or more atlas sub-bitstreams and one or more encoded video sub-bitstreams;

reconstruct, using a result of decoding the one or more atlas sub-bitstreams and a result of decoding the one or more encoded video sub-bitstreams, the 3-dimensional scene; and

render a target view of the 3-dimensional scene based on a desired viewing position and/or a desired viewing orientation, and

wherein the decoding of the bitstream includes:

decapsulating a group of volumetric visual tracks corresponding to an atlas group based on a first syntax element of a volumetric visual parameter track identified according to a first sample entry type, the atlas group including all atlases generated from a same view group from which one or more views of a volumetric visual data have been selected for the rendering of the target view; and

decoding the atlas group corresponding to the same view group, and

wherein each of volumetric visual tracks in the group of volumetric visual tracks is associated with a second syntax element with a track group type equal to a certain value indicating that a corresponding volumetric visual track belongs to the group of volumetric visual tracks that correspond to the atlas group.

20. A communication apparatus comprising a processor configured to:

generate a bitstream containing volumetric visual information for a 3-dimensional scene by representing the 3-dimensional scene using one or more atlas sub-bitstreams and one or more encoded video sub-bitstreams, and

include, in the bitstream, information enabling rendering of a target view of the 3-dimensional scene based on a desired viewing position and/or a desired viewing orientation,

wherein the generating of the bitstream comprises:

encoding an atlas group corresponding to a view group from which one or more views of a volumetric visual data are selectable has been selected for the rendering of the target view, the atlas group including all atlases generated from the view group;

encapsulating a group of volumetric visual tracks corresponding to the atlas group based on a first syntax element of a volumetric visual parameter track identified according to a first sample entry type, and

wherein each of volumetric visual tracks in the group of volumetric visual tracks is associated with a second syntax element with a track group type equal to a certain value indicating that a corresponding volumetric visual track belongs to the group of volumetric visual tracks that correspond to the atlas group.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2022
From: HUANG, CHENG; BAI, YAXIAN
To: ZTE CORPORATION
Reel/Frame 060610/0092 →
Continuity (2)
Continuation PCTCN2020084837 · Apr 15, 2020
Related Publication 20220360819A1 · Nov 10, 2022