IP Library Granted Patent US 12671855
Granted Patent B2
US 12671855 · App. 18/928,266 · Granted Jun 30, 2026

Media data processing method and apparatus, device, and readable storage medium

Inventor: Ying Hu (Shenzhen, CN)
Assignee: TENCENT TECHNOLOGY (SHENZHEN) COMPANY LIMITED
H04N21/236H04L65/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12671855
App. No.
18/928,266
Granted
Jun 30, 2026
Kind
B2
Abstract

A media data processing method, performed by a computer device, includes: acquiring point cloud media; determining saliency information of the point cloud media; encoding the point cloud media to acquire a point cloud code stream; and encapsulating the point cloud code stream and the saliency information into a media file, wherein the saliency information includes a saliency level parameter indicating a target range of the point cloud media, and wherein the target range includes at least one of a spatial range or a time range.

Claims (121)

1 . A media data processing method, performed by a computer device, comprising:

acquiring point cloud media;

determining saliency information of the point cloud media, the saliency information comprising a saliency level parameter indicating a target range of the point cloud media;

encoding the point cloud media to obtain a point cloud code stream; and

encapsulating the point cloud code stream and the saliency information into a media file,

wherein the target range comprises at least one of a spatial range or a time range,

wherein the saliency information is included in a saliency information data box within a sample entry of a point cloud track of the media file when the target range comprises the spatial range and not the time range,

wherein the saliency information is included in a saliency information metadata track different from the point cloud track when the target range comprises the time range and not the spatial range, and

wherein the saliency information metadata track comprises a spatial region identifier corresponding to the spatial range when the target range comprises both the spatial range and the time range.

2 . The method according to claim 1 , wherein the saliency information data box indicates the saliency information,

wherein a first saliency level parameter of the saliency information indicates one spatial range,

wherein the one spatial range comprises one spatial region,

wherein the point cloud media comprises one or more point cloud frames, and

wherein one or more of the one or more point cloud frames comprise the one spatial range.

3 . The method according to claim 2 , wherein the saliency information data box comprises a data structure quantity field, and

wherein the data structure quantity field indicates a total quantity of saliency information data structures.

4 . The method according to claim 3 , wherein a value of the data structure quantity field is S, representing S saliency information data structures,

wherein the S saliency information data structures comprise a saliency information data structure B c ,

wherein c indicates an index,

wherein S and c are both positive integers, and c is less than or equal to S,

wherein the saliency information data structure B c comprises a saliency level field,

wherein the saliency level field comprises a saliency level parameter D c and a target range field,

wherein the target range field is a first value,

wherein the saliency level parameter D c corresponds to a second saliency level parameter of the saliency information,

wherein the first value references the saliency level parameter D c , and

wherein the saliency level parameter D c indicates a saliency level of the spatial range.

5 . The method according to claim 4 , wherein the saliency information data structure B c further comprises a first spatial range field,

wherein, based on the first spatial range field being a second value, a first spatial range indicated by the saliency level parameter D c is determined based on a spatial region identifier,

wherein, based on the first spatial range field being a third value, the first spatial range is determined based on spatial region location information, and

wherein the third value is different than the second value.

6 . The method according to claim 5 , wherein based on the first spatial range field being the second value, the saliency information data structure B c further comprises a spatial region identifier field, and

wherein the spatial region identifier field indicates a first spatial region identifier of the first spatial range.

7 . The method according to claim 5 , wherein based on the first spatial range field being the third value, the saliency information data structure B c further comprises a spatial region location information field, and

wherein the spatial region location information field indicates spatial region location information of the first spatial range.

8 . The method according to claim 5 , wherein the saliency information data structure B c further comprises a point cloud slice information field,

wherein, based on the point cloud slice information field being a first information value, the first spatial range corresponds to a point cloud slice,

wherein, based on the point cloud slice information field being a second information value,

the first spatial range does not have an associated point cloud slice, and

the second information value is different from the first information value,

wherein, based on the point cloud slice information field is the first information value,

the saliency information data structure B c further comprises a point cloud slice quantity field and a point cloud slice identifier field,

the point cloud slice quantity field is configured for indicating a total quantity of associated point cloud slices, and

the point cloud slice identifier field is configured for indicating a point cloud slice identifier corresponding to the associated point cloud slice.

9 . The method according to claim 8 , wherein the saliency information data structure B c further comprises a spatial tile information field,

wherein, based on the spatial tile information field being a third information value, the first spatial range has an associated spatial tile,

wherein, based on the point cloud slice information field being a fourth information value,

the first spatial range has no associated spatial tile, and

the fourth information value is different from the third information value,

wherein, based on the spatial tile information field being the third information value,

the saliency information data structure B c further comprises a spatial tile quantity field and a spatial tile identifier field,

the spatial tile quantity field is configured for indicating a total quantity of associated spatial tiles, and

the spatial tile identifier field is configured for indicating a spatial tile identifier corresponding to the associated spatial tile.

10 . The method according to claim 2 , wherein the saliency information data box comprises a saliency algorithm type field,

wherein, based on the saliency algorithm type field being a first type value, the saliency information is determined by a saliency detection algorithm, and

wherein, based on the saliency algorithm type field being a second type value,

the saliency information is determined based on data statistics, and

the second type value is different from the first type value.

11 . A media data processing apparatus, comprising:

at least one memory configured to store computer program code; and

at least one processor configured to read the program code and operate as instructed by the program code, the program code comprising:

acquiring code configured to cause at least one of the at least one processor to acquire point cloud media;

determining code configured to cause at least one of the at least one processor to determine saliency information of the point cloud media, the saliency information comprising a saliency level parameter indicating a target range of the point cloud media;

encoding code configured to cause at least one of the at least one processor to encode the point cloud media to obtain a point cloud code stream; and

encapsulation code configured to cause at least one of the at least one processor to encapsulate the point cloud code stream and the saliency information into a media file,

wherein the target range comprises at least one of a spatial range or a time range,

wherein the saliency information is included in a saliency information data box within a sample entry of a point cloud track of the media file when the target range comprises the spatial range and not the time range,

wherein the saliency information is included in a saliency information metadata track different from the point cloud track when the target range comprises the time range and not the spatial range, and

wherein the saliency information metadata track comprises a spatial region identifier corresponding to the spatial range when the target range comprises both the spatial range and the time range.

12 . The apparatus according to claim 11 , wherein

the saliency information data box indicates the saliency information,

wherein a first saliency level parameter of the saliency information indicates one spatial range,

wherein the one spatial range comprises one spatial region,

wherein the point cloud media comprises one or more point cloud frames, and

wherein one or more of the one or more point cloud frames comprise the one spatial range.

13 . The apparatus according to claim 12 , wherein the saliency information data box comprises a data structure quantity field, and

wherein the data structure quantity field indicates a total quantity of saliency information data structures.

14 . The apparatus according to claim 13 , wherein a value of the data structure quantity field is S, representing S saliency information data structures,

wherein the S saliency information data structures comprise a saliency information data structure B c ,

wherein c indicates an index,

wherein S and c are both positive integers, and c is less than or equal to S,

wherein the saliency information data structure B c comprises a saliency level field,

wherein the saliency level field comprises a saliency level parameter D c and a target range field,

wherein the target range field is a first value,

wherein the saliency level parameter D c corresponds to a second saliency level parameter of the saliency information,

wherein the first value references the saliency level parameter D c , and

wherein the saliency level parameter D c indicates a saliency level of the spatial range.

15 . The apparatus according to claim 14 , wherein the saliency information data structure B c further comprises a first spatial range field,

wherein, based on the first spatial range field being a second value, a first spatial range indicated by the saliency level parameter D c is determined based on a spatial region identifier,

wherein, based on the first spatial range field being a third value, the first spatial range is determined based on spatial region location information, and

wherein the third value is different than the second value.

16 . The apparatus according to claim 15 , wherein based on the first spatial range field being the second value, the saliency information data structure B c further comprises a spatial region identifier field, and

wherein the spatial region identifier field indicates a first spatial region identifier of the first spatial range.

17 . The apparatus according to claim 15 , wherein based on the first spatial range field being the third value, the saliency information data structure B c further comprises a spatial region location information field, and

wherein the spatial region location information field indicates spatial region location information of the first spatial range.

18 . The apparatus according to claim 15 , wherein the saliency information data structure B c further comprises a point cloud slice information field,

wherein, based on the point cloud slice information field being a first information value, the first spatial range corresponds to a point cloud slice,

wherein, based on the point cloud slice information field being a second information value,

the first spatial range does not have an associated point cloud slice, and

the second information value is different from the first information value,

wherein, based on the point cloud slice information field is the first information value,

the saliency information data structure B c further comprises a point cloud slice quantity field and a point cloud slice identifier field,

the point cloud slice quantity field is configured for indicating a total quantity of associated point cloud slices, and

the point cloud slice identifier field is configured for indicating a point cloud slice identifier corresponding to the associated point cloud slice.

19 . The apparatus according to claim 18 , wherein the saliency information data structure B c further comprises a spatial tile information field,

wherein, based on the spatial tile information field being a third information value, the first spatial range has an associated spatial tile,

wherein, based on the point cloud slice information field being a fourth information value,

the first spatial range has no associated spatial tile, and

the fourth information value is different from the third information value,

wherein, based on the spatial tile information field being the third information value,

the saliency information data structure B c further comprises a spatial tile quantity field and a spatial tile identifier field,

the spatial tile quantity field is configured for indicating a total quantity of associated spatial tiles, and

the spatial tile identifier field is configured for indicating a spatial tile identifier corresponding to the associated spatial tile.

20 . A non-transitory computer readable storage medium, storing computer code which, when executed by at least one processor, causes the at least one processor to at least:

acquire point cloud media;

determine saliency information of the point cloud media, the saliency information comprising a saliency level parameter indicating a target range of the point cloud media;

encode the point cloud media to obtain a point cloud code stream; and

encapsulate the point cloud code stream and the saliency information into a media file,

wherein the target range comprises at least one of a spatial range or a time range,

wherein the saliency information is included in a saliency information data box within a sample entry of a point cloud track of the media file when the target range comprises the spatial range and not the time range,

wherein the saliency information is included in a saliency information metadata track different from the point cloud track when the target range comprises the time range and not the spatial range, and

wherein the saliency information metadata track comprises a spatial region identifier corresponding to the spatial range when the target range comprises both the spatial range and the time range.