IP Library Granted Patent US 12671841
Granted Patent B2
US 12671841 · App. 18/526,640 · Granted Jun 30, 2026

Combination of subpictures and scalability

Inventors: Ye-kui Wang (San Diego, CA); Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Zhipin Deng (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/70H04N19/105H04N19/172H04N19/187H04N19/46H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12671841
App. No.
18/526,640
Granted
Jun 30, 2026
Kind
B2
Abstract

Several techniques for video encoding and video decoding are described. One example method includes performing a conversion between a subpicture in a video picture of a video and a bitstream of the video according to a rule. The rule specifies that, in in case a subpicture is treated as a video picture for the conversion, a cross-layer alignment restriction is applied to less than all of the multiple layers including a current layer that includes the subpicture and a subset of layers associated with the current layer.

Claims (70)

1 . A method of processing video data, comprising:

performing a conversion between a current picture of a video and a bitstream of the video according to a rule,

wherein the rule specifies that a first syntax element specifying whether reference picture resampling is enabled is included in a sequence parameter set of the bitstream, and a second syntax element specifying whether a spatial resolution of a picture is allowed to change within a coded layer video sequence (CLVS) referring to the sequence parameter set is conditionally included in the sequence parameter set;

wherein the video comprises multiple layers, wherein the rule further specifies that, when a fourth syntax element included in the sequence parameter set indicates that number of subpictures in a video picture is greater than 1, and a subpicture with a first subpicture index is treated as one video picture for the conversion, a cross-layer alignment restriction is applied to a current layer that includes the subpicture and a subset of layers associated with the current layer, and

wherein the cross-layer alignment restriction includes a restriction of the current picture including the subpicture and a first picture in the subset of layers having a same value for at least one of following syntax elements:

a value of the fourth syntax element,

a value of a fifth syntax element specifying a dimension of a video picture,

a value of a sixth syntax element indicating a dimension of an i-th subpicture,

a value of a seventh syntax element indicating a location of the i-th subpicture, or

a value of the first subpicture index.

2 . The method of claim 1 , wherein the second syntax element is included in the sequence parameter set when the first syntax element specifies that the reference picture resampling is enabled, and

wherein, when the second syntax element is not included in the sequence parameter set, the second syntax element is inferred to be equivalent to a value indicating that the spatial resolution of the picture is disallowed from changing within the CLVS.

3 . The method of claim 1 , wherein a first general constraint flag corresponding to the first syntax element and a second general constraint flag corresponding to the second syntax element are included in the bitstream,

wherein, when the first general constraint flag has a first value, first syntax elements for all pictures in the sequence parameter set are equal to a value specifying that the reference picture resampling is disabled, and

wherein, when the second general constraint flag has the first value, the second syntax element for all pictures in the sequence parameter set is equal to a value specifying that the spatial resolution of the picture does not change within any CLVS referring to the sequence parameter set.

4 . The method of claim 1 , wherein the rule further specifies that when the second syntax element has a second value specifying that the spatial resolution of the picture is allowed to change within the CLVS, only one subpicture is allowed in each picture of the CLVS.

5 . The method of claim 1 , wherein the rule further specifies that when the first syntax element has a second value specifying that the reference picture resampling is enabled, more than one subpicture is allowed in each picture of the CLVS.

6 . The method of claim 1 , wherein the rule further specifies that a third syntax element is included in the bitstream,

wherein the third syntax element specifies whether scaling window offset parameters are present in a picture parameter set,

wherein, when the first syntax element has a value specifying that the reference picture resampling is disabled, the third syntax element is equal to a value specifying that the scaling window offset parameters are not present in the picture parameter set, and

wherein the first syntax element, the second syntax element, and the third syntax element are coded by using a unary coding method with one bit.

7 . The method of claim 1 , wherein the first picture is a reference picture of the current picture or the current picture is a reference picture of the first picture.

8 . The method of claim 1 , wherein the subset of layers associated with the current layer includes one or more higher layers that depend on the current layer, or

wherein the subset of layers associated with the current layer excludes all higher layers that do not depend on the current layer.

9 . The method of claim 1 , wherein the subset of layers associated with the current layer excludes all lower layers of the current layer.

10 . The method of claim 1 , wherein the subset of layers associated with the current layer is a subset of a dependency tree associated with the current layer,

wherein the dependency tree associated with the current layer includes the current layer, all layers that have the current layer as a reference layer, and all reference layers of the current layer, and

wherein the subset of layers associated with the current layer is the subset of the dependency tree, regardless of whether any of the subset of the dependency tree is an output layer in an output layer set.

11 . The method of claim 1 , wherein the cross-layer alignment restriction further includes a restriction regarding a value of an eighth syntax element,

wherein the current layer and the subset of layers associated with the current layer have a same value of the eighth syntax element, and

wherein the eighth syntax element specifies whether the subpicture of each coded picture in a CLVS is treated as a picture in a decoding process excluding in-loop filtering operations.

12 . The method of claim 1 , wherein the cross-layer alignment restriction excludes a value of a ninth syntax element, and

wherein the ninth syntax element specifies whether an in-loop filtering operation across subpicture boundaries is enabled.

13 . The method of claim 1 , wherein the rule further specifies that the cross-layer alignment restriction is not applied when the fourth syntax element indicates that the video picture includes a single subpicture.

14 . The method of claim 1 , wherein the rule further specifies that the cross-layer alignment restriction is not applied when a tenth syntax element included in the sequence parameter set indicates that subpicture information is not present.

15 . The method of claim 1 , wherein the rule further specifies that the cross-layer alignment restriction is applied to pictures in a target set of access units, and

wherein, for each CLVS of the current layer referring to the sequence parameter set, the target set of access units includes all access units starting from a first access unit that includes a second picture of the CLVS to a second access unit that includes a last picture of the CLVS according to a decoding order.

16 . The method of claim 1 , wherein the conversion comprises encoding the video into the bitstream.

17 . The method of claim 1 , wherein the conversion comprises decoding the bitstream to generate the video.

18 . An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

perform a conversion between a current picture of a video and a bitstream of the video according to a rule,

wherein the rule specifies that a first syntax element specifying whether reference picture resampling is enabled is included in a sequence parameter set of the bitstream, and a second syntax element specifying whether a spatial resolution of a picture is allowed to change within a coded layer video sequence (CLVS) referring to the sequence parameter set is conditionally included in the sequence parameter set;

wherein the video comprises multiple layers, wherein the rule further specifies that, when a fourth syntax element included in the sequence parameter set indicates that number of subpictures in a video picture is greater than 1, and a subpicture with a first subpicture index is treated as one video picture for the conversion, a cross-layer alignment restriction is applied to a current layer that includes the subpicture and a subset of layers associated with the current layer, and

wherein the cross-layer alignment restriction includes a restriction of the current picture including the subpicture and a first picture in the subset of layers having a same value for at least one of following syntax elements:

a value of the fourth syntax element,

a value of a fifth syntax element specifying a dimension of a video picture,

a value of a sixth syntax element indicating a dimension of an i-th subpicture,

a value of a seventh syntax element indicating a location of the i-th subpicture, or

a value of the first subpicture index.

19 . A non-transitory computer-readable storage medium storing instructions that cause a processor to:

perform a conversion between a current picture of a video and a bitstream of the video according to a rule,

wherein the rule specifies that a first syntax element specifying whether reference picture resampling is enabled is included in a sequence parameter set of the bitstream, and a second syntax element specifying whether a spatial resolution of a picture is allowed to change within a coded layer video sequence (CLVS) referring to the sequence parameter set is conditionally included in the sequence parameter set;

wherein the video comprises multiple layers, wherein the rule further specifies that, when a fourth syntax element included in the sequence parameter set indicates that number of subpictures in a video picture is greater than 1, and a subpicture with a first subpicture index is treated as one video picture for the conversion, a cross-layer alignment restriction is applied to a current layer that includes the subpicture and a subset of layers associated with the current layer, and

wherein the cross-layer alignment restriction includes a restriction of the current picture including the subpicture and a first picture in the subset of layers having a same value for at least one of following syntax elements:

a value of the fourth syntax element,

a value of a fifth syntax element specifying a dimension of a video picture,

a value of a sixth syntax element indicating a dimension of an i-th subpicture,

a value of a seventh syntax element indicating a location of the i-th subpicture, or

a value of the first subpicture index.

20 . A method for storing bitstream of a video, comprising:

generating the bitstream for a current picture of the video according to a rule,

storing the bitstream in a non-transitory computer-readable recording medium,

wherein the rule specifies that a first syntax element specifying whether reference picture resampling is enabled is included in a sequence parameter set of the bitstream, and a second syntax element specifying whether a spatial resolution of a picture is allowed to change within a coded layer video sequence (CLVS) referring to the sequence parameter set is conditionally included in the sequence parameter set;

wherein the video comprises multiple layers, wherein the rule further specifies that, when a fourth syntax element included in the sequence parameter set indicates that number of subpictures in a video picture is greater than 1, and a subpicture with a first subpicture index is treated as one video picture for the generating, a cross-layer alignment restriction is applied to a current layer that includes the subpicture and a subset of layers associated with the current layer, and

wherein the cross-layer alignment restriction includes a restriction of the current picture including the subpicture and a first picture in the subset of layers having a same value for at least one of following syntax elements:

a value of the fourth syntax element,

a value of a fifth syntax element specifying a dimension of a video picture,

a value of a sixth syntax element indicating a dimension of an i-th subpicture,

a value of a seventh syntax element indicating a location of the i-th subpicture, or

a value of the first subpicture index.