IP Library › Granted Patent US 12,750,481
Granted Patent B2
US 12,750,481 · App. 19/199,112 · Granted Sep 29, 2026

Sub-picture motion vectors in video coding

Inventors: Ye-Kui Wang (San Diego, CA); Jianle Chen (San Diego, CA); Fnu Hendry (San Diego, CA)
Assignee: Huawei Technologies Co., Ltd.
H04N19/117H04N19/105H04N19/119H04N19/132H04N19/137H04N19/159H04N19/172H04N19/174H04N19/176H04N19/184H04N19/186H04N19/46H04N19/52H04N19/593H04N19/70H04N19/82H04N19/86H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,750,481
App. No.
19/199,112
Granted
Sep 29, 2026
Kind
B2
Abstract

A video coding mechanism includes receiving a bitstream comprising a current picture including a sub-picture coded according to inter-prediction. Coded blocks contain candidate motion vectors for a current block of the sub-picture. The coded blocks include a collocated block from a different picture. A candidate list of candidate motion vectors for the current block are derived by excluding collocated motion vectors from the candidate list when the collocated motion vectors are included in the collocated block, when the collocated motion vectors point outside of the sub-picture, and when a flag is set to indicate the sub-picture is treated as a picture. A current motion vector for the current block is determined from the candidate list of candidate motion vectors. The current block is decoded based on the current motion vector. The current block is forwarded for display as part of a decoded video sequence.

Claims (83)

1 . A method implemented in a decoder, the method comprising:

receiving, by a receiver of the decoder, a bitstream comprising a coded sub-picture of a current picture;

obtaining, by one or more processors of the decoder, candidate motion vectors for a current block of the coded sub-picture;

deriving, by the one or more processors, a candidate list of candidate motion vectors for the current block according to temporal luma motion vector prediction, wherein the temporal luma motion vector prediction is performed according to:

when yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:

deriving variables mvLXCol and availableFlagLXCol;

otherwise, excluding collocated motion vectors of a collocated block of the current block from the candidate list by setting both components of the mvLXCol equal to zero and setting the availableFlagLXCol equal to zero;

wherein

xColBr = xCb + cbWidth;

yColBr = yCb + cbHeight;

rightBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicRightBoundaryPos : pic_width_in_luma_samples − 1; and

botBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicBotBoundaryPos : pic_height_in_luma_samples − 1,

where xColBr and yColBr specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the coded sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the coded sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma_samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the coded sub-picture, rightBoundaryPos is a computed position of the right boundary of the coded sub-picture, SubPicIdx is an index of the coded sub-picture, subpic_treated_as_pic_flag[SubPicIdx] is a flag indicating whether the coded sub-picture is treated as a picture during decoding excluding in-loop filtering operations, and CtbLog2SizeY indicates a size of a coding tree block;

determining, by the one or more processors, a current motion vector for the current block from the candidate list of candidate motion vectors; and

decoding, by the one or more processors, the current block based on the current motion vector.

2 . The method of claim 1 , further comprising obtaining, by the one or more processors, a flag from a sequence parameter set, SPS, wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the coded sub-picture, wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th coded sub-picture of each coded picture in a coded video sequence, CVS, is treated as a picture in a decoding process exclusive of in-loop filtering operations.

3 . The method of claim 1 , wherein the current block is a luma block of luma samples.

4 . The method of claim 1 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is decoded based on the reference luma samples.

5 . A method implemented in an encoder, the method comprising:

partitioning, by a one or more processors of the encoder, a current picture into a sub-picture, and the sub-picture into a current block;

obtaining, by the one or more processors, candidate motion vectors for the current block of the sub-picture;

deriving, by the one or more processors, a candidate list of candidate motion vectors for the current block according to temporal luma motion vector prediction, wherein the temporal luma motion vector prediction is performed according to:

when yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:

deriving variables mvLXCol and availableFlagLXCol;

otherwise, excluding collocated motion vectors of a collocated block of the current block from the candidate list by setting both components of the mvLXCol equal to zero and setting the availableFlagLXCol equal to zero;

wherein

xColBr = xCb + cb Width;

yColBr = yCb + cbHeight;

rightBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicRightBoundaryPos : pic_width_in_luma_samples − 1; and

botBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicBotBoundaryPos : pic_height_in_luma_samples − 1,

where xColBr and yColBR specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma_samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the sub-picture, rightBoundaryPos is a computed position of the right boundary of the sub-picture, SubPicIdx is an index of the sub-picture, subpic_treated_as_pic_flag[SubPicIdx] is a flag indicating whether the sub-picture is treated as a picture during decoding excluding in-loop filtering operations, and CtbLog2SizeY indicates a size of a coding tree block;

selecting, by the one or more processors, a current motion vector for the current block from the candidate list of candidate motion vectors;

encoding, by the one or more processors, the current block into a bitstream based on the current motion vector.

6 . The method of claim 5 , further comprising encoding, by the one or more processors, a flag into a sequence parameter set, SPS, in the bitstream, wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the sub-picture, wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th sub-picture of each coded picture in a coded video sequence, CVS, is treated as a picture in an encoding process exclusive of in-loop filtering operations.

7 . The method of claim 5 , wherein the current block is a luma block of luma samples.

8 . The method of claim 5 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is encoded based on the reference luma samples.

9 . A decoder comprising:

a receiver configured to receive a bitstream comprising a coded sub-picture of a current picture; and

one or more processors configured to:

obtain candidate motion vectors for a current block of the coded sub-picture;

derive a candidate list of candidate motion vectors for the current block according to temporal luma motion vector prediction, wherein the temporal luma motion vector prediction is performed according to:

when yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:

deriving variables mvLXCol and availableFlagLXCol;

otherwise, excluding collocated motion vectors of a collocated block of the current block from the candidate list by setting both components of the mvLXCol equal to zero and setting the availableFlagLXCol equal to zero;

wherein

xColBr = xCb + cbWidth;

yColBr = yCb + cbHeight;

rightBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicRightBoundaryPos : pic_width_in_luma_samples − 1; and

botBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

   SubPicBotBoundaryPos : pic_height_in_luma_samples − 1,

where xColBr and yColBr specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the coded sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the coded sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma_samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the coded sub-picture, rightBoundaryPos is a computed position of the right boundary of the coded sub-picture, SubPicIdx is an index of the coded sub-picture, subpic_treated_as_pic_flag[SubPicIdx] is a flag indicating whether the coded sub-picture is treated as a picture during decoding excluding in-loop filtering operations, and CtbLog2SizeY indicates a size of a coding tree block;

determine a current motion vector for the current block from the candidate list of candidate motion vectors; and

decode the current block based on the current motion vector.

10 . The decoder of claim 9 , wherein the one or more processors are further configured to obtain a flag from a sequence parameter set, SPS, wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the coded sub-picture, wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th coded sub-picture of each coded picture in a coded video sequence, CVS, is treated as a picture in a decoding process exclusive of in-loop filtering operations.

11 . The decoder of claim 9 , wherein the current block is a luma block of luma samples.

12 . The decoder of claim 9 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is decoded based on the reference luma samples.

13 . An encoder comprising:

one or more processors configured to:

partition a current picture into a sub-picture, and the sub-picture into a current block;

obtain candidate motion vectors for the current block of the sub-picture;

derive a candidate list of candidate motion vectors for the current block according to temporal luma motion vector prediction, wherein the temporal luma motion vector prediction is performed according to:

when yCb>>CtbLog2SizeY is equal to yColBr>>CtbLog2SizeY, yColBr is less than or equal to botBoundaryPos and xColBr is less than or equal to rightBoundaryPos, the following applies:

deriving variables mvLXCol and availableFlagLXCol;

otherwise, excluding collocated motion vectors of a collocated block of the current block from the candidate list by setting both components of the mvLXCol equal to zero and setting the availableFlagLXCol equal to zero;

wherein

   xColBr = xCb + cb Width;

   yColBr = yCb + cbHeight;

   rightBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

      SubPicRightBoundaryPos : pic_width_in_luma_

      samples − 1; and

   botBoundaryPos = subpic_treated_as_pic_flag[ SubPicIdx ] ?

      SubPicBotBoundaryPos : pic_height_in_luma_samples − 1,

where xColBr and yColBR specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma_samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the sub-picture, rightBoundaryPos is a computed position of the right boundary of the sub-picture, SubPicIdx is an index of the sub-picture, subpic_treated_as_pic_flag[SubPicIdx] is a flag indicating whether the sub-picture is treated as a picture during decoding excluding in-loop filtering operations, and CtbLog2SizeY indicates a size of a coding tree block;

select a current motion vector for the current block from the candidate list of candidate motion vectors; and

encode the current block into a bitstream based on the current motion vector.

14 . The encoder of claim 13 , wherein the one or more processors are further configured to encode a flag into a sequence parameter set, SPS, in the bitstream, wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the sub-picture, wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th sub-picture of each coded picture in a coded video sequence, CVS, is treated as a picture in an encoding process exclusive of in-loop filtering operations.

15 . The encoder of claim 13 , wherein the current block is a luma block of luma samples.

16 . The encoder of claim 13 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is encoded based on the reference luma samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 5, 2025
From: FUTUREWEI TECHNOLOGIES, INC.
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 071930/0941 →
Continuity (6)
Continuation 18519875 · Nov 27, 2023
Continuation 17470363 · Sep 9, 2021
Continuation PCTUS2020022082 · Mar 11, 2020
Provisional Application 62826659 · Mar 29, 2019
Provisional Application 62816751 · Mar 11, 2019
Related Publication 20250337894A1 · Oct 30, 2025
References Cited (49)
US 9462298B2 · Chong et al. · 2016 [cited by applicant]
US 11831816B2 · Wang et al. · 2023 [cited by applicant]
US 12267490B2 · Wang et al. · 2025 [cited by applicant]
US 12425582B2 · Wang et al. · 2025 [cited by applicant]
US 20070183676A1 · Hannuksela et al. · 2007 [cited by applicant]
US 20130107973A1 · Wang et al. · 2013 [cited by applicant]
US 20130202051A1 · Zhou · 2013 [cited by applicant]
US 20130208806A1 · Hu et al. · 2013 [cited by applicant]
US 20130208808A1 · Sasai et al. · 2013 [cited by applicant]
US 20140037011A1 · Lim et al. · 2014 [cited by applicant]
US 20140254666A1 · Rapaka et al. · 2014 [cited by applicant]
US 20140369415A1 · Naing et al. · 2014 [cited by applicant]
US 20150010091A1 · Hsu et al. · 2015 [cited by applicant]
US 20160080753A1 · Oh et al. · 2016 [cited by applicant]
US 20160255354A1 · Yamamoto · 2016 [cited by examiner]
US 20170085917A1 · Hannuksela · 2017 [cited by applicant]
US 20180091825A1 · Zhao et al. · 2018 [cited by applicant]
US 20190075328A1 · Huang et al. · 2019 [cited by applicant]
US 20190306515A1 · Shima · 2019 [cited by applicant]
US 20200120359A1 · Hanhart et al. · 2020 [cited by applicant]
US 20200260071A1 · Hannuksela · 2020 [cited by examiner]
US 20210092436A1 · Zhang · 2021 [cited by examiner]
US 20210152816A1 · Zhang · 2021 [cited by examiner]
US 20210185330A1 · Kim et al. · 2021 [cited by applicant]
US 20210409703A1 · Wang et al. · 2021 [cited by applicant]
EP 2728876A1 · 2014 [cited by applicant]
JP 2014534738A · 2014 [cited by applicant]
JP 2015504254A · 2015 [cited by applicant]
JP 2018107500A · 2018 [cited by applicant]
JP 2020517164A · 2020 [cited by applicant]
JP 2022507590A · 2022 [cited by applicant]
WO 2013168407A1 · 2013 [cited by applicant]
WO 2014051915A1 · 2014 [cited by applicant]
WO 2018191224A1 · 2018 [cited by applicant]
WO 2019009590A1 · 2019 [cited by applicant]
Document: JCTVC-10356, Coban, M., et al., “Support of independent sub-pictures,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 9th Meeting: Geneva, CH, Apr. 27-May 7,… [cited by applicant]
Document: JVET-Q0352, Jang, H., et al., “AHG9/AHG12: On subpicture boundary”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 17th Meeting: Brussels, BE, Jan. 7-17, 2020, 12 pages. [cited by applicant]
Hannuksela, M., et al., “Sub-Picture Video Coding for Unequal Error Protection”, Mar. 2002, 4 pages. [cited by applicant]
Huawei Technologies Co., Ltd. “[OMAF] On Signaling of some properties of sub-picture tracks”, ISO/IEC JTC1/SC29/WG11 MPEG2018/M42964, Ljubljana, SI, Jul. 2018, 3 pages. [cited by applicant]
“Line Transmission of Non-Telephone Signals, Video Codec for Audiovisual Services at p x 64 kbits,” ITU-T, H.261, Mar. 1993, 29 pages. [cited by applicant]
“Transmission of Non-Telephone Signals, Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Video,” H.262, Jul. 1995, 211 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Video coding for low bit rate communication,” ITU-T, H.263, Jan. 2005, 226 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services,” ITU-T, H.264, Jun. 2019, 836 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, High efficiency video coding,” H.265, Apr. 2013, 317 pages. [cited by applicant]
Bross, B., et al., “Versatile Video Coding (Draft 3),” JVET-L1001-v3, Oct. 3-12, 2018, 181 pages. [cited by applicant]
JVET-M1001-v6, “Versatile Video Coding (Draft 4),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, 298 pages. [cited by applicant]
JVET-M0261, “AHG12: On grouping of tiles,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, 11 pages. [cited by applicant]
JCTVC-AC1005-v2, “HEVC Additional Supplemental Enhancement Information (Draft 4),” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 29th Meeting: Macao, CN, Oct. 19-25… [cited by applicant]
Document: JCTVC-J0088, Zhou, M. et al., “AHG4: Enable parallel decoding with tiles,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 10th Meeting: Stockholm, Sweden, Jul. 1… [cited by applicant]