IP Library › Granted Patent US 12,294,699
Granted Patent B2
US 12,294,699 · App. 18/519,875 · Granted May 6, 2025

Sub-picture motion vectors in video coding

Inventors: Ye-Kui Wang (San Diego, CA); Jianle Chen (San Diego, CA); Fnu Hendry (San Diego, CA)
Assignee: Huawei Technologies Co., Ltd.
H04N19/117H04N19/105H04N19/119H04N19/132H04N19/137H04N19/159H04N19/172H04N19/174H04N19/176H04N19/184H04N19/186H04N19/46H04N19/52H04N19/593H04N19/70H04N19/82H04N19/86H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,294,699
App. No.
18/519,875
Granted
May 6, 2025
Kind
B2
Abstract

A video coding mechanism includes receiving a bitstream comprising a current picture including a sub-picture coded according to inter-prediction. Coded blocks contain candidate motion vectors for a current block of the sub-picture. The coded blocks include a collocated block from a different picture. A candidate list of candidate motion vectors for the current block are derived by excluding collocated motion vectors from the candidate list when the collocated motion vectors are included in the collocated block, when the collocated motion vectors point outside of the sub-picture, and when a flag is set to indicate the sub-picture is treated as a picture. A current motion vector for the current block is determined from the candidate list of candidate motion vectors. The current block is decoded based on the current motion vector. The current block is forwarded for display as part of a decoded video sequence.

Claims (47)

1. A method implemented by an encoder, the method comprising:

determining to encode a current block of a sub-picture of a current picture;

deriving a candidate list of candidate motion vectors for the current block based on candidate motion vectors for the current block without adding a collocated motion vector associated with a collocated block from a different picture than the current picture, when the collocated block, from a different picture than the current picture, is located outside the sub-picture, and when a flag indicates the sub-picture is treated as a picture;

determining a current motion vector for the current block from the candidate list of candidate motion vectors; and

encoding the current block into a bitstream based on the current motion vector.

2. The method of claim 1 , further comprising encoding the flag into a sequence parameter set (SPS), wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the sub-picture.

3. The method of claim 2 , wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th sub-picture of each coded picture in a coded video sequence (CVS) is treated as a picture in a decoding process excluding in-loop filtering operations.

4. The method of claim 1 , wherein deriving the candidate list of candidate motion vectors for the current block is performed according to temporal luma motion vector prediction.

5. The method of claim 4 , wherein the temporal luma motion vector prediction is performed according to:

xColBr=xCb+cbWidth;

yColBr=yCb+cbHeight;

rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicRightBoundaryPos: pic_width_in_luma_samples−1; and

botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos: pic_height_in_luma_samples−1,

where xColBr and yColBR specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma_samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the sub-picture, rightBoundaryPos is a computed position of the right boundary of the sub-picture, SubPicIdx is an index of the sub-picture, and wherein collocated motion vectors are excluded when yColBR is greater than botBoundaryPos or xColBr is greater than rightBoundaryPos.

6. The method of claim 1 , wherein the current block is a luma block of luma samples.

7. The method of claim 1 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is decoded based on the reference luma samples.

8. An encoder comprising:

a processor configured to:

determine to encode a current block of a sub-picture of a current picture;

derive a candidate list of candidate motion vectors for the current block based on candidate motion vectors for the current block without adding a collocated motion vector associated with a collocated block from a different picture than the current picture, when the collocated block, from a different picture than the current picture, is located outside the sub-picture, and when a flag indicates the sub-picture is treated as a picture; and

determine a current motion vector for the current block from the candidate list of candidate motion vectors; and

a memory coupled to the processor and configured to store a bitstream containing an encoding of the current block based on the current motion vector.

9. The encoder of claim 8 , wherein the processor is further configured to encode the flag into a sequence parameter set (SPS), wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the sub-picture.

10. The encoder of claim 9 , wherein the subpic_treated_as pic_flag[i] is set equal to one to specify that an i-th sub-picture of each coded picture in a coded video sequence (CVS) is treated as a picture in a decoding process excluding in-loop filtering operations.

11. The encoder of claim 8 , wherein deriving the candidate list of candidate motion vectors for the current block is performed according to temporal luma motion vector prediction.

12. The encoder of claim 11 , wherein the temporal luma motion vector prediction is performed according to:

xColBr=xCb+cbWidth;

yColBr=yCb+cbHeight;

rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicRightBoundaryPos: pic_width_in_luma_samples−1; and

botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos: pic_height_in_luma_samples−1,

where xColBr and yColBR specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the sub-picture, rightBoundaryPos is a computed position of the right boundary of the sub-picture, SubPicIdx is an index of the sub-picture, and wherein collocated motion vectors are excluded when yColBR is greater than botBoundaryPos or xColBr is greater than rightBoundaryPos.

13. The encoder of claim 8 , wherein the current block is a luma block of luma samples.

14. The encoder of claim 13 , wherein the current motion vector is a temporal luma motion vector pointing to reference luma samples in a reference block, and wherein the current block is decoded based on the reference luma samples.

15. A non-transitory storage medium storing an encoded bitstream for video signals,

wherein the bitstream comprises a current block of a sub-picture of a current picture,

wherein the encoded bitstream further comprises a syntax element indicating a motion vector based on a candidate list of candidate motion vectors,

wherein the candidate list of candidate motion vectors for the current block is derived based on candidate motion vectors for the current block without adding a collocated motion vector associated with a collocated block from a different picture than the current picture, when the collocated block, from a different picture than the current picture, is located outside the sub-picture, and when a flag indicates the sub-picture is treated as a picture.

16. The non-transitory storage medium of claim 15 , wherein the bitstream further comprises the flag in a sequence parameter set (SPS), wherein the flag is denoted as a subpic_treated_as_pic_flag[i], and wherein i is an index of the sub-picture.

17. The non-transitory storage medium of claim 16 , wherein the subpic_treated_as_pic_flag[i] is set equal to one to specify that an i-th sub-picture of each coded picture in a coded video sequence (CVS) is treated as a picture in a decoding process excluding in-loop filtering operations.

18. The non-transitory storage medium of claim 15 , wherein deriving the candidate list of candidate motion vectors for the current block is performed according to temporal luma motion vector prediction.

19. The non-transitory storage medium of claim 18 , wherein the temporal luma motion vector prediction is performed according to:

xColBr=xCb+cbWidth;

yColBr=yCb+cbHeight;

rightBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicRightBoundaryPos: pic_width_in_luma_samples−1; and

botBoundaryPos=subpic_treated_as_pic_flag[SubPicIdx]? SubPicBotBoundaryPos: pic_height_in_luma_samples−1,

where xColBr and yColBR specify a location of the collocated block, xCb and yCb specify a top left sample of the current block relative to a top left sample of the current picture, cbWidth is a width of the current block, cbHeight is a height of the current block, SubPicRightBoundaryPos is a position of a right boundary of the sub-picture, SubPicBotBoundaryPos is a position of a bottom boundary of the sub-picture, pic_width_in_luma_samples is a width of the current picture measured in luma samples, pic_height_in_luma samples is a height of the current picture measured in luma samples, botBoundaryPos is a computed position of the bottom boundary of the sub-picture, rightBoundaryPos is a computed position of the right boundary of the sub-picture, SubPicIdx is an index of the sub-picture, and wherein collocated motion vectors are excluded when yColBR is greater than botBoundaryPos or xColBr is greater than rightBoundaryPos.

20. The non-transitory storage medium of claim 15 , wherein the current block is a luma block of luma samples.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 27, 2024
From: FUTUREWEI TECHNOLOGIES, INC.
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 066579/0433 →
Continuity (5)
Continuation 17470363 · Sep 9, 2021
Continuation PCTUS2020022082 · Mar 11, 2020
Provisional Application 62826659 · Mar 29, 2019
Provisional Application 62816751 · Mar 11, 2019
Related Publication 20240171738A1 · May 23, 2024
References Cited (43)
US 9462298B2 · Chong et al. · 2016 [cited by applicant]
US 11831816B2 · Wang et al. · 2023 [cited by applicant]
US 20070183676A1 · Hannuksela et al. · 2007 [cited by applicant]
US 20130107973A1 · Wang et al. · 2013 [cited by applicant]
US 20130202051A1 · Zhou · 2013 [cited by applicant]
US 20130208806A1 · Hu et al. · 2013 [cited by applicant]
US 20130208808A1 · Sasai et al. · 2013 [cited by applicant]
US 20140037011A1 · Lim et al. · 2014 [cited by applicant]
US 20140254666A1 · Rapaka et al. · 2014 [cited by applicant]
US 20140369415A1 · Naing et al. · 2014 [cited by applicant]
US 20150010091A1 · Hsu et al. · 2015 [cited by applicant]
US 20160080753A1 · Oh et al. · 2016 [cited by applicant]
US 20170085917A1 · Hannuksela · 2017 [cited by applicant]
US 20180091825A1 · Zhao et al. · 2018 [cited by applicant]
US 20190075328A1 · Huang et al. · 2019 [cited by applicant]
US 20190306515A1 · Shima · 2019 [cited by applicant]
US 20200120359A1 · Hanhart et al. · 2020 [cited by applicant]
US 20200260071A1 · Hannuksela · 2020 [cited by examiner]
US 20210152816A1 · Zhang et al. · 2021 [cited by applicant]
US 20210185330A1 · Kim et al. · 2021 [cited by applicant]
EP 2728876A1 · 2014 [cited by applicant]
JP 2014534738A · 2014 [cited by applicant]
JP 2015504254A · 2015 [cited by applicant]
JP 2018107500A · 2018 [cited by applicant]
JP 2020517164A · 2020 [cited by applicant]
JP 2022507590A · 2022 [cited by applicant]
WO 2013168407A1 · 2013 [cited by applicant]
WO 2014051915A1 · 2014 [cited by applicant]
WO 2018191224A1 · 2018 [cited by applicant]
WO 2019009590A1 · 2019 [cited by applicant]
Document: JCTVC-10356, Coban, M., et al., “Support of independent sub-pictures,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 Wp 3 and ISO/IEC JTC 1/SC 29/WG 11 9th Meeting: Geneva, CH, Apr. 27-May 7,… [cited by applicant]
Document: JVET-Q0352, Jang, H., et al., “AHG9/AHG12: On subpicture boundary”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 17th Meeting: Brussels, BE, Jan. 7-17, 2020, 12 pages. [cited by applicant]
Hannuksela, M., et al., “Sub-Picture Video Coding for Unequal Error Protection”, Mar. 2002, 4 pages. [cited by applicant]
Huawei Technologies Co., Ltd. “[OMAF] On Signaling of some properties of sub-picture tracks”, ISO/IEC JTC1/SC29/WG11 MPEG2018/M42964, Ljubljana, SI, Jul. 2018, 3 pages. [cited by applicant]
“Line Transmission of Non-Telephone Signals, Video Codec for Audiovisual Services at p × 64 kbits,” ITU-T, H.261, Mar. 1993, 29 pages. [cited by applicant]
“Transmission of Non-Telephone Signals, Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Video,” H.262, Jul. 1995, 211 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Video coding for low bit rate communication,” ITU-T, H.263, Jan. 2005, 226 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, Advanced video coding for generic audiovisual services,” ITU-T, H.264, Jun. 2019, 836 pages. [cited by applicant]
“Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video, High efficiency video coding,” H.265, Apr. 2013, 317 pages. [cited by applicant]
Bross, B., et al., “Versatile Video Coding (Draft 3),” JVET-L1001-v3, Oct. 3-12, 2018, 181 pages. [cited by applicant]
JVET-M1001-v6, “Versatile Video Coding (Draft 4),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, 298 pages. [cited by applicant]
JVET-M0261, “AHG12: On grouping of tiles,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 13th Meeting: Marrakech, MA, Jan. 9-18, 2019, 11 pages. [cited by applicant]
JCTVC-AC1005-v2, “HEVC Additional Supplemental Enhancement Information (Draft 4),” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 29th Meeting: Macao, CN, Oct. 19-25… [cited by applicant]