IP Library › Granted Patent US 12,368,875
Granted Patent B2
US 12,368,875 · App. 18/472,353 · Granted Jul 22, 2025

Two-step cross-component prediction mode

Inventors: Li Zhang (San Diego, CA); Kai Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Jizheng Xu (San Diego, CA); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/50H04N19/117H04N19/157H04N19/176H04N19/184H04N19/186H04N19/593
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,368,875
App. No.
18/472,353
Granted
Jul 22, 2025
Kind
B2
Abstract

A method for video bitstream processing includes generating, using a first video block derived from a third video block of a first component and having a first size, a prediction block for a second video block of a video related to a second component, where the first component is different from the second component, and where the second video block has a second size that is different from the first size. The method also includes performing, using the prediction block, a conversion between the second video block and a bitstream representation of the video according to a two-step cross-component prediction mode (TSCPM).

Claims (35)

1. A method of video processing, comprising:

deriving a prediction block for a chroma block of a current video block of a video based on a temporary video block, wherein the temporary video block is derived from a reconstructed luma block of the current video block, both the reconstructed luma block and the temporary video block have a first size, and the chroma block has a second size which is different from the first size; and

performing, a conversion between the current video block and a bitstream of the video based on the prediction block for the chroma block according to a cross-component prediction mode,

wherein the prediction block is generated by applying one or more downsampling filters to samples in the temporary video block, and wherein the one or more downsampling filters are pre-defined, and

wherein a prediction sample of the chroma block only depends on samples located at (2*x, 2*y) and (2*x, 2*y+1), and wherein the prediction sample is located at (x, y), x and y are integers and a top-left sample's coordinates of the temporary video block is set to (0, 0).

2. The method of claim 1 , wherein the first size is (M′+W0)×(N′+H0), the second size is M×N, and wherein either M′ is unequal to M or N′ is unequal to N.

3. The method of claim 1 , wherein the first size is (M′+W0)×(N′+H0), wherein the second size is M×N, and wherein one or both of W0 and H0 are equal to a value of zero.

4. The method of claim 1 , wherein the temporary video block is associated with a first set of samples and the reconstructed luma block is associated with a second set of samples, wherein at least one sample S Temp (x0, y0) in the first set of samples is derived as a first function of a corresponding sample S (x0,y0) from the second set of samples,

wherein x0 is within a first range of zero to (M′−1), inclusive, and y0 is within a second range of zero to (N′−1), inclusive, and

wherein M′ represents a width of the first size, and N′ represents a height of the second size.

5. The method of claim 4 , wherein in the first function, the S Temp (x0, y0) is defined as ((S(x0, y0)*a)>>k)+b.

6. The method of claim 4 , the method further comprising clipping the S Temp (x0, y0) to an allowed range of chroma values.

7. The method of claim 1 , wherein a selection of the one or more downsampling filters is based on a relative position of a sample to be predicted.

8. The method of claim 1 , wherein the one or more downsampling filters includes a downsampling filter with coefficient [1 1] for a sample located at (0, 0) relative to the temporary video block.

9. The method of claim 1 , wherein the one or more downsampling filters includes a 6-tap filter for a sample not located at (0, 0) relative to the temporary video block.

10. The method of claim 9 , wherein the 6-tap filter includes coefficients [1 2 1; 1 2 1].

11. The method of claim 1 , wherein the conversion includes encoding the current video block into the bitstream.

12. The method of claim 1 , wherein the conversion includes decoding the current video block from the bitstream.

13. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

derive a prediction block for a chroma block of a current video block of a video based on a temporary video block, wherein the temporary video block is derived from a reconstructed luma block of the current video block, both the reconstructed luma block and the temporary video block have a first size, and the chroma block has a second size which is different from the first size; and

perform, a conversion between the current video block and a bitstream of the video based on the prediction block for the chroma block according to a cross-component prediction mode,

wherein the prediction block is generated by applying one or more downsampling filters to samples in the temporary video block, wherein the one or more downsampling filters are pre-defined, and

wherein a prediction sample of the chroma block only depends on samples located at (2*x, 2*y) and (2*x, 2*y+1), wherein the prediction sample is located at (x, y), x and y are integers and a top-left sample's coordinates of the temporary video block is set to (0, 0).

14. The apparatus of claim 13 , wherein the first size is (M′+W0)×(N′+H0), the second size is M×N, and wherein either M′ is unequal to M or N′ is unequal to N.

15. The apparatus of claim 13 , wherein the first size is (M′+W0)×(N′+H0), wherein the second size is M×N, and wherein one or both of W0 and H0 are equal to a value of zero.

16. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

derive a prediction block for a chroma block of a current video block of a video based on a temporary video block, wherein the temporary video block is derived from a reconstructed luma block of the current video block, both the reconstructed luma block and the temporary video block have a first size, and the chroma block has a second size which is different from the first size; and

perform, a conversion between the current video block and a bitstream of the video based on the prediction block for the chroma block according to a cross-component prediction mode,

wherein the prediction block is generated by applying one or more downsampling filters to samples in the temporary video block, and wherein the one or more downsampling filters are pre-defined, and

wherein a prediction sample of the chroma block only depends on samples located at (2*x, 2*y) and (2*x, 2*y+1), and wherein the prediction sample is located at (x, y), x and y are integers and a top-left sample's coordinates of the temporary video block is set to (0, 0).

17. A method for storing a bitstream of a video, comprising:

deriving a prediction block for a chroma block of a current video block of the video based on a temporary video block, wherein the temporary video block is derived from a reconstructed luma block of the current video block, both the reconstructed luma block and the temporary video block have a first size, and the chroma block has a second size which is different from the first size; generating the bitstream based on the prediction block for the chroma block according to a cross-component prediction mode; and

storing the bitstream in a non-transitory computer-readable recording medium,

wherein the prediction block is generated by applying one or more downsampling filters to samples in the temporary video block, and wherein the one or more downsampling filters are pre-defined, and

wherein a prediction sample of the chroma block only depends on samples located at (2*x, 2*y) and (2*x, 2*y+1), and wherein the prediction sample is located at (x, y), x and y are integers and a top-left sample's coordinates of the temporary video block is set to (0, 0).

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065093/0817 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 2, 2023
From: ZHANG, LI; ZHANG, KAI; XU, JIZHENG
To: BYTEDANCE INC.
Reel/Frame 065093/0951 →
Priority Claims (2)
WO PCT/CN2018/122955 · Dec 22, 2018 · international
WO PCT/CN2018/123394 · Dec 25, 2018 · international
Continuity (3)
Continuation 17353629 · Jun 21, 2021
Continuation PCTCN2019127377 · Dec 23, 2019
Related Publication 20240022754A1 · Jan 18, 2024
References Cited (116)
US 9294769B2 · Jeon · 2016 [cited by applicant]
US 9866842B2 · Kim · 2018 [cited by applicant]
US 10200700B2 · Zhang · 2019 [cited by applicant]
US 10419757B2 · Chen · 2019 [cited by applicant]
US 10469847B2 · Xiu · 2019 [cited by applicant]
US 10939128B2 · Zhang · 2021 [cited by applicant]
US 10979717B2 · Zhang · 2021 [cited by applicant]
US 11190790B2 · Jang · 2021 [cited by applicant]
US 20130136375A1 · Sasai · 2013 [cited by applicant]
US 20130148722A1 · Zhang · 2013 [cited by applicant]
US 20130223522A1 · Song · 2013 [cited by applicant]
US 20140098189A1 · Deng · 2014 [cited by examiner]
US 20140098869A1 · Su · 2014 [cited by applicant]
US 20140219336A1 · Jeon · 2014 [cited by applicant]
US 20140314142A1 · Oh · 2014 [cited by applicant]
US 20140328397A1 · Jeon · 2014 [cited by applicant]
US 20140348236A1 · Kim · 2014 [cited by applicant]
US 20150124876A1 · Lee · 2015 [cited by applicant]
US 20150264345A1 · Cohen · 2015 [cited by applicant]
US 20150334405A1 · Rosewarne et al. · 2015 [cited by applicant]
US 20150382016A1 · Cohen · 2015 [cited by applicant]
US 20160219283A1 · Chen · 2016 [cited by applicant]
US 20160373741A1 · Zhao · 2016 [cited by applicant]
US 20160373742A1 · Zhao · 2016 [cited by applicant]
US 20160373785A1 · Said · 2016 [cited by applicant]
US 20170150156A1 · Zhang · 2017 [cited by examiner]
US 20170150186A1 · Zhang · 2017 [cited by examiner]
US 20170214912A1 · Cote · 2017 [cited by applicant]
US 20170244975A1 · Huang · 2017 [cited by applicant]
US 20170324643A1 · Seregin · 2017 [cited by applicant]
US 20170347102A1 · Panusopone et al. · 2017 [cited by applicant]
US 20170347103A1 · Yu · 2017 [cited by applicant]
US 20180063527A1 · Chen · 2018 [cited by examiner]
US 20180077426A1 · Zhang · 2018 [cited by examiner]
US 20180084258A1 · Kim · 2018 [cited by applicant]
US 20180184083A1 · Panusopone · 2018 [cited by applicant]
US 20180205946A1 · Zhang · 2018 [cited by examiner]
US 20190158851A1 · Kim · 2019 [cited by applicant]
US 20200128272A1 · Jangwon · 2020 [cited by examiner]
US 20200177878A1 · Choi · 2020 [cited by examiner]
US 20200195976A1 · Zhao · 2020 [cited by examiner]
US 20200252619A1 · Zhang · 2020 [cited by applicant]
US 20200288135A1 · Laroche · 2020 [cited by examiner]
US 20200359051A1 · Zhang · 2020 [cited by applicant]
US 20200366896A1 · Zhang · 2020 [cited by applicant]
US 20200366910A1 · Zhang · 2020 [cited by applicant]
US 20200366933A1 · Zhang · 2020 [cited by applicant]
US 20200382769A1 · Zhang · 2020 [cited by applicant]
US 20200382800A1 · Zhang · 2020 [cited by applicant]
US 20200413062A1 · Onno · 2020 [cited by examiner]
US 20210076028A1 · Heo · 2021 [cited by applicant]
US 20210092395A1 · Zhang · 2021 [cited by applicant]
US 20210092396A1 · Zhang · 2021 [cited by applicant]
US 20210243457A1 · Ahn · 2021 [cited by examiner]
US 20210321131A1 · Zhang · 2021 [cited by applicant]
US 20220007012A1 · Srinivasan · 2022 [cited by applicant]
US 20220038683A1 · Choi · 2022 [cited by applicant]
AR 092495A1 · 2015 [cited by applicant]
CN 103220508A · 2013 [cited by applicant]
CN 103959782A · 2014 [cited by applicant]
CN 105247866A · 2016 [cited by applicant]
CN 105493505A · 2016 [cited by applicant]
CN 105594213A · 2016 [cited by applicant]
CN 106464885A · 2017 [cited by applicant]
CN 107079157A · 2017 [cited by applicant]
CN 107079166A · 2017 [cited by applicant]
CN 107211151A · 2017 [cited by applicant]
CN 107646195A · 2018 [cited by applicant]
CN 107707920A · 2018 [cited by applicant]
CN 107736022A · 2018 [cited by applicant]
CN 108605135A · 2018 [cited by applicant]
CN 109005408A · 2018 [cited by applicant]
CN 109076230A · 2018 [cited by applicant]
CN 109076237A · 2018 [cited by applicant]
CN 113273203B · 2021 [cited by applicant]
CN 113287311B · 2021 [cited by applicant]
CN 113261291B · 2024 [cited by applicant]
EP 3203746A1 · 2017 [cited by applicant]
JP 2017538381A · 2017 [cited by applicant]
KR 20130050902A · 2013 [cited by applicant]
KR 20150070848A · 2015 [cited by applicant]
TW 201811055A · 2018 [cited by applicant]
TW 201817236A · 2018 [cited by applicant]
TW 201818720A · 2018 [cited by applicant]
WO 2013029560A1 · 2013 [cited by applicant]
WO 2013069972A1 · 2013 [cited by applicant]
WO 2015196119A1 · 2015 [cited by applicant]
WO 2016065538A1 · 2016 [cited by applicant]
WO 2016205718A1 · 2016 [cited by applicant]
WO 2017059926A1 · 2017 [cited by applicant]
WO 2017209328A1 · 2017 [cited by applicant]
WO 2018053293A1 · 2018 [cited by applicant]
WO 2018132710A1 · 2018 [cited by applicant]
WO 2018174617A1 · 2018 [cited by applicant]
WO 2018191224A1 · 2018 [cited by applicant]
Document: JVET-L0191, Laroche, G., et al., “CE3-5.1: On cross-component linear model simplification,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting: Macao, CN, Oct. 3-12,… [cited by applicant]
Chinese Notice of Allowance from Chinese Patent Application No. 201980085389.5 dated May 10, 2024, 6 pages. [cited by applicant]
“High Efficiency Video Coding,” Series H: Audiovisual and Multimedia Systems: Infrastructure of Audiovisual Services Coding of Moving Video, ITU-T Telecommunication Standardization Sector of ITU, H.265, Feb. 2018. [cited by applicant]
Rosewarne et al. “High Efficiency Video Coding (HEVC) Test Model 16 (HM 16) Improved Encoder Description Update 7,” Joint Collaborative Team on Video Coding (JCT-VG) ITU-T SG 16 WP3 and ISO/IEC JTC1/SC29/WG11, 25th Meet… [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 7 (JEM 7),” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, document J… [cited by applicant]
JEM-7.0: https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/ HM-16.6-JEM-7.0, Sep. 9, 2021. [cited by applicant]
Bross et al. ““Versatile Video Coding (Draft 3),”” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting, Macao, CN, Oct. 3-12, 2018, document JVET-L1001, Oct. 2018. http:1/pheni… [cited by applicant]
https://vcgit.hhi.fraunhofer.de/jvet/VVCSoftware_VTM/tags/VTM-3.0, Sep. 9, 2021. [cited by applicant]
Ma et al. “CE3: Multi-directional LM (MDLM) (Test 5.4.1 and 5.4.2)” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and 1SO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, document JVET-L0338, 2018. [cited by applicant]
Laroche et al. “CE3-5.1: On Cross-Component Linear Model Simplification,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WVG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018, document JVET L0191… [cited by applicant]
Liao et al. “CE10.3.1.b: Triangular Prediction Unit Mode,” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018, document JVET-L0124, 2018. [cited by applicant]
Chen et al.CE4: Affine Merge Enhancement with Simplification (Test 4.2.2), Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, 3-12, Oct. 2018, document JVET-L0368… [cited by applicant]
Information Technology—High Efficiency Media Coding—Part 2 Video, GB/T 33475.2, Dec. 30, 2016. (with English equivalent). [cited by applicant]
Document: JVET-L1023-v3, Van Der Auwera, G., et al., “Description of Core Experiment 3 (CE3): Intra Prediction and Mode Coding,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Mee… [cited by applicant]
Zhang, T., “Research on High Efficiency Intra Coding in Video Compression” Dissertation for the Doctoral Degree in Engineering, School of Computer Science and Technology, Jan. 1, 2017, total 158 pages. [cited by applicant]
Notice of Allowance from U.S. Appl. No. 17/353,661 dated Oct. 7, 2022. [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/127377 dated Mar. 25, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/127379 dated Mar. 23, 2020 (9 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/127380 dated Mar. 25, 2020 (11 pages). [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/353,629 dated Feb. 14, 2023, 20 pages. [cited by applicant]
Zhang, T., et al., “Research on High Efficiency Intra Coding in Video Compression,” Dyssertation for the Doctoral Degree in Engineering, 2017, the whole passage, 158 pages. [cited by applicant]