IP Library › Granted Patent US 12,395,671
Granted Patent B2
US 12,395,671 · App. 18/486,641 · Granted Aug 19, 2025

Interaction between Intra Block Copy mode and inter prediction tools

Inventors: Kai Zhang (San Diego, CA); Li Zhang (San Diego, CA); Hongbin Liu (Beijing, CN); Yue Wang (Beijing, CN)
Assignees: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.; BYTEDANCE INC.
H04N19/52H04N19/159H04N19/176H04N19/184H04N19/51H04N19/593
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,395,671
App. No.
18/486,641
Granted
Aug 19, 2025
Kind
B2
Abstract

The present disclosure relates to interaction between intra block copy mode and inter prediction tools. A method for video processing includes determining that an intra block copy (IBC) mode is applied to a current video block of a video. In the IBC mode, at least one reference picture used by the current video block is a current picture where the current video block is located. The method also includes making a decision regarding a disabling of a specific coding mode for the current block. The specific coding mode uses a motion vector and a non-current picture to derive a prediction of a video block. The method further includes performing, based on the decision, a conversion between the current video block and the bitstream representation.

Claims (59)

1. A method of processing video data, comprising:

determining that an intra block copy (IBC) mode is applied to a first video block of a video, wherein in the IBC mode, reference samples from a video region including the first video block are used; and

performing, based on the determining, a first conversion between the first video block and a bitstream of the video,

wherein the first video block is a block with a dual coding tree in which a luma component and chroma components are coded with separate separated coding structure trees,

wherein the method further comprises:

constructing, during a second conversion between a second video block of the video and the bitstream of the video, a motion candidate list for the second video block based on an affine motion candidate,

wherein the second conversion is performed based on the motion candidate list, and

wherein the affine motion candidate is derived based on a spatial neighboring block of the second video block, and wherein, in response to the spatial neighboring block being the first video block, the spatial neighboring block is excluded from deriving the affine motion candidate.

2. The method of claim 1 , further comprising making a decision regarding a disabling of a specific coding mode for the first video block.

3. The method of claim 2 , wherein a flag for the specific coding mode is not included in the bitstream in response to the IBC mode being used in the first video block, and wherein when the flag is not included in the bitstream, the flag is inferred to be zero.

4. The method of claim 2 , wherein the specific coding mode comprises a bi-prediction with CU-level weights mode, wherein in the bi-prediction with CU-level weights mode, different weighting values relate with different reference pictures in a prediction derivation process.

5. The method of claim 4 , wherein a weighting index of the bi-prediction with CU-level weights mode is not included in the bitstream in response to the IBC mode being used in the first video block, and

wherein when the weighting index is not included in the bitstream, the weighting index is inferred to be 0.

6. The method of claim 2 , wherein the specific coding mode comprises a merge with motion vector difference (MMVD) mode, wherein in the MMVD mode, a motion vector of a video block is derived based on a merge motion candidate list and is further refined by at least one motion vector offset, or

wherein the specific coding mode comprises an affine mode and a combined inter-intra prediction mode, wherein the affine mode uses at least one control point motion vector, and wherein in the combined inter-intra prediction mode, a final prediction is generated at least based on a weighted sum of an intermediate intra prediction signal and an intermediate inter prediction signal, or

wherein the specific coding mode comprises a sub-block based temporal motion vector prediction mode, and wherein in the sub-block based temporal motion vector prediction mode, motion information is derived based on a collocated region that is determined by at least one temporal motion offset.

7. The method of claim 1 , wherein before the performing, the method further comprises:

deriving a block vector for the first video block; and

using at least one block vector difference included in the bitstream to refine the block vector.

8. The method of claim 1 , wherein a width of the first video block is greater than or equal to 2 and a height is greater than or equal to 2.

9. The method of claim 1 , wherein the motion candidate list is a subblock merge candidate list.

10. The method of claim 9 , wherein the spatial neighboring block is treated as unavailable based on determining that the spatial neighboring block uses the IBC mode, and wherein the subblock merge candidate list is constructed further based on a subblock-based temporal motion vector prediction candidate.

11. The method of claim 10 , wherein constructing the subblock merge candidate list comprises adding of the subblock-based temporal motion vector prediction candidate is disabled in case that a temporal motion vector prediction tool is disabled or a subblock-based temporal motion vector prediction tool is disabled.

12. The method of claim 10 , wherein the subblock-based temporal motion vector prediction candidate is derived by:

initializing a temporal motion information to a default motion information;

setting, in response to a specific neighboring block which is adjacent to a lower left corner of the second video block being available and a reference picture of the specific neighboring block being a collocated picture of the second video block, the temporal motion information to specific motion information related to the specific neighboring block; and

deriving the subblock-based temporal motion vector prediction candidate based on the temporal motion information,

wherein the specific neighboring block covers a luma location (xCb−1, yCb+cbHeight−1), wherein (xCb,yCb) is a luma location of a top-left sample of the second video block relative to a top-left luma sample of a current picture and cbHeight is a height of the second video block.

13. The method of claim 1 , wherein the motion candidate list is a motion vector prediction candidate list, and wherein performing the second conversion based on the motion candidate list comprises deriving a motion information for the second video block based on the affine motion candidate and at least one motion vector difference which is used to refine the motion information, and

wherein the spatial neighboring block is treated as unavailable based on determining that the spatial neighboring block uses the IBC mode.

14. The method of claim 1 , wherein the second conversion comprises generating the bitstream from the second video block, and the method further comprises:

storing the bitstream in a non-transitory computer-readable recording medium.

15. The method of claim 1 , wherein the first conversion comprises decoding the first video block from the bitstream.

16. The method of claim 1 , wherein the first conversion comprises encoding the first video block into the bitstream.

17. An apparatus for processing video data comprising a processor and a non-transitory memory with instructions thereon, wherein the instructions upon execution by the processor, cause the processor to:

determine that an intra block copy (IBC) mode is applied to a first video block of a video, wherein in the IBC mode, reference samples from a video region including the first video block are used; and

perform, based on the determination, a first conversion between the first video block and a bitstream of the video,

wherein the first video block is a block with a dual coding tree in which a luma component and chroma components are coded with separate separated coding structure trees,

wherein the instructions upon execution by the processor, further cause the processor to:

construct, during a second conversion between a second video block of the video and the bitstream of the video, a motion candidate list for the second video block based on an affine motion candidate,

wherein the second conversion is performed based on the motion candidate list, and

wherein the affine motion candidate is derived based on a spatial neighboring block of the second video block, and wherein in response to the spatial neighboring block being the first video block, the spatial neighboring block is excluded from deriving the affine motion candidate.

18. A non-transitory computer-readable storage medium storing instructions that cause a processor to:

determine that an intra block copy (IBC) mode is applied to a first video block of a video, wherein in the IBC mode, reference samples from a video region including the first video block are used; and

perform, based on the determination, a first conversion between the first video block and a bitstream of the video,

wherein the first video block is a block with a dual coding tree in which a luma component and chroma components are coded with separate separated coding structure trees,

wherein the instructions further cause the processor to:

construct, during a second conversion between a second video block of the video and the bitstream of the video, a motion candidate list for the second video block based on an affine motion candidate,

wherein the second conversion is performed based on the motion candidate list, and

wherein the affine motion candidate is derived based on a spatial neighboring block of the second video block, and wherein in response to the spatial neighboring block being the first video block, the spatial neighboring block is excluded from deriving the affine motion candidate.

19. A method for storing a bitstream of a video, comprising:

determining that an intra block copy (IBC) mode is applied to a first video block of the video, wherein in the IBC mode, reference samples from a video region including the first video block are used; and

generating the bitstream from the first video block based on the determining,

storing the bitstream in a non-transitory computer-readable recording medium;

wherein the first video block is a block with a dual coding tree in which a luma component and chroma components are coded with separate coding structure trees,

wherein the method further comprises:

constructing a motion candidate list for a second video block based on an affine motion candidate,

generating the bitstream from the second video block based on the motion candidate list, and

wherein the affine motion candidate is derived based on a spatial neighboring block of the second video block, and wherein in response to the spatial neighboring block being the first video block, the spatial neighboring block is excluded from deriving the affine motion candidate.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: ZHANG, KAI; ZHANG, LI
To: BYTEDANCE INC.
Reel/Frame 065407/0388 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 31, 2023
From: LIU, HONGBIN; WANG, YUE
To: BEIJING BYTEDANCE NETWORK TECHNOLOGY CO., LTD.
Reel/Frame 065407/0469 →
Priority Claims (1)
WO PCT/CN2018/118167 · Nov 29, 2018 · international
Continuity (4)
Continuation 17403707 · Aug 16, 2021
Continuation 17167266 · Feb 4, 2021
Continuation PCTCN2019122183 · Nov 29, 2019
Related Publication 20240056599A1 · Feb 15, 2024
References Cited (113)
US 9591325B2 · Li · 2017 [cited by applicant]
US 9883197B2 · Chen · 2018 [cited by applicant]
US 9918105B2 · Pang · 2018 [cited by applicant]
US 10178403B2 · Seregin · 2019 [cited by applicant]
US 10264290B2 · Xu · 2019 [cited by applicant]
US 10484686B2 · Xiu · 2019 [cited by applicant]
US 10638140B2 · Seregin · 2020 [cited by applicant]
US 11095917B2 · Zhang · 2021 [cited by applicant]
US 11115676B2 · Zhang · 2021 [cited by applicant]
US 20100241852A1 · Sela · 2010 [cited by applicant]
US 20150030073A1 · Chen · 2015 [cited by applicant]
US 20150085929A1 · Chen · 2015 [cited by applicant]
US 20150139296A1 · Yu · 2015 [cited by applicant]
US 20150264396A1 · Zhang · 2015 [cited by applicant]
US 20150373359A1 · He · 2015 [cited by applicant]
US 20160100189A1 · Pang · 2016 [cited by applicant]
US 20160241852A1 · Gamei · 2016 [cited by examiner]
US 20160241858A1 · Li · 2016 [cited by applicant]
US 20160241876A1 · Xu · 2016 [cited by applicant]
US 20160353117A1 · Seregin · 2016 [cited by applicant]
US 20160360234A1 · Tourapis · 2016 [cited by applicant]
US 20170034526A1 · Rapaka · 2017 [cited by examiner]
US 20170099490A1 · Seregin · 2017 [cited by applicant]
US 20170195677A1 · Ye · 2017 [cited by applicant]
US 20170280159A1 · Xu · 2017 [cited by applicant]
US 20170289566A1 · He · 2017 [cited by applicant]
US 20170332095A1 · Zou · 2017 [cited by applicant]
US 20180070105A1 · Jin · 2018 [cited by applicant]
US 20180109810A1 · Xu · 2018 [cited by applicant]
US 20180109814A1 · Chuang · 2018 [cited by applicant]
US 20180192069A1 · Chen · 2018 [cited by applicant]
US 20180270500A1 · Li · 2018 [cited by applicant]
US 20180288430A1 · Chen · 2018 [cited by applicant]
US 20180376149A1 · Zhang · 2018 [cited by examiner]
US 20190246128A1 · Xu · 2019 [cited by examiner]
US 20200036997A1 · Li · 2020 [cited by applicant]
US 20200177910A1 · Li · 2020 [cited by applicant]
US 20200195959A1 · Zhang · 2020 [cited by applicant]
US 20200396465A1 · Zhang · 2020 [cited by applicant]
US 20210160525A1 · Zhang · 2021 [cited by applicant]
US 20210160533A1 · Zhang · 2021 [cited by applicant]
CN 105392008A · 2016 [cited by applicant]
CN 105493505A · 2016 [cited by applicant]
CN 106416253A · 2017 [cited by applicant]
CN 106797476A · 2017 [cited by applicant]
CN 107079161A · 2017 [cited by applicant]
CN 107409220A · 2017 [cited by applicant]
CN 107646195A · 2018 [cited by applicant]
CN 107690809A · 2018 [cited by applicant]
CN 107852490A · 2018 [cited by applicant]
CN 108012153A · 2018 [cited by applicant]
CN 108432250A · 2018 [cited by applicant]
CN 108702509A · 2018 [cited by applicant]
CN 108886619A · 2018 [cited by applicant]
CN 113196772B · 2024 [cited by applicant]
IN 547928 · 2024 [cited by applicant]
JP 2017508345A · 2017 [cited by applicant]
JP 2017522790A · 2017 [cited by applicant]
JP 2018526881A · 2018 [cited by applicant]
KR 2016059513A · 2016 [cited by applicant]
KR 102695787B1 · 2024 [cited by applicant]
WO 2015035449A1 · 2015 [cited by applicant]
WO 2015192353A1 · 2015 [cited by applicant]
WO 2016057938A1 · 2016 [cited by applicant]
WO 2016123068A1 · 2016 [cited by applicant]
WO 2016123081A1 · 2016 [cited by applicant]
WO 2017118409A1 · 2017 [cited by applicant]
WO 2017148345A1 · 2017 [cited by applicant]
WO 2017157259A1 · 2017 [cited by applicant]
WO 2017206804A1 · 2017 [cited by applicant]
WO 2018192574A1 · 2018 [cited by applicant]
WO 2018200960A1 · 2018 [cited by applicant]
WO 2018205954A1 · 2018 [cited by applicant]
WO 2018210315A1 · 2018 [cited by applicant]
Lee et al. “CE4: Simplification of the Common Base for Affine Merge (Test 4.2.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and 1SO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macau, CN, Oct. 8-12, 2012, document JV… [cited by applicant]
Chen et al. “Crosscheck of JVET-L0142 (CE4: Simplification of the Common Base for Affine Merge (Test 4.2.6)),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, … [cited by applicant]
Huang et al. “CE4.2.5: Simplification of Affine Merge List Construction and Move ATMVP lo Affine Merge List,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, O… [cited by applicant]
Su et al. “CE4-Related: Generalized Bi-Prediction Improvements Combined from JVET-L0197 and JVET-L0296,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3… [cited by applicant]
Chen et al. “Generalized Bi-Prediction for Inter Coding,” Joint Video Exploration Team (JVET) of ITU-T Sg 16 WP 3 and ISO-IEC JTC 1/SC 29/WG 11, 3rd Meeting, Geneva, CH, 26 May to Jun. 1, 2016, document JVET-C0047, 2016. [cited by applicant]
Su et al. “CE4.4.1: Generalized Bi-Prediction for Intercoding,” Joint Video Exploration Team of ISO/IEC JTC 1/SC 29/ WG 11 and ITU-T SG 16, Ljubljana, Jul. 10-18, 2018, document No. JVET-K0248, 2018. [cited by applicant]
Su et al.“CE4-Related: Generalized Bi-Prediction Improvements,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018, documentJVET-L0197, 2018. [cited by applicant]
He et al. “CE4-Related: Encoder Speed-Up and Bug Fix for Generalized Bi-Prediction in BMS-2.1,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018… [cited by applicant]
Chiang et al. “CE10.1.1: Multi-Hypothesis Prediction for Improving AMVP Mode, Skip or Merge Mode, and Intra Mode,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, … [cited by applicant]
“High Efficiency Video Coding” Series H: Audiovisual and Multimedia Systems: Infrastructure of Audiovisual Services—Coding of Moving Video, ITU-T, H.265, 2018. [cited by applicant]
Rosewarne et al. “High Efficiency Video Coding (HEVC) Test Model 16 (HM 16) Improved Encoder Description Update 7,” Joint Collaborative Team on Video Coding (JCT-VG) ITU-T SG 16 WP3 and 1SO/IEC JTC1/SC29/WG11, 25th Meet… [cited by applicant]
Chen et al. “Algorithm Description of Joint Exploration Test Model 7 (JEM 7),” Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 7th Meeting: Torino, IT, Jul. 13-21, 2017, document J… [cited by applicant]
JEM-7.0, Retrieved from the internet: https://jvet.hhi.fraunhofer.de/svn/svn_HMJEMSoftware/tags/ HM-16.6-JEM-7.0, Aug. 16, 2021. [cited by applicant]
Chen et al. “CE4: Affine Merge Enhancement with Simplification (Test 4.2.2),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting: Macao, CN, Oct. 3-12, 2018, document JVET L0… [cited by applicant]
Li et al. “Non-SCCE1: Unification of Intra BC and Inter Modes,” Joint Collaborative Team on Video Coding (JCT-VG) of ITU-T SG 16 WP 3 and 1SO/IEC JTC 1/SC 29/WG 11 18th Meeting, Sapporo, JP, Jun. 30-Jul. 9, 2014, docume… [cited by applicant]
Chien et al. “Methodology and Reporting Templace for Tool Testing,” Joint Video Exploration Team {JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 12th Meeting, Macao, CN, Oct. 3-12, 2018, document JVET-L1005, 2… [cited by applicant]
Bross et al. “Versatile Video Coding (Draft 3),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting: Macao, CN, Oct. 3-12, 2018, document JVET-L1001, 2018. [cited by applicant]
Lin et al. “CE4.2.3: Affine Merge Mode,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 12th Meeting: Macao, CN, Oct. 3-12, 2018, document JVET-L0088, 2018. [cited by applicant]
Sullivan et al. “Meeting Report of the 21st Meeting of the Joint Collaborative Team on Video Coding (JCT-VC), Warsaw, PL, Jun. 19-26, 2015,” Joint Collaborative Team on Video Coding (JCT-VC) of ITU-T SG16 WP3 and ISO/IE… [cited by applicant]
Document: JVET-J0029-v1, Li, X., et al., “Description of SDR video coding technology proposal by Tencent,” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 10th Meeting: San Diego, US, A… [cited by applicant]
Document: JVET-P2002-v1, Chen, J., et al., “Algorithm description for Versatile Video Coding and Test Model 7 (VTM 7),” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 16th Meeting: Gen… [cited by applicant]
Document: JVET-Q0095, Li, G., et al., “CE5-2.2: Multiplication removal for CCALF with coefficient range in [-8, 8],” Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 17th Meeting: Brusse… [cited by applicant]
Extended European Search Report from European Patent Application No. 19888914.9 dated Nov. 19, 2021 (11 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/122183 dated Feb. 28, 2020 (13 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/122194 dated Mar. 6, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/122195 dated Mar. 2, 2020 (10 pages). [cited by applicant]
International Search Report and Written Opinion from International Patent Application No. PCT/CN2019/122198 dated Feb. 28, 2020 (11 pages). [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/167,303 dated Apr. 1, 2021. [cited by applicant]
Non-Final Office Action from U.S. Appl. No. 17/167,266 dated Mar. 23, 2021. [cited by applicant]
Office Action from Japanese Patent Application No. 2021-528964 dated Dec. 13, 2022. [cited by applicant]
Document: JVET-L0198, Wang, S., et al., “Simplification of ATMVP Candidate Derivation,” Institute of Digital Media, Peking University, Oct. 5, 2018, 12 pages. [cited by applicant]
Document: JCTVC-S0172, He, Y., et al., “Non-CE2: Unification of IntraBC mode with inter mode,” Joint Video Experts Team on Video Coding (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 19th Meeting: Strasbourg, … [cited by applicant]
Document: JVET-K1002-v2, Chen, J., et al., “Algorithn description for Versatile Video Coding and Test Model 2 (VTM 2),” Joint Video Experts Team on Video Coding (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11 1… [cited by applicant]
Extended European Search Report from European Application No. 23213886.7 dated Mar. 6, 2024, 12 pages. [cited by applicant]
Chinese Notice of Allowance from Chinese Patent Application No. 201980077207.X dated Feb. 24, 2024, 14 pages. With English Translation. [cited by applicant]
Document: JCTVC-C403-v1, Wiegand, T., et al., “WDI: Working Draft 1 of High-Efficiency Video Coding,” Joint Collaborative Team on Video Coding (.TCT-VC) of' ITU-T SG16 WP3 and ISO/IEC JTC1/SC29/WG11 3rd Meeting: Guangzh… [cited by applicant]
Japanese Office Action from Japanese Patent Application No. 2023-002737 dated Jul. 2, 2024, 14 pages. [cited by applicant]
Chinese Notice of Allowance from Chinese Patent Application No. 201980077195.0 dated Sep. 16, 2024, 5 pages. [cited by applicant]
Chinese Office Action from Chinese Patent Application No. 201980077195.0 dated Jun. 28, 2024, 16 pages. [cited by applicant]