IP Library › Granted Patent US 12,284,344
Granted Patent B2
US 12,284,344 · App. 16/771,105 · Granted Apr 22, 2025

Deep learning based image partitioning for video compression

Inventors: Fabrice Leleannec (Cesson-Sevigne, FR); Franck Galpin (Cesson-Sevigne, FR); Sunil Jaiswal (Rennes, FR); Fabien Racape (Palo Alto, CA)
Assignee: InterDigital VC Holdings, Inc.
H04N19/119G06N3/08H04N19/105H04N19/176H04N19/96
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,284,344
App. No.
16/771,105
Granted
Apr 22, 2025
Kind
B2
Abstract

A block of video data is split using one or more of several possible partition operations by using the partitioning choices obtained through use of a deep learning-based image partitioning. In at least one embodiment, the block is split in one or more splitting operations using a convolutional neural network. In another embodiment, inputs to the convolutional neural network come from pixels along the block's causal borders. In another embodiment, boundary information, such as the location of partitions in spatially neighboring blocks, is used by the convolutional neural network. Methods, apparatus, and signal embodiments are provided for encoding.

Claims (32)

1. A method for video processing, comprising:

determining, using a convolutional neural network, a vector of split possibilities based on a block of image data and a plurality of pixels adjacent to the block, wherein the plurality of pixels comprises a causal pixel, wherein the causal pixel is adjacent to the block of image data, wherein the causal pixel is a pixel of a causal border of the block of image data, wherein the causal border comprises a row of pixels on top of the block of image data and a column of pixels to the left of the block of image data, wherein the convolutional neural network comprises a convolutional layer and a fully connected layer, wherein an output of the convolutional layer comprises a vector associated with a level of splitting, wherein the output of the convolutional layer is concatenated with quantization information associated with the block of image data, wherein the output of the convolution layer, concatenated with the quantization information, is an input of the fully connected layer, and wherein an output of the fully connected layer, based on a dimension reduction, comprises the vector of split possibilities;

partitioning the block of image data into a plurality of smaller blocks based on the vector of split possibilities; and

predicting the plurality of smaller blocks.

2. A non-transitory computer readable medium containing program instructions that, when executed by a processor, cause the processor to perform a method comprising:

determining, using a convolutional neural network, a vector of split possibilities based on a block of image data and a plurality of pixels adjacent to the block, wherein the plurality of pixels comprises a causal pixel, wherein the causal pixel is adjacent to the block of image data, wherein the causal pixel is a pixel of a causal border of the block of image data, wherein the causal border comprises a row of pixels on top of the block of image data and a column of pixels to the left of the block of image data, wherein the convolutional neural network comprises a convolutional layer and a fully connected layer, wherein an output of the convolutional layer comprises a vector associated with a level of splitting, wherein the output of the convolutional layer is concatenated with quantization information associated with the block of image data, wherein the output of the convolution layer, concatenated with the quantization information, is an input of the fully connected layer, and wherein an output of the fully connected layer, based on a dimension reduction, comprises the vector of split possibilities;

partitioning the block of image data into a plurality of smaller blocks based on the vector of split possibilities; and

predicting the plurality of smaller blocks.

3. A non-transitory computer program product comprising program instructions which, when executed by a computer, cause the computer to decode or encode data generated according to the method of claim 1 .

4. An apparatus for video processing, comprising:

a memory, and

at least one processor, configured to:

determine, using a convolutional neural network, a vector of split possibilities based on a block of image data and a plurality of pixels adjacent to the block, wherein the plurality of pixels comprises a causal pixel, wherein the causal pixel is adjacent to the block of image data, wherein the causal pixel is a pixel of a causal border of the block of image data, wherein the causal border comprises a row of pixels on top of the block of image data and a column of pixels to the left of the block of image data, wherein the convolutional neural network comprises a convolutional layer and a fully connected layer, wherein an output of the convolutional layer comprises a vector associated with a level of splitting, wherein the output of the convolutional layer is concatenated with quantization information associated with the block of image data, wherein the output of the convolutional layer, concatenated with the quantization information, is an input of the fully connected layer, and wherein an output of the fully connected layer, based on a dimension reduction, comprises the vector of split possibilities;

partition the block of image data into a plurality of smaller blocks based on the vector of split possibilities; and

predict the plurality of smaller blocks.

5. A non-transitory computer program product comprising program instructions which, when executed by a processor, cause the processor to decode or encode data generated by the apparatus of claim 4 .

6. The method of claim 1 , wherein the quantization information comprises a quantization parameter of the block of image data.

7. The apparatus of claim 4 , wherein the quantization information comprises a quantization parameter of the block of image data.

8. The method of claim 1 , wherein the causal border of the block of image data is associated with a reconstructed image.

9. The method of claim 1 , wherein information relating to a spatial neighboring block of the block of image data is concatenated to an input and further input to said convolutional neural network.

10. The apparatus of claim 4 , wherein the block of image data is associated with an image, and the causal border of the block of image data is associated with the image.

11. The apparatus of claim 4 , wherein the causal border of the block of image data is associated with a reconstructed image.

12. The apparatus of claim 4 , wherein information relating to a spatial neighboring block of the block of image data is concatenated to an input and further input to the convolutional neural network.

13. The apparatus of claim 4 , wherein the determination of the vector of split possibilities using the convolutional neural network is based on location information of a partition boundary associated with a neighboring block, and at least one vector representing the location information of the partition boundary associated with the neighboring block is reduced in dimension for use in the convolutional neural network, and wherein the neighboring block is located above or to the left of the block of image data.

14. The apparatus of claim 13 , wherein the location information of the partition boundary associated with the neighboring block is contained in a vector.

15. The apparatus of claim 4 , wherein the at least one processor is further configured to:

generate a residual based on the predicted plurality of smaller blocks; and

include the residual in a bitstream.

16. The apparatus of claim 4 , wherein the vector of split possibilities is a single-component vector.

17. The apparatus of claim 4 , wherein the block of image data has a size of 64×64 pixels.

18. The method of claim 1 , wherein the vector of split possibilities is a single-component vector.

19. The method of claim 1 , wherein the block of image data has a size of 64×64 pixels.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2020
From: LELEANNEC, FABRICE; GALPIN, FRANCK; JAISWAL, SUNIL; RACAPE, FABIEN
To: INTERDIGITAL VC HOLDINGS, INC.
Reel/Frame 052933/0501 →
Priority Claims (3)
EP 17306773 · Dec 14, 2017 · regional
EP 18305069 · Jan 26, 2018 · regional
EP 18305070 · Jan 26, 2018 · regional
Continuity (1)
Related Publication 20200344474A1 · Oct 29, 2020
References Cited (64)
US 8515193B1 · Han · 2013 [cited by examiner]
US 8665951B2 · Ameres et al. · 2014 [cited by applicant]
US 9706219B2 · Henry et al. · 2017 [cited by applicant]
US 9883203B2 · Chien et al. · 2018 [cited by applicant]
US 10057586B2 · Gu et al. · 2018 [cited by applicant]
US 10708589B2 · Mishurovskiy et al. · 2020 [cited by applicant]
US 20040165765A1 · Sung et al. · 2004 [cited by applicant]
US 20150189269A1 · Han et al. · 2015 [cited by applicant]
US 20160065959A1 · Stobaugh · 2016 [cited by examiner]
US 20160065964A1 · Zhang et al. · 2016 [cited by applicant]
US 20160134857A1 · An et al. · 2016 [cited by applicant]
US 20160174902A1 · Georgescu et al. · 2016 [cited by applicant]
US 20170118486A1 · Rusanovskyy · 2017 [cited by applicant]
US 20170214913A1 · Zhang et al. · 2017 [cited by applicant]
US 20170280144A1 · Dvir · 2017 [cited by examiner]
US 20180089493A1 · Nirenberg · 2018 [cited by examiner]
US 20180227585A1 · Wang · 2018 [cited by examiner]
US 20190007686A1 · Galpin et al. · 2019 [cited by applicant]
US 20190166380A1 · Chen · 2019 [cited by examiner]
US 20190230354A1 · Kim · 2019 [cited by examiner]
US 20200145661A1 · Jeon · 2020 [cited by examiner]
US 20210150767A1 · Ikai · 2021 [cited by examiner]
CN 100479527 · 2009 [cited by applicant]
CN 101789124A · 2010 [cited by applicant]
CN 103999465A · 2014 [cited by applicant]
CN 104041039 · 2014 [cited by applicant]
CN 106162167A · 2016 [cited by applicant]
CN 106464855A · 2017 [cited by applicant]
CN 106651886A · 2017 [cited by applicant]
CN 107197260A · 2017 [cited by applicant]
FR 3035760A1 · 2016 [cited by applicant]
JP 2017155903 · 2017 [cited by examiner]
KR 1020170059040A · 2017 [cited by applicant]
TW 201724853 · 2017 [cited by applicant]
WO 2010039731A2 · 2010 [cited by applicant]
WO 2015006884A1 · 2015 [cited by applicant]
WO 2015142070A1 · 2015 [cited by applicant]
WO 2015190839A1 · 2015 [cited by applicant]
WO 2015200820A1 · 2015 [cited by applicant]
WO WO2019009449A1 · 2019 [cited by examiner]
IEEE Transactions on Image Processing, vol. 25, No. 11, Nov. 2016 CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network Zhenyu Liu, Member, IEEE, Xianyu Yu, Yuan Gao, Shaolin Chen,… [cited by examiner]
Liu et al., CU Partition Mode Decision for HEVC Hardwired Intra Encoder Using Convolution Neural Network, IEEE Transactions on Image Processing IEEE Service Center, vol. 25, No. 11, Nov. 1, 2016, pp. 5088-5103. [cited by applicant]
Li et al., A Deep Convolutional Neural Network Approach for Complexity Reduction on Intra-Mode HEVC, 2017 IEEE International Conference on Multimedia and Expo (ICME), IEEE, Jul. 10, 2017, pp. 1255-1260. [cited by applicant]
Ruiz et al., Fast CU Partitioning Algorithm for HEVC Intra Coding Using Data Mining, Multimedia Tools and Applications, vol. 76, No. 1, 861-94, (2017). [cited by applicant]
Jin et al., CNN Oriented Fast QTBT Partition Algorithm for JVET Intra Coding, 2017 IEEE Visual Communications and Image Processing (VCIP), IEEE, Dec. 10, 2017, pp. 1-4. [cited by applicant]
Laude et al. Deep Learning-Based Intra Prediction Mode Decision for HEVC, 2016 Picture Coding Symposium (PCS), IEEE, Dec. 4, 2016, pp. 1-5. [cited by applicant]
Galpin et al., AHG9: CNN-Based Driving of Block Partitioning for Intra Slices Encoding, 10. JVET Meeting, Oct. 4, 2018-Apr. 20, 2018, San Diego, (The Joint Video Exploration Team of ISO/IEC JTC1/SC29/WG11 and ITU-T SG.1… [cited by applicant]
Jie et al., Content Based Hierarchical Fast Coding Unit Decision Algorith for HEVC, Multimedia and Signal Processing (CMSP), 2011 International Conference on, IEEE, May 14, 2011, pp. 56-59. [cited by applicant]
Xiaolin et al., CU Splitting Eady Termination Based on Weighted SVM, Eurasip Journal on Image and Video Processing, vol. 2013, No. 1, Jan. 1, 2013, p. 4. [cited by applicant]
Cassa et al.. Fast Rate Distortion Optimization for the Emerging HEVC Standard, 2012 Picture Coding Symposium, May 7-9, 2012, Krakow, Poland. [cited by applicant]
Sun et al., Efficient coding unit partition strategy for HEVC intracoding, Journal of Electronic Imaging, vol. 26, No. 4, pp. 043023-1-043023-8, (Jul./Aug. 2017). [cited by applicant]
ITU-T, “Advanced Video Coding for Generic Audiovisual Services”, Recommendation H.264, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Feb. 2014, 790 pages. [cited by applicant]
ITU-T, “High Efficiency Video Coding”, Recommendation H.265, Series H: Audiovisual and Multimedia Systems, Infrastructure of Audiovisual Services—Coding of Moving Video, Oct. 2014, 540 pages. [cited by applicant]
CN 106162167 (A) Cited in Office Action dated Oct. 11, 2023, in related Chinese Patent application No. 201880080429.2. [cited by applicant]
CN 107197260 (A) Cited in Office Action dated Oct. 11, 2023, in related Chinese Patent application No. 201880080429.2. [cited by applicant]
Chen et al., “Algorithm description of Joint Exploration Test Model 7 (JEM7)”, JVET-G1001, Aug. 19, 2017. [cited by applicant]
Leannec et al., “Asymmetric Coding Units in QTBT”, JVET-D0064, Technicolor, Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 15-21, 2016, 10 pages. [cited by applicant]
Li et al., “Multi-Type-Tree”, JVET-D0117, Qualcomm Inc., Joint Video Exploration Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, 4th Meeting: Chengdu, CN, Oct. 15-21, 2016, 3 pages. [cited by applicant]
Shen et al., “CU Splitting Early Termination Based on Weighted SVM”, Eurasip Journal on Image and Video Processing, vol. 2013, No. 4, 2013, 11 pages. [cited by applicant]
CN 101789124 (A), Cited in notice of allowance dated Aug. 1, 2024 in related Chinese Patent Application No. 201880080802.4. [cited by applicant]
CN 106651886 (A), Cited in Office Action dated May 15, 2024 in related Chinese Patent Application No. 201880080429.2. [cited by applicant]
FR 3035760 (A1), Cited in Office Action dated May 15, 2024 in related Chinese Patent Application No. 201880080429.2. [cited by applicant]
KR 10-2017-0059040 (A), Cited in notice of allowance dated Aug. 1, 2024 in related. [cited by applicant]
WO 2015/142070 (A1), U.S. Pat. No. 10,708,589 (B2). [cited by applicant]