IP Library Granted Patent US 12,335,539
Granted Patent B2
US 12,335,539 · App. 17/800,002 · Granted Jun 17, 2025

Neural network-based intra prediction for video encoding or decoding

Inventors: Thierry Dumas (Rennes, FR); Franck Galpin (Thorigne-Fouillard, FR); Philippe Bordes (Laille, FR); Fabrice Leleannec (Betton, FR)
Assignee: InterDigital Madison Patent Holdings, SAS
H04N19/88H04N19/159H04N19/176H04N19/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,335,539
App. No.
17/800,002
Granted
Jun 17, 2025
Kind
B2
Abstract

A video coding system is provided that performs intra prediction in a mode using a neural network for block of only a set of specific block sizes. The signaling of this mode is designed to be efficient in terms of rate-distortion under this constraint. Different transformations of the context of a block and the neural network prediction of this block are introduced in order to use one single neural network for predicting blocks of several sizes, as well as the corresponding signaling. The neural network-based prediction mode considers both luminance blocks and chrominance blocks. The video coding system comprises encoder and decoder apparatuses, encoding, decoding and signal generation methods and a signal carrying information corresponding to the described coding mode.

Claims (222)

1. A video encoding method comprising:

performing intra prediction for at least one block in a picture or video using a neural network based intra prediction by feeding a block context into a neural network selected based on a size of the at least one block, the block context comprising pixels of blocks located at a top side, at a left side, at a diagonal top left side, at a diagonal top right side and at a diagonal bottom left side of the at least one block, wherein the size of the block context is based on the size of the at least one block, and wherein the block context comprises n l columns and n a rows, wherein n l and n a are selected as:

n

a

=

α

H

,

n

l

=

β

W

,

α

1

4

,

1

2

,

3

4

,

1

,

2

,

β

1

4

,

1

2

,

3

4

,

1

,

2

where H is a height of the block and W is a width of the block;

generating signaling information representative that the intra prediction mode is a neural network based intra prediction; and

encoding at least information representative of the at least one block and the neural network-based intra prediction mode.

2. The method of claim 1 , wherein the signaling information is encoded in a bitstream and comprises a flag indicating that a neural network-based intra prediction mode is selected for the at least one block, the flag being based on a set of flags representing a plurality of intra prediction modes arranged in a binary tree for being encoded in a bitstream and wherein the flag indicating that neural network-based intra prediction mode is selected is located at a first level of the tree and encoded with a single bit.

3. A video decoding method comprising:

obtaining, for at least one block in a picture or video, at least information representative that an intra prediction is a neural network-based prediction and a block context, the block context comprising pixels of blocks located at a top side, at a left side, at a diagonal top left side, at a diagonal top right side and at a diagonal bottom left side of the at least one block, wherein the size of the block context is based on the size of the at least one block, and wherein the block context comprises n l columns and n a rows, wherein n l and n a are selected as:

n

a

=

α

H

,

n

l

=

β

W

,

α

1

4

,

1

2

,

3

4

,

1

,

2

,

β

1

4

,

1

2

,

3

4

,

1

,

2

where H is a height of the block and W is a width of the block; and

performing intra prediction for the at least one block in a picture or video by feeding the block context into a neural network based on the size of the at least one block.

4. The method of claim 3 , wherein the neural network-based intra prediction is performed based on a position of the at least one block.

5. The method of claim 3 , wherein the block context is down-sampled prior to performing the intra prediction and the at least one predicted block resulting from the neural network based intra prediction is interpolated after the intra prediction.

6. The method of claim 3 , wherein the block context is transposed prior to performing the intra prediction and the at least one predicted block resulting from the neural network based intra prediction is transposed back after the intra prediction.

7. The method of claim 3 , wherein the block context is down-sampled and transposed prior to performing the intra prediction and the at least one predicted block resulting from the neural network based intra prediction is transposed back and interpolated after the intra prediction.

8. The method of claim 3 , wherein the neural network-based intra prediction is done in both luminance and chrominance of the at least one block.

9. An apparatus, comprising an encoder for encoding a current block in a picture or video wherein the encoder is configured to:

perform intra prediction for at least one block in a picture or video using a neural network based intra prediction by feeding a block context into a neural network selected based on a size of the at least one block, the block context comprising pixels of blocks located at a top side, at a left side, at a diagonal top left side, at a diagonal top right side and at a diagonal bottom left side of the at least one block, wherein the size of the block context is based on the size of the at least one block, and wherein the block context comprises n l columns and n a rows, wherein n l and n a are selected as:

n

a

=

α

H

,

n

l

=

β

W

,

α

1

4

,

1

2

,

3

4

,

1

,

2

,

β

1

4

,

1

2

,

3

4

,

1

,

2

where H is a height of the block and W is a width of the block;

generate signaling information representative that the intra prediction mode is a neural network based intra prediction; and

encode at least information representative of the at least one block and the neural network-based intra prediction mode.

10. An apparatus, comprising a decoder for decoding picture data for a current block in a picture or video wherein the decoder is configured to:

obtain, for at least one block in a picture or video, at least information representative that an intra prediction is a neural network-based prediction and a block context, the block context comprising pixels of blocks located at a top side, at a left side, at a diagonal top left side, at a diagonal top right side and at a diagonal bottom left side of the at least one block, wherein the size of the block context is based on the size of the at least one block, and wherein the block context comprises n l columns and n a rows, wherein n l and n a are selected as:

n

a

=

α

H

,

n

l

=

β

W

,

α

1

4

,

1

2

,

3

4

,

1

,

2

,

β

1

4

,

1

2

,

3

4

,

1

,

2

where H is a height of the block and W is a width of the block; and

perform intra prediction for the at least one block in a picture or video by feeding the block context into a neural network selected based on the size of the at least one block.

11. The apparatus of claim 10 , wherein the neural network-based intra prediction is performed based on a position of the at least one block.

12. The apparatus of claim 10 , wherein the block context is down-sampled prior to performing the intra prediction and the at least one predicted block resulting from the neural network-based intra prediction is interpolated after the intra prediction.

13. The apparatus of claim 10 , wherein the block context is transposed prior to performing the intra prediction and the at least one predicted block resulting from the neural network-based intra prediction is transposed back after the intra prediction.

14. The apparatus of claim 10 , wherein the block context is down-sampled and transposed prior to performing the intra prediction and the at least one predicted block resulting from the neural network-based intra prediction is transposed back and interpolated after the intra prediction.

15. A non-transitory computer readable medium comprising program code instructions for implementing the steps of a method according to claim 3 when executed by a processor.

16. The method of claim 1 , wherein the neural network is selected among a set of at least two neural networks, and wherein at least one size of a block is associated to each neural network of the set of at least two neural networks.

17. The method of claim 3 , wherein the neural network is selected among a set of at least two neural networks, and wherein at least one size of a block is associated to each neural network of the set of at least two neural networks.

18. The apparatus of claim 9 , wherein the signaling information is encoded in a bitstream and comprises a flag indicating that a neural network-based intra prediction mode is selected for the at least one block, the flag being based on a set of flags representing a plurality of intra prediction modes arranged in a binary tree for being encoded in a bitstream and wherein the flag indicating that neural network-based intra prediction mode is selected is located at a first level of the tree and encoded with a single bit.

19. The apparatus of claim 9 , wherein the neural network is selected among a set of at least two neural networks, and wherein at least one size of a block is associated to each neural network of the set of at least two neural networks.

20. The apparatus of claim 10 , wherein the neural network is selected among a set of at least two neural networks, and wherein at least one size of a block is associated to each neural network of the set of at least two neural networks.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 6, 2023
From: INTERDIGITAL CE PATENT HOLDINGS, SAS
To: INTERDIGITAL MADISON PATENT HOLDINGS, SAS
Reel/Frame 065465/0293 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: INTERDIGITAL VC HOLDINGS FRANCE, SAS
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 064460/0921 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2023
From: LELEANNEC, FABRICE
To: INTERDIGITAL VC HOLDINGS FRANCE
Reel/Frame 062745/0044 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2022
From: DUMAS, THIERRY; GALPIN, FRANCK; BORDES, PHILIPPE; LELEANNEC, FABRICE
To: INTERDIGITAL VC HOLDINGS FRANCE
Reel/Frame 060819/0247 →
Priority Claims (1)
EP 20305169 · Feb 21, 2020 · regional
Continuity (1)
Related Publication 20230095387A1 · Mar 30, 2023
References Cited (16)
US 20180184123A1 · Terada · 2018 [cited by examiner]
US 20180332282A1 · He et al. · 2018 [cited by applicant]
US 20190387222A1 · Kim et al. · 2019 [cited by applicant]
US 20200186796A1 · Mukherjee · 2020 [cited by examiner]
US 20200304832A1 · Ramasubramonian · 2020 [cited by examiner]
US 20220360785A1 · Huo · 2022 [cited by examiner]
WO WO2019072921A1 · 2019 [cited by applicant]
WO WO2019185808A1 · 2019 [cited by examiner]
WO WO2019185883A1 · 2019 [cited by examiner]
Hu et al., “Progressive Spatial Recurrent Neural Network for Intra Prediction”, Institute of Electronics and Electronical Engineers (IEEE), IEEE Transactions on MultiMedia, vol. 21, Issue: 12, Dec. 2019, 14 pages. [cited by applicant]
Li et al., “Fully-Connected Network-Based Intra Prediction for Image Coding”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 27, No. 7, Jul. 2018, 12 pages. [cited by applicant]
Dumas et al., “Context-adaptive neural network-based prediction for image compression”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 29, Aug. 16, 2019, 15 pages. [cited by applicant]
Bross et al., “Versatile Video Coding (Draft 8)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-Q2001-vD, 17th Meeting, Brussels, Belgium, Jan. 7, 2020, 511 pages. [cited by applicant]
Anonymous, “High Efficiency Video Coding”, International Telecommunication Union, Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems—Infrastructure of audiovisual services—Codi… [cited by applicant]
Dumas et al., “AHG11: Neural Network-based intra prediction with transform selection in VVC”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, Document: JVET-T0073, 20th Meeting, Oct. 7, 2020… [cited by applicant]
Pfaff et al., “Neural Network based intra prediction for video coding”, International Society for Optics and Photonics (SPIE), SPIE Proceedings, vol. 10752, Applications of Digital Image Processing XLI, 1075213, Sep. 17… [cited by applicant]
Cited By (2)
US 12,555,201 US 12,587,640