IP Library Granted Patent US 12,355,942
Granted Patent B2
US 12,355,942 · App. 18/011,184 · Granted Jul 8, 2025

Adapting the transform process to neural network-based intra prediction mode

Inventors: Thierry Dumas (Rennes, FR); Franck Galpin (Thorigne-Fouillard, FR); Philippe Bordes (Laille, FR); Fabrice Le Leannec (Betton, FR)
Assignee: InterDigital CE Patent Holdings, SAS
H04N19/105H04N19/12H04N19/176H04N19/593
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,355,942
App. No.
18/011,184
Granted
Jul 8, 2025
Kind
B2
Abstract

At least a method and an apparatus are presented for efficiently encoding or decoding video. For example, an intra prediction of an image block using at least one neural network from a context comprising pixels surrounding the image block is determined and an information relative to a transform method to apply for decoding the image block is also determined. The transform method is adapted to the neural network intra prediction mode of the block to encode or decode. The information relative to the transform method is inferred from the at least one neural network used in intra prediction of the image block at the encoding and either signaled or also inferred at the decoding.

Claims (75)

1. A method comprising:

determining an intra prediction of an image block and a vector by applying at least one neural network to a context comprising pixels adjacent to the image block, wherein an ith element of the vector indicates a relative relevance that a respective group of transforms corresponding to an index i includes a transform that is applied to the image block and wherein the index i identifies the respective group of transforms from among a plurality of groups of transforms;

inferring an index corresponding to a higher relative relevance of the vector, wherein the index identifies a particular group of transforms from among a plurality of groups of transforms;

obtaining a block of residue for the image block by at least applying an inverse transform to transform coefficients of the image block, the inverse transform belonging to the particular group of transforms identified by the index; and

decoding the image block based on the intra prediction and the block of residue.

2. The method of claim 1 , further comprising:

decoding an explicit transform index, wherein the explicit transform index indicates a particular transform applied to the image block among the particular group of transforms.

3. The method of claim 1 , wherein the at least one neural network comprises one or more input data among:

a Quantization Parameter used to decode some of the pixels that belongs to the context;

an index of an intra prediction mode used to predict the image block that is located at a top-right of the context if the image block at the top-right of the context is predicted in intra; and

an index of an intra prediction mode used to predict an image block that is located at a bottom-left of the context if the image block at the bottom-left of the context is predicted in intra.

4. The method of claim 1 , wherein obtaining a block of residue for the image block further comprises:

applying an inverse secondary transform to the transform coefficients; and

applying an inverse primary transform after the inverse secondary transform to obtain the block of residue, wherein the inverse secondary transform belongs to the particular group of transforms identified by the index.

5. The method of claim 1 , wherein obtaining a block of residue for the image block further comprises:

applying an inverse secondary Low Frequency Non-Separable Transform LFNST to the transform coefficients; and

applying an inverse primary transform to the inverse secondary transform coefficients, wherein a group of transforms includes 2 LFNST matrices with a decision of transposing primary transform coefficients.

6. The method of claim 5 , further comprising:

decoding an index of a Low Frequency Non-Separable Transform in the group of 2 LFNST matrices.

7. An apparatus comprising at least one processor configured to:

determine an intra prediction of an image block and a vector by applying at least one neural network to a context comprising pixels adjacent to the image block, wherein an ith element of the vector indicates a relative relevance that a respective group of transforms corresponding to an index i includes a transform that is applied to the image block and wherein the index i identifies the respective group of transforms from among a plurality of groups of transforms;

infer an index corresponding to a higher relative relevance of the vector, wherein the index identifies a particular group of transforms from among a plurality of groups of transforms;

obtain a block of residue for the image block by at least applying a particular inverse transform to transform coefficients of the image block, the particular inverse transform belonging to the particular group of transforms identified by the index; and

decode the image block based on the intra prediction and the block of residue.

8. The apparatus of claim 7 , the at least one processor further configured to:

decode an explicit transform index, wherein the explicit transform index indicates a particular transform applied to the image block among the particular group of transforms.

9. The apparatus of claim 7 , wherein the at least one neural network comprises one or more input data among:

a Quantization Parameter used to decode some of the pixels that belongs to the context;

an index of an intra prediction mode used to predict the image block that is located at a top-right of the context if the image block at the top-right of the context is predicted in intra; and

an index of an intra prediction mode used to predict an image block that is located at a bottom-left of the context if the image block at the bottom-left of the context is predicted in intra.

10. The apparatus of claim 7 , the at least one processor further configured to:

apply an inverse secondary transform to the transform coefficients; and

apply an inverse primary transform after the inverse secondary transform to obtain the block of residue, wherein the inverse secondary transform belongs to the particular group of transforms identified by the index.

11. The apparatus of claim 7 , the at least one processor further configured to:

apply an inverse secondary Low Frequency Non-Separable Transform LFNST to the transform coefficients; and

apply an inverse primary transform to the inverse secondary transform coefficients to obtain the block of residue, wherein the group of transforms includes 2 LFNST matrices with a decision of transposing primary transform coefficients.

12. The apparatus of claim 11 , the at least one processor further configured to:

decode an index of a Low Frequency Non-Separable Transform in the group of 2 LFNST matrices.

13. A method comprising:

determining an intra prediction of an image block and a vector by applying at least one neural network to a context comprising pixels adjacent to the image block, wherein an ith element of the vector indicates a relative relevance that a respective group of transforms corresponding to an index i includes a transform that is applied to the image block and wherein the index i identifies the respective group of transforms from among a plurality of groups of transforms;

inferring an index corresponding to a higher relative relevance of the vector, wherein the index identifies a particular group of transforms from among a plurality of groups of transforms;

at least applying a particular transform to a block of residue for the image block to obtain transform coefficients of the image block, the particular transform belonging to the particular group of transforms identified by the index; and

encoding the block of transform coefficients.

14. The method of claim 13 , further comprising:

encoding an explicit transform index, wherein the explicit transform index indicates the particular transform applied to the image block among the particular group of transforms.

15. The method of claim 13 , wherein the at least one neural network comprises one or more input data among:

a Quantization Parameter used to encode the image block that either partially or fully belongs to the context;

an index of an intra prediction mode used to predict the image block that is located at a top-right of the context if the image block at the top-right of the context is predicted in intra; and

an index of an intra prediction mode used to predict an image block that is located at a bottom-left of the context if the image block at the bottom-left of the context is predicted in intra.

16. The method of claim 13 , wherein obtaining a block of transform coefficient for the image block further comprises:

applying a primary transform to the block of residue; and

applying a secondary transform after the primary transform to obtain to the block of transform coefficients, wherein the secondary transform belongs to the particular group of transforms identified by the index.

17. The method of claim 13 , wherein obtaining a block of transform coefficient for the image block further comprises:

applying a primary transform to the block of residue; and

applying a secondary Low Frequency Non-Separable Transform LFNST to the block of primary transform coefficients, wherein a group of transforms includes 2 LFNST matrices with a decision of transposing primary transform coefficients.

18. The method of claim 17 , further comprising:

encoding an index of a Low Frequency Non-Separable Transform in the group of 2 LFNST matrices.

19. An apparatus comprising at least one processor configured to:

determine an intra prediction of an image block and a vector by applying least one neural network applied to a context comprising pixels adjacent to the image block, wherein an ith element of the vector indicates a relative relevance that a respective group of transforms corresponding to an index i includes a transform that is applied to the image block and wherein the index i identifies the respective group of transforms from among a plurality of groups of transforms;

infer an index corresponding to a higher relative relevance of the vector, wherein the index identifies a particular group of transforms from among a plurality of groups of transforms;

at least apply a particular transform to a block of residue to obtain transform coefficients of the image block, the particular transform belonging to the particular group of transforms identified by the index; and

encode the block of transform coefficients.

20. The apparatus of claim 19 , the at least one processor further configured to:

encode an explicit transform index, wherein the explicit transform index indicates the particular transform applied to the image block among the particular group of transforms.

21. The apparatus of claim 19 , wherein the at least one neural network comprises one or more input data among:

a Quantization Parameter used to encode the image block that either partially or fully belongs to the context;

an index of an intra prediction mode used to predict the image block that is located at a top-right of the context if the image block at the top-right of the context is predicted in intra; and

an index of an intra prediction mode used to predict an image block that is located at a bottom-left of the context if the image block at the bottom-left of the context is predicted in intra.

22. The apparatus of claim 19 , the at least one processor further configured to:

apply a primary transform to the block of residue; and

apply a secondary transform after the primary transform to obtain to the block of transform coefficients, wherein the secondary transform belongs to the particular group of transforms identified by the index.

23. The apparatus of claim 19 , the at least one processor further configured to:

apply a primary transform to the block of residue; and

apply a secondary Low Frequency Non-Separable Transform LFNST to the block of primary transform coefficients, wherein the group of transforms includes 2 LFNST matrices with a decision of transposing primary transform coefficients.

24. The apparatus of claim 23 , the at least one processor further configured to encode an index of a Low Frequency Non-Separable Transform in the group of 2 LFNST matrices.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 1, 2023
From: INTERDIGITAL VC HOLDINGS FRANCE, SAS
To: INTERDIGITAL CE PATENT HOLDINGS, SAS
Reel/Frame 064460/0921 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 18, 2022
From: DUMAS, THIERRY; GALPIN, FRANCK; BORDES, PHILIPPE; LE LEANNEC, FABRICE
To: INTERDIGITAL VC HOLDINGS FRANCE, SAS
Reel/Frame 062133/0338 →
Priority Claims (3)
EP 20305668 · Jun 18, 2020 · regional
EP 20306137 · Sep 30, 2020 · regional
EP 21305378 · Mar 26, 2021 · regional
Continuity (1)
Related Publication 20230224454A1 · Jul 13, 2023
References Cited (17)
US 20200186808A1 · Joshi · 2020 [cited by examiner]
US 20200389661A1 · Zhao · 2020 [cited by examiner]
US 20200396455A1 · Liu · 2020 [cited by examiner]
US 20220070482A1 · Kang · 2022 [cited by examiner]
US 20220201316A1 · Coelho · 2022 [cited by examiner]
EP 3310058A1 · 2018 [cited by applicant]
Dumas et al., “Iterative Training of Neural Networks for Intra Prediction”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 30, Nov. 23, 2020, 15 pages. [cited by applicant]
Pfaff et al., “Intra prediction Modes based on Neural Networks”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-J0037-v1, 10th Meeting: San Diego, California, USA, Apr.… [cited by applicant]
Dumas et al., “Context-Adaptive Neural Network-Based Prediction for Image Compression”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 29, Aug. 16, 2019, 16 pages. [cited by applicant]
Dumas et al., “AHG11: Neural Network-based intra prediction with transform selection in VVC”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, Document: JVET-T0073, 20th Meeting, Oct. 7, 2020… [cited by applicant]
Anonymous, “Reference software for ITU-T H.265 high efficiency video coding”, International Telecommunication Union, ITU-T Telecommunication Standardization Sector of ITU, Series H: Audiovisual and Multimedia Systems, I… [cited by applicant]
Anonymous, “Series H: Audiovisual and Multimedia Systems, Infrastructure of audiovisual services—Coding of moving video—Information Technology—Generic coding of moving pictures and associated audio information: Video”, … [cited by applicant]
Li et al., “Fully-Connected Network-Based Intra Prediction for Image Coding”, Institute of Electrical and Electronics Engineers (IEEE), IEEE Transactions on Image Processing, vol. 27, No. 7, Jul. 2018, 12 pages. [cited by applicant]
Bross et al, “Versatile video coding (draft 9)”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29/WG 11, Document: JVET-R2001-vA, 18th Meeting, by teleconference, Apr. 15, 2020, 524 pages. [cited by applicant]
Pfaff et al., “Neural Network based intra prediction for video coding”, International Society for Optics and Photonics (SPIE), SPIE Proceedings, vol. 10752, Applications of Digital Image Processing XLI, 1075213, Sep. 17… [cited by applicant]
Anonymous, “Transmission of Non-Telephone Signals: Information Technology—Generic Coding of Moving Pictures and Associated Audio Information: Systems”, International Telecommunication Union, ITU-T Telecommunication Stan… [cited by applicant]
Pfaff et al., “Intra Prediction Modes based on Neural Networks”, Joint Video Experts Team (JVET) of ITU-T SG 16 WP 3 and ISO/IEC JTC 1/SC 29, Fraunhofer Heinreich Hertz Institute, Document: JVET0037-v2, 10th Meeting, Sa… [cited by applicant]
Cited By (1)
US 12,641,222