IP Library › Granted Patent US 10,999,606
Granted Patent B2
US 10,999,606 · App. 16/418,803 · Granted May 4, 2021

Method and system of neural network loop filtering for video coding

Inventors: Hujun Yin (Saratoga, CA); Shoujiang Ma (Shanghai, CN); Xiaoran Fang (Shanghai, CN); Rongzhen Yang (Shanghai, CN)
Assignee: Intel Corporation
H04N19/82G06N3/0454H04N19/117H04N19/172H04N19/70
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,999,606
App. No.
16/418,803
Granted
May 4, 2021
Kind
B2
Abstract

A method, system, medium, and article provide neural network loop filtering for video coding with multiple alternative neural networks.

Claims (46)

1. A computer implemented method of video coding comprising:

obtaining compressed image data of at least one frame of a video sequence;

decoding the at least one frame to form a reconstructed version of the frame;

applying multiple alternative convolutional neural networks to the same pixel locations of the same blocks of the reconstructed version of the at least one frame;

selecting one of the convolutional neural networks based on at least one criterion, and

refining the image data of the part comprising using the output of the selected convolutional neural network.

2. The method of claim 1 wherein the multiple alternative convolutional neural networks at least partly establish an adaptable neural network in-loop filter on a decoding loop of an encoder.

3. The method of claim 1 comprising indicating a selection among the alternative convolutional neural networks in syntax data transmitted from an encoder to a remote decoder.

4. The method of claim 1 wherein the refining occurs at a decoder remote from an encoder and according to a selection indicated by the encoder so that the decoder does not need to perform the selecting.

5. The method of claim 1 comprising receiving the multiple alternative convolutional neural networks and the identification of the selected convolutional neural networks at a decoder remote from an encoder, and the decoder performing the refining.

6. The method of claim 5 wherein the encoder transmits all alternative convolutional neural networks to a decoder without checking which alternative neural networks were selected for a block.

7. The method of claim 1 wherein the encoder trains the multiple alternative convolutional neural networks before transmitting the multiple alternative convolutional neural networks to the decoder.

8. The method of claim 1 wherein each of the multiple alternative convolutional neural networks has only two convolutional layers.

9. The method of claim 8 wherein a rectified linear operation is performed on the output of a first layer of the two convolutional layers.

10. The method of claim 8 wherein the two convolutional layers comprises a first 1×1 filter layer and a second 3×3 filter layer.

11. The method of claim 1 wherein the selecting is performed during a run-time to complete an encode or decode of the at least one frame comprising forming a dataset to train the convolutional neural networks with image data of a set of previous frames already reconstructed.

12. The method of claim 1 comprising generating and training the multiple alternative convolutional neural networks during a run-time of an encoder and before applying the multiple alternative convolutional neural networks to the reconstructed version of the at least one frame comprising applying an initial neural network to a full training dataset to obtain an output dataset; and partitioning the output dataset by at least one criterion to form separate datasets to train separate neural networks.

13. A computer-implemented system comprising:

at least one display;

memory to store image data of at least one frame of a video sequence;

at least one processor communicatively coupled to the memory and display, and the at least one processor to operate by:

obtaining compressed image data of at least one current frame of a video sequence;

decoding the at least one current frame to form a reconstructed version of the current frame;

during a run-time of an encoder, training multiple alternative convolutional neural networks to output data used to refine image data of the reconstructed version of the frame and comprising establishing an initial training dataset comprising image data of a set of frames decoded previously to the decoding of the current frame; and

applying the multiple alternative convolutional neural networks to the reconstructed version of the current frame to refine the image data of the current frame.

14. The system of claim 13 comprising:

applying the multiple alternative convolutional neural networks to at least the same part of the reconstructed version of the at least one frame;

selecting one of the convolutional neural networks based on at least one criterion, and

refining the image data of the part comprising using the output of the selected convolutional neural network.

15. The system of claim 13 wherein the training dataset comprises data only of one or more I-frames and frames that use the I-frame as a reference frame.

16. The system of claim 13 wherein the training dataset comprises data only of the same random access segment or group of pictures.

17. The system of claim 13 wherein the training dataset comprises data of a predetermined number of frames before the current frame regardless of frame location in a particular random access segment and group of pictures.

18. The system of claim 13 wherein the at least one processor to operate by training the multiple alternative convolutional neural networks before applying the multiple alternative convolutional neural networks to the reconstructed version of the at least one frame comprising applying an initial neural network to a full training dataset to obtain an output dataset; and partitioning the output dataset by at least one criterion to form separate datasets to train separate neural networks.

19. At least one non-transitory computer-readable medium having stored thereon instructions that when executed cause a computing device to operate by:

obtaining compressed image data of at least one current frame of a video sequence;

decoding the at least one current frame to form a reconstructed version of the current frame;

during a run-time of an encoder, training multiple alternative convolutional neural networks to output data used to refine image data of the reconstructed version of the frame and comprising establishing an initial training dataset comprising image data of a set of frames decoded previously to the decoding of the current frame; and

applying the multiple alternative convolutional neural networks to the reconstructed version of the current frame to refine the image data of the current frame.

20. The medium of claim 19 wherein the training comprises applying an initial neural network to the initial training dataset, partitioning the output data of the initial neural network into subsets based on at least one criterion, using at least one of the subsets to train a separate neural network, and repeating the partitioning and using of subsets until a desired number of multiple alternative neural networks is reached.

21. The medium of claim 20 wherein the criterion is whether values of the output data indicate a gain versus a loss, wherein gain refers to output image data becoming closer in value to original image data of the same pixel or block location than the input reconstructed image data, and wherein loss refers to output image data becoming farther in value from original image data of the same pixel or block location than the input reconstructed image data.

22. The medium of claim 21 wherein only a loss-associated subset is used to train a new alternative neural network after two alternative neural networks are trained at least once.

23. The medium of claim 20 wherein after three or more neural networks are established, the instructions cause the computing device to operate by training the neural networks on the highest gain output data subset among output subsets from the three or more neural networks resulting from applying the three or more neural networks to the initial training dataset.

24. The medium of claim 19 wherein the instructions cause the computing device to operate by applying the multiple alternative convolutional neural networks to at least the same part of the reconstructed version of the at least one frame;

selecting one of the convolutional neural networks based on at least one criterion, and

refining the image data of the part comprising using the output of the selected convolutional neural network.

25. The medium of claim 19 wherein the initial training dataset comprises data of a predetermined number of frames before a current frame and not after the current frame being reconstructed and in encoding order.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2019
From: YIN, HUJUN; MA, SHOUJIANG; FANG, XIAORAN; YANG, RONGZHEN
To: INTEL CORPORATION
Reel/Frame 049349/0156 →
Continuity (2)
Provisional Application 62789952 · Jan 8, 2019
Related Publication 20190273948A1 · Sep 5, 2019
Cited By (4)
US 12,231,646 US 12,526,457 US 12,647,614 US 12,707,099