IP Library Granted Patent US 11,792,391
Granted Patent B2
US 11,792,391 · App. 17/494,238 · Granted Oct 17, 2023

Neural network based coefficient sign prediction

Inventors: Xin Zhao (Palo Alto, CA); Yixin Du (Palo Alto, CA); Liang Zhao (Palo Alto, CA); Madhu Peringassery Krishnan (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
H04N19/105G06N3/08H04N19/129H04N19/132H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,792,391
App. No.
17/494,238
Granted
Oct 17, 2023
Kind
B2
Abstract

A method, computer program, and computer system is provided for coding video data. Reference samples and magnitudes of transform coefficients corresponding to a current block of video data from an input to a neural network are identified. Sign values associated with the transform coefficients are predicted based on at least the identified reference samples. The video data is encoded/decoded based on the predicted sign values.

Claims (34)

1. A method of video coding, executable by a processor, the method comprising:

identifying reference samples and magnitudes of transform coefficients corresponding to a current block of video data from inputs to a neural network, the inputs including a quantization parameter, a reconstruction parameter, and at least one POC distance;

extracting a set of features from the reference samples, wherein the extracting includes performing an integer-approximation of an ELU function;

predicting sign values associated with the transform coefficients based at least on the identified reference samples and the set of features; and

coding the video data based on the predicted sign values.

2. The method of claim 1 , wherein the reference samples comprise reconstructed samples from spatially neighboring blocks corresponding to the current block.

3. The method of claim 1 , wherein the reference samples include reconstructed samples specified by motion vectors associated with a reference image from the video data.

4. The method of claim 1 , wherein the transform coefficients are de-quantized transform coefficients.

5. The method of claim 1 , wherein the predicted sign values correspond to a limited set of transform coefficients.

6. The method of claim 5 , wherein the limited set of transform coefficients comprises low frequency coefficients at a pre-defined top left area of the current block.

7. The method of claim 5 , wherein the limited set of transform coefficients comprises a pre-defined number of coefficients along a forward scanning order.

8. The method of claim 1 , wherein the sign values are predicted using a convolutional neural network.

9. The method of claim 8 , wherein based on reconstructed samples being associated with a current image and/or multiple reference images, a phase-only correlation difference between a current image and the reference images is used as an input to the convolutional neural network.

10. The method of claim 8 , wherein indicators of a primary transform type or kernel are an input to the convolutional neural network.

11. The method of claim 10 , wherein indicators of a secondary transform type or kernel are an input to the convolutional neural network.

12. The method of claim 8 , further comprising obtaining estimated reconstructed samples of the current block based on a candidate of the predicted sign values, wherein the estimated reconstructed samples of the current block and neighboring reconstructed values are inputs to the convolutional neural network.

13. The method of claim 12 , wherein score values associated with the predicted sign values are outputs of the convolutional neural network, and wherein a sign value is selected from among the predicted sign values by the current block based on the score values.

14. The method of claim 8 , wherein a number of transform coefficients used as an input to the convolutional neural network varies based on a secondary transform being applied.

15. The method of claim 14 , wherein a number of transform coefficients in a forward scanning order less than a pre-defined number of coefficients are used based on the secondary transform being applied to the current block.

16. The method of claim 1 , wherein a subset of reconstructed samples of the current block is used based on the predicted sign values.

17. The method of claim 16 , wherein the subset of reconstructed samples includes boundary samples of current block.

18. The method of claim 17 , wherein the boundary samples comprise a first pre-determined number of top rows of the current block and a second pre-determined number of left columns of the current block, and the first and the second pre-determined numbers are based on a size associated with the current block.

19. A computer system for coding video data, the computer system comprising:

one or more computer-readable non-transitory storage media configured to store computer program code; and

one or more computer processors configured to access the computer program code stored in the one or more computer-readable non-transitory storage media and operate as instructed by the computer program code, the computer program code including:

identifying code configured to cause the one or more computer processors to identify reference samples and magnitudes of transform coefficients corresponding to a current block of video data from inputs to a neural network, the inputs including a quantization parameter, a reconstruction parameter, and at least one POC distance;

extracting code configured to cause the one or more computer processors to extract a set of features from the reference samples, wherein the extracting includes performing an integer-approximation of an ELU function;

predicting code configured to cause the one or more computer processors to predict sign values associated with the transform coefficients based on at least the identified reference samples and the set of features; and

coding code configured to cause the one or more computer processors to code the video data based on the predicted sign values.

20. A non-transitory computer readable medium storing a computer program for coding video data, the computer program configured to cause one or more computer processors to at least:

identify reference samples and magnitudes of transform coefficients corresponding to a current block of video data from inputs to a neural network, the inputs including a quantization parameter, a reconstruction parameter, and at least one POC distance;

extract a set of features from the reference samples, wherein the extracting includes performing an integer-approximation of an ELU function;

predict sign values associated with the transform coefficients based on at least the identified reference samples and the set of features; and

code the video data based on the predicted sign values.

Continuity (2)
Continuation 17071582 · Oct 15, 2020
Related Publication 20220124311A1 · Apr 21, 2022