IP Library Granted Patent US 12,143,573
Granted Patent B2
US 12,143,573 · App. 18/244,768 · Granted Nov 12, 2024

Neural network based coefficient sign prediction field

Inventors: Xin Zhao (Palo Alto, CA); Yixin Du (Palo Alto, CA); Liang Zhao (Palo Alto, CA); Madhu Peringassery Krishnan (Palo Alto, CA); Shan Liu (Palo Alto, CA)
Assignee: TENCENT AMERICA LLC
H04N19/105G06N3/08H04N19/129H04N19/132H04N19/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,143,573
App. No.
18/244,768
Granted
Nov 12, 2024
Kind
B2
Abstract

A method, computer program, and computer system is provided for coding video data. Reference samples and magnitudes of transform coefficients corresponding to a current block of video data from an input to a neural network are identified. Sign values associated with the transform coefficients are predicted using neural networks. The video data is encoded/decoded based on the predicted sign values.

Claims (31)

1. A method of video decoding, executable by a processor, the method comprising:

receiving a video bitstream comprising a current block in a current picture of video data of the video bitstream;

identifying reference samples and magnitudes of transform coefficients, corresponding to the current block, from an input to a neural network;

predicting sign values associated with the transform coefficients based at least on the identified reference samples, wherein the sign values are predicted using a convolutional neural network; and

decoding the video data based on the predicted sign values.

2. The method of claim 1 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.

3. The method of claim 1 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.

4. The method of claim 1 , wherein the transform coefficients are de-quantized transform coefficients.

5. The method of claim 1 , wherein the predicted sign values for correspond to a limited set of transform coefficients.

6. The method of claim 5 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.

7. The method of claim 5 , wherein the limited set of transform coefficients includes a pre-defined number of coefficients along a forward scanning order.

8. A method of video encoding, executable by a processor, the method comprising:

receiving video data;

identifying reference samples and magnitudes of transform coefficients, corresponding to a current block of a picture of the video data, from an input to a neural network;

determining sign values associated with the transform coefficients based at least on the identified reference samples, wherein the sign values are predicted using a convolutional neural network; and

encoding the video data based on the determined sign values.

9. The method of claim 8 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.

10. The method of claim 8 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.

11. The method of claim 8 , wherein the transform coefficients are de-quantized transform coefficients.

12. The method of claim 8 , wherein the predicted sign values for correspond to a limited set of transform coefficients.

13. The method of claim 12 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.

14. The method of claim 12 , wherein the limited set of transform coefficients includes a pre-defined number of coefficients along a forward scanning order.

15. A method of processing visual media data, executable by a processor, the method comprising:

performing a conversion between a visual media and a bitstream of visual media data according to a format rule, wherein the bitstream comprises video data including a current block in a current picture, and wherein the format rule specifies:

identifying reference samples and magnitudes of transform coefficients, corresponding to the current block of the current picture, from an input to a neural network, and

predicting sign values associated with the transform coefficients using a convolutional neural network based at least on the identified reference samples.

16. The method of claim 15 , wherein the reference sample includes reconstructed samples from spatially neighboring blocks corresponding to the current block.

17. The method of claim 15 , wherein the reference sample includes reconstructed samples specified by motion vectors associated with a reference image from the video data.

18. The method of claim 15 , wherein the transform coefficients are de-quantized transform coefficients.

19. The method of claim 15 , wherein the predicted sign values for correspond to a limited set of transform coefficients.

20. The method of claim 19 , wherein the limited set of transform coefficients includes low frequency coefficients at a pre-defined top left area of the current block.

Continuity (3)
Continuation 17494238 · Oct 5, 2021
Continuation 17071582 · Oct 15, 2020
Related Publication 20230421754A1 · Dec 28, 2023