IP Library Granted Patent US 10,021,383
Granted Patent B2
US 10,021,383 · App. 14/123,009 · Granted Jul 10, 2018

Method and system for structural similarity based perceptual video coding

Inventors: Zhou Wang (Waterloo, CA); Abdul Rehman (Kitchener, CA)
Assignee: SSIMWAVE INC.
H04N19/00096H04N19/126H04N19/154H04N19/18H04N19/19H04N19/61
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,021,383
App. No.
14/123,009
Granted
Jul 10, 2018
Kind
B2
Abstract

The present invention is a system and method for video coding. The video coding system may involve a structural similarity-based divisive normalization approach, wherein the frame prediction residual of the current frame may be transformed to form a set of coefficients and a divisive normalization mechanism may be utilized to normalize each coefficient. The normalization factor may be designed to reflect or approximate the normalization factor in a structural similarity definition. The Lagrange parameter for RDO for divisive normalization coefficients may be determined by both the quantization step and a prior distribution function of the coefficients. The present invention may generally be utilized to improve the perceptual quality of decoded video without increasing data rate, or to reduce the data rate of compressed video stream without sacrificing the perceived quality of decoded video. The present invention has shown to significantly improve the coding efficiency of MPEG4/H.264 AVC and HEVC coding schemes. The present invention may be utilized to create video codes compatible with prior art and state-of-the-art video coding standards such as MPEG4/H.264 AVC and HEVC. The present invention may also be utilized to create video codecs incompatible with existing standards, so as to further improve the coding gain.

Claims (35)

1. A computer-implemented method of perceptual video coding utilizing a structural similarity-based divisive normalization approach, comprising:

producing a prediction residual by subtracting a current frame of video footage from a prediction from one or more previously coded frames while coding the current frame;

transforming the prediction residual to form a set of transform coefficients;

utilizing a transform domain structural similarity index to determine a divisive normalization factor for each transform coefficient;

adjusting a quantization parameter by adding a factor proportional to the logarithm of the divisive normalization factor;

quantizing the transform coefficients using the adjusted quantization parameter to obtain normalized and quantized coefficients; and

performing a rate-distortion optimization and entropy coding on the normalized and quantized coefficients.

2. The method of claim 1 , further comprising approximating the structural similarity divisive normalization factor based on estimating the energy of AC coefficients in the current frame by applying a scale factor to the energy of the corresponding coefficients in one or more previously coded frames that are neighboring frames to the current frame.

3. The method of claim 1 , further comprising quantizing the quantization parameter value to an integer number, so as to make the codec compatible with MPEG4/H.264 AVC and HEVC standards.

4. The method of claim 1 , further comprising performing rate-distortion optimization on normalized coefficients, wherein a Lagrange parameter is determined by utilizing an approximation model or a lookup table comprising one or more input arguments that are at least one of the following: a quantization step; and one or more parameters of a prior distribution of a normalized coefficient.

5. The method of claim 1 , further comprising adjusting the divisive normalization factors based on local content of the video frame, where the local content may be characterized by a local complexity measure computed as local contrast, local energy or local signal activities.

6. The method of claim 5 , further comprising spatially adapting the divisive normalization factor for each transform unit (TU), which may be blocks with variable sizes across space.

7. The method of claim 6 , further comprising dividing the TU to smaller blocks of equal size in the whole frame and then average the divisive normalization factors for all small blocks within the TU.

8. The method of claim 7 , further comprising normalizing local divisive normalization factor for each TU by the expected value of local divisive normalization factors of the whole frame being encoded.

9. A non-transient computer readable medium storing computer code that when executed on a computer device adapts the device to perform the method of claim 1 .

10. The method of claim 1 , further comprising:

performing an inverse divisive normalization transform on de-quantized coefficients; and

utilizing inverse divisive normalization transformed coefficients of previously coded frames to determine the divisive normalization factors of the coefficients in the current frame.

11. A computer-implemented system for perceptual video coding utilizing a structural similarity-based divisive normalization approach, wherein the system is adapted to:

produce a prediction residual by subtracting a current frame of video footage from a prediction from one or more previously coded frames while coding the current frame;

transform the prediction residual to form a set of transform coefficients;

utilize a transform domain structural similarity index to determine a divisive normalization factor for each transform coefficient;

adjust a quantization parameter by adding a factor proportional to the logarithm of the divisive normalization factor;

quantize the transform coefficients using the adjusted quantization parameter to obtain normalized and quantized coefficients; and

perform a rate-distortion optimization and entropy coding on the normalized coefficients.

12. The system of claim 11 , wherein the system is further adapted to: approximate structural similarity divisive normalization factor based on estimating the energy of AC coefficients in the current frame by applying a scale factor to the energy of the corresponding coefficients in one or more previously coded frames that are neighboring frames to the current frame.

13. The system of claim 11 , wherein the system is further adapted to quantize the QP value to an integer number, so as to make the codec compatible with MPEG4/H.264 AVC and HEVC standards.

14. The system of claim 11 , wherein the system is further adapted to perform rate-distortion optimization on normalized coefficients, wherein a Lagrange parameter is determined by utilizing an approximation model or a lookup table comprising one or more input arguments that are at least one of the following: a quantization step; and one or more parameters of a prior distribution of a normalized coefficient.

15. The system of claim 11 , wherein the system is further adapted to adjust the divisive normalization factors based on local content of the video frame, where the local content may be characterized by a local complexity measure computed as local contrast, local energy or local signal activities.

16. The system of claim 15 , wherein the system is further adapted to spatially adapt the divisive normalization factor for each transform unit (TU), which may be blocks with variable sizes across space.

17. The system of claim 16 , wherein the system is further adapted to divide the TU to smaller blocks of equal size in the whole frame and then average the divisive normalization factors for all small blocks within the TU.

18. The system of claim 17 , wherein the system is further adapted to normalize local divisive normalization factor for each TU by the expected value of local divisive normalization factors of the whole frame being encoded.

19. The system of claim 11 , wherein the system is further adapted to:

perform an inverse divisive normalization transform on de-quantized coefficients; and

utilize inverse divisive normalization transformed coefficients of previously coded frames to determine the divisive normalization factors of the coefficients in the current frame.

Assignments (5)
RELEASE OF SECURITY INTEREST Recorded Jul 29, 2025
From: SILICON VALLEY BANK
To: SSIMWAVE INC.
Reel/Frame 071866/0803 →
SECURITY INTEREST Recorded Jul 14, 2025
From: IMAX CORPORATION
To: WELLS FARGO BANK, NATIONAL ASSOCIATION, AS ADMINISTRATIVE AGENT
Reel/Frame 071935/0813 →
MERGER Recorded Feb 21, 2024
From: SSIMWAVE INC.
To: IMAX CORPORATION
Reel/Frame 066643/0122 →
SECURITY INTEREST Recorded Oct 28, 2020
From: SSIMWAVE INC.
To: SILICON VALLEY BANK
Reel/Frame 054196/0709 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 27, 2018
From: WANG, ZHOU; REHMAN, ABDUL
To: SSIMWAVE INC.
Reel/Frame 045358/0100 →
Continuity (3)
Provisional Application 61492081 · Jun 1, 2011
Provisional Application 61523610 · Aug 15, 2011
Related Publication 20140140396A1 · May 22, 2014