Video encoding using pre-processing
There is provided a technique for video encoding. The technique comprises downsampling at a downsampler ( 820 ), an input video stream at a first resolution ( 805 ) to a second resolution ( 825 ), the second resolution being lower than the first resolution. The technique uses a set of encoders to encode signals derived from the input video stream at the first spatial resolution and the second spatial resolution. There is also provided a pre-processing stage ( 830 ) to pre-process the input video stream at the first resolution ( 805 ) prior to the downsampling at downsampler ( 820 ). The pre-processing comprises an application of a blurring filter ( 810 ) and a sharpening filter ( 815 ).
1 . A method for video encoding using a hierarchical coding format, the method comprising:
downsampling, at a downsampler, an input video stream from a first spatial resolution to a second spatial resolution, the second spatial resolution being lower than the first spatial resolution, wherein said downsampling includes a number of downsampling operations, and wherein the number of downsampling operations is selected to be one less than a number of echelon indices that are to be used as a part of the hierarchical coding format;
encoding, at a set of encoders, a signal derived from the input video stream at the first spatial resolution and a signal derived from the downsampled input video stream at the second spatial resolution; and
pre-processing, at a pre-processing stage, the input video stream prior to the downsampling, wherein the pre-processing comprises the application of:
a blurring filter; and
a sharpening filter,
wherein the blurring filter and the sharpening filter, which are applied prior to the downsampling, are cascaded in that order prior to the downsampling, and wherein the pre-processing is configured to modify characteristics of residual data generated by the hierarchical coding format such that residual encoding efficiency is modified independently of perceptual image quality at the second spatial resolution.
2 . The method of claim 1 , wherein the pre-processing and the downsampling implement a non-linear modification of the input video stream.
3 . The method of claim 1 , wherein the pre-processing at the pre-processing stage is controllably enabled or disabled.
4 . The method of claim 1 , wherein the blurring filter is a Gaussian filter.
5 . The method of claim 1 , wherein the sharpening filter comprises an unsharp mask.
6 . The method of claim 1 , wherein the sharpening filter is a 2D N×N filter, where N is an integer value.
7 . The method of claim 1 , wherein the sharpening filter uses adjustable coefficient values.
8 . The method of claim 1 , wherein the set of encoders implement a bitrate ladder.
9 . The method of claim 1 , wherein the encoding at the set of encoders comprises encoding the signal derived from the input video stream at the first spatial resolution using a first encoding method and the signal derived from the downsampled input video stream at the second spatial resolution using a second method, wherein the first encoding method and the second encoding method are different.
10 . The method of claim 9 , wherein the encoded signals from the first and second methods are output as an LCEVC encoded data stream.
11 . The method of claim 1 , wherein the encoding at the set of encoders comprise encoding the signal derived from the input video stream at the first spatial resolution using a first encoding method and the signal derived from the downsampled input video stream at the second spatial resolution using a second method, wherein the first encoding method and the second encoding method are the same.
12 . The method of claim 11 , wherein the first encoding method and the second encoding method generate at least part of a VC-6 encoded data stream.
13 . The method of claim 1 , wherein the encoding at the set of encoders comprise encoding a residual stream, the residual stream being generated based on a comparison of a reconstruction of the input video stream at the first spatial resolution with the input video stream at the first spatial resolution, the reconstruction of the video stream at the first spatial resolution being derived from a reconstruction of the video stream at the second spatial resolution.
14 . The method of claim 1 , wherein the encoding at the set of encoders comprise encoding the input video stream at the second spatial resolution or lower, and wherein the encoding at the set of encoders further comprise encoding a second residual stream, the second residual stream being generated based on a comparison of a reconstruction of the input video stream at the second spatial resolution with the input video stream at the second spatial resolution, the reconstruction of the input video stream at the second spatial resolution being derived from a decoding of the encoded input video stream at the second spatial resolution or lower.
15 . The method of claim 1 , wherein the method further comprises a second downsampling at a second downsampler to convert the input video stream from the second spatial resolution to a third spatial resolution, the third spatial resolution being lower than the second spatial resolution, and applying the pre-processing at a second pre-processing stage before the second downsampler.
16 . The method of claim 15 , wherein the pre-processing at the pre-processing stage and at the second pre-processing stage are enabled or disabled in different combinations.
17 . The method of claim 1 , wherein one or more image metrics used by one or more of the set of encoders are disabled when the pre-processing is enabled, the one or more image metrics optionally comprising PSNR or SSIM image metrics.
18 . The method of claim 1 , wherein the number of echelon indices is 4, and wherein the number of downsampling operations is 3.
19 . A system for video encoding using a hierarchical coding format, said system comprising:
one or more processors; and
one or more computer hardware storage devices having stored thereon executable instructions that are executable by the one or more processors to cause the system to:
downsample, at a downsampler, an input video stream from a first spatial resolution to a second spatial resolution, the second spatial resolution being lower than the first spatial resolution, wherein said downsampling includes a number of downsampling operations, and wherein the number of downsampling operations is selected to be one less than a number of echelon indices that are to be used as a part of the hierarchical coding format;
encode, at a set of encoders, a signal derived from the input video stream at the first spatial resolution and a signal derived from the downsampled input video stream at the second spatial resolution; and
pre-process, at a pre-processing stage, the input video stream prior to the downsampling, wherein the pre-processing comprises the application of:
a blurring filter; and
a sharpening filter,
wherein the blurring filter and the sharpening filter, which are applied prior to the downsampling, are cascaded in that order prior to the downsampling, and wherein the pre-processing is configured to modify characteristics of residual data generated by the hierarchical coding format such that residual encoding efficiency is modified independently of perceptual image quality at the second spatial resolution.
20 . A non-transitory computer-readable storage medium comprising instructions that are executable by a processor to cause the processor to:
downsample, at a downsampler, an input video stream from a first spatial resolution to a second spatial resolution, the second spatial resolution being lower than the first spatial resolution, wherein said downsampling includes a number of downsampling operations, and wherein the number of downsampling operations is selected to be one less than a number of echelon indices that are to be used as a part of the hierarchical coding format;
encode, at a set of encoders, a signal derived from the input video stream at the first spatial resolution and a signal derived from the downsampled input video stream at the second spatial resolution; and
pre-process, at a pre-processing stage, the input video stream prior to the downsampling, wherein the pre-processing comprises the application of:
a blurring filter; and
a sharpening filter,
wherein the blurring filter and the sharpening filter, which are applied prior to the downsampling, are cascaded in that order prior to the downsampling, and wherein the pre-processing is configured to modify characteristics of residual data generated by the hierarchical coding format such that residual encoding efficiency is modified independently of perceptual image quality at the second spatial resolution.