SYSTEMS AND METHODS FOR CONTENT ADAPTIVE MULTI-SCALE FEATURE LAYER FILTERING AND REDUNDANT CHANNEL PROCESSING
Systems and methods are provided for encoding and decoding video for machine consumption in which bandwidth is reduced by filtering feature layers and filtering channels at the encoder site that are determined to be redundant or of reduced relevance. A video encoder includes a neural network front end which receives image data and generates a plurality of feature layers. The relevance of the plurality of feature layers to a machine task at the decoder site is determined and redundant layers can be removed. Channels in at least one feature layer can be evaluated for redundancy and redundant channels also removed prior to encoding.
1 . A video encoder in a system for video coding for machines, comprising:
a neural network front end, the neural network front end receiving image data and generating a plurality of feature layers;
layer context processor evaluating context of objects in the image data which impacts the significance of the layers of the feature map for a machine task;
a redundant layer identifier, the redundant layer identifier applying the output of the context processor and determining the relevance of the plurality of feature layers to a machine task;
a layer filter, the layer filter receiving the plurality of feature layers from the neural network front end and the output of the redundant layer identifier and performing at least one of removing a redundant layer and scaling a layer identified as having low relevance to a machine task to generate a filtered layer set;
a redundant channel identifier; the redundant channel identifier receiving the filtered layer set and identifying at least one set of correlated channels in at least one feature layer
a channel filter selecting at least one channel of the at least on set of correlated channels for encoding and removing the remaining correlated channels to generate a filtered channel set; and
an encoder receiving the filtered layer set and filtered channel set and generating a coded bitstream of the filtered layer and channel set.
2 . The video encoder of claim 1 further comprising signaling information in the coded bitstream indicating which layers of the plurality of layers are removed or modified by the layer filter and which channels of the plurality of channels are removed or modified by the channel filter.
3 . The video encoder of claim 1 , wherein the at least one set of correlated channels comprises a two sets of correlated channels.
4 . The video encoder of claim 1 , wherein the neural network front end comprises a feature pyramid network.
5 . The video encoder of claim 1 , wherein the redundant layer identifier further comprises a lightweight object detector.
6 . A decoder in a system for video coding for machines, comprising circuitry configured to:
receive an encoded bitstream generated by an encoder that generates a plurality of feature layers and selectively filters the plurality of plurality of feature layers prior to encoding, the bitstream comprising the filtered feature layers and signaling information identifying which layers were impacted by filtering;
decompress the encoded bitstream;
applying the signaling information to the decompressed bitstream and generating layers and channels removed by filtering at the encoder;
applying the reconstructed feature layers and channels to a neural network trained for a machine task.
7 . The decoder of claim 6 , wherein the signaling information explicitly signals which channels are removed during encoding.
8 . The decoder of claim 6 , wherein the signaling information implicitly signals which channels are removed prior to encoding.
9 . The decoder of claim 8 , wherein the channels removed prior to encoding are among a plurality of correlated channels in which one channel among the plurality was selected for encoding, and wherein the decoder generates the channels removed by filtering by applying values from the selected channel to the plurality of correlated channels removed at the encoder.
10 . A method of transmitting an encoded bitstream for video coding for machines, comprising:
applying a neural network front end, the neural network front end receiving image data and generating a plurality of feature layers;
applying a layer context processor evaluating context of objects in the image data which impacts the significance of the layers of the feature map for a machine task;
applying a redundant layer identifier, the redundant layer identifier applying the output of the context processor and determining the relevance of the plurality of feature layers to a machine task;
applying a layer filter, the layer filter receiving the plurality of feature layers from the neural network front end and the output of the redundant layer identifier and performing at least one of removing a redundant layer and scaling a layer identified as having low relevance to a machine task to generate a filtered layer set;
applying a redundant channel identifier, the redundant channel identifier determining a correlation among a plurality of channels in at least one feature layer a channel filter selecting one of a plurality of correlated channels to encode and removing the remaining plurality of correlated channels; and
applying an encoder receiving the filtered layer set and channels and generating the coded bitstream for transmission.
11 . The method of claim 9 , wherein the coded bitstream further comprises signaling information in the coded bitstream indicating which layers of the plurality of layers are removed or modified by the layer filter.
12 . The method of claim 10 , wherein the bitstream explicitly signals which channels are removed by the layer filter.
13 . The method of claim 10 , wherein the bitstream implicitly signals which channels are removed by the layer filter.