Encoding device and method for utility-driven video compression
View Patent ↗An encoding device for utility-driven video compression, includes circuitry configured to accept an input video having a first data volume, identify at least a feature of interest in the input video, generate an output video, wherein the output video contains a second data volume that is less than the first data volume and the output video preserves the at least a feature of interest, and encode a bitstream using the output video.
1 . An encoding device for utility-driven video compression, the encoding device comprising circuitry configured to:
accept an input video having a first data volume comprising at least one frame, each frame comprising a plurality of blocks;
identify at least one region of interest in a first spatial region of the first data volume, the first spatial region comprising a first plurality of included blocks;
identify at least a feature reflecting a structural or content attribute of the at least one region of interest;
identify a region of exclusion identified as a region containing blocks associated with at least one object detected in the input video, the at least one object including at least one of a human face, a person, a vehicle, or a license plate, the at least one object being a feature to be selectively excluded from the bitstream;
generate an output video, wherein:
the output video contains a second data volume that is less than the first data volume; and
the output video preserves the at least a feature of the region of interest; and
generate an encoded bitstream including the output video and signaling information including a header indicating a block-level association of blocks to regions in the second data volume.
2 . The encoding device of claim 1 , wherein the encoding device is further configured to accept the input video by:
receiving an input bitstream; and
decoding the input video from the input bitstream.
3 . The encoder of claim 1 , further configured to identify the at least a feature of interest by:
receiving at least a supervised annotation indicating the at least a feature of interest; and
identifying the at least a feature of interest using the at least a supervised annotation.
4 . The encoder of claim 1 , further configured to identify the at least a feature of interest using a neural network.
5 . The encoder of claim 4 , further configured to:
receive an output bitstream recipient characteristic; and
select the neural network from a plurality of neural networks as a function of the output bitstream recipient characteristic.
6 . The encoding device of claim 1 , wherein the at least a feature of interest includes at least one of an audio feature, a visual feature, and an element of metadata.
7 . The encoding device of claim 1 , wherein the header includes a privacy flag for each block and the flag for the blocks of a region of exclusion are set to indicate exclusion from the encoded bitstream.
8 . The encoding device of claim 1 , wherein the-block-level association includes a spatial region identifier.
9 . The encoding device of claim 1 , wherein encoding the bitstream further comprises compressing the output video.
10 . The encoding device of claim 1 , wherein the header is a slice header.
11 . The encoding device of claim 1 , wherein the signaling information is provided in a sequence parameter set (SPS) in a slice header of the bitstream.
12 . The encoding device of claim 1 , wherein the feature describes at least one of spatial and temporal characteristics of a frame or group of frames in the output video.
13 . The encoding device of claim 1 , wherein signaling information is provided as supplemental information in the bitstream.
14 . The encoding device of claim 13 , wherein the second feature is motion information of blocks sufficient to perform motion analysis.
15 . A method for utility-driven video compression, the method comprising:
accepting, by an encoding device, an input video having a first data volume comprising at least one frame, each frame comprising a plurality of blocks;
identify, by the encoding device, at least one region of interest in a first spatial region of the first data volume, the first spatial region comprising a first plurality of included blocks;
identifying, by the encoding device, at least a feature reflecting a structural or content attribute of the at least one region of interest;
identifying, by the encoding device, a region of exclusion identified as a region containing blocks associated with at least one object detected in the input video, the at least one object including at least one of a human face, a person, a vehicle, or a license plate, the at least one object being a feature to be selectively excluded from the bitstream;
generating, by the encoding device, an output video, wherein:
the output video contains a second data volume that is less than the first data volume; and
the output video preserves the at least a feature of the region of interest; and
generating, by the encoding device, an encoded bitstream including the output video and signaling information including a header indicating a block-level association of blocks to regions in the second data volume.
16 . The method of claim 15 , wherein accepting the input video further comprises:
receiving an input bitstream; and
decoding the input video from the input bitstream.
17 . The encoder of claim 15 , wherein identifying the at least a feature of interest further comprises:
receiving at least a supervised annotation indicating the at least a feature of interest; and
identifying the at least a feature of interest using the at least a supervised annotation.
18 . The encoder of claim 15 , wherein identifying the at least a feature of interest further comprises identifying the at least a feature of interest using a neural network.
19 . The encoder of claim 18 further comprising:
receiving an output bitstream recipient characteristic; and
selecting the neural network from a plurality of neural networks as a function of the output bitstream recipient characteristic.
20 . The method of claim 15 , wherein the at least a feature of interest includes at least one of an audio feature, a visual feature, and an element of metadata.
21 . The method of claim 15 , wherein the header includes a privacy flag for each block and the flag for the blocks of a region of exclusion are set to indicate exclusion from the encoded bitstream.
22 . The method of claim 15 , wherein the block-level association includes a spatial region identifier.
23 . The method of claim 15 , wherein encoding the bitstream further comprises compressing the output video.
24 . An encoding device for utility-driven video compression, the encoding device comprising circuitry configured to:
accept an input video comprising a plurality of frames and having a first data volume;
identify in the input video at least a first feature relevant to a first application and a second feature relevant to a second application, wherein the first and second feature each concerning a specific structural or content attribute of the input video;
generate a first encoded application specific sub-bitstream for human viewing consisting essentially of a subset of frames including the first feature and having a volume less than the first data volume; and
generate a second encoded application specific sub-bitstream for a machine-based application consisting essentially of a subset of frames including the second feature and having a volume less than the first data volume, the second encoded application specific sub-bitstream comprising motion information of blocks in the input video and excluding texture data and color data of the blocks, the second encoded application specific sub-bitstream being incapable of reconstructing pixels of the input video necessary for human viewing.