SYSTEMS AND METHODS FOR OBJECT AND EVENT DETECTION AND FEATURE-BASED RATE-DISTORTION OPTIMIZATION FOR VIDEO CODING
Systems and methods for event and object detection and annotation in the video streams may include extracting a plurality of features in a picture in a video frame, grouping at least a portion of the plurality of features into at least one object, determining a region for the at least one object, assigning object identifiers to the at least one object and encoding the object identifiers into the bitstream. Feature-based rate distortion optimization may be employed for video coding including extracting a set of features from a picture in the video, generating a relevance map for the extracted features, determining a relevance score for portions of the picture using the relevance map, and encoding the portion of the picture with a bit rate determined at least in part by the relevance score.
1 . A method of encoding video comprising:
extract a plurality of features in a picture in a video frame;
group at least a portion of the plurality of features into at least one object;
determine a region for the at least one object;
assign object identifiers to the at least one object; and
encode the object identifiers into the bitstream.
2 . The method of claim 1 , wherein a feature model is used to extract the plurality of features.
3 . The method of claim 1 , wherein the region is represented by a geometric representation.
4 . The method of claim 3 , wherein the geometric representation is one of a bounding box or a contour.
5 . The method of claim 4 , wherein the object identifiers comprise a region identifier and a label.
6 . The method of claim 5 , wherein when the geometric representation is a bounding box, the bounding box is identified by the coordinates of a specific corner and the width and height of the bounding box.
7 . The method of claim 1 , wherein an object is further evaluated over a sequence of frames to determine an event, an event identifier is associated with an object and the event identifier is encoded into the bitstream.
8 . The method of claim 1 , wherein the object identifiers are inserted into the bitstream as supplemental enhancement information.
9 . The method of claim 1 , wherein the bitstream includes a slice header and the sliced header is used to signal the presence of an object in a given slice.
10 . The method of claim 1 , further comprising:
generate a relevance map for the extracted features;
determine a relevance score for portions of the picture using the relevance map; and
encode the portion of the picture with a bit rate determined at least in part by the relevance score.
11 . A method for encoding video comprising:
extracting a set of features from a picture in the video;
generate a relevance map for the extracted features;
determine a relevance score for portions of the picture using the relevance map; and
encode the portion of the picture with a bit rate determined at least in part by the relevance score.
12 . The method of claim 11 , wherein the picture is represented by a plurality of coding units and the relevance map is determined at the coding unit level with each coding unit having a coding unit relevance score.
13 . The method of claim 12 , wherein the encoding includes allocating bit rate for each coding unit.
14 . The method of claim 13 , wherein the relevance score includes a relative relevance score for each coding unit.
15 . The method of claim 14 , wherein the encoding includes at least one of intra prediction, motion estimation, and transform quantization, and wherein the relative relevance score is used in an explicit rate distortion optimization mode to alter the encoding during at least one of the intra prediction, motion estimation, and transform quantization processes.
16 . The method of claim 15 , wherein the relative relevance score is used in a rate distortion function to determine an adjusted bitrate for each coding unit.
17 . The method of claim 11 , further comprising:
grouping at least a portion of the extracted features into at least one object;
determining a region for the at least one object;
assigning object identifiers to the at least one object; and
encoding the object identifiers into the bitstream.
18 . An encoded video bitstream comprising:
encoded video content data, the video content including at least one object identified by an encoder extracting a plurality of features of a picture in the video content;
at least one object identifier and associated object annotation; and
at least one event identifier and associated event annotation.
19 . The bitstream of claim 18 , further comprising a supplemental enhancement information (SEI) message, wherein information related to the at least one object and at least one event is signaled in the SEI message.
20 . The bitstream of claim 19 , further comprising a slice header, wherein information related to the at least one object and at least one event in a video slice is signaled in the slice header.