Electronic device, contents searching system and searching method thereof
Optical flow information is determined and used to identify a video clip from among a sequence of frames. The video clip may be identified based on motion features derived in part from the optical flow information. In some embodiments, semantic information is concatenated with the motion features derived from the optical flow information.
1 . A video processing method, comprising:
acquiring first optical flow information of at least one object, surface or edge in a video, wherein the video comprises a sequence of frame images;
extracting motion features of the at least one object, surface or edge based on the sequence of frame images and the first optical flow information;
extracting semantic features of at least one level of any first frame image of the sequence of frame images via first cascaded feature extraction layers, wherein the sequence of frame images comprises the first frame image;
determining a first video clip in which an event occurs, based on the extracted semantic features; and
generating, based on the motion features of the at least one object, surface or edge, a second video clip based on the first video clip,
wherein the second video clip is a reduced video clip of the determined first video clip such that the second video clip has a shorter length than the first video clip, and
wherein the second video clip comprises the occurred event.
2 . The method of claim 1 , wherein the extracting the motion features of the at least one object, surface, or edge comprises:
extracting motion features of at least one level of the any second frame image via second cascaded feature extraction layers, based on second optical flow information of the first frame image, wherein the sequence of frame images comprises the second frame image; and
merging, by at least one second feature extraction layer of the second cascaded feature extraction layers, motion features and semantic features of corresponding levels to provide merged features of the corresponding levels.
3 . The method of claim 2 , wherein a first number of output channels of the first cascaded feature extraction layers is less than a second number of output channels of the second cascaded feature extraction layers of a corresponding level.
4 . The method of claim 2 , wherein the merging comprises:
determining first weight information corresponding to the merged features via a weight determination network;
weighting the merged features according to the first weight information to obtain weighted merged features; and
outputting the weighted merged features.
5 . The method of claim 4 , wherein the first weight information includes second weight information of the motion features and third weight information of the semantic features, and
the weighting the merged features comprises weighting the merged features according to the second weight information and the third weight information to obtain the weighted merged features.
6 . The method of claim 1 , further comprising:
acquiring global optical flow features corresponding to the at least one object, surface, or edge based on the first optical flow information; and
merging the motion features of the at least one object, surface, or edge with the global optical flow features, to obtain second motion features of the at least one object, surface, or edge.
7 . The method of claim 1 , wherein the acquiring the first optical flow information comprises:
determining a motion area of a first frame image of the sequence of frame images;
performing a dense optical flow calculation based on the motion area and an image area of a second frame image, wherein the second frame image is a next frame image with respect to the first frame image, to obtain a dense optical flow of the any frame image; or
A) calculating the dense optical flow corresponding to the first frame image based on the first frame image and the second frame image, and
B) using a third optical flow of the motion area of the first frame image as the dense optical flow of the first frame image.
8 . The method of claim 7 , wherein the determining the motion area of the first frame image of the sequence of frame images comprises determining the motion area of the first frame image based on first brightness information of the first frame image and second brightness information of a predetermined number of frame images after the first frame image.
9 . The method of claim 8 , wherein the determining the motion area based on the first brightness information and the second brightness information comprises:
dividing the first frame image and the predetermined number of frame images into a plurality of image blocks; and
determining the motion area based on the first brightness information of the first frame image evaluated at a first position of the first frame image and based on the second brightness information evaluated at the first position in each of the predetermined number of frame images after the first frame image.
10 . The method of claim 9 , wherein the determining the motion area of the first frame image based on the first position of the first frame image and the first position in each of the predetermined number of frame images comprises:
for at least one image block, determining a first brightness mean and a first brightness variance of a Gaussian background model corresponding to the at least one image block based on the first brightness information of the at least one image block in the first frame image and the second brightness information of the predetermined number of frame images after the first frame image; and
determining whether a first pixel in the at least one image block is a motion pixel in the first frame image or whether the at least one image block corresponds to the motion area of the first frame image, based on third brightness information of the first pixel of the at least one image block in the first frame image and the Gaussian background model.
11 . The method of claim 1 , wherein the determining the first video clip in which the event occurs comprises:
determining motion state representation information of the at least one object, surface, or edge based on the motion features of the at least one object, surface, or edge; and
determining the first video clip based on the motion state representation information of the at least one object, surface, or edge.
12 . The method of claim 11 , wherein the determining the first video clip based on the motion state representation information of the at least one object, surface, or edge further comprises:
performing range merging on the motion state representation information of the at least one object, surface, or edge at least once to obtain a pyramid candidate event range including at least two levels;
determining the motion state representation information of each level of the at least two levels based on the motion state representation information of each frame image corresponding to each level of the at least two levels; and
determining the first video clip based on the motion state representation information of each level of the at least two levels.
13 . The method of claim 12 , wherein the determining the first video clip based on the motion state representation information of the at least one object, surface, or edge further comprises:
for a candidate event range of each level of the at least two levels, determining a target candidate range of a current level;
determining a first video clip range in which the event occurs in video clip ranges corresponding to the target candidate range of the current level, based on the motion state representation information of the target candidate range of the current level; and
retaining a second video clip range, wherein the second video clip range is determined in a lowest level as the first video clip in which the event occurs.
14 . The method of claim 13 , wherein for the determining the target candidate range of the current level comprises:
for a highest level in the candidate event range, determining the candidate event range of the current level as the target candidate range; and
for levels other than the highest level, determining, the candidate event range in the current level corresponding to a third video clip range in which the event occurs in a higher level, as the target candidate range.
15 . The method according to claim 1 , wherein the first optical flow information of the at least one object includes a velocity of pixel motion of at least one object.
16 . A video processing device, comprising:
one or more processors configured to:
acquire optical flow information of at least one object, surface or edge in a video, wherein the video comprises a sequence of frame images;
extract motion features of the at least one object, surface or edge based on the sequence of frame images and the optical flow information;
extract semantic features of at least one level of any first frame image of the sequence of frame images via first cascaded feature extraction layers, wherein the sequence of frame images comprises the first frame image;
determine a first video clip in which an event occurs, based on the extracted semantic features; and
generate, based on the motion features of the at least one object, surface or edge, a second video clip based on the first video clip,
wherein the second video clip is a reduced video clip of the determined first video clip such that the second video clip has a shorter length than the first video clip, and
wherein the second video clip comprises the occurred event.
17 . The video processing device according to claim 16 , wherein the optical flow information of the at least one object includes a velocity of pixel motion of at least one object.
18 . A non-transitory computer readable medium having instructions stored therein, which when executed by one or more processors of a computing device cause the computing device to:
acquire optical flow information of at least one object, surface or edge in a video, wherein the video comprises a sequence of frame images;
extract motion features of the at least one object, surface, or edge based on the sequence of frame images and the optical flow information;
extract semantic features of at least one level of any first frame image of the sequence of frame images via first cascaded feature extraction layers, wherein the sequence of frame images comprises the first frame image;
determine a first video clip in which an event occurs, based on the extracted semantic features; and
generate, based on the motion features of the at least one object, surface or edge, a second video clip based on the first video clip,
wherein the second video clip is a reduced video clip of the determined first video clip such that the second video clip has a shorter length than the first video clip, and
wherein the second video clip comprises the occurred event.
19 . The non-transitory computer readable medium according to claim 18 , wherein the optical flow information of the at least one object includes a velocity of pixel motion of at least one object.