Joint count and flow analysis for video crowd scenes
View Patent ↗A system for flow-count includes one or more cameras and a processing system. The one more cameras are configured to capture image data comprising a sequence of frames. The processing system is configured to extract spatial features and temporal features based on the sequence of frames; construct, based on the spatial features and temporal features, a density map and/or a flow map corresponding to a frame of the sequence of frames, wherein the density map comprises density values for the plurality of pixels, wherein the flow map comprises two-dimensional (2D) vectors for the plurality of pixels, and wherein a 2D vector indicates a speed and direction for a pixel of the plurality of pixels in the frame; and determine, based on the density map and/or the flow map, a count, a distribution, and/or movement of objects corresponding to the frame.
1 . A system, comprising:
one or more cameras configured to capture image data, the image data comprising a video stream that comprises a sequence of frames, each frame comprising a plurality of pixels; and
a processing system configured to:
extract spatial features and temporal features based on the sequence of frames;
construct, based on the spatial features and temporal features, a flow map corresponding to a frame of the sequence of frames, wherein the flow map comprises two-dimensional (2D) vectors for the plurality of pixels, wherein each pixel of the flow map corresponds to a respective 2D vector, wherein each respective 2D vector indicates a speed and direction for a respective pixel of the plurality of pixels in the frame, and wherein the speed and direction for the respective pixel represents a speed and direction of a crowd movement;
determine, based on the flow map, movement of objects corresponding to the frame, including the crowd movement; and
perform crowd management based on the crowd movement contained in the frame of the sequence of frames.
2 . The system of claim 1 , wherein to extract the spatial features and temporal features based on the sequence of frames, the processing system is further configured to:
extract, using an encoder comprising a transformer with a plurality of transformer layers and a convolutional network connected in series, the spatial features at different scales for each frame of the sequence of frames, wherein the spatial features at a given scale are produced by a corresponding transformer layer of the plurality of transformer layers; and
extract the temporal features from the sequence of frames, wherein the convolutional network is configured to process the spatial features output by the last transformer layer of the plurality of transformer lavers so as to produce the temporal features.
3 . The system of claim 2 , wherein the construction, based on the spatial features and temporal features, of the flow map corresponding to the frame of the sequence of frames is based on the spatial features at the different scales and the temporal features.
4 . The system of claim 2 , wherein the spatial features at the different scales and the temporal features for the sequence of frames are extracted in sequence.
5 . The system of claim 2 , wherein the spatial features at the different scales and the temporal features for the sequence of frames are extracted concurrently.
6 . The system of claim 2 , wherein the processing system is further configured to:
construct, based on the spatial features at the different scales and the temporal features, a density map corresponding to the frame of the sequence of frames, wherein the density map comprises density values for the plurality of pixels; and
determine, based on the flow map and the density map, at least one of a count or a distribution of the objects corresponding to the frame.
7 . The system of claim 6 , wherein the processing system is further configured to:
filter, based on the spatial features at the different scales and the temporal features, data irrelevant to the objects contained in the sequence of video frames,
wherein the constructed density map and flow map associated with the frame do not include the irrelevant data.
8 . The system of claim 7 , wherein the objects are persons contained in the sequence of video frames, wherein the irrelevant data includes non-human objects, and wherein the constructed density map and flow map do not include non-human objects.
9 . The system of claim 1 , wherein the processing system is further configured to:
compute a first loss between the constructed flow map and a ground-truth flow map for the frame of the sequence of frames; and
train, based on the first loss, a model for determining the count, distribution, and movement of the objects contained in the frame of the sequence of frames.
10 . The system of claim 9 , wherein the processing system is further configured to:
construct, based on the spatial features and temporal features, a density map corresponding to the frame of the sequence of frames, wherein the density map comprises density values for the plurality of pixels;
compute a second loss between the constructed density map and a ground-truth density map for the frame of the sequence of frames; and
train, based on the first loss and the second loss, the model for determining at least one of a count, a distribution, or the movement of the objects contained in the frame of the sequence of frames.
11 . The system of claim 1 , wherein the processing system is further configured to:
perform at least one of crowd management, service optimization, or security monitoring based on the count, distribution, and movement of the objects contained in the frame of the sequence of frames.
12 . A method, comprising:
obtaining, by a processing system, a video stream comprising a sequence of frames, each frame comprising a plurality of pixels;
extracting, by the processing system, spatial features and temporal features based on the sequence of frames;
constructing, by the processing system, based on the spatial features and the temporal features, a flow map corresponding to a frame of the sequence of frames, wherein the flow map comprises two-dimensional (2D) vectors for the plurality of pixels, wherein each pixel of the flow map corresponds to a respective 2D vector, wherein each respective 2D vector indicates a speed and direction for a pixel of the plurality of pixels in the frame, and wherein the speed and direction for the respective pixel represents a speed and direction of a crowd movement;
determining, by the processing system, based on the flow map, movement of objects corresponding to the frame, including the crowd movement; and
performing, by the processing system, crowd management based on the crowd movement contained in the frame of the sequence of frames.
13 . The method of claim 12 , wherein extracting the spatial features and temporal features based on the sequence of frames further comprises:
extracting, by the processing system implementing an encoder comprising a transformer with a plurality of transformer lavers and a convolutional network connected in series, the spatial features at different scales for each frame of the sequence of frames, wherein the spatial features at a given scale are produced by a corresponding transformer layer of the plurality of transformer layers; and
extracting, by the processing system, the temporal features from the sequence of frames, wherein the convolutional network is configured to process the spatial features output by the last transformer layer of the plurality of transformer lavers so as to produce the temporal features.
14 . The method of claim 13 , wherein the constructing, by the processing system, based on the spatial features and temporal features, the flow map corresponding to the frame of the sequence of frames is based on the spatial features at the different scales and the temporal features.
15 . The method of claim 13 , wherein the spatial features at the different scales and the temporal features for the sequence of frames are extracted in sequence.
16 . The method of claim 13 , wherein the spatial features at the different scales and the temporal features for the sequence of frames are extracted concurrently.
17 . The method of claim 13 , further comprising:
constructing, based on the spatial features at the different scales and the temporal features, a density map corresponding to the frame of the sequence of frames, wherein the density map comprises density values for the plurality of pixels, and
determining, based on the flow map and the density map, at least one of a count or a distribution of the objects corresponding to the frame.
18 . The method of claim 17 , further comprising:
filtering, by the processing system, based on the spatial features at the different scales and the temporal features, data irrelevant to the objects contained in the sequence of video frames, wherein the constructed density map and flow map associated with the frame do not include the irrelevant data.
19 . The method of claim 18 , wherein the objects are persons contained in the sequence of video frames, wherein the irrelevant data includes non-human objects, and wherein the constructed density map and flow map do not include non-human objects.
20 . A non-transitory computer-readable medium having processor-executable instructions stored thereon, wherein the processor-executable instructions, when executed, facilitate performance of the following:
obtaining a video stream comprising a sequence of frames, each frame comprising a plurality of pixels;
extracting spatial features and temporal features based on the sequence of frames;
constructing, based on the spatial features and temporal features, a flow map corresponding to a frame of the sequence of frames, wherein the flow map comprises two-dimensional (2D) vectors for the plurality of pixels, wherein each pixel of the flow map corresponds to a respective 2D vector, wherein each respective 2D vector indicates a speed and direction for a pixel of the plurality of pixels in the frame, and wherein the speed and direction for the respective pixel represents a speed and direction of a crowd movement;
determining, based on the flow map, a count, a distribution, and movement of objects corresponding to the frame, including the crowd movement; and
performing crowd management based on the crowd movement contained in the frame of the sequence of frames.