Transmission of a collage of detected objects in a video
A computer-implemented method in a processor device of a camera, the method comprising: acquiring image frames comprising video data, communicating with a receiver device for continuously transmitting a video stream to the receiver device over a communication network; detecting at least one object in the image frames of the video data, the detected at least one object belong to at least one predetermined object class selected as surveillance target; cropping sub-areas in the image frames of the video data, the sub-areas including the at least one detected object, and adding the cropped sub-areas to image frames of the video stream being continuously transmitted as a single video stream to the receiver device.
1 . A computer-implemented method in a processor device of a network surveillance camera mounted to a building or a pole and arranged to monitor a scene, the method comprising:
acquiring, by the camera, one or more image frames comprising video data;
communicating, by a processor device of the camera, with a receiver device for continuously transmitting a video stream comprising predetermined content to the receiver device over a communication network, the receiver device comprising more powerful processing resources than the processor device;
detecting, by the processor device, more than one object in the same image frame comprising the video data, the detected more than one object belonging to at least one predetermined object class selected as surveillance target;
cropping, by the processor device, one or more sub-areas in the one or more image frames comprising the video data to provide one or more cropped sub-areas each including one of the detected objects;
determining that the total size of the one or more cropped sub-areas exceeds that size of a present image frame of the video stream;
downscaling at least the largest of the cropped sub-areas so that the resulting total size of the one or more cropped sub-areas is below the total size of the image frame of the video stream; and
adding, by the processor device, the one or more cropped sub-areas to the present image frame being continuously transmitted as a single video stream to the receiver device,
wherein the locations of the one or more cropped sub-areas in the image frames of the video stream for a given detected object is fixed and predetermined based on a priority list of positions,
wherein a remaining area of the video stream image frames, which is an area other than the one or more cropped sub-areas added to the video stream image frames, includes the predetermined content, and
wherein, in the absence of detected objects in the one or more image frames comprising the video data, the video stream image frames include only the predetermined content.
2 . The computer-implemented method of claim 1 , comprising:
performing at least one of color adjustment and tone mapping to the one or more cropped sub-areas before adding them to the video stream.
3 . The computer-implemented method of claim 1 , wherein the at least one predetermined object class comprises a class with moving objects.
4 . The computer-implemented method of claim 1 , wherein the at least one predetermined object class comprises at least one of a people class, a vehicle class, and a biometric object class.
5 . The computer-implemented method of claim 1 , wherein a resolution of the video stream image frames is fixed.
6 . The computer-implemented method of claim 1 , wherein the one or more image frames comprising video data are continuously acquired from a camera 200 and the one or more cropped sub-areas are added as a collage of crops in the video stream image frames that are continuously transmitted as the single video stream to the receiver device to reduce a required bandwidth.
7 . The computer-implemented method of claim 1 , wherein the remaining area comprises a single color, background, or static pattern to provide computational efficiency for encoding the single video stream.
8 . A control unit comprising processing circuitry in association with a network surveillance camera mounted to a building or a pole and arranged to monitor a scene, the control unit comprising:
an interface acquiring one or more image frames comprising video data,
a transmitter for communicating with a receiver device for continuously transmitting a video stream comprising predetermined content to the receiver device over a communication network, the receiver device comprising more powerful processing resources than the control unit;
a processor for:
detecting more than one object in the same image frame of the video data, the detected more than one object belonging to at least one predetermined object class selected as surveillance target,
cropping one or more sub-areas in the image frames of the video data, the one or more sub-areas each including one of the detected objects,
determining that the total size of the one or more cropped sub-areas exceeds that size of a present image frame of the video stream,
downscaling at least the largest of the one or more cropped sub-areas so that the resulting total size of the one or more cropped sub-areas is below the total size of the image frame of the video stream, and
adding the one or more cropped sub-areas to the present image frame of the video stream being continuously transmitted as a single video stream to the receiver device, wherein the locations of the one or more cropped sub-areas in the image frames of the video stream for a given detected object is fixed and predetermined based on a priority list of positions,
wherein a remaining area of the image frames of the video stream, other than the added one or more cropped sub-areas include the predetermined content, and
wherein, in the absence of detected objects in the one or more image frames comprising the video data, the image frames of the video stream include only the predetermined content.
9 . The control unit of claim 8 , further comprising:
an input and output interface to communicate with a receiver device over the communication network.
10 . A computer-implemented method on a server, the computer-implemented method comprising:
communicating with a network surveillance camera over a communication network for continuously receiving a video stream from the camera;
receiving the video stream from the camera, the video stream comprising a set of image frames including one or more cropped sub-areas of objects on a background with predetermined content,
wherein, in the absence of the one or more cropped sub-areas of the objects, the set of image frames of the video stream include only the predetermined content,
wherein the largest of the cropped sub-areas are downscaled so that the resulting total size of the cropped sub-areas is below the total size of the image frame of the video stream,
wherein the locations of the cropped sub-areas in the image frames of the video stream for a given detected object is fixed and predetermined based on a priority list of positions,
identifying the one or more cropped sub-areas in the set of image frames of the video stream;
identifying the objects in the one or more cropped areas; and
providing a signal indicating the identified objects.
11 . A camera comprising an input and output interface to communicate with a receiver device over a communication network, and a control unit of claim 8 .