Image capture device with face blurring
An image capture device captures video frames using a capture rate. The image capture device performs face detection on the video frames using a face detection rate slower than the capture rate. Temporal filtering is applied to determine the positions and sizes of the faces within the video frames in which face detection has not been performed. Tracking filtering is applied to determine the positions and sizes of the faces within the video frames in which face detection has failed. One or more of the faces within the video frames are blurred, and the blurred video frames are transmitted to a computing device.
1 . An image capture device comprising:
a housing;
an image sensor carried by the housing and configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content;
an optical element carried by the housing and configured to guide light within a field of view to the image sensor; and
one or more physical processors carried by the housing, the one or more physical processors configured by machine-readable instructions to:
capture video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;
detect the faces depicted within the video frames, the detection of the faces depicted within the video frames including determination of positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;
apply a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;
apply a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;
blur the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and
transmit the blurred video frames to a computing device;
wherein:
the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;
the detection of the faces is performed on the first video frame and the third video frame;
the detection of the faces is not performed on the second video frame;
the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame;
the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame;
the video frames includes a fourth video frame and a fifth video frame subsequent to the fourth video frame;
the detection of the faces is performed on the fourth video frame and the fifth video frame, wherein the detection of the faces fails in the fifth video frame;
the tracking filtering determines a position of a particular face in the fifth video frame based on a prior position of the particular face in the fourth video frame; and
the tracking filtering determines a size of the particular face in the fifth video frame based on a prior size of the particular face in the fourth video frame,
wherein:
a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;
a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and
the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and
the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.
2 . An image capture device comprising:
a housing;
an image sensor carried by the housing and configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content;
an optical element carried by the housing and configured to guide light within a field of view to the image sensor; and
one or more physical processors carried by the housing, the one or more physical processors configured by machine-readable instructions to:
capture video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;
detect the faces depicted within the video frames, the detection of the faces depicted within the video frames including determination of positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;
apply a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;
apply a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;
blur the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and
transmit the blurred video frames to a computing device,
wherein:
a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;
a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and
the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and
the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.
3 . The image capture device of claim 2 , wherein:
the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;
the detection of the faces is performed on the first video frame and the third video frame;
the detection of the faces is not performed on the second video frame;
the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame; and
the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame.
4 . The image capture device of claim 2 , wherein:
the video frames include a first video frame and a second video frame subsequent to the first video frame;
the detection of the faces is performed on the first video frame and the second video frame, wherein the detection of the faces fails in the second video frame;
the tracking filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame; and
the tracking filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame.
5 . The image capture device of claim 2 , wherein a target face is not blurred in a given blurred video frame.
6 . The image capture device of claim 5 , wherein a given face is identified as the target face based on a given size of the given face.
7 . The image capture device of claim 5 , wherein a given face is identified as the target face based on proximity of the given face to a bottom edge of the given blurred video frame.
8 . The image capture device of claim 5 , wherein a given face is identified as the target face based on a user selection of the given face.
9 . The image capture device of claim 2 , wherein the transmission of the blurred video frames to the computing device includes live streaming of the blurred video frames to the computing device.
10 . A method for face blurring, the method performed by an image capture device including an image sensor, an optical element, and one or more processors, the image sensor configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content, the optical element configured to guide light within a field of view to the image sensor, the method comprising:
capturing video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;
detecting the faces depicted within the video frames, detecting the faces depicted within the video frames including determining positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;
applying a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;
applying a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;
blurring the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and
transmitting the blurred video frames to a computing device,
wherein:
a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;
a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and
the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and
the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.
11 . The method of claim 10 , wherein:
the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;
the detection of the faces is performed on the first video frame and the third video frame;
the detection of the faces is not performed on the second video frame;
the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame; and
the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame.
12 . The method of claim 10 , wherein:
the video frames include a first video frame and a second video frame subsequent to the first video frame;
the detection of the faces is performed on the first video frame and the second video frame, wherein the detection of the faces fails in the second video frame;
the tracking filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame; and
the tracking filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame.
13 . The method of claim 10 , wherein a target face is not blurred in a given blurred video frame.
14 . The method of claim 13 , wherein a given face is identified as the target face based on a given size of the given face.
15 . The method of claim 13 , wherein a given face is identified as the target face based on proximity of the given face to a bottom edge of the given blurred video frame.
16 . The method of claim 13 , wherein a given face is identified as the target face based on a user selection of the given face.
17 . The method of claim 10 , wherein transmitting the blurred video frames to the computing device includes live streaming the blurred video frames to the computing device.