IP Library Granted Patent US 12688559
Granted Patent B1
US 12688559 · App. 18/758,315 · Granted Jul 21, 2026

Image capture device with face blurring

Inventors: Marc Lebrun (Issy-les-Moulineaux, FR); Moctar Mounirou Arouna LM (Montrouge, FR); Jean-Marc Thiesse (Saint-Cyr-l'école, FR); Maxime Bichon (San Mateo, CA); Martin Raoul (Elancourt, FR)
Assignee: GoPro, Inc.
G06T5/70G06T7/74G06V40/161G06T2207/10016G06T2207/20092G06T2207/30201
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12688559
App. No.
18/758,315
Granted
Jul 21, 2026
Kind
B1
Abstract

An image capture device captures video frames using a capture rate. The image capture device performs face detection on the video frames using a face detection rate slower than the capture rate. Temporal filtering is applied to determine the positions and sizes of the faces within the video frames in which face detection has not been performed. Tracking filtering is applied to determine the positions and sizes of the faces within the video frames in which face detection has failed. One or more of the faces within the video frames are blurred, and the blurred video frames are transmitted to a computing device.

Claims (86)

1 . An image capture device comprising:

a housing;

an image sensor carried by the housing and configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content;

an optical element carried by the housing and configured to guide light within a field of view to the image sensor; and

one or more physical processors carried by the housing, the one or more physical processors configured by machine-readable instructions to:

capture video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;

detect the faces depicted within the video frames, the detection of the faces depicted within the video frames including determination of positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;

apply a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;

apply a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;

blur the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and

transmit the blurred video frames to a computing device;

wherein:

the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;

the detection of the faces is performed on the first video frame and the third video frame;

the detection of the faces is not performed on the second video frame;

the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame;

the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame;

the video frames includes a fourth video frame and a fifth video frame subsequent to the fourth video frame;

the detection of the faces is performed on the fourth video frame and the fifth video frame, wherein the detection of the faces fails in the fifth video frame;

the tracking filtering determines a position of a particular face in the fifth video frame based on a prior position of the particular face in the fourth video frame; and

the tracking filtering determines a size of the particular face in the fifth video frame based on a prior size of the particular face in the fourth video frame,

wherein:

a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;

a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and

the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and

the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.

2 . An image capture device comprising:

a housing;

an image sensor carried by the housing and configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content;

an optical element carried by the housing and configured to guide light within a field of view to the image sensor; and

one or more physical processors carried by the housing, the one or more physical processors configured by machine-readable instructions to:

capture video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;

detect the faces depicted within the video frames, the detection of the faces depicted within the video frames including determination of positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;

apply a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;

apply a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;

blur the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and

transmit the blurred video frames to a computing device,

wherein:

a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;

a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and

the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and

the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.

3 . The image capture device of claim 2 , wherein:

the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;

the detection of the faces is performed on the first video frame and the third video frame;

the detection of the faces is not performed on the second video frame;

the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame; and

the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame.

4 . The image capture device of claim 2 , wherein:

the video frames include a first video frame and a second video frame subsequent to the first video frame;

the detection of the faces is performed on the first video frame and the second video frame, wherein the detection of the faces fails in the second video frame;

the tracking filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame; and

the tracking filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame.

5 . The image capture device of claim 2 , wherein a target face is not blurred in a given blurred video frame.

6 . The image capture device of claim 5 , wherein a given face is identified as the target face based on a given size of the given face.

7 . The image capture device of claim 5 , wherein a given face is identified as the target face based on proximity of the given face to a bottom edge of the given blurred video frame.

8 . The image capture device of claim 5 , wherein a given face is identified as the target face based on a user selection of the given face.

9 . The image capture device of claim 2 , wherein the transmission of the blurred video frames to the computing device includes live streaming of the blurred video frames to the computing device.

10 . A method for face blurring, the method performed by an image capture device including an image sensor, an optical element, and one or more processors, the image sensor configured to generate a visual output signal conveying visual information based on light that becomes incident thereon, the visual information defining visual content, the optical element configured to guide light within a field of view to the image sensor, the method comprising:

capturing video frames based on the visual content, the video frames depicting faces, wherein a capture rate defines a rate at which the video frames are captured;

detecting the faces depicted within the video frames, detecting the faces depicted within the video frames including determining positions and sizes of the faces, wherein a face detection rate defines a rate at which the faces are detected within the video frames, the face detection rate slower than the capture rate;

applying a temporal filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has not been performed;

applying a tracking filtering to determine the positions and the sizes of the faces within the video frames in which the detection of the faces has failed;

blurring the video frames based on the positions and the sizes of the faces within the video frames, wherein one or more of the faces within the video frames are blurred; and

transmitting the blurred video frames to a computing device,

wherein:

a first face is detected in a first video frame, the first face having a first position and a first size in the first video frame;

a second face is detected in a second video frame, the second face having a second position and a second size in the second video frame; and

the first face is determined to be the same as the second face based on a distance between the first position of the first face in the first video frame and the second position of the second face in the second video frame being less than a threshold distance; and

the threshold distance is determined based on the first size of the first face in the first video frame and the second size of the second face in the second video frame.

11 . The method of claim 10 , wherein:

the video frames include a first video frame, a second video frame subsequent to the first video frame, and a third video frame subsequent to the second video frame;

the detection of the faces is performed on the first video frame and the third video frame;

the detection of the faces is not performed on the second video frame;

the temporal filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame and a subsequent position of the given face in the third video frame; and

the temporal filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame and a subsequent size of the given face in the third video frame.

12 . The method of claim 10 , wherein:

the video frames include a first video frame and a second video frame subsequent to the first video frame;

the detection of the faces is performed on the first video frame and the second video frame, wherein the detection of the faces fails in the second video frame;

the tracking filtering determines a position of a given face in the second video frame based on a prior position of the given face in the first video frame; and

the tracking filtering determines a size of the given face in the second video frame based on a prior size of the given face in the first video frame.

13 . The method of claim 10 , wherein a target face is not blurred in a given blurred video frame.

14 . The method of claim 13 , wherein a given face is identified as the target face based on a given size of the given face.

15 . The method of claim 13 , wherein a given face is identified as the target face based on proximity of the given face to a bottom edge of the given blurred video frame.

16 . The method of claim 13 , wherein a given face is identified as the target face based on a user selection of the given face.

17 . The method of claim 10 , wherein transmitting the blurred video frames to the computing device includes live streaming the blurred video frames to the computing device.