Enhancing deepfake detection using forensic models
There is provided a computer implemented method of detecting a deepfake video, comprising: feeding a video into at least one forensic model, obtaining at least one indication of at least one property of a person or object depicted in the video as an outcome of the at least one forensic model, comparing the at least one indication to at least one ground truth obtained by applying the at least one forensic model to at least one authentic video depicting the person or object and/or by searching a ground truth dataset, in response to a mismatch between the at least one indication and the at least one ground truth, increasing likelihood that the video is deepfake, and in response to a match between the at least one indication and the at least one ground truth, increasing likelihood that the video is authentic.
1 . A computer implemented method of detecting a deepfake video, comprises:
feeding a video into at least one forensic model;
obtaining at least one indication of at least one property of a person or object depicted in the video as an outcome of the at least one forensic model;
comparing the at least one indication to at least one ground truth obtained by applying the at least one forensic model to at least one authentic video depicting the person or object and/or by searching a ground truth dataset;
in response to a mismatch between the at least one indication and the at least one ground truth, increasing likelihood that the video is deepfake; and
in response to a match between the at least one indication and the at least one ground truth, increasing likelihood that the video is authentic,
wherein the at least one forensic model is not trained to detect whether the video is a deepfake video.
2 . The computer implemented method of claim 1 , wherein the at least one ground truth for comparison is found by searching an index of the ground truth dataset of records according to the person or object to find a matching record that includes the ground truth.
3 . The computer implemented method of claim 2 , wherein the at least one ground truth comparison is found by searching to find the matching record without applying the at least one forensic model to the at least one authentic video to obtain the at least one ground truth.
4 . The computer implemented method of claim 2 , further comprising creating the ground truth dataset by:
searching for a plurality of authentic videos that include different people and/or objects;
applying the forensic model to each authentic video to obtain a ground truth for each person and/or object;
creating a record for each person and/or object that includes the ground truth corresponding to the person and/or object; and
adding the record to the index according to the person and/or object corresponding to the ground truth of the record.
5 . The computer implemented method of claim 1 , wherein the at least one forensic model comprises at least one voice forensic model trained to analyze voice to determine at least one property of the speaker.
6 . The computer implemented method of claim 5 , wherein the at least one property is selected from: height, weight, office size, medical condition, native tongue, and tiredness.
7 . The computer implemented method of claim 1 , wherein the at least one property of the person or object comprises at least one physical property of the person or object.
8 . The computer implemented method of claim 1 , wherein a forensic model is assigned a classification according to predicted changes in outcome over time, and the at least one property outcome of the forensic model is evaluated according to the classification.
9 . A computer implemented method of detecting a deepfake video, comprises:
feeding a video into at least one forensic model;
obtaining at least one indication of at least one property of a person or object depicted in the video as an outcome of the at least one forensic model;
comparing the at least one indication to at least one ground truth obtained by applying the at least one forensic model to at least one authentic video depicting the person or object and/or by searching a ground truth dataset;
in response to a mismatch between the at least one indication and the at least one ground truth, increasing likelihood that the video is deepfake; and
in response to a match between the at least one indication and the at least one ground truth, increasing likelihood that the video is authentic;
wherein a forensic model is assigned a classification according to predicted changes in outcome over time, and the at least one property outcome of the forensic model is evaluated according to the classification;
wherein the classification and evaluation are selected from: (i) predicted to stay the same, and at least one property used as is, (ii) predicted to vary within a range, and verifying whether the at least one property falls within the range of the ground truth, and (iii) predicted to last a predefined amount of time, and not using the at least one property when the predefined amount of time has elapsed.
10 . The computer implemented method of claim 1 , further comprising generating a report indicating at least one of: the match or mismatch, a value of the increase in likelihood that the video is deepfake or authentic, and in response to a mismatch providing the at least one indication outcome of the at least one forensic model and the at least one ground truth.
11 . A system for detecting a deepfake video, comprises:
at least one processor executing a code for:
feeding a video into at least one forensic model;
obtaining at least one indication of at least one property of a person or object depicted in the video as an outcome of the at least one forensic model;
comparing the at least one indication to at least one ground truth obtained by applying the at least one forensic model to at least one authentic video depicting the person or object and/or by searching a ground truth dataset;
in response to a mismatch between the at least one indication and the at least one ground truth, increasing likelihood that the video is deepfake; and
in response to a match between the at least one indication and the at least one ground truth, increasing likelihood that the video is authentic,
wherein the at least one forensic model is not trained to detect whether the video is a deepfake video.
12 . A system for detecting a deepfake video, comprises:
at least one processor executing a code for:
feeding a video into at least one forensic model;
obtaining at least one indication of at least one property of a person or object depicted in the video as an outcome of the at least one forensic model;
comparing the at least one indication to at least one ground truth obtained by applying the at least one forensic model to at least one authentic video depicting the person or object and/or by searching a ground truth dataset;
in response to a mismatch between the at least one indication and the at least one ground truth, increasing likelihood that the video is deepfake; and
in response to a match between the at least one indication and the at least one ground truth, increasing likelihood that the video is authentic,
wherein the at least one ground truth for comparison is found by searching an index of the ground truth dataset of records according to the person or object to find a matching record that includes the ground truth, without applying the at least one forensic model to the at least one authentic video to obtain the at least one ground truth.
13 . The system of claim 12 , further comprising code for creating the ground truth dataset by:
searching for a plurality of authentic videos that include different people and/or objects;
applying the forensic model to each authentic video to obtain a ground truth for each person and/or object;
creating a record for each person and/or object that includes the ground truth corresponding to the person and/or object; and
adding the record to the index according to the person and/or object corresponding to the ground truth of the record.
14 . The system of claim 11 , wherein the at least one forensic model comprises at least one voice forensic model trained to analyze voice to determine at least one property of the speaker.
15 . The system of claim 11 , wherein the at least one property of the person or object comprises at least one physical property of the person or object.
16 . The system of claim 11 , wherein a forensic model is assigned a classification according to predicted changes in outcome over time, and the at least one property outcome of the forensic model is evaluated according to the classification.
17 . The system of claim 11 , further comprising code for generating a report indicating at least one of: the match or mismatch, a value of the increase in likelihood that the video is deepfake or authentic, and in response to a mismatch providing the at least one indication outcome of the at least one forensic model and the at least one ground truth.