IP Library Granted Patent US 12711793
Granted Patent B1
US 12711793 · App. 19/330,538 · Granted Aug 18, 2026

Deepfake detection in data at rest and real-time network traffic

Inventors: Siying Yang (Saratoga, CA); Krishna Narayanaswamy (Saratoga, CA); Sanjay Beri (Los Altos, CA); Yihua Liao (Palo Alto, CA); Jason Bryslawskyj (San Diego, CA); Xinjun Zhang (Fremont, CA)
Assignee: NETSKOPE, INC.
G06V20/95G06V10/32G06V10/70G06V20/46
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12711793
App. No.
19/330,538
Granted
Aug 18, 2026
Kind
B1
Abstract

A detecting deepfake system for detecting deepfake in live network traffic communications and stored media. Visual data and audio data are extracted from a plurality of users of an enterprise. Digital fingerprints are generated from the extracted visual data and audio data. The digital fingerprints are matched with the pre-stored fingerprints to produce an output. Further, the visual data and audio data are analyzed by Artificial Intelligence (AI)-detection model. A deepfake detection score is generated based on the analysis by the AI-detection model. The output from the matching of the digital fingerprints with the pre-stored fingerprints and the deepfake detection score are combined to determine a level of deepfake in the extracted visual data and audio data. The policies are applied based on the level of deepfake detected in the visual data and the audio data.

Claims (80)

1 . A method for detecting deepfake in live network traffic communications and stored media, the method comprising:

extracting visual data and audio of a plurality of users of an enterprise;

generating digital fingerprints from the extracted visual data and the audio;

matching the generated digital fingerprints with pre-stored normalized fingerprints to produce an output of the matching;

analyzing the extracted visual data and the audio using an artificial intelligence (AI)-detection model;

generating a deepfake detection score based on the analysis by the AI-detection model;

generating a fingerprint mismatch flag based on the output of the matching;

generating a deepfake machine learning flag based on the deepfake detection score;

combining the output of the matching and the deepfake detection score to determine a level of deepfake in the extracted visual data and audio;

applying a plurality of policies in real-time based on the level of deepfake in the extracted visual data and the audio, wherein applying the plurality of policies includes blocking access of enterprise content to a user based on a user role in the enterprise, upon triggering of the fingerprint mismatch flag and the deepfake machine learning flag; and

where in the plurality of policies are predefined in a policy store of the enterprise based on the user role and when the user role corresponds to a Trainee in the enterprise and the level of deepfake is indicated by:

(i) the fingerprint mismatch flag, a policy comprises holding access of the enterprise content to the user and re-authenticating the user,

(ii) the deepfake machine learning flag, the policy comprises holding access of the enterprise content until an administrator of the enterprise contacts the user, or

(iii) both the fingerprint mismatch flag and the deepfake machine learning flag, the policy comprises blocking user access of the enterprise content and sending an alert to a user device of the user and an administration of the enterprise.

2 . The method for detecting deepfake in live network traffic communications and stored media of claim 1 , further comprising:

pre-storing the normalized fingerprints includes:

receiving images and audio recordings of the plurality of users;

pre-processing the received images and the audio recordings using a trained model;

generating fingerprints for the pre-processed images and the audio recordings;

normalizing the fingerprints with corresponding images and the audio recordings of the plurality of users; and

pre-storing the normalized fingerprints of the plurality of users in a database.

3 . The method for detecting deepfake in live network traffic communications and stored media of claim 2 , wherein the pre-processing is performed by:

cropping and resampling the received images, and

extracting one or more audio features from the audio recordings.

4 . The method for detecting deepfake in live network traffic communications and stored media of claim 3 , wherein extracting of the one or more audio features is performed by segmenting the audio recordings into key audio characteristics.

5 . The method for detecting deepfake in live network traffic communications and stored media of claim 2 , wherein the trained model is a custom object detection model.

6 . The method for detecting deepfake in live network traffic communications and stored media of claim 1 , wherein extracting the visual data and the audio of the plurality of users includes extracting video frames and the audio by monitoring the live network traffic, and the live network traffic includes a live communication between two or more of the plurality of users.

7 . The method for detecting deepfake in live network traffic communications and stored media of claim 1 , wherein extracting the visual data and the audio of the plurality of users is performed by extracting images and the audio from an enterprise data or the stored media, and the enterprise data comprises documents and videos uploaded by the plurality of users of the enterprise.

8 . The method for detecting deepfake in live network traffic communications and stored media of claim 1 , wherein applying the plurality of policies comprises generating real-time security alerts and reports, and flagging the plurality of users.

9 . The method for detecting deepfake in live network traffic communications and stored media of claim 1 , wherein the plurality of policies are predefined in a policy store of the enterprise based on the user role and when the user role corresponds to Senior Management in the enterprise and the level of deepfake is indicated by triggering of (i) the fingerprint mismatch flag, (ii) the deepfake machine learning flag, or (iii) both the fingerprint mismatch flag and the deepfake machine learning flag, a policy comprises blocking access of the enterprise content to the user and contacting the user immediately.

10 . A deepfake detection system for detecting deepfake in live network traffic communications and stored media, the deepfake detection system comprising:

one or more processors, and

memory coupled with the one or more processors, the memory configured to store instructions that when executed by the one or more processors cause the one or more processors to:

extract, by a mid-link server, a visual data and an audio of a plurality of users of an enterprise;

generate, by the mid-link server, digital fingerprints from the extracted visual data and the audio;

match, by the mid-link server, the generated digital fingerprints with pre-stored normalized fingerprints to produce an output of the match;

analyze, by the mid-link server, the extracted visual data and the audio using an artificial intelligence (AI)-detection model;

generate, by the mid-link server, a deepfake detection score based on the analysis by the AI-detection model;

generate, by the mid-link server, a fingerprint mismatch flag based on the output of the matching;

generate, by the mid-link server, a deepfake machine learning flag based on the deepfake detection score;

combine, by the mid-link server, the output of the match and the deepfake detection score to determine a level of deepfake in the extracted visual data and the audio;

apply, by the mid-link server, a plurality of policies in real-time based on the level of deepfake in the extracted visual data and the audio, wherein applying the plurality of policies includes blocking access of enterprise content to a user based on a user role in the enterprise, upon triggering of the fingerprint mismatch flag and the deepfake machine learning flag; and

where in the plurality of policies are predefined in a policy store of the enterprise based on the user role and when the user role corresponds to a Trainee in the enterprise and the level of deepfake is indicated by:

(i) the fingerprint mismatch flag, a policy comprises holding access of the enterprise content to the user and re-authenticating the user,

(ii) the deepfake machine learning flag, the policy comprises holding access of the enterprise content until an administrator of the enterprise contacts the user, or

(iii) both the fingerprint mismatch flag and the deepfake machine learning flag, the policy comprises blocking user access of the enterprise content and sending an alert to a user device of the user and an administration of the enterprise.

11 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 10 , wherein the one or more processors are further configured to:

receive images and audio recordings of the plurality of users;

pre-process the received images and the audio recordings using a trained model;

generate fingerprints for the pre-processed images and the audio recordings;

normalize the fingerprints with corresponding images and the audio recordings of the plurality of users; and

pre-store the normalized fingerprints of the plurality of users in a database of the mid-link server.

12 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 11 , wherein the pre-processing the images and the audio recordings is performed by cropping and resampling the received images and extracting one or more audio features from the audio recordings.

13 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 12 , wherein the extracting of the one or more audio features is performed by segmenting the audio recordings into key audio characteristics.

14 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 11 , wherein the trained model is a custom object detection model.

15 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 10 , wherein extraction of the visual data and the audio of the plurality of users is performed by extracting video frames and the audio by monitoring live network traffic, and the live network traffic includes a live communication between two or more of the users.

16 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 10 , wherein extraction of the visual data and the audio of the plurality of users is performed by extracting images and the audio from enterprise data and/or the stored media, and the enterprise data comprises documents and videos uploaded by the plurality of users of the enterprise.

17 . The deepfake detection system for detecting deepfake in live network traffic communications and stored media of claim 10 , wherein the plurality of policies includes generating real-time security alerts and reports, and by flagging the plurality of users.

18 . A non-transitory computer-readable storage medium having stored thereon instructions for causing at least one computer to facilitate a method for detecting deepfake in live network traffic communications and stored media, the method comprising:

extracting visual data and audio of a plurality of users of an enterprise;

generating digital fingerprints from the extracted visual data and the audio;

matching the generated digital fingerprints with pre-stored normalized fingerprints to produce an output of the matching;

analyzing the extracted visual data and the audio using an artificial intelligence (AI)-detection model;

generating a deepfake detection score based on the analysis by the AI-detection model;

generating a fingerprint mismatch flag based on the output of the matching;

generating a deepfake machine learning flag based on the deepfake detection score;

combining the output of the matching and the deepfake detection score to determine a level of deepfake in the extracted visual data and the audio;

applying a plurality of policies in real-time based on the level of deepfake in the extracted visual data and the audio, wherein applying the plurality of policies includes blocking access of enterprise content to a user based on a user role in the enterprise, upon triggering of the fingerprint mismatch flag and the deepfake machine learning flag; and

wherein the plurality of policies are predefined in a policy store of the enterprise based on the user role and when the user role corresponds to a Trainee in the enterprise and the level of deepfake is indicated by:

(i) the fingerprint mismatch flag, a policy comprises holding access of the enterprise content to the user and re-authenticating the user,

(ii) the deepfake machine learning flag, the policy comprises holding access of the enterprise content until an administrator of the enterprise contacts the user, or

(iii) both the fingerprint mismatch flag and the deepfake machine learning flag, the policy comprises blocking user access of the enterprise content and sending an alert to a user device of the user and an administration of the enterprise.

19 . The non-transitory computer-readable storage medium having stored thereon instructions for causing at least one computer to facilitate a method for detecting deepfake in live network traffic communications and stored media of claim 18 , wherein the method further comprises:

receiving images and audio recordings of the plurality of users;

pre-processing the received images and the audio recordings using a trained model;

generating fingerprints for the pre-processed images and the audio recordings;

normalizing the fingerprints with corresponding images and the audio recordings of the plurality of users; and

pre-storing the normalized fingerprints of the plurality of users in a database.

20 . The non-transitory computer-readable storage medium having stored thereon instructions for causing at least one computer to facilitate a method for detecting deepfake in live network traffic communications and stored media of claim 18 , wherein extracting the visual data and the audio of the plurality of users includes extracting video frames and the audio by monitoring live network traffic, and the live network traffic includes a live communication between two or more of the plurality of users.

21 . The non-transitory computer-readable storage medium having stored thereon instructions for causing at least one computer to facilitate a method for detecting deepfake in live network traffic communications and stored media of claim 18 , wherein extracting the visual data and the audio of the plurality of users includes extracting images and the audio from enterprise data and/or the stored media, and the enterprise data comprises documents and videos uploaded by the plurality of users of the enterprise.