IP Library Patent Application 18388364
Patent Application
App. No. 18/388,364

DEEPFAKE DETECTION

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/388,364
Abstract

Disclosed are systems and methods including software processes executed by a server that detect audio-based synthetic speech (“deepfakes”) in a call conversation. The server applies an NLP engine to transcribe call audio and analyze the text for anomalous patterns to detect synthetic speech. Additionally or alternatively, the server executes a voice “liveness” detection system for detecting machine speech, such as synthetic speech or replayed speech. The system performs phrase repetition detection, background change detection, and passive voice liveness detection in call audio signals to detect liveness of a speech utterance. An automated model update module allows the liveness detection model to adapt to new types of presentation attacks, based on the human provided feedback.

Claims (29)

1 . A computer-implemented method for detecting machine-based speech in calls, comprising:

obtaining, by a computer, inbound audio data for a call, including a plurality of speech segments corresponding to a dialogue between a caller and an agent;

detecting, by the computer, from the plurality of speech segments of the inbound audio data, a speech region including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller;

determining, by the computer, a response delay between the first segment corresponding to the agent and the second segment corresponding to the caller;

identifying, by the computer, the caller as a deepfake, in response to determining that the response delay fails to satisfy an expected response time for a human speaker.

2 . The method of claim 1 , wherein determining the response delay includes identifying, by the computer, based on the plurality of timestamps, the response delay corresponding to a time difference between the first speech segment corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller.

3 . The method of claim 1 , wherein detecting the speech region further includes detecting, by the computer, a plurality of timestamps defining the first speech segment and the second speech segment in the speech region.

4 . The method of claim 1 , wherein determining the response delay includes determining, by the computer, a plurality of statistical measures of response delays based on a plurality of timestamps derived from the plurality of speech segments of the inbound audio data.

5 . The method of claim 4 , wherein the plurality of statistical measures further comprises at least one of: (i) a running variance, (ii) a running inter-quartile range, or (iii) a running mean.

6 . The method of claim 1 , wherein determining the response delay includes adjusting, by the computer, the response delay based on a context of the dialogue between the caller and the agent.

7 . The method of claim 1 , further comprising extracting, by the computer, a transcription containing text in chronological sequence of each caller speech segment and each agent speech segment.

8 . The method of claim 1 , further comprising determining, by the computer, the expected response time for the human speaker based on historical data of audio data between callers and agents.

9 . The method of claim 1 , further comprising identifying, by the computer, the caller as human, in response to determining that the response delay satisfies an expected response time for a human speaker.

10 . The method of claim 1 , further comprising generating, by the computer, an indication, for a user interface, indicating the caller as one of the deepfake or human.

11 . A system for detecting machine-based speech in calls, comprising:

a computer having one or more processors and configured to:

obtain inbound audio data for a call, including a plurality of speech segments corresponding to a dialogue between a caller and an agent;

detect from the plurality of speech segments of the inbound audio data, a speech region including a first speech segment corresponding to the agent and a second speech segment corresponding to the caller;

determine a response delay between the first segment corresponding to the agent and the second segment corresponding to the caller; and

identify the caller as a deepfake, in response to determining that the response delay fails to satisfy an expected response time for a human speaker.

12 . The system of claim 11 , wherein, when determining the response delay, the computer is further configured to identify, based on the plurality of timestamps, the response delay corresponding to a time difference between the first speech segment corresponding to a question by the agent and the second speech segment corresponding to an answer by the caller.

13 . The system of claim 11 , wherein, where detecting the speech region, the computer is further configured to detect a plurality of timestamps defining the first speech segment and the second speech segment in the speech region.

14 . The system of claim 11 , wherein, when determining the response delay, the computer is further configured to determine a plurality of statistical measures of response delays based on a plurality of timestamps derived from the plurality of speech segments of the inbound audio data.

15 . The system of claim 14 , wherein, where the plurality of statistical measures, the computer is further configured to include at least one of: (i) a running variance, (ii) a running inter-quartile range, or (iii) a running mean.

16 . The system of claim 11 , wherein, where determining the response delay the computer is further configured to adjust the response delay based on a context of the dialogue between the caller and the agent.

17 . The system of claim 11 , the computer is further configured to extract a transcription containing text in chronological sequence of each caller speech segment and each agent speech segment.

18 . The system of claim 11 , the computer is further configured to determine the expected response time for the human speaker based on historical data of audio data between callers and agents.

19 . The system of claim 11 , the computer is further configured to identify the caller as human, in response to determining that the response delay satisfies an expected response time for a human speaker.

20 . The system of claim 11 , the computer is further configured to provide an indication of an identification of the caller as one of the deepfake or human.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 4, 2024
From: ALTAF, UMAIR; PERI, SAI PRADEEP; PHATELA, LAKSHAY; GUPTA, PAYAS; SUN, YITAO; AFANASEVA, SVETLANA; PATIL, KAILASH; KHOURY, ELIE; MAGNETTA, BRADLEY; BALASUBRAMANIYAN, VIJAY; CHEN, TIANXIANG
To: PINDROP SECURITY, INC.
Reel/Frame 067618/0883 →