IP Library Patent Application 18598595
Patent Application
App. No. 18/598,595

PRESENTATION ATTACKS IN REVERBERANT CONDITIONS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/598,595
Abstract

Embodiments include a computing device that executes software routines and/or one or more machine-learning architectures including obtaining training audio signals having corresponding training impulse responses associated with reverberation degradation, training a machine-learning model of a presentation attack detection engine to generate one or more acoustic parameters by executing the presentation attack detection engine using the training impulse responses of the training audio signals and a loss function, obtaining an audio signal having an acoustic impulse response associated with reverberation degradation caused by one or more rooms, generating the one or more acoustic parameters for the audio signal by executing the machine-learning model using the audio signal as input, and generating an attack score for the audio signal based upon the one or more parameters generated by the machine-learning model.

Claims (28)

1 . A computer-implemented method comprising:

obtaining, by a computer, a plurality of training audio signals having corresponding training acoustic impulse responses including at least one single-room training acoustic impulse response and at least one multi-room training acoustic impulse response;

training, by the computer, a parameter estimation machine-learning model of a presentation attack detection (PAD) engine to estimate one or more acoustic parameters by executing the parameter estimation machine-learning model of the PAD engine using the training acoustic impulse responses of the plurality of training audio signals and a loss function;

obtaining, by the computer, an audio signal having an acoustic impulse response caused by one or more rooms; and

generating, by the computer, the one or more acoustic parameters for the audio signal by executing the parameter estimation machine-learning model using the audio signal as input.

2 . The method according to claim 1 , further comprising generating, by the computer, an attack score for the audio signal based upon the one or more acoustic parameters by executing a PAD scoring machine-learning model of the PAD engine, the attack score indicating a likelihood that the acoustic impulse response of the audio signal is caused by two or more rooms.

3 . The computer-implemented method of claim 2 , further comprising detecting, by the computer, that the audio signal is a presentation attack in response to determining that the attack score satisfies an attack score threshold.

4 . The computer-implemented method of claim 2 , wherein generating the attack score for the audio signal includes determining, by the computer, whether the acoustic impulse response of the audio signal is consistent throughout the audio signal.

5 . The computer-implemented method of claim 1 , further comprising segmenting, by the computer, the audio signal into a plurality of frames, wherein the computer executes the PAD engine using each frame of the audio signal as input to the presentation attack engine.

6 . The computer-implemented method of claim 1 , further comprising transforming, by the computer, the audio signal into a spectral domain representation, wherein the computer executes the PAD engine using the spectral representation of the audio signal as input to the PAD engine.

7 . The computer-implemented method of claim 1 , wherein the one or more acoustic parameters include at least one of: spectral standard deviation, late reverberation onset, SNR, or energy decay curve.

8 . The computer-implemented method of claim 1 , further comprising extracting, by the computer, a feature vector for the one or more acoustic parameters of the acoustic impulse response of the audio signal.

9 . The computer-implemented method of claim 1 , wherein obtaining the plurality of training audio signals includes generating, by the computer, one or more of the plurality of training audio signals according to a simulated environment.

10 . The computer-implemented method of claim 1 , wherein the acoustic impulse response of the audio signal is a type of acoustic impulse response not represented in the training audio signals.

11 . A non-transitory computer-readable medium comprising machine-executable instructions which, when executed by one or more processors, cause the one or more processors to:

obtain a plurality of training audio signals having corresponding training acoustic impulse responses including at least one single-room training acoustic impulse response and at least one multi-room training acoustic impulse response;

train a parameter estimation machine-learning model of a presentation attack detection (PAD) engine to estimate one or more acoustic parameters by executing the parameter estimation machine-learning model of the PAD engine using the training acoustic impulse responses of the plurality of training audio signals and a loss function;

obtain an audio signal having an acoustic impulse response caused by one or more rooms; and

generate the one or more acoustic parameters for the audio signal by executing the machine-learning model using the audio signal as input.

12 . The non-transitory computer-readable medium of claim 11 , wherein the instructions further cause the one or more processors to generate an attack score for the audio signal based upon the one or more acoustic parameters by executing a PAD scoring machine-learning model of the PAD engine, the attack score indicating a likelihood that the acoustic impulse response of the audio signal is caused by two or more rooms.

13 . The non-transitory computer-readable medium of claim 12 , wherein the instructions further cause the one or more processors to detect that the audio signal is a presentation attack in response to determining that the attack score satisfies an attack score threshold.

14 . The non-transitory computer-readable medium of claim 11 , wherein, when generating the attack score for the audio signal, the instructions further cause the one or more processors to determine whether the acoustic impulse response of the audio signal is consistent throughout the audio signal.

15 . The non-transitory computer-readable medium of claim 11 , wherein the instructions further cause the one or more processors to segment the audio signal into a plurality of frames, wherein the computer executes the PAD engine using each frame of the audio signal as input to the presentation attack engine.

16 . The non-transitory computer-readable medium of claim 11 , wherein the instructions further cause the one or more processors to transform the audio signal into a spectral domain representation, wherein the computer executes the PAD engine using the spectral representation of the audio signal as input to the PAD engine.

17 . The non-transitory computer-readable medium of claim 11 , wherein the one or more acoustic parameters include at least one of: spectral standard deviation, late reverberation onset, SNR, or energy decay curve.

18 . The non-transitory computer-readable medium of claim 11 , wherein the instructions further cause the one or more processors to extract a feature vector for the one or more acoustic parameters of the acoustic impulse response of the audio signal.

19 . The non-transitory computer-readable medium of claim 11 , wherein, when obtaining the plurality of training audio signals, the instructions further cause the one or more processors to generate one or more of the plurality of training audio signals in a simulated environment.

20 . The non-transitory computer-readable medium of claim 11 , wherein the acoustic impulse response of the audio signal is a type of acoustic impulse response not represented in the training audio signals.

Assignments (2)
SECURITY INTEREST Recorded Jun 26, 2024
From: PINDROP SECURITY, INC.
To: HERCULES CAPITAL, INC., AS AGENT
Reel/Frame 067867/0860 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 7, 2024
From: GAUBITCH, NIKOLAY; LOONEY, DAVID
To: PINDROP SECURITY, INC.
Reel/Frame 066686/0478 →