IP Library Granted Patent US 11,631,402
Granted Patent B2
US 11,631,402 · App. 16/869,176 · Granted Apr 18, 2023

Detection of replay attack

Inventors: John Paul Lesso (Edinburgh, GB); César Alonso (Madrid, ES)
Assignee: Cirrus Logic, Inc.
G10L15/20G06F21/32G10L15/22G10L21/0208G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,631,402
App. No.
16/869,176
Granted
Apr 18, 2023
Kind
B2
Abstract

A method of detecting a replay attack comprises: receiving an audio signal representing speech; identifying speech content present in at least a portion of the audio signal; obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified; and, for each portion of the audio signal for which speech content is identified: retrieving information about an expected frequency spectrum of the audio signal; comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum; and determining that the audio signal may result from a replay attack if a measure of a difference between the frequency spectrum of the portions of the audio signal for which speech content is identified and the respective expected frequency spectrum exceeds a threshold level.

Claims (24)

1. A method of detecting a replay attack, the method comprising:

receiving an audio signal representing speech;

identifying speech content present in at least a portion of the audio signal;

obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified, the frequency spectrum in one of a frequency band of 20-200 Hz and an ultrasonic frequency band; and

for each portion of the audio signal for which speech content is identified, providing the information about the frequency spectrum in one of the frequency band of 20-200 Hz and the ultrasonic frequency band to a trained neural network to determine a score indicative of a likelihood that the speech content is live speech.

2. A method according to claim 1 , comprising:

removing effects of a channel and/or noise from the received audio signal; and

using the audio signal after removing the effects of the channel and/or noise when obtaining the information about the frequency spectrum of each portion of the audio signal for which speech content is identified.

3. A method according to claim 1 , wherein the at least one acoustic class comprises one or more specific phonemes.

4. A method according to claim 3 , wherein the at least one acoustic class comprises plosives.

5. A method according to claim 1 , wherein the at least one acoustic class comprises fricatives.

6. A method according to claim 5 , wherein the at least one acoustic class comprises sibilants.

7. A method according to claim 1 , wherein the score indicative of the likelihood that the speech content is live speech is based on an identified acoustic class of the speech content.

8. A method according to claim 1 , wherein identifying speech content present in at least the portion of the audio signal comprises identifying speech content of at least one test acoustic class.

9. A method according to claim 8 , wherein identifying speech content of the at least one acoustic class comprises identifying a location of occurrences of the acoustic class in known speech content.

10. A method according to claim 9 , wherein the known speech content comprises a pass phrase.

11. A system for detecting a replay attack, the system comprising:

an input, for receiving an audio signal representing speech; and

a processor, wherein the processor is configured for:

identifying speech content present in at least a portion of the audio signal;

obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified, the frequency spectrum in one of a frequency band of 20-200 Hz and an ultrasonic frequency band; and

for each portion of the audio signal for which speech content is identified, providing the information about the frequency spectrum in one of the frequency band of 20-200 Hz and the ultrasonic frequency band to a trained neural network to determine a score indicative of a likelihood that the speech content is live speech.

12. A device comprising the system as claimed in claim 11 , wherein the device comprises one of: a smartphone, a tablet or laptop computer, a games console, a home control system, a home entertainment system, an in-vehicle entertainment system, or a domestic appliance.

13. A computer program product, comprising a tangible, non-transitory computer-readable medium, storing code for causing a suitable programmed processor to perform the method as claimed in claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 17, 2023
From: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
To: CIRRUS LOGIC, INC.
Reel/Frame 062390/0710 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 7, 2020
From: LESSO, JOHN PAUL; ALONSO, CÉSAR
To: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
Reel/Frame 052603/0246 →
Continuity (2)
Continuation 16050593 · Jul 31, 2018
Related Publication 20200265834A1 · Aug 20, 2020