IP Library Granted Patent US 10,692,490
Granted Patent B2
US 10,692,490 · App. 16/050,593 · Granted Jun 23, 2020

Detection of replay attack

Inventors: John Paul Lesso (Edinburgh, GB); César Alonso (Madrid, ES)
Assignee: Cirrus Logic, Inc.
G10L15/20G06F21/32G10L15/22G10L21/0208G10L15/063
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,692,490
App. No.
16/050,593
Granted
Jun 23, 2020
Kind
B2
Abstract

A method of detecting a replay attack comprises: receiving an audio signal representing speech; identifying speech content present in at least a portion of the audio signal; obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified; and, for each portion of the audio signal for which speech content is identified: retrieving information about an expected frequency spectrum of the audio signal; comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum; and determining that the audio signal may result from a replay attack if a measure of a difference between the frequency spectrum of the portions of the audio signal for which speech content is identified and the respective expected frequency spectrum exceeds a threshold level.

Claims (35)

1. A method of detecting a replay attack, the method comprising:

receiving an audio signal representing speech;

identifying speech content present in at least a portion of the audio signal;

obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified;

for each portion of the audio signal for which speech content is identified, retrieving information about an expected frequency spectrum of the audio signal;

comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum in one of a frequency band of 20-200 Hz and an ultrasonic frequency band; and

determining that the audio signal results from a replay attack if a measure of a difference between the frequency spectrum of the portions of the audio signal for which speech content is identified and the respective expected frequency spectrum exceeds a threshold level.

2. A method according to claim 1 , comprising:

removing effects of a channel and/or noise from the received audio signal; and

using the audio signal after removing the effects of the channel and/or noise when obtaining the information about the frequency spectrum of each portion of the audio signal for which speech content is identified.

3. A method according to claim 1 , wherein identifying speech content present in at least a portion of the audio signal comprises identifying at least one test acoustic class.

4. A method according to claim 3 , wherein the at least one test acoustic class comprises one or more specific phonemes.

5. A method according to claim 4 , wherein the at least one test acoustic class comprises fricatives.

6. A method according to claim 5 , wherein the at least one test acoustic class comprises sibilants.

7. A method according to claim 4 , wherein the at least one test acoustic class comprises plosives.

8. A method according to claim 3 , wherein identifying at least one test acoustic class comprises identifying a location of occurrences of the test acoustic class in known speech content.

9. A method according to claim 8 , wherein the known speech content comprises a pass phrase.

10. A method according to claim 1 , wherein comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum comprises:

comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum in a frequency band in the range of 5-20 kHz.

11. A method according to claim 1 , wherein comparing the identified parts of the audio signal with the respective retrieved information for the corresponding test acoustic class comprises:

comparing a power level in at least one frequency band of the identified parts of the audio signal with a power level in at least one corresponding frequency band of the expected spectrum of the audio signal.

12. A method according to claim 11 , wherein the measure of the difference between the identified parts of the audio signal and the respective retrieved information for the corresponding test acoustic class comprises a difference in power of greater than 1 dB.

13. A method according to claim 1 , further comprising:

performing a speaker identification process on the received audio signal; and

for each test acoustic class, retrieving information about an expected spectrum of the audio signal for a speaker identified by said speaker identification process.

14. A system for detecting a replay attack, the system comprising:

an input, for receiving an audio signal representing speech; and

a processor, wherein the processor is configured for:

identifying speech content present in at least a portion of the audio signal;

obtaining information about a frequency spectrum of each portion of the audio signal for which speech content is identified;

for each portion of the audio signal for which speech content is identified, retrieving information about an expected frequency spectrum of the audio signal;

comparing the frequency spectrum of portions of the audio signal for which speech content is identified with the respective expected frequency spectrum in one of a frequency band of 20-200 Hz and an ultrasonic frequency band; and

determining that the audio signal results from a replay attack if a measure of a difference between the frequency spectrum of the portions of the audio signal for which speech content is identified and the respective expected frequency spectrum exceeds a threshold level.

15. A device comprising a system as claimed in claim 14 , wherein the device comprises one of: a smartphone, a tablet or laptop computer, a games console, a home control system, a home entertainment system, an in-vehicle entertainment system, or a domestic appliance.

16. A computer program product, comprising a non-transitory computer-readable medium, storing code for causing a suitable programmed processor to perform a method as claimed in claim 1 .

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2020
From: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
To: CIRRUS LOGIC, INC.
Reel/Frame 052568/0880 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 16, 2018
From: LESSO, JOHN PAUL; ALONSO, CÉSAR
To: CIRRUS LOGIC INTERNATIONAL SEMICONDUCTOR LTD.
Reel/Frame 047186/0478 →
Continuity (1)
Related Publication 20200043484A1 · Feb 6, 2020