IP Library Granted Patent US 12,586,592
Granted Patent B2
US 12,586,592 · App. 17/882,447 · Granted Mar 24, 2026

Methods and apparatus for generating audio fingerprints for calls using power spectral density values

Inventors: Shrirang Jangi (Acton, MA); Vilas Bhade (Acton, MA)
Assignee: Ribbon Communications Operating Company, Inc.
G10L19/02H03M7/3073
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,586,592
App. No.
17/882,447
Granted
Mar 24, 2026
Kind
B2
Abstract

The present invention relates to methods, systems, and apparatus for processing audio signals. An exemplary method embodiment includes the steps of: removing silence from an audio signal; determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1; identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks; and generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks. In various embodiments, audio fingerprints are generated from an audio signal of call and then used to determine if the call is a robocall or SPAM call.

Claims (113)

1 . A method of processing an audio signal comprising:

removing silence from the audio signal, said audio signal being from a first call;

after the silence has been removed from the audio signal, determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1;

identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks;

generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks, said first audio fingerprint being an ordered set of information including: a first frequency bin value, a second frequency bin value, and a delta time, said first frequency bin value corresponding to a first dominant frequency peak, said first dominant frequency peak being one of the identified dominant frequency peaks, said second frequency bin value corresponding to a second dominant frequency peak, said second dominant frequency peak being one of the identified dominant frequency peaks, said first dominant frequency peak and said second dominant frequency peak being different dominant frequency peaks, said delta time being a time difference between a second identified location in the audio signal corresponding to the second dominant frequency peak and the first identified location in the audio signal corresponding to the first dominant frequency peak;

generating a first set of fuzzy audio fingerprints from the first audio fingerprint, said generating the first set of fuzzy audio fingerprints from the first audio fingerprint including modifying one or more of: the first frequency bin value or the second frequency bin value, said first frequency bin value identifying a first range of frequencies corresponding to a first frequency bin, said second frequency bin value identifying a second range of frequencies corresponding to a second frequency bin; and

using the first audio fingerprint and the first set of fuzzy audio fingerprints to determine whether the first call is a robocall; and

wherein said modifying includes performing one of the following:

(i) changing the first frequency bin value to a third frequency bin value, said third frequency bin value identifying a third range of frequencies corresponding to a third frequency bin, said third range of frequencies being different than said first range of frequencies,

(ii) changing the second frequency bin value to a fourth frequency bin value, said fourth frequency bin value identifying a fourth range of frequencies corresponding to a fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies, or

(iii) changing the first frequency bin value to the third frequency bin value, said third frequency bin value identifying the third range of frequencies corresponding to the third frequency bin, said third range of frequencies being different than said first range of frequencies and changing the second frequency bin value to the fourth frequency bin value, said fourth frequency bin value identifying the fourth range of frequencies corresponding to the fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies; and

wherein said plurality of N different frequency bins includes said first frequency bin, said second frequency bin, said third frequency bin, and said fourth frequency bin.

2 . The method of claim 1 ,

wherein said identifying a plurality of dominant frequency peaks based on the determined power spectral density values includes: identifying for each of the plurality of time segments of the audio signal a set of frequency bins with the highest power spectral density values above a first threshold value, said set of frequency bins having M or fewer entries, where M is less than N, and where M is an integer; and

wherein said identified positions in the audio signal corresponding to the identified peaks are times corresponding to the time segments in which the identified peaks appear.

3 . The method of claim 2 ,

wherein said first audio fingerprint further includes a first time;

wherein said first time is a first identified location in the audio signal corresponding to the first dominant frequency peak, said first time being a time corresponding to a first time segment of the plurality of time segments, said first dominant frequency peak appearing in said first time segment.

4 . The method of claim 2 , further comprising:

generating a first fingerprint-set for the first call, said generating a first fingerprint-set for the first call including generating a plurality of audio fingerprints from the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks, said first audio fingerprint being one of said plurality of audio fingerprints.

5 . The method of claim 4 , further comprising:

generating a fingerprint-set dictionary for the first call without using a hash function for audio fingerprints in the fingerprint-set dictionary, said fingerprint-set dictionary including a key value identifying individual fingerprints for the first call, and a list of time entries identifying individual fingerprints in the fingerprint-set for the call by the time in the audio signal to which the individual fingerprint corresponds.

6 . The method of claim 1 , further comprising:

performing, prior to said identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks, a filtering operation on the audio signal to remove high frequency signals above a first frequency threshold level.

7 . The method of claim 1 , further comprising:

quantizing the determined power spectral density (PSD) values of the audio signal.

8 . The method of claim 7 ,

wherein said step of identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks includes:

generating a spectrogram of power spectral density values based on: (i) said determined power spectral density values of the audio signal, (ii) the plurality of N different frequency bins, and (iii) the plurality of time segments; and

applying a maximal filter to said spectrogram of power spectral density values to locate frequency peaks in said spectrogram.

9 . The method of claim 8 ,

wherein said step of identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks further includes:

applying an erosion filter to said spectrogram of power spectral density values after applying said maximal filter.

10 . The method of claim 1 ,

wherein said audio signal is digitally encoded audio; and

wherein said method of processing said audio signal further includes prior to determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins:

decoding said digitally encoded audio; and

converting a sampling rate for said audio to an 8 KHz sampling rate when said sampling rate is not 8 KHz.

11 . The method of claim 1 ,

wherein said step of removing silence from the audio signal includes: (i) using voice activation detection to determine portions of the audio signal with a signal level less than a first threshold value, said portions of the audio signal being less than the first threshold value being determined to be silence; and (ii) removing portions of the audio signal determined to be silence; and

wherein said step of removing silence from the audio signal is performed using a voice activated detector.

12 . The method of claim 1 , further compromising:

receiving, by a Session Border Controller, the audio signal of the first call;

wherein said first audio fingerprint is generated in real-time by the Session Border Controller as said audio signal passes through said Session Border Controller.

13 . The method of claim 1 , wherein said generating the first set of fuzzy audio fingerprints from the first audio fingerprint further includes:

modifying the delta time of the first audio fingerprint.

14 . The method of claim 1 , wherein said generating the first set of fuzzy audio fingerprints from the first audio fingerprint including modifying one or more of: the first frequency bin value or the second frequency bin value includes:

modifying the first frequency bin value in a logarithmic manner based on the first frequency range to which the first frequency bin value corresponds.

15 . A system for processing an audio signal comprising:

an audio fingerprinting device including a first processor, said first processor controlling the audio fingerprinting device to perform the following operations:

removing silence from the audio signal, said audio signal being from a first call;

after the silence has been removed from the audio signal, determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1;

identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks;

generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks, said first audio fingerprint being an ordered set of information including: a first frequency bin value, a second frequency bin value, and a delta time, said first frequency bin value corresponding to a first dominant frequency peak, said first dominant frequency peak being one of the identified dominant frequency peaks, said second frequency bin value corresponding to a second dominant frequency peak, said second dominant frequency peak being one of the identified dominant frequency peaks, said first dominant frequency peak and said second dominant frequency peak being different dominant frequency peaks, said delta time being a time difference between a second identified location in the audio signal corresponding to the second dominant frequency peak and the first identified location in the audio signal corresponding to the first dominant frequency peak;

generating a first set of fuzzy audio fingerprints from the first audio fingerprint by modifying one or more of: the first frequency bin value or the second frequency bin value, said first frequency bin value identifying a first range of frequencies corresponding to a first frequency bin, said second frequency bin value identifying a second range of frequencies corresponding to a second frequency bin; and

using the first audio fingerprint and the first set of fuzzy audio fingerprints to determine whether the first call is a robocall; and

wherein said modifying includes performing one of the following:

(i) changing the first frequency bin value to a third frequency bin value, said third frequency bin value identifying a third range of frequencies corresponding to a third frequency bin, said third range of frequencies being different than said first range of frequencies,

(ii) changing the second frequency bin value to a fourth frequency bin value, said fourth frequency bin value identifying a fourth range of frequencies corresponding to a fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies, or

(iii) changing the first frequency bin value to the third frequency bin value, said third frequency bin value identifying the third range of frequencies corresponding to the third frequency bin, said third range of frequencies being different than said first range of frequencies and changing the second frequency bin value to the fourth frequency bin value, said fourth frequency bin value identifying the fourth range of frequencies corresponding to the fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies; and

wherein said plurality of N different frequency bins includes said first frequency bin, said second frequency bin, said third frequency bin, and said fourth frequency bin.

16 . The system of claim 15 ,

wherein said identifying a plurality of dominant frequency peaks based on the determined power spectral density values includes: identifying for each of the plurality of time segments of the audio signal a set of frequency bins with the highest power spectral density values above a first threshold value, said set of frequency bins having M or fewer entries, where M is less than N, and where M is an integer; and

wherein said identified positions in the audio signal corresponding to the identified peaks are times corresponding to the time segments in which the identified peaks appear.

17 . The system of claim 15 ,

wherein said first processor further controls the audio fingerprinting device to perform the following operations:

performing, prior to said identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks, a filtering operation on the audio signal to remove high frequency signals above a first frequency threshold level.

18 . The system of claim 15 , wherein said first processor further controls the audio fingerprinting device to perform the following operations:

quantizing the determined power spectral density (PSD) values of the audio signal.

19 . The system of claim 18 ,

wherein said operation of identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks includes:

generating a spectrogram of power spectral density values based on: (i) said determined power spectral density values of the audio signal, (ii) the set of frequency bins, and (iii) the plurality of time segments; and

applying a maximal filter to said spectrogram of power spectral density values to locate frequency peaks in said spectrogram.

20 . The system of claim 19 ,

wherein said operation of identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks further includes:

applying an erosion filter to said spectrogram of power spectral density values after applying said maximal filter.

21 . The system of claim 15 ,

wherein said audio signal is digitally encoded audio; and

wherein said first processor controls the audio fingerprinting device prior to determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins to perform the following operations:

decoding said digitally encoded audio; and

converting a sampling rate for said audio to an 8 KHz sampling rate when said sampling rate is not 8 KHz.

22 . A non-transitory computer readable medium including a first set of computer executable instructions which when executed by a processor of a computing device cause the computing device to:

remove silence from an audio signal, said audio signal being from a first call;

determine, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1;

identify (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks;

generate a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks, said first audio fingerprint being an ordered set of information including: a first frequency bin value, a second frequency bin value, and a delta time, said first frequency bin value corresponding to a first dominant frequency peak, said first dominant frequency peak being one of the identified dominant frequency peaks, said second frequency bin value corresponding to a second dominant frequency peak, said second dominant frequency peak being one of the identified dominant frequency peaks, said first dominant frequency peak and said second dominant frequency peak being different dominant frequency peaks, said delta time being a time difference between a second identified location in the audio signal corresponding to the second dominant frequency peak and the first identified location in the audio signal corresponding to the first dominant frequency peak;

generate a first set of fuzzy audio fingerprints from the first audio fingerprint by modifying one or more of: the first frequency bin value or the second frequency bin value, said first frequency bin value identifying a first range of frequencies corresponding to a first frequency bin, said second frequency bin value identifying a second range of frequencies corresponding to a second frequency bin; and

use the first audio fingerprint and the first set of fuzzy audio fingerprints to determine whether the first call is a robocall; and

wherein said modifying includes performing one of the following:

(i) changing the first frequency bin value to a third frequency bin value, said third frequency bin value identifying a third range of frequencies corresponding to a third frequency bin, said third range of frequencies being different than said first range of frequencies,

(ii) changing the second frequency bin value to a fourth frequency bin value, said fourth frequency bin value identifying a fourth range of frequencies corresponding to a fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies, or

(iii) changing the first frequency bin value to the third frequency bin value, said third frequency bin value identifying the third range of frequencies corresponding to the third frequency bin, said third range of frequencies being different than said first range of frequencies and changing the second frequency bin value to the fourth frequency bin value, said fourth frequency bin value identifying the fourth range of frequencies corresponding to the fourth frequency bin, said fourth range of frequencies being different than said second range of frequencies; and

wherein said plurality of N different frequency bins includes said first frequency bin, said second frequency bin, said third frequency bin, and said fourth frequency bin.

23 . A method of processing an audio signal comprising:

removing silence from the audio signal, said audio signal being from a first call;

after the silence has been removed from the audio signal, determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins, N being an integer greater than 1;

identifying (i) a plurality of dominant frequency peaks based on the determined power spectral density values, and (ii) positions in the audio signal corresponding to the identified peaks;

generating a first audio fingerprint from at least some of the identified plurality of dominant frequency peaks and the identified positions in the audio signal corresponding to the identified peaks, said first audio fingerprint being an ordered set of information including: a first frequency bin value, a second frequency bin value, and a delta time, said first frequency bin value corresponding to a first dominant frequency peak, said first dominant frequency peak being one of the identified dominant frequency peaks, said second frequency bin value corresponding to a second dominant frequency peak, said second dominant frequency peak being one of the identified dominant frequency peaks, said first dominant frequency peak and said second dominant frequency peak being different dominant frequency peaks, said delta time being a time difference between a second identified location in the audio signal corresponding to the second dominant frequency peak and the first identified location in the audio signal corresponding to the first dominant frequency peak;

generate a first set of fuzzy audio fingerprints from the first audio fingerprint, said generating the first set of fuzzy audio fingerprints from the first audio fingerprint including modifying one or more of the following of the first audio fingerprint: the first frequency bin value or the second frequency bin value; and

using the first audio fingerprint and the first set of fuzzy audio fingerprints to determine whether the first call is a robocall; and

wherein said generating the first set of fuzzy audio fingerprints from the first audio fingerprint including modifying one or more of the following of the first audio fingerprint: the first frequency bin value or the second frequency bin value includes:

generating a first fuzzy audio fingerprint and a second fuzzy audio fingerprint, said generating the first fuzzy audio fingerprint and the second fuzzy audio fingerprint including:

generating the first fuzzy audio fingerprint by adding 1 to the first frequency bin value when the first frequency bin value is a value greater than 1 and less than 64; and

generating the second fuzzy audio fingerprint by subtracting 1 from the first frequency bin value when the first frequency bin value is a value greater than 1 and less than 64; and

generating the first fuzzy audio fingerprint by adding 2 to the first frequency bin value when the first frequency bin value is a value equal to or greater than 64 and less than 128; and

generating the second fuzzy audio fingerprint by subtracting 2 from the first frequency bin value when the first frequency bin value is a value equal to or greater than 64 and less than 128; and

generating the first fuzzy audio fingerprint by adding 4 to the first frequency bin value when the first frequency bin value is a value equal to or greater than 128 and less than 256; and

generating the second fuzzy audio fingerprint by subtracting 4 from the first frequency bin value when the first frequency bin value is a value equal to or greater than 128 and less than 256.

24 . The method of claim 23 ,

wherein a 1024 point Fast Fourier Transform (FFT) is utilized for said determining, for a plurality of time segments of the audio signal, power spectral density values of the audio signal for each of a plurality of N different frequency bins;

wherein each of the time segments of the audio signal have been sampled at an 8000 Hz sample rate;

wherein the first frequency bin value is a quantized frequency value in the range of (0, 256) corresponding to 0-2000 Hz; and

wherein the second frequency bin value is a quantized frequency value in the range of (0, 256) corresponding to 0-2000 Hz.

Assignments (3)
SHORT-FORM PATENTS SECURITY AGREEMENT Recorded Sep 5, 2024
From: RIBBON COMMUNICATIONS OPERATING COMPANY, INC.
To: HPS INVESTMENT PARTNERS, LLC, AS ADMINISTRATIVE AGENT
Reel/Frame 068857/0351 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 16, 2022
From: BHADE, VILAS
To: RIBBON COMMUNICATIONS OPERATING COMPANY, INC.
Reel/Frame 060820/0364 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 15, 2022
From: JANGI, SHRIRANG
To: RIBBON COMMUNICATIONS OPERATING COMPANY, INC.
Reel/Frame 060812/0047 →
Continuity (2)
Provisional Application 63346989 · May 30, 2022
Related Publication 20230386484A1 · Nov 30, 2023
References Cited (50)
US 8160877B1 · Nucci · 2012 [cited by examiner]
US 8934383B1 · McConkey · 2015 [cited by examiner]
US 9065409B2 · Sandgren · 2015 [cited by examiner]
US 9437204B2 · Grancharov · 2016 [cited by examiner]
US 9762731B1 · Cohen · 2017 [cited by applicant]
US 10089994B1 · Radzishevsky · 2018 [cited by examiner]
US 10110741B1 · Cohen · 2018 [cited by examiner]
US 10121488B1 · Drews · 2018 [cited by examiner]
US 10666792B1 · Marzuoli · 2020 [cited by examiner]
US 10798241B1 · Quilici · 2020 [cited by examiner]
US 10878838B1 · Mayol Cuevas · 2020 [cited by examiner]
US 11153436B2 · Aravena · 2021 [cited by examiner]
US 11343374B1 · Rolia et al. · 2022 [cited by applicant]
US 11978461B1 · Radzishevsky · 2024 [cited by examiner]
US 20030176934A1 · Gopalan · 2003 [cited by examiner]
US 20060143190A1 · Haitsma et al. · 2006 [cited by applicant]
US 20070192390A1 · Wang · 2007 [cited by examiner]
US 20080178288A1 · Alperovitch et al. · 2008 [cited by applicant]
US 20110173208A1 · Vogel · 2011 [cited by examiner]
US 20120215853A1 · Sundaram et al. · 2012 [cited by applicant]
US 20130259211A1 · Vlack · 2013 [cited by examiner]
US 20150279381A1 · Goesnar · 2015 [cited by examiner]
US 20180007199A1 · Quilici et al. · 2018 [cited by applicant]
US 20180294959A1 · Traynor · 2018 [cited by examiner]
US 20200257722A1 · Zhang et al. · 2020 [cited by applicant]
US 20200296510A1 · Li · 2020 [cited by examiner]
US 20210092223A1 · Gallagher et al. · 2021 [cited by applicant]
US 20210136200A1 · Li · 2021 [cited by examiner]
US 20220086175A1 · Bharrat et al. · 2022 [cited by applicant]
US 20220247866A1 · Xiao-Devins et al. · 2022 [cited by applicant]
US 20220270017A1 · Singh · 2022 [cited by examiner]
US 20220399945A1 · Frenkel · 2022 [cited by examiner]
US 20230388414A1 · Bharrat et al. · 2023 [cited by applicant]
CN 110602303A · 2019 [cited by applicant]
EP 1667106A1 · 2006 [cited by examiner]
EP 3023884A1 · 2016 [cited by examiner]
EP 3324607A1 · 2018 [cited by applicant]
EP 3937474A1 · 2022 [cited by examiner]
GB 2455505A · 2009 [cited by applicant]
“Spectral density”, Wikipedia, 11 Pages, downloaded Sep. 4, 2024. [cited by examiner]
“Spectral density”, Wikipedia, 11 Pages, downloaded Mar. 19, 2025 (Year: 2025). [cited by examiner]
Richard Harter, The minimum on a sliding window algorithm, downloaded from Internet address http://richardhartersworld.com/slidingmin/ on Aug. 5, 2022, 4 pages. [cited by applicant]
scipy.ndimage.maximum_filter, SciPy.org, SciPy v1.9.0 Manual, downloaded from Internet address https://docs.scipy.org/doc/scipy/reference/generated/scipy.ndimage.maximum_filter.html on Aug. 5, 2022, 3 pages. [cited by applicant]
scipy.ndimage.maximum_filter1d, SciPy.org, SciPy v1.9.0 Manual, downloaded from Internet address https://docs.scipy.org/doc/scipy/reference/generated/scipy.ndimage.maximum_filter1d.html on Aug. 5, 2022, 3 pages. [cited by applicant]
scipy.ndimage.filters.maximum_filter, SciPy.org, SciPy v0.14.0 Reference Guide, downloaded from Internet address https://docs.scipy.org/doc/scipy-0.14.0/reference/generated/scipy.ndimage.filters.maximum_filter.html on A… [cited by applicant]
Multidimensional image processing, SciPy.org, SciPy v1.9.0 Manual, downloaded from Internet address https://docs.scipy.org/doc/scipy/tutorial/ndimage.html on Aug. 5, 2022, 27 pages. [cited by applicant]
Kim Hyoung-Gook et al., Robust audio fingerprinting using peak-pair-based hash of non-repeating foreground audio in a real. environment, Cluster Computing, Baltzer Science Publishers, Bussum, NL, vol. 19, No. 1, Jan. 2,… [cited by applicant]
Notification of Transmittal of the International Search Report and the Written Opinion of the International Searching Authority, International Search Report and Written Opinion of the International Searching Authority f… [cited by applicant]
Multidimensional image processing, SciPy.org, SciPy v0.14.0 Reference Guide, downloaded from Internet address https://docs.scipy.org/doc/scipy-0.14.0/reference/tutorial/ndimage.html on Aug. 16, 2022, 28 pages. [cited by applicant]
Will Drevo, Audio Fingerprinting with Python and Numpy, Nov. 15, 2013, downloaded from Internet address https://willdrevo.com/fingerprinting-and-audio-recognition-with-python/ on Aug. 8, 2022, 15 pages. [cited by applicant]