IP Library Granted Patent US 12,598,253
Granted Patent B2
US 12,598,253 · App. 18/501,134 · Granted Apr 7, 2026

Systems and methods for media analysis for call state detection

Inventors: Kosta Pribić (Zagreb, HR); Sead Delalić (Sarajevo, BA); Hadžem Hadžić (Sarajevo, BA); Kemal Altwlkany (Sarajevo, BA); Amar Kurić (Sarajevo, BA); Luka Čižmek (Zagreb, HR); Ivan Ðumlija (Velika Gorica, HR)
Assignee: Infobip Ltd.
H04M3/4365H04M3/4936
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,598,253
App. No.
18/501,134
Granted
Apr 7, 2026
Kind
B2
Abstract

A method includes: initiating a call to a telephone number based on a request from a call originator; generating a fingerprint of an audio sample of the call; matching the fingerprint to a record in a fingerprint database; determining, based on the matched record, whether the recorded call is an actionable machine response, a non-actionable machine response, or a human response, as a detected response; providing the detected response to the call originator; receiving, based on the providing the detected response, an action from the call originator; and performing the action associated with the call.

Claims (69)

1 . A method comprising:

initiating a call to a telephone number based on a request from a call originator;

generating a fingerprint of an audio sample of the call based on determining whether the audio sample includes silence that exceeds a threshold;

matching the fingerprint to a record in a fingerprint database, the fingerprint database including records of non-actionable machine responses representing voicemails;

determining, based on the matched record, whether the call includes an actionable machine response, a non-actionable machine response, or a human response, as a detected response;

providing the detected response to the call originator;

receiving, based on the providing the detected response, an action from the call originator; and

performing the action associated with the call.

2 . The method of claim 1 , wherein the generating the fingerprint of the audio sample of the call includes:

generating a first fingerprint of a first audio sample of the call, and

generating a second fingerprint of a second audio sample of the call.

3 . The method of claim 2 , wherein the second audio sample of the call includes the first audio sample of the call, such that the second audio sample is a longer duration sample than the first audio sample.

4 . The method of claim 1 , wherein the generating the fingerprint of the audio sample of the call includes:

detecting the silence in the audio sample of the call;

determining that the detected silence in the audio sample of the call exceeds the threshold;

retrieving a different audio sample of the call; and

generating the fingerprint of the different audio sample of the call.

5 . The method of claim 1 , wherein the generating the fingerprint of the audio sample of the call includes:

classifying the audio sample of the call as an announcement, among classifications including music, ringing, and announcement; and

generating the fingerprint of the audio sample of the call based on classifying the audio sample.

6 . The method of claim 5 , wherein the classifying the audio sample includes using one or more machine learning models.

7 . The method of claim 1 , wherein the matching the fingerprint to the record in the fingerprint database includes:

determining that a similarity of the fingerprint to the record exceeds a threshold.

8 . The method of claim 1 , wherein the matching the fingerprint to the record in the fingerprint database includes:

determining that a similarity of the fingerprint, as a first fingerprint, to the record is below a threshold;

generating a second fingerprint of an extended audio sample of the call, wherein the extended audio sample includes the audio sample, as a first audio sample, used for the first fingerprint merged with a second audio sample of the call subsequent to the first audio sample; and

determining that a similarity of the second fingerprint to the record is above the threshold.

9 . The method of claim 1 , wherein the providing the detected response to the call originator includes:

providing the detected response to the call originator as a status code.

10 . The method of claim 1 , wherein the performing the action associated with the call includes:

disconnecting the call, retrying the call on one of a same channel, route, or gateway, or retrying the call on a different channel, route, or gateway.

11 . A method comprising:

recording audio samples of calls to telephone numbers as a dataset;

generating corresponding text transcriptions of the recorded audio samples;

generating corresponding fingerprints of the recorded audio samples;

clustering the recorded audio samples based on the fingerprints and text transcriptions, as groups;

assigning a category to each group;

generating a fingerprint database with the groups and the assigned categories, wherein the assigned categories include a non-actionable machine response representing voicemail; and

incorporating fingerprints of new audio samples of calls to telephone numbers into the generated fingerprint database based on determining whether at least one of the new audio samples includes silence that exceeds a threshold.

12 . The method of claim 11 , further comprising:

classifying the audio samples of the calls as music, ringing, or announcement; and

removing recorded audio samples from the dataset with a classification of music or ringing.

13 . The method of claim 11 , wherein the clustering the recorded audio samples based on the fingerprints and text transcriptions includes:

grouping recorded audio samples with similar text transcriptions into a single group.

14 . The method of claim 11 , further comprising:

comparing the text transcriptions of the recorded audio samples; and

removing redundant recorded audio samples from the dataset where the comparing text transcriptions indicates that a first recorded audio sample is included in a second recorded audio sample, by removing the first recorded audio sample and associated text transcription from the dataset.

15 . The method of claim 11 , further comprising:

comparing the fingerprints of the recorded audio samples; and

removing redundant recorded audio samples from the dataset where the comparing fingerprints indicates that a first recorded audio sample is similar to a second recorded audio sample having a longer duration than the first recorded audio sample, by removing the first recorded audio sample and associated fingerprint from the dataset.

16 . The method of claim 11 , wherein:

the categories further include an actionable machine response or a human response, and

the actionable machine response includes a Private Branch exchange announcement requesting extension input.

17 . A system comprising one or more processors configured to execute a method including:

initiating a call to a telephone number based on a request from a call originator;

generating a fingerprint of an audio sample of the call based on determining whether the audio sample includes silence that exceeds a threshold;

matching the fingerprint to a record in a fingerprint database, the fingerprint database including records of non-actionable machine responses representing voicemails;

determining, based on the matched record, whether the call includes an actionable machine response, a non-actionable machine response, or a human response, as a detected response;

providing the detected response to the call originator;

receiving, based on the providing the detected response, an action from the call originator; and

performing the action associated with the call.

18 . The system of claim 17 , wherein the generating the fingerprint of the audio sample of the call includes:

classifying the audio sample of the call as an announcement, among classifications including music, ringing, and announcement; and

generating the fingerprint of the audio sample of the call based on classifying the audio sample.

19 . The system of claim 17 , wherein the performing the action associated with the call includes:

disconnecting the call, retrying the call on one of a same channel, route, or gateway, or retrying the call on a different channel, route, or gateway.

20 . The system of claim 17 , wherein generating the fingerprint of the audio sample of the call includes:

generating a first fingerprint of a first audio sample of the call; and

generating a second fingerprint of a second audio sample of the call.

Assignments (2)
SECURITY INTEREST Recorded Jul 17, 2025
From: INFOBIP LIMITED
To: ALTER DOMUS (US) LLC, AS COLLATERAL AGENT
Reel/Frame 071745/0615 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2023
From: PRIBIC, KOSTA; DELALIC, SEAD; HADZIC, HADZEM; ALTWLKANY, KEMAL; KURIC, AMAR; CIZMEK, LUKA; DUMLIJA, IVAN
To: INFOBIP LTD.
Reel/Frame 065445/0567 →
Continuity (1)
Related Publication 20250150530A1 · May 8, 2025
References Cited (24)
US 8065146B2 · Acero · 2011 [cited by examiner]
US 9100479B2 · Bouzid · 2015 [cited by examiner]
US 9571994B2 · Yagey · 2017 [cited by examiner]
US 9596578B1 · Clark · 2017 [cited by examiner]
US 10757252B1 · Quilici · 2020 [cited by examiner]
US 10848618B1 · Quilici · 2020 [cited by examiner]
US 11050876B2 · Balasubramaniyan · 2021 [cited by examiner]
US 11558506B1 · Cardillo · 2023 [cited by examiner]
US 12244762B2 · Horton · 2025 [cited by examiner]
US 20130259211A1 · Vlack · 2013 [cited by examiner]
US 20140369479A1 · Siminoff · 2014 [cited by examiner]
US 20170229133A1 · Bilobrov · 2017 [cited by examiner]
US 20180007199A1 · Quilici · 2018 [cited by examiner]
US 20190391788A1 · Blake · 2019 [cited by examiner]
US 20210136200A1 · Li · 2021 [cited by examiner]
US 20210141879A1 · Calahan · 2021 [cited by examiner]
US 20230041266A1 · Earman · 2023 [cited by examiner]
US 20230082094A1 · Keret · 2023 [cited by examiner]
US 20230085012A1 · Baughman · 2023 [cited by examiner]
US 20230386484A1 · Jangi · 2023 [cited by examiner]
US 20230388414A1 · Bharrat · 2023 [cited by examiner]
US 20240161753A1 · Dentel · 2024 [cited by examiner]
US 20240363125A1 · Khoury · 2024 [cited by examiner]
PCT International Application No. PCT/EP2024/079168, Partial International Search Report and Provisional Opinion of the International Searching Authority, dated Jan. 7, 2025, 13 pages. [cited by applicant]