IP Library Granted Patent US 12682732
Granted Patent B2
US 12682732 · App. 18/478,507 · Granted Jul 14, 2026

Method and apparatus for identifying and relaying audio segments of interest in a monitored security environment

Inventors: Bhaskar Thadisetty (Boca Raton, FL); Isabel Fernandez (Lauderdale by the Sea, FL); Steven Alfonse Zaccardi (Coral Springs, FL)
Assignee: The ADT Security Corporation
G08B13/1672G10L17/22G10L25/84
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12682732
App. No.
18/478,507
Granted
Jul 14, 2026
Kind
B2
Abstract

An apparatus and method are disclosed. In an embodiment, an audio analysis device is provided. The audio analysis device is configured to receive administrator criteria that specifies an alarm event handling response, obtain an audio file associated with a monitored security environment. detect an audio anomaly in the audio file, and in response to detecting the audio anomaly, generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly. The audio analysis device is further configured to generate metadata comprising a characteristic associated with the audio segment, determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal, and in response to the administrator criteria, encode the audio segment and the metadata for rendering in the agent portal.

Claims (90)

1 . An audio analysis device, the audio analysis device

comprising:

at least one processor; and

memory comprising a plurality of instructions that, when executed by the at least one processor, are configured to cause the at least one processor to:

receive administrator criteria that specifies an alarm event handling response;

obtain an audio file associated with an alarm event at a monitored security environment;

detect, based on a machine learning model, an audio anomaly in the audio file;

in response to detecting the audio anomaly:

generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;

generate metadata based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment, at least one of the plurality of characteristics comprising:

a description of the audio anomaly;

a secret phrase;

a description of at least one other sound in the audio segment;

a location where the audio file was recorded;

an identity of at least one person present at the location;

a number of distinct speakers having voices present in the audio segment;

a distance between a source of the audio anomaly and the audio recording device; or

an identity of at least one background noise in the audio segment;

determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal for a monitoring agent assigned to the alarm event;

select, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and

encode the audio segment and the subset of the metadata for rendering in the agent portal.

2 . The audio analysis device of claim 1 , wherein the plurality of instructions is further configured to cause the at least one processor to:

select, based on the metadata, a second ML model for analyzing the audio segment; and

update the metadata based on the analysis by the second ML model.

3 . An audio analysis device, the audio analysis device comprising:

at least one processor; and

memory comprising a plurality of instructions that, when executed by the at least one processor, are configured to cause the at least one processor to:

receive administrator criteria that specifies an alarm event handling response;

obtain an audio file associated with a monitored security environment;

detect an audio anomaly in the audio file;

in response to detecting the audio anomaly:

generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;

generate metadata based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment;

determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal;

select, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and

encode the audio segment and the subset of the metadata for rendering in the agent portal.

4 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to detect the audio anomaly in the audio file using the first ML model.

5 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to:

receive resolution information relating to the audio anomaly; and

train, based on the resolution information, the first ML model.

6 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to:

receive at least one change to the first ML model;

perform, based on the at least one change to the first ML model and on determining not to transmit, a further analysis of the audio file; and

based on the further analysis, one of transmit the audio file to an agent device or mark the audio file for deletion.

7 . The audio analysis device of claim 3 , wherein at least one of the characteristics comprises at least one of:

a description of the audio anomaly;

a secret phrase;

a description of at least one other sound in the audio segment;

a location where the audio file was recorded;

an identity of at least one person present at the location;

a number of distinct speakers having voices present in the audio segment;

a distance between a source of the audio anomaly and the audio recording device; or

an identity of at least one background noise in the audio segment.

8 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to:

select, based on the metadata, a second ML model for analyzing the audio segment; and

update the metadata based on the analysis by the second ML model.

9 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to select a subset of the metadata by at least selecting individual ones of the characteristics that have a respective priority weight above a priority threshold, the priority threshold being based on the administrator criteria.

10 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to increase, based on the metadata, a length of the audio segment.

11 . A method performed by an audio analysis device, the method comprising:

receiving administrator criteria that specifies an alarm event handling response;

obtaining an audio file associated with a monitored security environment;

detecting an audio anomaly in the audio file;

in response to detecting the audio anomaly;

generating an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;

generating metadata, based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment;

determining that the administrator criteria specifies providing the audio segment and the metadata in an agent portal;

selecting, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and

encoding the audio segment and the subset of the metadata for rendering in the agent portal.

12 . The method of claim 11 , further comprising detecting the audio anomaly in the audio file using the first ML model.

13 . The method of claim 11 , further comprising:

receiving resolution information relating to the audio anomaly; and

training, based on the resolution information, the first ML model.

14 . The method of claim 11 , further comprising:

receiving at least one change to the first ML model;

performing, based on the at least one change to the first ML model and on determining not to transmit, a further analysis of the audio file; and

based on the further analysis, one of transmitting the audio file to an agent device or marking the audio file for deletion.

15 . The method of claim 11 , wherein at least one of the characteristics comprises at least one of:

a description of the audio anomaly;

a secret phrase;

a description of at least one other sound in the audio segment;

a location where the audio file was recorded;

an identity of at least one person present at the location;

a number of distinct speakers having voices present in the audio segment;

a distance between a source of the audio anomaly and the audio analysis device; or

an identity of at least one background noise in the audio segment.

16 . The method of claim 11 , further comprising:

selecting, based on the metadata, a second ML model for analyzing the audio segment; and

updating the metadata based on the analysis by the second ML model.

17 . The method of claim 16 , further comprising selecting a subset of the metadata by at least selecting individual ones of the characteristics that have a respective priority weight above a priority threshold, the priority threshold being based on the administrator criteria.

18 . The method of claim 11 , further comprising increasing, based on the metadata, a length of the audio segment.