Method and apparatus for identifying and relaying audio segments of interest in a monitored security environment
An apparatus and method are disclosed. In an embodiment, an audio analysis device is provided. The audio analysis device is configured to receive administrator criteria that specifies an alarm event handling response, obtain an audio file associated with a monitored security environment. detect an audio anomaly in the audio file, and in response to detecting the audio anomaly, generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly. The audio analysis device is further configured to generate metadata comprising a characteristic associated with the audio segment, determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal, and in response to the administrator criteria, encode the audio segment and the metadata for rendering in the agent portal.
1 . An audio analysis device, the audio analysis device
comprising:
at least one processor; and
memory comprising a plurality of instructions that, when executed by the at least one processor, are configured to cause the at least one processor to:
receive administrator criteria that specifies an alarm event handling response;
obtain an audio file associated with an alarm event at a monitored security environment;
detect, based on a machine learning model, an audio anomaly in the audio file;
in response to detecting the audio anomaly:
generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;
generate metadata based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment, at least one of the plurality of characteristics comprising:
a description of the audio anomaly;
a secret phrase;
a description of at least one other sound in the audio segment;
a location where the audio file was recorded;
an identity of at least one person present at the location;
a number of distinct speakers having voices present in the audio segment;
a distance between a source of the audio anomaly and the audio recording device; or
an identity of at least one background noise in the audio segment;
determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal for a monitoring agent assigned to the alarm event;
select, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and
encode the audio segment and the subset of the metadata for rendering in the agent portal.
2 . The audio analysis device of claim 1 , wherein the plurality of instructions is further configured to cause the at least one processor to:
select, based on the metadata, a second ML model for analyzing the audio segment; and
update the metadata based on the analysis by the second ML model.
3 . An audio analysis device, the audio analysis device comprising:
at least one processor; and
memory comprising a plurality of instructions that, when executed by the at least one processor, are configured to cause the at least one processor to:
receive administrator criteria that specifies an alarm event handling response;
obtain an audio file associated with a monitored security environment;
detect an audio anomaly in the audio file;
in response to detecting the audio anomaly:
generate an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;
generate metadata based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment;
determine that the administrator criteria specifies providing the audio segment and the metadata in an agent portal;
select, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and
encode the audio segment and the subset of the metadata for rendering in the agent portal.
4 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to detect the audio anomaly in the audio file using the first ML model.
5 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to:
receive resolution information relating to the audio anomaly; and
train, based on the resolution information, the first ML model.
6 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to:
receive at least one change to the first ML model;
perform, based on the at least one change to the first ML model and on determining not to transmit, a further analysis of the audio file; and
based on the further analysis, one of transmit the audio file to an agent device or mark the audio file for deletion.
7 . The audio analysis device of claim 3 , wherein at least one of the characteristics comprises at least one of:
a description of the audio anomaly;
a secret phrase;
a description of at least one other sound in the audio segment;
a location where the audio file was recorded;
an identity of at least one person present at the location;
a number of distinct speakers having voices present in the audio segment;
a distance between a source of the audio anomaly and the audio recording device; or
an identity of at least one background noise in the audio segment.
8 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to:
select, based on the metadata, a second ML model for analyzing the audio segment; and
update the metadata based on the analysis by the second ML model.
9 . The audio analysis device of claim 3 , wherein the plurality of instructions are further configured to cause the at least one processor to select a subset of the metadata by at least selecting individual ones of the characteristics that have a respective priority weight above a priority threshold, the priority threshold being based on the administrator criteria.
10 . The audio analysis device of claim 3 , wherein the plurality of instructions is further configured to cause the at least one processor to increase, based on the metadata, a length of the audio segment.
11 . A method performed by an audio analysis device, the method comprising:
receiving administrator criteria that specifies an alarm event handling response;
obtaining an audio file associated with a monitored security environment;
detecting an audio anomaly in the audio file;
in response to detecting the audio anomaly;
generating an audio segment based on the audio anomaly, the audio segment comprising the audio anomaly;
generating metadata, based on a first machine learning (ML) model and by assigning, based on the administrator criteria, a respective priority weight to each of a plurality of characteristics associated with the audio segment;
determining that the administrator criteria specifies providing the audio segment and the metadata in an agent portal;
selecting, based on the administrator criteria and the respective priority weight of each of the plurality of characteristics, a subset of the metadata for rendering in the agent portal; and
encoding the audio segment and the subset of the metadata for rendering in the agent portal.
12 . The method of claim 11 , further comprising detecting the audio anomaly in the audio file using the first ML model.
13 . The method of claim 11 , further comprising:
receiving resolution information relating to the audio anomaly; and
training, based on the resolution information, the first ML model.
14 . The method of claim 11 , further comprising:
receiving at least one change to the first ML model;
performing, based on the at least one change to the first ML model and on determining not to transmit, a further analysis of the audio file; and
based on the further analysis, one of transmitting the audio file to an agent device or marking the audio file for deletion.
15 . The method of claim 11 , wherein at least one of the characteristics comprises at least one of:
a description of the audio anomaly;
a secret phrase;
a description of at least one other sound in the audio segment;
a location where the audio file was recorded;
an identity of at least one person present at the location;
a number of distinct speakers having voices present in the audio segment;
a distance between a source of the audio anomaly and the audio analysis device; or
an identity of at least one background noise in the audio segment.
16 . The method of claim 11 , further comprising:
selecting, based on the metadata, a second ML model for analyzing the audio segment; and
updating the metadata based on the analysis by the second ML model.
17 . The method of claim 16 , further comprising selecting a subset of the metadata by at least selecting individual ones of the characteristics that have a respective priority weight above a priority threshold, the priority threshold being based on the administrator criteria.
18 . The method of claim 11 , further comprising increasing, based on the metadata, a length of the audio segment.