Contextual automated audio talkdown for remote guarding
An electronic device and method for contextual automated audio talkdown for remote guarding is provided. The electronic device detects a movement of an object in a physical area inside or in a vicinity of a built environment and device receives, based on the detected movement, a sequence of images of the physical area that include the object. The electronic device detects physical activities of the object that are associated with a behavior of an intruder based on application of a first AI model on the sequence of images and generates information that includes a textual description of the physical activities. The electronic device generates an audio alert based on the information and controls a playback of the audio alert via an audio reproduction device installed in a vicinity of the built environment. The playback includes a recitation of the textual description included in the information.
1 . An electronic device, comprising:
circuitry configured to:
detect a movement of an object in a physical area that is inside or in a vicinity of a built environment;
receive, based on the detected movement, a sequence of images of the physical area that include the object;
detect one or more physical activities of the object that are associated with a behavior of an intruder, based on application of a first Artificial Intelligence (AI) model on the sequence of images;
generate information that includes a textual description of the detected one or more physical activities, wherein the generation of the information including the textual description is based on application of a second AI model on an output of the first AI model, and the second AI model is configured to translate one or more text labels for the detected one or more physical activities and one or more attributes of the object to a natural language description;
generate an audio alert based on the generated information, wherein the generation of the audio alert comprises:
converting the textual description to an audio message;
selecting an audio template from a set of audio templates based on the detected one or more physical activities, wherein
the selected audio template includes a set of audio slots, and
the set of audio slots includes an object slot associated with a type of the object, an attribute slot associated with the one or more attributes of the object, and a warning slot associated with the detection of the one or more physical activities; and
inserting the audio message in a corresponding audio slot of the set of audio slots of the selected audio template; and
control a playback of the audio alert via an audio reproduction device that is installed in a vicinity of the built environment, wherein the playback includes a recitation of the textual description included in the generated information.
2 . The electronic device according to claim 1 , wherein the circuitry is further configured to determine a true alarm probability (TAP) based on the detected one or more physical activities of the object.
3 . The electronic device according to claim 1 , wherein the circuitry is further configured to recognize the object based on the application of the first AI model on the sequence of images, wherein the object is recognized as a person or a vehicle.
4 . The electronic device according to claim 1 , wherein the circuitry is further configured to execute one or more of an object detection task and an activity recognition task based on an application of the first AI model on the sequence of images, wherein the detection of the one or more physical activities of the object is based on the execution, and the one or more physical activities include an interaction between the object and one or more items in the physical area.
5 . The electronic device according to claim 1 , wherein the output of the first AI model includes the detected one or more physical activities of the object.
6 . The electronic device according to claim 1 , wherein the circuitry is further configured to:
determine, based on the received sequence of images, the object as a person;
extract a set of features of the person from the received sequence of images; and
classify the person as one of a whitelisted person, a blacklisted person, or an unrecognized person based on whether the extracted set of features is present in a feature database, wherein the audio alert is generated based on a determination that the person is classified as one of the blacklisted person or the unrecognized person.
7 . The electronic device according to claim 1 , wherein the circuitry is further configured to:
detect one or more instances of a damage to a property that includes the physical area and the built environment, based on application of a third AI model on the sequence of images; and
control a display device associated with a user of the built environment to render images that include the detected one or more instances of the damage, wherein the generated information includes a description of the detected one or more instances, and the audio alert is generated further based on the description of the detected one or more instances.
8 . A method comprising:
in an electronic device:
detecting a movement of an object in a physical area that is inside or in a vicinity of a built environment;
receiving, based on the detected movement, a sequence of images of the physical area that include the object;
detecting one or more physical activities of the object that are associated with a behavior of an intruder, based on application of a first Artificial Intelligence (AI) model on the sequence of images;
generating information that includes a textual description of the detected one or more physical activities, wherein the generation of the information including the textual description is based on application of a second AI model on an output of the first AI model, and the second AI model is configured to translate one or more text labels for the detected one or more physical activities and one or more attributes of the object to a natural language description;
generating an audio alert based on the generated information, wherein the generation of the audio alert comprises:
converting the textual description to an audio message;
selecting an audio template from a set of audio templates based on the detected one or more physical activities, wherein
the selected audio template includes a set of audio slots, and
the set of audio slots includes an object slot associated with a type of the object, an attribute slot associated with the one or more attributes of the object, and a warning slot associated with the detection of the one or more physical activities; and
inserting the audio message in a corresponding audio slot of the set of audio slots of the selected audio template; and
controlling a playback of the audio alert via an audio reproduction device that is installed in a vicinity of the built environment, wherein the playback includes a recitation of the textual description included in the generated information.
9 . The method according to claim 8 , further comprising determining a true alarm probability (TAP) based on the detected one or more physical activities of the object.
10 . The method according to claim 8 , further comprising recognizing the object based on the application of the first AI model on the sequence of images, wherein the object is determined as a person or a vehicle.
11 . The method according to claim 8 , further comprising executing one or more of an object detection task and an activity recognition task based on an application of the first AI model on the sequence of images, wherein the detection of the one or more physical activities of the object is based on the execution, and the one or more physical activities include an interaction between the object and one or more items in the physical area.
12 . The method according to claim 8 , wherein the output of the first AI model includes the detected one or more physical activities of the object.
13 . The method according to claim 8 , further comprising:
determining, based on the received sequence of images, the object as a person;
extracting a set of features of the person from the received sequence of images; and
classifying the person as one of a whitelisted person, a blacklisted person, or an unrecognized person based on whether the extracted set of features is present in a feature database, wherein the audio alert is generated based on a determination that the person is classified as one of the blacklisted person or the unrecognized person.
14 . The method according to claim 8 , further comprising:
detecting one or more instances of a damage to a property that includes the physical area and the built environment, based on application of a third AI model on the sequence of images; and
controlling a display device associated with a user of the built environment to render images that include the detected one or more instances of the damage, wherein the generated information includes a description of the detected one or more instances, and the audio alert is generated further based on the description of the detected one or more instances.
15 . A non-transitory computer-readable storage medium having stored thereon, computer-executable instructions that when executed by an electronic device, causes the electronic device to execute operations, the operations comprise:
detecting a movement of an object in a physical area that is inside or in a vicinity of a built environment;
receiving, based on the detected movement, a sequence of images of the physical area that include the object;
detecting one or more physical activities of the object that are associated with a behavior of an intruder, based on application of a first Artificial Intelligence (AI) model on the sequence of images;
generating information that includes a textual description of the detected one or more physical activities, wherein the generation of the information including the textual description is based on application of a second AI model on an output of the first AI model, and the second AI model is configured to translate one or more text labels for the detected one or more physical activities and one or more attributes of the object to a natural language description;
generating an audio alert based on the generated information, wherein the generation of the audio alert comprises:
converting the textual description to an audio message;
selecting an audio template from a set of audio templates based on the detected one or more physical activities, wherein
the selected audio template includes a set of audio slots, and
the set of audio slots includes an object slot associated with a type of the object, an attribute slot associated with the one or more attributes of the object, and a warning slot associated with the detection of the one or more physical activities; and
inserting the audio message in a corresponding audio slot of the set of audio slots of the selected audio template; and
controlling a playback of the audio alert via an audio reproduction device that is installed in a vicinity of the built environment, wherein the playback includes a recitation of the textual description included in the generated information.
16 . The non-transitory computer-readable storage medium according to claim 15 , wherein the operations further comprise determining a true alarm probability (TAP) based on the detected one or more physical activities of the object.
17 . The non-transitory computer-readable storage medium according to claim 15 , wherein the operations further comprise recognizing the object based on the application of the first AI model on the sequence of images, wherein the object is recognized as a person or a vehicle.
18 . The non-transitory computer-readable storage medium according to claim 15 , wherein the operations further comprise executing one or more of an object detection task and an activity recognition task based on an application of the first AI model on the sequence of images, wherein the detection of the one or more physical activities of the object is based on the execution, and the one or more physical activities include an interaction between the object and one or more items in the physical area.