METHOD/SYSTEM TO IDENTIFY, EXTRACT, AND CONVEY KEY WORDS, PHRASES, AND MEANINGS IN SERVICE OF EMERGENCY COMMUNICATIONS (E.G., VOICE, TEXT, AND VIDEO)
A method includes receiving visual content of a physical location, the visual content identified by a content timestamp, the physical location identified by a spatial identifier; receiving audio content of a call; performing voice recognition on the call to extract a first audio symbol; receiving a first timestamp of the call, the first timestamp indicating a time at which the call was initiated or a time extracted from the call by voice recognition; and determining a feature in the visual content, at least in part based on the first audio symbol, the feature defined by a person, object, or situation.
1 - 20 . (canceled)
21 . A method, comprising:
receiving visual content of a physical location, the visual content identified by a content timestamp, the physical location identified by a spatial identifier;
receiving audio content of a call;
performing voice recognition on the call to extract a first audio symbol;
receiving a first timestamp of the call, the first timestamp indicating a time at which the call was initiated or a time extracted from the call by voice recognition;
identifying the visual content from among a plurality of visual content, at least in part based on the content timestamp and the first timestamp; and
determining a feature in the visual content, at least in part based on the first audio symbol, the feature defined by a person, object, or situation.
22 . The method of claim 21 , further comprising:
determining a relative timestamp, at least in part based on the first timestamp and a spoken timestamp extracted from the call by voice recognition, wherein the identifying is at least in part based on the relative timestamp.
23 . The method of claim 21 , further comprising:
receiving a location of the call, the location being a GPS location of a device that initiated the call or a location identified in the call by voice recognition, wherein the identifying is further at least in part based on the location of the call and the spatial identifier.
24 . The method of claim 21 , further comprising:
standardizing the first audio symbol to a standardized word, wherein the determining the feature is at least in part based on the standardized word.
25 . The method of claim 24 , wherein the standardizing is at least in part based on a thesaurus.
26 . The method of claim 21 , further comprising:
annotating the visual content with an identifier of the feature to produce an annotation; and
transmitting the visual content and the annotation to an emergency system.
27 . The method of claim 21 , further comprising:
determining, based on entropy of the visual content, a portion of the visual content that includes meaningful information.
28 . An apparatus, comprising:
a network interface that receives visual content of a physical location and audio content of a call, the visual content identified by a content timestamp, the physical location identified by a spatial identifier; and
a processing unit configured to:
perform voice recognition on the call to extract a first audio symbol,
receive a first timestamp of the call, the first timestamp indicating a time at which the call was initiated or a time extracted from the call by voice recognition,
identify the visual content from among a plurality of visual content, at least in part based on the content timestamp and the first timestamp, and
determine a feature in the visual content, at least in part based on the first audio symbol, the feature defined by a person, object, or situation.
29 . The apparatus of claim 28 , wherein the processing unit further is configured to standardize the first audio symbol to a standardized word, and the determining the feature is at least in part based on the standardized word.
30 . The apparatus of claim 28 , wherein the processing unit further is configured to assign a relevance score to the first audio symbol.
31 . The apparatus of claim 30 , wherein the processing unit further is configured to determine a confidence score of a scene represented by the visual content, at least in part based on the relevance score.
32 . The apparatus of claim 28 , wherein the processing unit further is configured to contribute the feature and the first audio symbol to a learning model, the learning model configured to learn a fit of a situation to a modeled assessment.
33 . The apparatus of claim 32 , wherein the learning model is further configured to receive feedback from a human trainer regarding accuracy of the modeled assessment.
34 . The apparatus of claim 28 , wherein the processing unit further is configured to create a statement about a scene represented by the visual content, and the network interface transmits the statement to an emergency communications center or a first responder.
35 . A method, comprising:
receiving audio content of a call;
performing voice recognition on the call to extract a first audio symbol;
standardizing the first audio symbol to a standardized word based on a thesaurus;
receiving visual content of a physical location;
determining a feature in the visual content, at least in part based on the standardized word, the feature defined by a person, object, or situation;
assigning a relevance score to the first audio symbol; and
contributing the feature and the relevance score to a learning model.
36 . The method of claim 35 , further comprising:
receiving caller location information comprising an automatic location identification (ALI);
comparing at least one of an initiating location of the call or a spoken location extracted from the call to the ALI; and
producing an identified location based on the comparing.
37 . The method of claim 35 , further comprising:
extracting a spoken telephone number from the call by voice recognition;
comparing the spoken telephone number to an automatic number identification (ANI) of the call; and
adding the spoken telephone number to an annotation of the visual content.
38 . The method of claim 35 , wherein the performing voice recognition comprises performing voice recognition on a dialogue between an emergency calltaker and a caller, and the method further comprises resolving a relative timestamp spoken by the caller based on a time referenced by the calltaker.
39 . The method of claim 35 , further comprising:
storing information associated with the call in an immutable data repository, the information comprising an incident ID, a source type, a time, a date, and a location.
40 . The method of claim 35 , wherein the performing voice recognition comprises situationally ignoring a negation in the call based on a determined importance of a subject of the negation.