Enhanced computing device representation of audio
An example method includes receiving, by one or more processors of a computing device, audio data recorded by one or more microphones of the computing device; and generating, based on the audio data and by the one or more processors, one or more structured sound records, a first structured sound record of the one or more structured sound records including: a description of a first sound, the description including a descriptive label of the first sound, the descriptive label different than a text transcription of the first sound, and a time stamp indicating a time at which the first sound occurred; and outputting a graphical user interface including a timeline representation of the one or more structured sound records.
1 . A method comprising:
receiving, by one or more processors of a computing device, audio data recorded by one or more microphones of the computing device; and
generating, based on the audio data and by the one or more processors, one or more structured sound records, a first structured sound record of the one or more structured sound records including:
a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and
a time stamp indicating a time at which the first sound occurred; and
outputting a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.
2 . The method of claim 1 , wherein the timeline representation indicates a sequence in which sounds of the one or more structured sound records occurred.
3 . The method of claim 1 , wherein a second structured sound record of the one or more sound records includes:
a description of a second sound, the description including a text transcription of the second sound, and
a time stamp indicating a time at which the second sound occurred.
4 . The method of claim 1 , wherein outputting the graphical user interface comprises outputting a first graphical user interface including a current timeline representation of the one or more structured sound records, the method further comprising:
outputting a second graphical user interface including a past timeline representation of the one or more structured sound records.
5 . The method of claim 4 , wherein outputting the first graphical user interface comprises outputting the first graphical user interface for display at a first display, and wherein outputting the second graphical user interface comprises outputting the second graphical user interface at a second display that is included in a device that is different than a device that includes the first display.
6 . The method of claim 1 , wherein outputting the graphical user interface comprises outputting the graphical user interface in response to determining that a newly generated sound record of the one or more structured sound records has a descriptive label included in a pre-determined set of descriptive labels.
7 . The method of claim 6 , wherein the pre-determined set of descriptive labels includes emergency category labels, priority category labels, and other category labels.
8 . The method of claim 7 , wherein:
the emergency category labels include one or more of a smoke alarm label, a fire alarm label, a carbon monoxide label, a siren label, and a shouting label;
the priority category labels include one or more of a baby crying label, a doorbell label, a door knocking label, an animal alerting label, and a glass breaking label; and
the other category labels include one or more of a water running label, a landline phone ringing label, and one or more appliance beep labels.
9 . The method of claim 1 , wherein outputting the graphical user interface including the timeline representation comprises outputting the graphical user interface including the timeline representation at a display of the computing device.
10 . The method of claim 1 , wherein the computing device is a first computing device, wherein outputting the graphical user interface including the timeline representation comprises causing a second computing device to output, at a display of the second computing device, the graphical user interface including the timeline representation, and wherein the second computing device is different than the first computing device.
11 . The method of claim 10 , further comprising:
responsive to receiving, at the computing device, user input indicating viewing of the timeline representation at a particular time, modifying output of the timeline representation at the second computing device to indicate previous viewing.
12 . The method of claim 10 , wherein the second computing device comprises a wearable computing device.
13 . The method of claim 10 , wherein the first computing device does not include a display.
14 . The method of claim 1 , further comprising: determining that the classification of the non-speech environmental sound corresponds to an emergency category; and outputting, concurrently with the graphical user interface, a non-audio alert via at least one of a haptic output device or a light device of the computing device.
15 . The method of claim 1 , further comprising: transmitting a signal to a wearable computing device communicatively coupled to the computing device, wherein the signal causes the wearable computing device to output a haptic alert corresponding to the classification of the non-speech environmental sound.
16 . The method of claim 1 , further comprising: receiving user input to scroll the timeline representation backwards in time; and updating the graphical user interface to display a past timeline representation indicating a sequence of previously generated structured sound records corresponding to previously classified non-speech environmental sounds.
17 . The method of claim 1 , wherein the timeline representation further includes a graphical plot indicating an amplitude of the first sound over a duration of the first sound.
18 . A computing device comprising:
one or more microphones configured to record audio data; and
one or more processors configured to:
generate, based on the audio data, one or more structured sound records, a first structured sound record of the one or more structured sound records including:
a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and
a time stamp indicating a time at which the first sound occurred; and
output a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.
19 . The computing device of claim 18 , wherein the timeline representation indicates a sequence in which sounds of the one or more structured sound records occurred.
20 . The computing device of claim 18 , wherein a second structured sound record of the one or more sound records includes:
a description of a second sound, the description including a text transcription of the second sound, and
a time stamp indicating a time at which the second sound occurred.
21 . The computing device of claim 18 , wherein, to output the graphical user interface, the one or more processors are configured to output a first graphical user interface including a current timeline representation of the one or more structured sound records, and wherein the one or more processors are further configured to:
output a second graphical user interface including a past timeline representation of the one or more structured sound records.
22 . The computing device of claim 18 , wherein outputting the graphical user interface including the timeline representation comprises causing another computing device to output the graphical user interface including the timeline representation, and wherein the other computing device is different than the computing device.
23 . The computing device of claim 22 , wherein the one or more processors are further configured to:
modify, responsive to receiving user input indicating viewing of the timeline representation at a particular time, output of the timeline representation at the other computing device to indicate previous viewing.
24 . A computer-readable storage medium storing instructions that, when executed, cause one or more processors of a computing device to:
receive audio data recorded by one or more microphones of the computing device;
generate, based on the audio data, one or more structured sound records, a first structured sound record of the one or more structured sound records including:
a description of a first sound, the description including a descriptive label of the first sound, the descriptive label representing a classification of a non-speech environmental sound and being different than a text transcription of the first sound, and
a time stamp indicating a time at which the first sound occurred; and
output a graphical user interface including a timeline representation of the one or more structured sound records, wherein the timeline representation visually indicates a sequence of a plurality of structured sound records corresponding to a plurality of different classified non-speech environmental sounds.