SYSTEM AND METHOD FOR VISUALLY PRESENTING AUDITORY INFORMATION
System and method for analyzing audio data are provided. The audio data may be analyzed to visually present auditory information. For example, the audio data may be analyzed to obtain textual information. Speaker information may be obtained, for example by analyzing input data, such as audio data and image data. Presentation parameters may be selected based on the speaker information. The textual information may be visually present, for example according to the presentation parameters.
1 . A wearable apparatus for processing audio and visually presenting information, the wearable apparatus comprising:
one or more wearable audio sensors configured to capture audio data from an environment of a user;
one or more visual output devices; and
at least one processing unit configured to:
analyze the audio data to obtain textual information;
obtain speaker information associated with one or more speakers that produced speech in the audio data; and
based on the speaker information, present the textual information to the user using the one or more visual output devices.
2 . The wearable apparatus of claim 1 , wherein the speaker information comprises an association of one or more portions of the textual information with the user; and the at least one processing unit is further configured to forgo presenting at least one of the one or more portions of the textual information associated with the user based on said association.
3 . The wearable apparatus of claim 1 , wherein the at least one processing unit is further configured to:
analyze the audio data to obtain auxiliary information associated with one or more properties of a voice associated with a portion of the textual information; and
present, using the one or more visual output devices, information based on the auxiliary information in conjunction with the presentation of the portion of the textual information.
4 . The wearable apparatus of claim 3 , wherein at least one of the one or more properties is at least one of: pitch, intensity, tempo, rhythm, prosody, and flatness.
5 . The wearable apparatus of claim 1 , wherein the at least one processing unit is further configured to:
analyze the audio data to obtain at least one non-textual graphical symbol associated with at least a portion of the audio data; and
present the non-textual graphical symbol to the user using the one or more visual output devices.
6 . The wearable apparatus of claim 1 , wherein the speaker information comprises an association of a portion of the textual information with a first speaker of the one or more speakers; and wherein the at least one processing unit is further configured to:
present, using the one or more visual output devices, information associated with the first speaker in conjunction with the presentation of the portion of the textual information.
7 . The wearable apparatus of claim 1 , wherein the speaker information comprises an association of a portion of the textual information with a first speaker of the one or more speakers; and wherein the at least one processing unit is further configured to:
select at least one display parameter based on said association; and
present the portion of the textual information using the at least one display parameter.
8 . The wearable apparatus of claim 7 , wherein the at least one display parameter comprises a presentation region; and wherein the presentation region is selected based on orientation information associated with the first speaker.
9 . The wearable apparatus of claim 8 , wherein the one or more wearable audio sensors is two or more wearable audio sensors; and wherein the at least one processing unit is further configured to:
analyze the audio data to obtain the orientation information associated with the first speaker.
10 . The wearable apparatus of claim 8 , further comprising one or more image sensors configured to capture one or more images from the environment of the user; and wherein the at least one processing unit is further configured to:
analyze the one or more images to obtain the orientation information associated with the first speaker.
11 . A method for processing audio and visually presenting information, the method comprising:
obtaining audio data captured using one or more wearable audio sensors;
analyzing the audio data to obtain textual information;
obtaining speaker information associated with one or more speakers that produced speech in the audio data; and
based on the speaker information, visually presenting the textual information to a user.
12 . The method of claim 11 , wherein the speaker information comprises an association of one or more portions of the textual information with the user; and wherein the method further comprising:
forgo presenting at least one of the one or more portions of the textual information associated with the user.
13 . The method of claim 11 , wherein the method further comprising:
analyzing the audio data to obtain auxiliary information associated with one or more properties of a voice associated with a portion of the textual information; and
visually presenting information based on the auxiliary information in conjunction with the presentation of the portion of the textual information.
14 . The method of claim 11 , further comprising:
analyzing the audio data to obtain at least one non-textual graphical symbol associated with at least a portion of the audio data; and
visually presenting the non-textual graphical symbol to the user.
15 . The method of claim 11 , wherein the speaker information comprises an association of a portion of the textual information with a first speaker of the one or more speakers; and wherein the method further comprise:
visually presenting information associated with the first speaker in conjunction with the presentation of the portion of the textual information.
16 . The method of claim 11 , wherein the speaker information comprises an association of a portion of the textual information with a first speaker of the one or more speakers; and wherein the method further comprise:
selecting at least one display parameter based on the first speaker; and
visually presenting the portion of the textual information using the at least one display parameter.
17 . The method of claim 16 , wherein the at least one display parameter comprises a presentation region; and wherein the presentation region is selected based on orientation information associated with the first speaker.
18 . The method of claim 17 , further comprising:
analyzing the audio data to obtain orientation information associated with the first speaker.
19 . The method of claim 18 , further comprising:
obtaining one or more images using one or more wearable image sensors; and
analyzing the one or more images to obtain orientation information associated with the first speaker.
20 . A non-transitory computer readable medium storing data and computer implementable instructions for carrying out the method of claim 11 .