Latency handling for point-to-point communications
Aspects of the subject technology provide improved point-to-point audio communications based on human variable sensitivity to latency differences in multipath communications. In aspects, improved techniques may include measuring a level of ambient noise, and then selecting processing for a received electronic audio based on the measured level of ambient noise before emitting the processed audio signal at a loudspeaker worn by a listener.
1 . A method of audio processing, comprising:
receiving, by a first device from a second device, an electronic audio signal corresponding to an audio source within a first proximity of a first user;
measuring a level of ambient noise with respect to the first user;
determining that the audio source is also present in the ambient noise such that a multipath latency difference exists between an ambient acoustic path of the audio source and an electronic path of the audio source that corresponds to the electronic audio signal;
processing the electronic audio signal based on the level of ambient noise to reduce an effect of the multipath latency difference; and
causing the processed electronic audio signal to be emitted from a loudspeaker in a device communicatively coupled to the first device and configured to be worn by the first user.
2 . The method of claim 1 , wherein the audio source corresponds to a second user, the electronic audio signal is captured at a microphone of the second device, the microphone located proximate to the second user, and the measuring of the level of ambient noise is based on a signal captured at a microphone in the first device configured to be worn by the first user.
3 . The method of claim 1 , further comprising:
selecting, when the level of ambient noise is high, a first speech-processing pipeline having a first end-to-end latency, and when the level of ambient noise is low, a second speech-processing pipeline having a second end-to-end latency to reduce a latency discontinuity due to speech-processing pipeline delays, the second end-to-end latency being shorter than the first end-to-end latency.
4 . The method of claim 1 , wherein the level of ambient noise is a signal-to-noise ratio based on a source signal captured at a microphone of the second device located proximate to the audio source and an ambient noise signal captured at a microphone in the first device configured to be worn by the first user.
5 . The method of claim 1 , wherein the processing of the electronic audio signal includes:
raising a noise floor in the electronic audio signal based on the level of ambient noise to mask the multipath latency difference so that a latency discontinuity due to the multipath latency difference has a reduced perceptibility to the first user.
6 . The method of claim 1 , wherein the processing of the electronic audio signal includes:
when the level of ambient noise is low, raising a noise floor to a higher level to mask the multipath latency difference so that a latency discontinuity due to the multipath latency difference has a reduced perceptibility to the first user; and
when the level of ambient noise is high, lowering the noise floor to a lower level.
7 . The method of claim 5 , wherein the raising the noise floor includes adding an artificial noise to the electronic audio signal.
8 . The method of claim 5 , wherein the raising the noise floor includes capturing an ambient noise, and adding the captured ambient noise to the electronic audio signal.
9 . The method of claim 5 , wherein the raising the noise floor includes reducing a noise cancelling effect at the loudspeaker in the device configured to be worn by the first user.
10 . The method of claim 5 , wherein the raising the noise floor in the electronic audio signal is further based on an estimate of a physical distance between the first user and the audio source.
11 . The method of claim 3 , wherein the processing of the electronic audio signal includes a speech enhancement processing of the electronic audio signal for the first user according to the first speech-processing pipeline or the second speech-processing pipeline.
12 . The method of claim 11 , wherein the speech enhancement processing includes language translation.
13 . The method of claim 11 , wherein the speech enhancement processing is based on an indication of a hearing limitation of the first user.
14 . The method of claim 1 , wherein the audio source is a second user, the second device is a user device of the second user, and the first proximity of the first user includes distances between the first user and the second user that are within human audible hearing range via sound waves traveling through air.
15 . A system for audio processing, comprising:
a processor; and
a memory storing instructions, that when executed by the processor, cause the system to:
receive, by a first device from a second device, an electronic audio signal corresponding to an audio source within a first proximity of a first user;
measure a level of ambient noise with respect to the first user;
determine that the audio source is also present in the ambient noise such that a multipath latency difference exists between an ambient acoustic path of the audio source and an electronic path of the audio source that corresponds to the electronic audio signal;
process the electronic audio signal based on the level of ambient noise to reduce an effect of the multipath latency difference; and
cause the processed electronic audio signal to be emitted from a loudspeaker in a device communicatively coupled to the first device and configured to be worn by the first user.
16 . The system of claim 15 , wherein the processing of the electronic audio signal includes:
raising a noise floor in the electronic audio signal based on the level of ambient noise to mask the multipath latency difference so that a latency discontinuity due to the multipath latency difference has a reduced perceptibility to the first user.
17 . The system of claim 15 , wherein the processing of the electronic audio signal includes a speech enhancement processing of the electronic audio signal for the first user, the instructions further causing the system to:
select, when the level of ambient noise is high, a first speech-processing pipeline for the speech enhancement processing having a first end-to-end latency, and when the level of ambient noise is low, a second speech-processing pipeline for the speech enhancement processing having a second end-to-end latency to reduce a latency discontinuity due to speech-processing pipeline delays, the second end-to-end latency being shorter than the first end-to-end latency.
18 . A non-transitory computer readable memory storing instructions that, when executed by a processor, cause the processor to:
receive, by a first device from a second device, an electronic audio signal corresponding to an audio source within a first proximity of a first user;
measure a level of ambient noise with respect to the first user;
determine that the audio source is also present in the ambient noise such that a multipath latency difference exists between an ambient acoustic path of the audio source and an electronic path of the audio source that corresponds to the electronic audio signal;
process the electronic audio signal based on the level of ambient noise to reduce an effect of the multipath latency difference; and
cause the processed electronic audio signal to be emitted from a loudspeaker in a device communicatively coupled to the first device and configured to be worn by the first user.
19 . The computer readable memory of claim 18 , wherein the processing of the electronic audio signal includes:
raising a noise floor in the electronic audio signal based on the level of ambient noise to mask the multipath latency difference so that a latency discontinuity due to the multipath latency difference has a reduced perceptibility to the first user.
20 . The computer readable memory of claim 18 , wherein the processing of the electronic audio signal includes a speech enhancement processing of the electronic audio signal for the first user, the instructions further causing the processor to:
select, when the level of ambient noise is high, a first speech-processing pipeline for the speech enhancement processing having a first end-to-end latency, and when the level of ambient noise is low, a second speech-processing pipeline for the speech enhancement processing having a second end-to-end latency to reduce a latency discontinuity due to speech-processing pipeline delays, the second end-to-end latency being shorter than the first end to end latency.