Spatial audio system for videoconferencing
Systems and methods for spatial audio generation for videoconferencing are provided. For example, a computing device detects the head pose of a user of a video conference. The device further determines a first relative direction of a participant user interface (UI) element for an active speaker of the video conference with respect to the user based on a first position of the participant UI element and the head pose of the user. The device generates spatial audio for an audio signal of the active speaker based on the first relative direction and outputs the spatial audio. The device further generates an updated spatial audio based on a second relative direction towards a second location closer to a center location than the first relative direction. The center location is determined based on the head pose of the user. The device outputs the updated spatial audio.
1 . A method performed by a computing device, the method comprising:
detecting a head pose of a user in front of the computing device that has joined a video conference;
determining a first relative direction of a participant user interface (UI) element for an active speaker of the video conference with respect to the user based on a first position of the participant UI element and the head pose of the user, the participant UI element comprising a window and a representation of a corresponding participant presented in a UI for the video conference displayed on a screen of the computing device;
generating spatial audio for at least an audio signal of the active speaker based on the first relative direction;
outputting the spatial audio;
generating an updated spatial audio based on a second relative direction, wherein the second relative direction is towards a second location closer to a center location on the screen that the user is facing than the first relative direction, the center location determined based on the head pose of the user; and
outputting the updated spatial audio.
2 . The method of claim 1 , further comprising:
determining a third relative direction of a participant UI element for a second active speaker with respect to the user based on a third location of the participant UI element for the second active speaker and the head pose of the user,
wherein generating the spatial audio further comprising:
generating a first individual spatial audio for the audio signal of the active speaker based on the first relative direction and a second individual spatial audio for a second audio signal of the second active speaker based on the second relative direction, and
combining the first individual spatial audio and the second individual spatial audio.
3 . The method of claim 1 , further comprising:
modifying the UI for the video conference by moving the participant UI element for the active speaker to the second location; and
presenting the modified UI on the computing device.
4 . The method of claim 1 , wherein generating the spatial audio is performed using head-related transfer functions (HRTFs).
5 . The method of claim 1 , wherein moving the participant UI element for the active speaker to the second location comprises switching the participant UI element for the active speaker with another participant UI element for another participant positioned at the second location.
6 . The method of claim 1 , further comprising stopping generating the updated spatial audio and moving the participant UI element for the active speaker towards the center location based on determining that the active speaker is inactive for a threshold time period.
7 . The method of claim 1 , further comprising stopping generating the spatial audio based on determining that the participant UI element for the active speaker has reached the center location.
8 . A computing device, comprising:
a non-transitory computer-readable medium; and
a processor communicatively coupled to the non-transitory computer-readable medium, the processor configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to:
detect a head pose of a user in front of the computing device that has joined a video conference;
determine a first relative direction of a participant user interface (UI) element for an active speaker of the video conference with respect to the user based on a first location of the participant UI element and the head pose of the user, the participant UI element comprising a window and a representation of a corresponding participant presented in a UI for the video conference displayed on a screen of the computing device;
generate a spatial audio for at least an audio signal of the active speaker based on the first relative direction;
output the spatial audio;
generate an updated spatial audio based on a second relative direction, wherein the second relative direction is towards a second location closer to a center location on the screen that the user is facing than the first relative direction, the center location determined based on the head pose of the user; and
output the updated spatial audio.
9 . The computing device of claim 8 , wherein the processor is configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to further:
determine a third relative direction of a participant UI element for a second active speaker with respect to the user based on a third location of the participant UI element for the second active speaker and the head pose of the user,
wherein generating the spatial audio further comprising:
generating a first individual spatial audio for the audio signal of the active speaker based on the first relative direction and a second individual spatial audio for a second audio signal of the second active speaker based on the second relative direction, and
combining the first individual spatial audio and the second individual spatial audio.
10 . The computing device of claim 8 , wherein the processor is configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to further:
modifying the UI for the video conference by moving the participant UI element for the active speaker to the second location; and
presenting the modified UI on the computing device.
11 . The computing device of claim 8 , wherein generating the spatial audio is performed using head-related transfer functions (HRTFs).
12 . The computing device of claim 8 , wherein moving the participant UI element for the active speaker to the second location comprises switching the participant UI element for the active speaker with another participant UI element for another participant positioned at the second location.
13 . The computing device of claim 8 , wherein the processor is configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to further stop generating the updated spatial audio and moving the participant UI element for the active speaker towards the center location based on determining that the active speaker is inactive for a threshold time period.
14 . The computing device of claim 8 , wherein the processor is configured to execute processor-executable instructions stored in the non-transitory computer-readable medium to further stop generating the spatial audio based on determining that the participant UI element for the active speaker has reached the center location.
15 . A non-transitory computer-readable medium comprising processor-executable instructions configured to cause one or more processors to:
detect a head pose of a user in front of a computing device that has joined a video conference;
determine a first relative direction of a participant user interface (UI) element for an active speaker of the video conference with respect to the user based on a first location of the participant UI element and the head pose of the user, the participant UI element comprising a window and a representation of a corresponding participant presented in a UI for the video conference displayed on a screen of the computing device;
generate a spatial audio for at least an audio signal of the active speaker based on the first relative direction;
output the spatial audio;
generate an updated spatial audio based on a second relative direction, wherein the second relative direction is towards a second location closer to a center location on the screen that the user is facing than the first relative direction, the center location determined based on the head pose of the user; and
play the updated spatial audio.
16 . The non-transitory computer-readable medium of claim 15 , wherein the processor-executable instructions are configured to further cause the one or more processors to:
determine a third relative direction of a participant UI element for a second active speaker with respect to the user based on a third location of the participant UI element for the second active speaker and the head pose of the user,
wherein generating the spatial audio further comprising:
generating a first individual spatial audio for the audio signal of the active speaker based on the first relative direction and a second individual spatial audio for a second audio signal of the second active speaker based on the second relative direction, and
combining the first individual spatial audio and the second individual spatial audio.
17 . The non-transitory computer-readable medium of claim 15 , wherein the processor-executable instructions are configured to further cause the one or more processors to:
modifying the UI for the video conference by moving the participant UI element for the active speaker to the second location; and
presenting the modified UI on the computing device.
18 . The non-transitory computer-readable medium of claim 15 , wherein generating the spatial audio is performed using head-related transfer functions (HRTFs).
19 . The non-transitory computer-readable medium of claim 15 , wherein moving the participant UI element for the active speaker to the second location comprises switching the participant UI element for the active speaker with another participant UI element for another participant positioned at the second location.
20 . The non-transitory computer-readable medium of claim 15 , wherein the processor-executable instructions are configured to further cause the one or more processors to: stop generating the updated spatial audio and moving the participant UI element for the active speaker towards the center location based on determining that the active speaker is inactive for a threshold time period.