AUDIO SOURCE POSITIONING USING A CAMERA
Audio source positioning technique embodiments are presented that are employed in a video teleconference or telepresence session between a local site and one or more remote sites. Each of these sites has one participant, and a virtual scene is constructed and displayed at each site that depicts each of the participants from the other sites in the constructed scene. However, rather than simply playing audio captured at the other site or sites in the viewing participant's site, audio source positioning is used to make it seem to a participant viewing a rendering of the virtual scene that the voice of another participant is emanating from a location on the display device where the remote participant is depicted.
1 . A computer-implemented process for audio source positioning in a video teleconference or telepresence session between a local site and one or more remote sites, each of said sites having one or more participants, comprising for the local site:
using a computing device to perform the following process actions:
receiving from each remote site, scene proxies representing successive scene proxy frames transmitted by a remote site over a data communication network;
receiving from at least one remote site, along with each frame of scene proxies received from the site,
audio data representing each remote site participant's voice captured, if any, during the time period between the currently received frame and the next frame of scene proxies to be received from the remote site, and
a 3D point representing the location of each participant in the remote site;
for each frame of scene proxies received from a remote site if there is only one remote site sending frames, or for each group of frames of scene proxies contemporaneously received from remote sites if there are multiple remote sites sending frames,
rendering a frame of a virtual scene comprising a depiction of each of the remote site participants from the last-received frame or frames of scene proxies, and
displaying the rendered frame to the local site participant or participants via a display device;
for each remote site participant depicted in the last-rendered frame of the virtual scene that is resident at a remote site that sent audio data representing the remote site participant's voice and the 3D point representing the location of the participant in the remote site, employing a spatial audio technique to make it seem to each local site participant that the voice of the remote site participant is emanating from a location on the display device where the remote participant is depicted using the audio data and the 3D point representing the location of the participant in the remote site that was received from the remote site in conjunction with the last-received frame of scene proxies.
2 . The process of claim 1 , wherein said 3D point representing the location of a participant in a remote site is a 3D point representing the location of the participant's mouth in that remote site for each remote site participant depicted in the last-rendered frame of a virtual scene whose mouth is visible.
3 . The process of claim 1 , wherein said 3D point representing the location of a participant in a remote site is a 3D point representing the location of the participant's head in that remote site for each remote site participant depicted in the last-rendered frame of a virtual scene whose mouth is not visible.
4 . The process of claim 1 , wherein the process action of rendering a frame of the virtual scene, comprises, for each remote site, computing a first transform that converts 3D locations in the remote site to points in the frame of the virtual scene, and wherein the process action of displaying the rendered frame to the local site participant via a display device, comprises computing a second transform that converts points in a frame of the virtual scene to screen coordinates on the display device.
5 . The process of claim 4 , wherein the process action of employing a spatial audio technique to make it seem to the local site participant that the voice of a remote site participant is emanating from a location on the display device where the remote participant is depicted using the audio data and the 3D point representing the location of a remote participant in the remote site that was received in conjunction with the last-received frame of scene proxies from the remote site, comprises the actions of:
employing the first transform computed to convert 3D locations in the remote site to points in the last-rendered frame of the virtual scene, to convert the 3D point representing the location of the remote participant in the remote site to a point in the last-rendered frame of the virtual scene;
employing the second transform computed to convert points in a frame of the virtual scene to screen coordinates on the display device, to convert the point in the last-rendered frame of the virtual scene representing the remote participant location to screen coordinates on the display device;
employing a third transform that converts screen coordinates in the display device to 3D points in the local site to compute the 3D point in the local site of the screen coordinates representing the location of the remote participant depicted on the display device; and
employing said spatial audio technique and a plurality of audio speakers resident in the local site to make it seem to the local site participant that the voice of the remote site participant is emanating from the computed 3D point in the local site of the screen coordinates representing the location of the remote participant depicted on the display device.
6 . The process action of claim 1 , wherein the process action of employing a spatial audio technique to make it seem to a local site participant that the voice of the remote site participant is emanating from a location on the display device where the remote participant is depicted, further comprises the actions of:
tracking the head of the local site participant and periodically computing a 3D point representative of the location of the local site participant's head in the local site; and
each time a 3D point representative of the location of the local site participant's head in the local site is computed, employing the spatial audio technique to make it seem to the local site participant that the voice of the remote site participant is emanating from a location on the display device where the remote participant is depicted taking into consideration the last-computed 3D point representative of the location of the local site participant's head.
7 . The process of claim 6 , wherein the process action of periodically computing a 3D point representative of the location of a local site participant's head in the local site, comprises computing a 3D point representative of the location of the local site participant's head in the local site at a rate that exceeds the rate at which frames of the virtual scene are computed.
8 . The process of claim 1 , wherein said audio data representing a remote site participant's voice received from a remote site has been modified so as to suppress reverberations and noise in the audio captured at that remote site.
9 . The process of claim 8 , wherein the process action of rendering a frame of the virtual scene, comprises for each remote site, computing a first transform that converts 3D locations in the remote site to points in the frame of the virtual scene.
10 . The process of claim 9 , further comprising the process actions of:
for each remote site participant depicted in the last-rendered frame of the virtual scene that is resident at a remote site that sent audio data representing the remote site participant's voice and the 3D point representing the location of the participant in the remote site,
employing the first transform computed to convert 3D locations in the remote site to points in the last-rendered frame of the virtual scene, to convert the 3D point representing the location of the remote participant in the remote site to a point in the last-rendered frame of the virtual scene, wherein said 3D point representing the location of the remote participant in the remote site corresponds to a 3D point representing the location of the remote participant's mouth in the remote site,
identifying the orientation of the remote site participant's face in the virtual scene as depicted in the last-rendered virtual scene frame,
computing the direction from the point in the last-rendered frame of the virtual scene that corresponds to the 3D point representing the location of the remote participant's mouth in the remote site that the remote participant's voice projects in the virtual space based on the orientation of the remote site participant's face in the virtual scene,
estimating the reverberation characteristics of the virtual scene as depicted in the last-rendered virtual scene frame,
computing reverberation audio data that when added to the received audio data simulates the reverberations of the remote participant's voice in the virtual scene as spoken from the point representing the location of the remote participant's mouth in the virtual scene in the computed direction, and
adding the computed reverberation audio data into audio played in the local site in conjunction with the display of the virtual scene frame.
11 . A computer-implemented process for facilitating audio source positioning at a remote site in a video teleconference or telepresence session between a local site and the remote site, each of said sites having one or more participants, comprising for the local site:
using a computing device to perform the following process actions:
inputting streams of sensor data generated from an arrangement of sensors that capture participant data, said arrangement comprising a plurality of video and audio devices which generate a plurality of streams of sensor data, each video capture device of which captures the participant from a different geometric perspective, and each audio capture device of which captures the voice of the participant at the local site;
generating scene proxies from the streams of sensor data which geometrically describes the local site including the participant on a frame by frame basis;
employing the streams of sensor data and a face tracking technique to identify a 3D point representing the location of the participant in the local site for each frame of the scene proxies; and
transmitting the scene proxies representing each frame in the order generated over a data communication network to the remote site, along with,
audio data representing each local site participant's voice captured, if any, during the time period between the frame currently being transmitted and next frame of scene proxies to be transmitted, and
the 3D point coordinates representing the location of each participant in the local site identified for the frame currently being transmitted.
12 . The process of claim 11 , wherein said 3D point representing the location of a participant in the local site is a 3D point representing the location of the participant's head in the local site.
13 . The process of claim 11 , wherein said 3D point representing the location of a participant in the local site is a 3D point representing the location of the participant's mouth in the local site.
14 . The process of claim 11 , wherein prior to performing the process action of transmitting audio data representing a local site participant's voice, performing an action of suppressing reverberations and noise in the audio data.
15 . A computer-implemented process for audio source positioning in a video teleconference or telepresence session between two non co-located sites, each of said sites having one participant, comprising for a first of the two sites:
using a computing device to perform the following process actions:
receiving from the other site, scene proxies representing successive scene proxy frames transmitted by the other site over a data communication network, along with for each scene proxy frame received,
audio data representing the other site participant's voice captured, if any, during the time period between the currently received frame and the next frame of scene proxies to be received from the other site, and
a 3D point representing the location of the participant in the other site;
for each frame of scene proxies received from the other site, rendering a frame of a virtual scene comprising a depiction of the other site's participant from the last-received frame of scene proxies and displaying the rendered frame to the first site participant via a display device; and
whenever audio data representing the other site participant's voice is received, employing a spatial audio technique to make it seem to the first site participant that the voice of the other site participant is emanating from a location on the display device where the other site participant is depicted using the audio data and the 3D point representing the location of the participant in the other site that was received from the other site in conjunction with the last-received frame of scene proxies.
16 . The process of claim 15 , wherein said 3D point representing the location of the participant in the other site is a 3D point representing the location of the participant's mouth in the other site whenever the other site participant's mouth is visible in the last-rendered frame of a virtual scene.
17 . The process of claim 15 , wherein said 3D point representing the location of the participant in the other site is a 3D point representing the location of the participant's head in the other site whenever the other site participant's mouth is not visible in the last-rendered frame of a virtual scene.
18 . The process of claim 15 , wherein the process action of rendering a frame of the virtual scene, comprises, computing a first transform that converts 3D locations in the other site to points in the frame of the virtual scene, and wherein the process action of displaying the rendered frame to the first site participant via a display device, comprises computing a second transform that converts points in a frame of the virtual scene to screen coordinates on the display device.
19 . The process of claim 18 , wherein the process action of employing a spatial audio technique to make it seem to the first site participant that the voice of the other site participant is emanating from a location on the display device where the other site participant is depicted using the audio data and the 3D point representing the location of the other participant in the other site that was received in conjunction with the last-received frame of scene proxies, comprises the actions of:
employing the first transform computed to convert 3D locations in the other site to points in the last-rendered frame of the virtual scene, to convert the 3D point representing the location of the other site participant in the other site to a point in the last-rendered frame of the virtual scene;
employing the second transform computed to convert points in a frame of the virtual scene to screen coordinates on the display device, to convert the point in the last-rendered frame of the virtual scene representing the other participant's location to screen coordinates on the display device;
employing a third transform that converts screen coordinates in the display device to 3D points in the first site to compute the 3D point in the first site of the screen coordinates representing the location of the other site participant depicted on the display device; and
employing said spatial audio technique and a plurality of audio speakers resident in the first site to make it seem to the first site participant that the voice of the other site participant is emanating from the computed 3D point in the first site of the screen coordinates representing the location of the other participant depicted on the display device.
20 . The process action of claim 15 , wherein the process action of employing a spatial audio technique to make it seem to the first site participant that the voice of the other site participant is emanating from a location on the display device where the other participant is depicted, further comprises the actions of:
tracking the head of the first site participant and periodically computing a 3D point representative of the location of the first site participant's head in the first site; and
each time a 3D point representative of the location of the first site participant's head in the first site is computed, employing the spatial audio technique to make it seem to the first site participant that the voice of the other site participant is emanating from a location on the display device where the other participant is depicted taking into consideration the last-computed 3D point representative of the location of the first site participant's head.