Systems and methods for presence-aware repositioning and reframing in video conferencing
Systems and methods are described herein for automatically reframing a video for a conference participant. Video of the participant is captured, and a first position of the participant is detected. An offset for the first position of the participant is then calculated to determine a relative distance from the center of the video frame. The captured video is modified based on the offset and is then presented in the video conference.
1. A method for automatically reframing a video conference participant in a video stream, the method comprising:
capturing video of the participant;
generating a pixel coordinate array for each respective frame of the video;
determining respective coordinates of a respective absolute center of each respective frame of the video;
detecting, within a first frame of the video, a first position of the participant within the pixel coordinate array;
calculating a participant center represented by participant center coordinates of an area of the pixel coordinate array of the first frame corresponding to the first position of the participant;
calculating an offset for the first position of the participant, wherein the offset is based at least in part on a difference between respective coordinates of a respective absolute center of the first frame and the participant center coordinates of the first frame;
modifying subsequent frames of the video based on the offset, wherein the subsequent frames are to be displayed after the first frame of the video; and
presenting the modified subsequent frames of the video of the participant.
2. The method of claim 1 , wherein modifying the subsequent frames of the video based on the offset further comprises translating the first position of the participant, based on the offset, to a second position corresponding to updated participant center coordinates.
3. The method of claim 2 , wherein translating the first position of the participant, based on the offset, to a second position further comprises:
calculating a translation vector; and
for each frame of the captured video, applying the translation vector to each pixel of a plurality of pixels of the respective frame that form an image of the participant.
4. The method of claim 1 , wherein modifying the subsequent frames of the video based on the offset further comprises cropping the video so that the first position of the participant is centered in the video.
5. The method of claim 1 , further comprising:
encoding a media stream including the captured video with the modified subsequent frames; and
transmitting, to a video conference server, the media stream.
6. The method of claim 1 , further comprising:
encoding a media stream including the captured video and the offset;
modifying the captured video to incorporate the modified subsequent frames; and
transmitting, to a video conference server, the media stream.
7. The method of claim 6 , wherein presenting the modified subsequent frames of the video of the participant further comprises:
retrieving, at the video conference server, from the media stream, the offset;
cropping, at the video conference server, the video based on the offset;
reencoding, at the video conference server, the cropped video in a second media stream; and
transmitting, from the video conference server, the second media stream to client devices associated with each participant in the video conference.
8. The method of claim 1 , further comprising:
detecting, within the video, a second position of a second participant; and
calculating a second offset for the second position of the second participant;
wherein modifying the video is further based on the second offset.
9. The method of claim 1 , further comprising:
determining a first resolution of the video;
determining a second resolution to which to scale the video to fit in a video conference layout;
scaling the video to the second resolution; and
adjusting the offset based on the scaling.
10. The method of claim 1 , further comprising:
monitoring the offset;
detecting a change in the offset; and
in response to detecting a change in the offset:
determining whether the offset has changed by at least a threshold amount; and
in response to determining that the offset has changed by at least the threshold amount, altering how the video is modified.
11. The method of claim 1 , further comprising:
monitoring the offset;
detecting a change in the offset; and
in response to detecting a change in the offset:
determining whether change in the offset is temporally stable; and
in response to determining that the change in the offset is temporally stable, altering how the video is modified.
12. A system for automatically reframing a video conference participant in a video stream, the system comprising:
video capture circuitry configured to capture video of the participant;
input/output circuitry; and
control circuitry configured to:
generate a pixel coordinate array for each respective frame of the video;
determine respective coordinates of a respective absolute center of each respective frame of the video;
detect, within a first frame of the video, a first position of the participant within the pixel coordinate array;
calculate a participant center represented by participant center coordinates of an area of the pixel coordinate array of the first frame corresponding to the first position of the participant;
calculate an offset for the first position of the participant, wherein the offset is based at least in part on a difference between respective coordinates of a respective absolute center of the first frame and the participant center coordinates of the first frame;
modify subsequent frames of the video based on the offset, wherein the subsequent frames are to be displayed after the first frame of the video; and
present, using the input/output circuitry, the modified subsequent frames of the video of the participant.
13. The system of claim 12 , wherein the control circuitry configured to modify the subsequent frames of the video based on the offset is further configured to translate the first position of the participant, based on the offset, to a second position corresponding to updated participant center coordinates.
14. The system of claim 13 , wherein the control circuitry configured to translate the first position of the participant, based on the offset, to a second position is further configured to:
calculate a translation vector; and
for each frame of the captured video, apply the translation vector to each pixel of a plurality of pixels of the respective frame that form an image of the participant.
15. The system of claim 12 , wherein the control circuitry configured to modify the subsequent frames of the video based on the offset is further configured to crop the video so that the first position of the participant is centered in the video.
16. The system of claim 12 , wherein the control circuitry is further configured to:
encode a media stream including the captured video with the modified subsequent frames; and
transmit, to a video conference server, the media stream.
17. The system of claim 12 , wherein the control circuitry is further configured to:
encode a media stream including the captured video and the offset;
modify the captured video to incorporate the modified subsequent frames; and
transmit, to a video conference server, the media stream.
18. The system of claim 12 , wherein the control circuitry is further configured to:
detect, within the video, a second position of a second participant; and
calculate a second offset for the second position of the second participant;
wherein the control circuitry configured to modify the video is further configured to do so based on the second offset.
19. The system of claim 12 , wherein the control circuitry is further configured to:
determine a first resolution of the video;
determine a second resolution to which to scale the video to fit in a video conference layout;
scale the video to the second resolution; and
adjust the offset based on the scaling.
20. The system of claim 12 , wherein the control circuitry is further configured to:
monitor the offset;
detect a change in the offset; and
in response to detecting a change in the offset:
determine whether the offset has changed by at least a threshold amount; and
in response to determining that the offset has changed by at least the threshold amount, alter how the video is modified.
21. The system of claim 12 , wherein the control circuitry is further configured to:
monitor the offset;
detect a change in the offset; and
in response to detecting a change in the offset:
determine whether change in the offset is temporally stable; and
in response to determining that the change in the offset is temporally stable, alter how the video is modified.