Video-conference device and method
A video-conference device for providing an augmented view of in-room participants in a video-conference. The image sensor is configured to capture images comprising the room and the in-room participants. The device comprises a depth sensor configured for measuring the distance to each of the in-room participants having a processing unit configured for providing the augmented view of the in-room. The processing unit is configured to performing an image processing of the captured images from the image sensor to virtually identify each of the in-room participants. The processing unit is configured to determining a desired virtual position for each of the in-room participants in the augmented view. The processing unit is configured to performing a virtual scaling of each of the in-room participants to the desired virtual size.
1 . A video-conference device for providing an augmented view of in-room participants in a video-conference, wherein the device is configured to be arranged in a room, a first number of participants includes a number of in-room participants present while the video-conference is held, and a second number of participants includes a number of far-end participants present at one or more different locations other than the room, wherein the device comprises:
an image sensor configured to capture images comprising the room and the in-room participants;
a depth sensor configured to measure a distance from the device to each of the in-room participants;
processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;
wherein the processing unit is configured to:
obtain the measurements of the distance from the device to each of the in-room participants by the depth sensor;
obtain the captured images by the images sensor;
perform an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;
determine a desired virtual position for each of the in-room participants in the augmented view;
determine a scaling factor for a desired virtual size of each of the in-room participants;
perform a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;
perform a virtual positioning of each of the in-room participants in the augmented view;
determine a distance A to each of the in-room participants thereby obtaining an actual position (x1,y1,z1) of each in-room participant;
determine a distance B between the in-room participants;
determine a relative location of the in-room participants based on the determined distance A and distance B;
determine the scaling factor N for the desired virtual size of each of the in-room participants based on the determined distance A and distance B; and
determine the virtual position (x2,y2,z2) of each of the in-room participants in the augmented view based on the determined relative location of the in-room participants.
2 . The device of claim 1 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the in-room participants are virtually positioned closer to each other and closer to the image sensor capturing the images.
3 . The device of claim 1 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the relative positions of the in-room participants relative to each other are maintained.
4 . The device of claim 1 , wherein the device is configured to be located at a non-central end of the room and the image sensor, the depth sensor and the processing unit of the device are also located in the room.
5 . The device of claim 1 , wherein the depth sensor comprises one or more of:
a time of flight (ToF) sensor configured for determining path lengths;
an infra-red sensor configured for determining distances based on reflected light; and
an acoustic sensor configured for determining distances based on reflected audio.
6 . The device according to claim 1 , wherein the image processing of the captured images to identify the in-room participants is performed by using one or more of:
virtual image segmentation;
virtual cut-out to virtually separate the in-room participants from the background;
virtual image resizing; and
virtual seam carving.
7 . The device of claim 1 , wherein the image processing is performed in two-dimensions (2D) and/or in three-dimensions (3D).
8 . The device to of claim 1 ,
wherein the depth sensor is configured to measure the distance to objects/parts of the room, and
wherein the image sensor is configured to capture images comprising the room with no in-room participants, and
wherein the processing unit is configured to:
map the room with no in-room participants, by determining distances C between the depth sensor and one or more objects/parts of the room.
9 . The device of claim 1 , wherein the processing unit is further configured to define a suitable illumination in the augmented view based on the time of day of the video-conference.
10 . The device of claim 1 , wherein the processing unit is further configured to perform blurring of objects in the augmented view of the room.
11 . A video-conference device for providing an augmented view of in-room participants in a video-conference, wherein the device is configured to be arranged in a room, a first number of participants includes a number of in-room participants present while the video-conference is held, and a second number of participants includes a number of far-end participants present at one or more different locations other than the room, wherein the device comprises:
an image sensor configured to capture images comprising the room and the in-room participants;
a depth sensor configured to measure a distance from the device to each of the in-room participants;
processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;
wherein the processing unit is configured to:
obtain the measurements of the distance from the device to each of the in-room participants by the depth sensor;
obtain the captured images by the images sensor;
perform an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;
determine a desired virtual position for each of the in-room participants in the augmented view;
determine a scaling factor for a desired virtual size of each of the in-room participants;
perform a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;
perform a virtual positioning of each of the in-room participants in the augmented view; and
in accordance with a determination that a criterion for the virtual positioning of an in-room participant is not satisfied:
forgo performing the virtual positioning of the in-room participant in the augmented view, and
virtually place the in-room participant in a dedicated picture/box in the augmented view.
12 . The device of claim 1 , wherein the device further comprises at least one of:
a display configured to display the augmented view and the far-end participants; or
an audio output transducer configured to transmit audio from the far-end participants to the room.
13 . A system comprising at least two video-conference devices according to claim 1 .
14 . A method, performed in a video-conference device, for providing an augmented view of in-room participants in the video-conference, wherein:
the device is configured to be arranged in a room,
a first number of participants is the in-room participants present in the room while the video-conference is held,
a second number of participants is far-end participants present at one or more different locations than the room,
the device comprises;
an image sensor configured to capture images comprising the room and the in-room participants;
a depth sensor configured to measure the distance to each of the in-room participants; and
a processing unit configured to provide the augmented view of the in-room participants by a processed/augmented video based on the captured images from the image sensor and the measurements by the depth sensor;
wherein the method comprises, by the processing unit:
obtaining the measurements of the distance to each of the in-room participants by the depth sensor;
obtaining the captured images by the images sensor;
performing an image processing of the captured images from the image sensor to virtually identify each of the in-room participants;
determining a desired virtual position for each of the in-room participants in the augmented view;
determining a scaling factor for a desired virtual size of each of the in-room participants;
performing a virtual scaling of each of the in-room participants to the desired virtual size in the augmented view;
performing a virtual positioning of each of the in-room participants in the augmented view;
determining a distance A to each of the in-room participants thereby obtaining an actual position (x1,y1,z1) of each in-room participant;
determining a distance B between the in-room participants;
determining a relative location of the in-room participants based on the determined distance A and distance B;
determining the scaling factor N for the desired virtual size of each of the in-room participants based on the determined distance A and distance B; and
determining the virtual position (x2,y2,z2) of each of the in-room participants in the augmented view based on the determined relative location of the in-room participants.
15 . The device of claim 11 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the in-room participants are virtually positioned closer to each other and closer to the image sensor capturing the images.
16 . The device of claim 11 , wherein the processing unit is configured to perform the virtual positioning of each of the in-room participants in the augmented view such that the relative positions of the in-room participants relative to each other are maintained.
17 . The device of claim 11 , wherein the device is configured to be located at a non-central end of the room and the image sensor, the depth sensor and the processing unit of the device are also located in the room.
18 . The device of claim 11 , wherein the depth sensor comprises one or more of:
a time of flight (ToF) sensor configured for determining path lengths;
an infra-red sensor configured for determining distances based on reflected light; and
an acoustic sensor configured for determining distances based on reflected audio.
19 . The device according to claim 11 , wherein the image processing of the captured images to identify the in-room participants is performed by using one or more of:
virtual image segmentation;
virtual cut-out to virtually separate the in-room participants from the background;
virtual image resizing;
virtual seam carving.