Augmenting foreground of a conference participant
A foreground portion and a background portion are obtained from an image of a conference participant. Using a background replacement image, a replacement background portion corresponding to the background portion is obtained. An indication of an image portion that is included in the image and is outside the foreground portion is obtained. An output image that includes the replacement background portion, the foreground portion, and the image portion is generated such that the foreground portion and the image portion are visible in the output image. The output image is transmitted.
1 . A method, comprising:
obtaining a foreground portion and a background portion from an image of a conference participant;
obtaining, using a background replacement image, a replacement background portion corresponding to the background portion;
obtaining, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein obtaining the indication of the image portion comprises:
determining that the gesture persists in images of the conference participant for a predetermined duration of time;
generating an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and
transmitting the output image.
2 . The method of claim 1 , further comprising:
receiving, from the conference participant, a bounding box surrounding another image portion; and
generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the output image.
3 . The method of claim 1 , further comprising:
obtaining an indication of another image portion verbally from the conference participant; and
generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the output image.
4 . The method of claim 1 , further comprising:
obtaining an indication of another image portion based on a pixel location received from a pointing device; and
generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the another output image.
5 . The method of claim 1 , further comprising:
receiving an indication of the background replacement image to blur the background portion,
wherein generating the output image comprises:
unblurring the image portion based on the indication of the image portion.
6 . The method of claim 1 , further comprising:
presenting a preceding image on a display of the conference participant, wherein the preceding image precedes the image in an image stream of the conference participant; and
obtaining another indication of another object based on a markup of the preceding image by the conference participant.
7 . The method of claim 1 , further comprising:
applying a filter to the image portion to obtain a filtered image portion, wherein generating the output image comprises:
generating the output image by combining the background portion, the foreground portion, and the filtered image portion.
8 . A device, comprising:
a memory; and
a processor, the processor configured to execute instructions stored in the memory to:
obtain a foreground portion and a background portion from an image of a conference participant;
obtain, using a background replacement image, a replacement background portion corresponding to the background portion;
obtain, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein to obtain the indication of the image portion comprises to:
determine that the gesture persists in images of the conference participant for a predetermined duration of time;
generate an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and
transmit the output image.
9 . The device of claim 8 , wherein to obtain the indication of the image portion, the processor is configured to execute the instructions to:
obtain the indication from a preceding image of the conference participant, wherein the preceding image precedes the image in an image stream of the conference participant.
10 . The device of claim 8 , wherein the processor is further configured to execute instructions to:
obtain another indication of another image portion by determining that the another indication is identified in a predetermined number of images of the conference participant.
11 . The device of claim 8 , wherein the processor is configured to execute the instructions to:
obtain, using a machine learning model for object detection, a list of objects from the image of the conference participant;
obtain, from the conference participant, a selection of at least one object of the list of objects; and
identify another image portion based on the at least one object.
12 . The device of claim 8 , wherein the processor is further configured to execute the instructions to:
receive a request from the conference participant to stop displaying the image portion; and
responsive to the request, generate subsequent output images from respective subsequent images of the conference participant, wherein each subsequent output image consists of a foreground portion obtained from a corresponding subsequent image and the replacement background portion.
13 . The device of claim 8 , wherein the processor is configured to execute the instructions to:
identify, in another image captured by a camera, another gesture of the conference participant;
determine that the conference participant is holding an object based on the another gesture; and
configure the camera to focus on the object.
14 . The device of claim 8 , wherein the processor is further configured to execute the instructions to:
perform optical character processing on the image portion to extract text included in the image portion; and
transmit the text to participants of the conference.
15 . The device of claim 8 , wherein to generate the output image, the processor is configured to execute the instructions to:
apply a special effect to the image portion to obtain a modified image portion; and
generate the output image by combining the replacement background portion, the foreground portion, and the modified image portion.
16 . The device of claim 8 , wherein to generate the output image, the processor is configured to execute the instructions to:
apply an anti-glare filter to the image portion to obtain a filtered image portion; and
generate the output image by combining the replacement background portion, the foreground portion, and the filtered image portion.
17 . A non-transitory computer readable medium that stores instructions operable to cause one or more processors to perform operations including:
obtaining a foreground portion and a background portion from an image of a conference participant;
obtaining, using a background replacement image, a replacement background portion corresponding to the background portion;
obtaining, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein obtaining the indication of the image portion comprises:
determining that the gesture persists in images of the conference participant for a predetermined duration of time;
generating an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and
transmitting the output image.
18 . The non-transitory computer readable medium of claim 17 , wherein the output image further comprises a highlight of the image portion.
19 . The method of claim 1 , further comprising:
determining that the conference participant is holding an object; and
configuring a camera that captured the image of the conference participant to focus on the object.
20 . The non-transitory computer readable medium of claim 17 , wherein the operations further comprise:
configuring a camera to focus on an object in response to determining that the conference participant is holding the object.