Iterative background generation for video streams
Systems and methods for iterative background generation for video streams are provided. A first background layer and a first foreground layer of a first frame of a video stream are determined. A second background layer and a second foreground layer of a second frame of the video stream are determined. The first and second background layers are combined. The combined background layer includes a region obscured by both the first and the second foreground layers. An inpainting of the obscured region is performed to obtain an image of the obscured region.
1 . A method comprising:
determining a first background layer and a first foreground layer of a first frame of a video stream provided by a client device associated with a first participant of a plurality of participants of a video conference;
determining, during the video conference, a second background layer and a second foreground layer of a second frame of the video stream;
combining, during the video conference, the first background layer and the second background layer to obtain a combined background layer, wherein the combined background layer comprises a region obscured by both the first foreground layer and the second foreground layer;
performing, during the video conference and using a generative machine learning model, inpainting of the obscured region to obtain an image of the obscured region;
modifying, during the video conference and using the image of the obscured region and the combined background layer, background layers of subsequent frames of the video stream;
performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed, wherein the iterative modifications are performed until a fidelity level of the combined background layer satisfies a fidelity level criterion, wherein the fidelity level is based on an area of the combined background layer compared to an area of the image of the obscured region combined with the area of the combined background layer; and
providing, during the video conference, the video stream with modified background layers for presentation on one or more client devices of one or more of the plurality of participants of the video conference.
2 . The method of claim 1 , wherein determining, during the video conference, the first background layer and the first foreground layer of the video stream comprises:
providing the first frame of the video stream as input to a machine learning model, wherein the machine learning model is trained to predict, based on a given frame, segmentation labels for the given frame that represent foreground and background regions of the given frame;
obtaining a plurality of outputs from the machine learning model, wherein the plurality of outputs comprises one or more background regions and one or more foreground regions;
combining the one or more background regions to obtain the first background layer; and
combining the one or more foreground regions to obtain the first foreground layer.
3 . The method of claim 1 , wherein performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed comprises:
determining a third background layer of a third frame of the video stream;
determining a shared region of the image that shares a common area with the third background layer; and
modifying the image to replace a portion of the image corresponding to the shared region with a portion of the third background layer corresponding to the shared region.
4 . The method of claim 3 , further comprising ceasing the iterative modifications on the image in response to satisfying one or more criteria.
5 . The method of claim 4 , wherein the one or more criteria comprise the fidelity level criterion that is satisfied when the fidelity level exceeds a threshold fidelity level.
6 . The method of claim 4 , wherein the one or more criteria comprise at least one of exceeding a threshold amount of time or a threshold number of frames of the video stream.
7 . The method of claim 4 , further comprising resuming the iterative modifications on the image in response to detecting movement within the video stream.
8 . A system comprising:
a memory device; and
a processing device coupled to the memory device, the processing device to perform operations comprising:
determining a first background layer and a first foreground layer of a first frame of a video stream provided by a client device associated with a participant of a plurality of participants of a video conference;
determining, during the video conference, a second background layer and a second foreground layer of a second frame of the video stream;
combining, during the video conference, the first background layer and the second background layer to obtain a combined background layer, wherein the combined background layer comprises a region obscured by both the first foreground layer and the second foreground layer;
performing, during the video conference and using a generative machine learning model, inpainting of the obscured region to obtain an image of the obscured region;
modifying, during the video conference and using the image of the obscured region and the combined background layer, background layers of subsequent frames of the video stream;
performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed, wherein the iterative modifications are performed until a fidelity level of the combined background layer satisfies a fidelity level criterion, wherein the fidelity level is based on an area of the combined background layer compared to an area of the image of the obscured region combined with the area of the combined background layer; and
providing, during the video conference, the video stream with modified background layers for presentation on one or more client devices of one or more of the plurality of participants of the video conference.
9 . The processing device of claim 8 , wherein determining, during the video conference, the first background layer and the first foreground layer of the video stream comprises:
providing the first frame of the video stream as input to a machine learning model, wherein the machine learning model is trained to predict, based on a given frame, segmentation labels for the given frame that represent foreground and background regions of the given frame;
obtaining a plurality of outputs from the machine learning model, wherein the plurality of outputs comprises one or more background regions and one or more foreground regions;
combining the one or more background regions to obtain the first background layer; and
combining the one or more foreground regions to obtain the first foreground layer.
10 . The processing device of claim 8 , wherein performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed comprises:
determining a third background layer of a third frame of the video stream;
determining a shared region of the image that shares a common area with the third background layer; and
modifying the image to replace a portion of the image corresponding to the shared region with a portion of the third background layer corresponding to the shared region.
11 . The processing device of claim 10 , further comprising ceasing the iterative modifications on the image in response to satisfying one or more criteria.
12 . The processing device of claim 11 , wherein the one or more criteria comprise the fidelity level criterion that is satisfied when the fidelity level exceeds a threshold fidelity level.
13 . The processing device of claim 11 , wherein the one or more criteria comprise at least one of exceeding a threshold amount of time or a threshold number of frames of the video stream.
14 . The processing device of claim 11 , further comprising resuming the iterative modifications on the image in response to detecting movement within the video stream.
15 . A non-transitory computer-readable storage medium comprising instructions for a server that, when executed by a processing device, cause the processing device to perform operations comprising:
determining a first background layer and a first foreground layer of a first frame of a video stream provided by a client device associated with a participant of a plurality of participants of a video conference;
determining, during the video conference, a second background layer and a second foreground layer of a second frame of the video stream;
combining, during the video conference, the first background layer and the second background layer to obtain a combined background layer, wherein the combined background layer comprises a region obscured by both the first foreground layer and the second foreground layer;
performing, during the video conference and using a generative machine learning model, inpainting of the obscured region to obtain an image of the obscured region;
modifying, during the video conference and using the image of the obscured region and the combined background layer, background layers of subsequent frames of the video stream;
performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed, wherein the iterative modifications are performed until a fidelity level of the combined background layer satisfies a fidelity level criterion, wherein the fidelity level is based on an area of the combined background layer compared to an area of the image of the obscured region combined with the area of the combined background layer; and
providing, during the video conference, the video stream with modified background layers for presentation on one or more client devices of one or more of the plurality of participants of the video conference.
16 . The non-transitory computer-readable storage medium of claim 15 , wherein determining, during the video conference, the first background layer and the first foreground layer of the video stream comprises:
providing the first frame of the video stream as input to a machine learning model, wherein the machine learning model is trained to predict, based on a given frame, segmentation labels for the given frame that represent foreground and background regions of the given frame;
obtaining a plurality of outputs from the machine learning model, wherein the plurality of outputs comprises one or more background regions and one or more foreground regions;
combining the one or more background regions to obtain the first background layer; and
combining the one or more foreground regions to obtain the first foreground layer.
17 . The non-transitory computer-readable storage medium of claim 15 , wherein performing iterative modifications on the image for subsequent frames of the video stream as portions of the obscured region are revealed comprises:
determining a third background layer of a third frame of the video stream;
determining a shared region of the image that shares a common area with the third background layer; and
modifying the image to replace a portion of the image corresponding to the shared region with a portion of the third background layer corresponding to the shared region.