ENHANCED MULTI-VIEW BACKGROUND MATTING FOR VIDEO CONFERENCING
This disclosure describes systems, methods, and devices related to video background matte generation. A method may include receiving, by a first neural network trained to generate alpha mattes and foreground multi-camera images, first inputs generated by a second neural network; receiving, by the first neural network, second inputs comprising grayscale images and depth maps of the multi-camera images; and generating, by the first neural network, based on the first inputs and the second inputs, multi-view alpha mattes and multi-view foreground estimates for the multi-camera images.
1 . A method for generating multi-camera background mattes for video, the method comprising:
receiving, by a first neural network trained to generate alpha mattes and foreground estimates of multi-camera images, first inputs generated by a second neural network;
receiving, by the first neural network, second inputs comprising grayscale images and depth maps of the multi-camera images; and
generating, by the first neural network, based on the first inputs and the second inputs, multi-view alpha mattes and multi-view foreground estimates for the multi-camera images.
2 . The method of claim 1 , wherein the multi-view alpha mattes and the multi-view foreground estimates are generated based on a loss function.
3 . The method of claim 2 , wherein the loss function minimizes a difference between predicted alpha mattes and ground truth alpha mattes for the multi-camera images.
4 . The method of claim 3 , wherein the first neural network is trained using the ground truth alpha mattes.
5 . The method of claim 2 , wherein the loss function minimizes a difference between predicted foreground estimates and ground truth foreground estimates for the multi-camera images.
6 . The method of claim 5 , wherein the first neural network is trained using the ground truth foreground estimates.
7 . The method of claim 1 , wherein the first neural network is a deep learning-based neural network.
8 . The method of claim 1 , wherein the first inputs comprise alpha mattes for each camera view of the multi-camera images.
9 . A non-transitory computer-readable medium storing computer-executable instructions, associated with video background matte generation, which when executed by one or more processors result in performing operations comprising:
receive, using a first neural network trained to generate alpha mattes and foreground estimates of multi-camera images, first inputs generated by a second neural network;
receive, using the first neural network, second inputs comprising grayscale images and depth maps of the multi-camera images; and
generate, using the first neural network, based on the first inputs and the second inputs, multi-view alpha mattes and multi-view foreground estimates for the multi-camera images.
10 . The non-transitory computer-readable medium of claim 9 , wherein the multi-view alpha mattes and the multi-view foreground estimates are generated based on a loss function.
11 . The non-transitory computer-readable medium of claim 10 , wherein the loss function minimizes a difference between predicted alpha mattes and ground truth alpha mattes for the multi-camera images.
12 . The non-transitory computer-readable medium of claim 11 , wherein the first neural network is trained using the ground truth alpha mattes.
13 . The non-transitory computer-readable medium of claim 10 , wherein the loss function minimizes a difference between predicted foreground estimates and ground truth foreground estimates for the multi-camera images.
14 . The non-transitory computer-readable medium of claim 13 , wherein the first neural network is trained using the ground truth foreground estimates.
15 . The non-transitory computer-readable medium of claim 9 , wherein the first neural network is a deep learning-based neural network.
16 . The non-transitory computer-readable medium of claim 9 , wherein the first inputs comprise alpha mattes for each camera view of the multi-camera images.
17 . A device for video background matte generation, the device comprising memory storing instructions associated with the video background matte generation, the memory coupled to at least one processor configured to:
receive, using a first neural network trained to generate alpha mattes and foreground estimates of multi-camera images, first inputs generated by a second neural network;
receive, using the first neural network, second inputs comprising grayscale images and depth maps of the multi-camera images; and
generate, using the first neural network, based on the first inputs and the second inputs, multi-view alpha mattes and multi-view foreground estimates for the multi-camera images.
18 . The device of claim 17 , wherein the multi-view alpha mattes and the multi-view foreground estimates are generated based on a loss function.
19 . The device of claim 18 , wherein the loss function minimizes a difference between predicted alpha mattes and ground truth alpha mattes for the multi-camera images.
20 . The device of claim 19 , wherein the first neural network is trained using the ground truth alpha mattes.