Extraneous video element detection and modification
A computer implemented method includes receiving images from a device camera during a video conference and processing the received images via a machine learning model trained on labeled image training data to detect extraneous image portions. Extraneous image pixels associated with the extraneous image portions are identified and replaced with replacement image pixels from previously stored image pixels to form modified images. The modified images may be transmitted during the video conference.
1 . A computer implemented method comprising:
receiving images of a user from a device camera during a video conference;
processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;
identifying extraneous image pixels associated with the extraneous image portions capturing extraneous motions of the user; and
replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.
2 . The method of claim 1 and further comprising transmitting the modified images to other devices on the video conference.
3 . The method of claim 1 wherein the received images and labeled image training data have a field of view wider than the modified images.
4 . The method of claim 1 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising unnecessary arm motion or eating activities of the user.
5 . The method of claim 1 and further comprising:
receiving audio from the video conference;
detecting that the user of the device is being addressed by processing the received audio via a natural language model trained to recognize a name of the user; and
alerting the user to discontinue activity resulting in extraneous image portions.
6 . The method of claim 1 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising a changed direction of gaze of the user.
7 . The method of claim 1 and further comprising:
receiving audio signals from the device;
processing the audio signals to detect that the user is speaking; and
suspending replacing of the extraneous image pixels in response to detecting that the user is speaking.
8 . The method of claim 1 and further comprising:
detecting that the user device has unmuted the device; and
suspending replacing of the extraneous image pixels in response to detecting that the user device is unmuted.
9 . The method of claim 1 wherein the motion detection labeled video image training data is identified via Motion Picture Expert's Group (MPEG) processing techniques.
10 . The method of claim 1 wherein the previously stored image pixels comprise reference images without extraneous image pixels captured during the video conference.
11 . The method of claim 1 wherein the previously stored image pixels comprise reference images without extraneous image pixels captured prior to the video conference.
12 . The method of claim 1 wherein the previously stored image pixels comprise reference images and further comprising:
comparing the received images having extraneous image pixels to the reference images;
identifying for each received image, a closest reference image; and
identifying the replacement image pixels from such closest reference images.
13 . The method of claim 1 and further comprising receiving a user-initiated replacement signal that triggers replacing the image pixels with replacement image pixels until a stop replacement signal is received.
14 . The method of claim 1 wherein the previously stored image pixels comprise reference images and further comprising suspending replacing the extraneous image pixels with replacement image pixels by transitioning between a reference video and a live video.
15 . The method of claim 14 wherein transitioning is performed using a generator network.
16 . A machine-readable storage device having instructions for execution by a processor of a machine to cause the processor to perform operations to perform a method, the operations comprising:
receiving images of a user from a device camera during a video conference;
processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;
identifying extraneous image pixels associated with the extraneous image portions capturing extraneous motions of the user; and
replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.
17 . The device of claim 16 and further comprising transmitting the modified images to other devices on the video conference.
18 . The device of claim 16 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising unnecessary arm motion or eating activities of the user.
19 . The device of claim 16 and further comprising:
receiving audio from the video conference;
detecting that the user of the device is being addressed by processing the received audio via a natural language model trained to recognize a name of the user, and
alerting the user to discontinue activity resulting in extraneous image portions.
20 . A device comprising:
a processor; and
a memory device coupled to the processor and having a program stored thereon for execution by the processor to perform operations comprising:
receiving images of a user from a device camera during a video conference;
processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;
identifying extraneous image pixels associated with the extraneous image portions of the user; and
replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.