IP Library Granted Patent US 12,646,358
Granted Patent B2
US 12,646,358 · App. 17/543,579 · Granted Jun 2, 2026

Extraneous video element detection and modification

Inventors: Howard J. Locker (Cary, NC); John Weldon Nicholson (Cary, NC); Daryl C Cromer (Raleigh, NC); David Alexander Schwarz (Morrisville, NC); Mounika Vanka (Durham, NC)
Assignee: Lenovo (United States) Inc.
G06V40/28G06F3/013G06N20/00G06V20/46G06V40/176G10L15/18G10L15/24
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,358
App. No.
17/543,579
Granted
Jun 2, 2026
Kind
B2
Abstract

A computer implemented method includes receiving images from a device camera during a video conference and processing the received images via a machine learning model trained on labeled image training data to detect extraneous image portions. Extraneous image pixels associated with the extraneous image portions are identified and replaced with replacement image pixels from previously stored image pixels to form modified images. The modified images may be transmitted during the video conference.

Claims (48)

1 . A computer implemented method comprising:

receiving images of a user from a device camera during a video conference;

processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;

identifying extraneous image pixels associated with the extraneous image portions capturing extraneous motions of the user; and

replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.

2 . The method of claim 1 and further comprising transmitting the modified images to other devices on the video conference.

3 . The method of claim 1 wherein the received images and labeled image training data have a field of view wider than the modified images.

4 . The method of claim 1 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising unnecessary arm motion or eating activities of the user.

5 . The method of claim 1 and further comprising:

receiving audio from the video conference;

detecting that the user of the device is being addressed by processing the received audio via a natural language model trained to recognize a name of the user; and

alerting the user to discontinue activity resulting in extraneous image portions.

6 . The method of claim 1 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising a changed direction of gaze of the user.

7 . The method of claim 1 and further comprising:

receiving audio signals from the device;

processing the audio signals to detect that the user is speaking; and

suspending replacing of the extraneous image pixels in response to detecting that the user is speaking.

8 . The method of claim 1 and further comprising:

detecting that the user device has unmuted the device; and

suspending replacing of the extraneous image pixels in response to detecting that the user device is unmuted.

9 . The method of claim 1 wherein the motion detection labeled video image training data is identified via Motion Picture Expert's Group (MPEG) processing techniques.

10 . The method of claim 1 wherein the previously stored image pixels comprise reference images without extraneous image pixels captured during the video conference.

11 . The method of claim 1 wherein the previously stored image pixels comprise reference images without extraneous image pixels captured prior to the video conference.

12 . The method of claim 1 wherein the previously stored image pixels comprise reference images and further comprising:

comparing the received images having extraneous image pixels to the reference images;

identifying for each received image, a closest reference image; and

identifying the replacement image pixels from such closest reference images.

13 . The method of claim 1 and further comprising receiving a user-initiated replacement signal that triggers replacing the image pixels with replacement image pixels until a stop replacement signal is received.

14 . The method of claim 1 wherein the previously stored image pixels comprise reference images and further comprising suspending replacing the extraneous image pixels with replacement image pixels by transitioning between a reference video and a live video.

15 . The method of claim 14 wherein transitioning is performed using a generator network.

16 . A machine-readable storage device having instructions for execution by a processor of a machine to cause the processor to perform operations to perform a method, the operations comprising:

receiving images of a user from a device camera during a video conference;

processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;

identifying extraneous image pixels associated with the extraneous image portions capturing extraneous motions of the user; and

replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.

17 . The device of claim 16 and further comprising transmitting the modified images to other devices on the video conference.

18 . The device of claim 16 wherein the machine learning model is trained on video images labeled as including extraneous image portions comprising unnecessary arm motion or eating activities of the user.

19 . The device of claim 16 and further comprising:

receiving audio from the video conference;

detecting that the user of the device is being addressed by processing the received audio via a natural language model trained to recognize a name of the user, and

alerting the user to discontinue activity resulting in extraneous image portions.

20 . A device comprising:

a processor; and

a memory device coupled to the processor and having a program stored thereon for execution by the processor to perform operations comprising:

receiving images of a user from a device camera during a video conference;

processing the received images via a machine learning model trained on motion detection labeled video image training data to detect extraneous image portions capturing extraneous motions of the user;

identifying extraneous image pixels associated with the extraneous image portions of the user; and

replacing the extraneous image pixels with replacement image pixels of the user from previously stored image pixels that do not contain extraneous image portions to form modified images.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: LENOVO (UNITED STATES) INC.
To: LENOVO (SINGAPORE) PTE. LTD.
Reel/Frame 061880/0110 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 6, 2021
From: LOCKER, HOWARD J.; NICHOLSON, JOHN WELDON; SCHWARZ, DAVID ALEXANDER; CROMER, DARYL C.; VANKA, MOUNIKA
To: LENOVO (UNITED STATES) INC.
Reel/Frame 058312/0747 →
Continuity (1)
Related Publication 20230177884A1 · Jun 8, 2023
References Cited (9)
US 7564476B1 · Coughlan · 2009 [cited by examiner]
US 9888211B1 · Browne · 2018 [cited by examiner]
US 20120327176A1 · Kee · 2012 [cited by examiner]
US 20150030314A1 · Skarakis · 2015 [cited by examiner]
US 20220060525A1 · Chavez · 2022 [cited by examiner]
US 20220141396A1 · Ruan · 2022 [cited by examiner]
US 20240121358A1 · Wang · 2024 [cited by examiner]
WO WO2022191848A1 · 2022 [cited by examiner]
Crowley, J. L., Coutaz, J., and Berard, F. Things that see. Communications of the ACM, 43(3), 54-64, 2000. (Year: 2000). [cited by examiner]