IP Library Granted Patent US 12,340,489
Granted Patent B2
US 12,340,489 · App. 17/697,242 · Granted Jun 24, 2025

Object removal during video conferencing

Inventors: John Weldon Nicholson (Cary, NC); Howard J. Locker (Cary, NC); Daryl C Cromer (Raleigh, NC)
Assignee: Lenovo (United States) Inc.
G06T5/77G06T7/50G06V20/40G06V20/50G06T2207/10016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,340,489
App. No.
17/697,242
Granted
Jun 24, 2025
Kind
B2
Abstract

A video conferencing system includes an image of a participant in a video conference and a depth map of the image. The system identifies objects in the background of the image, identifies objects in the foreground of the image, and identifies objects in the middle-ground of the image. The system removes the objects from the middle-ground, and replaces the removed objects from the middle-ground with the objects from the background that are located behind the removed objects. The system then uses the image with the removed and replaced objects in a video stream of the video conference.

Claims (47)

1. A process comprising:

receiving into a computer processor an image of a participant in a video conference;

receiving a depth map of the image;

identifying one or more objects forming a background in the image using the depth map;

identifying one or more objects forming a foreground in the image using the depth map;

identifying one or more objects forming a middle-ground in the image using the depth map, wherein the middle-ground comprises an area between an object or a person in the foreground and the background;

removing the one or more objects from the middle-ground;

replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects; and

using the image with the removed and replaced one or more objects in a video stream associated with the video conference;

wherein the identifying the one or more objects forming a middle-ground in the image using the depth map, the removing the one or more objects from the middle-ground, the replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects, and the using the image with the removed and replaced one or more objects in a video stream associated with the video conference are executed a plurality of times during the video conference.

2. The process of claim 1 , wherein the image is received into the computer processor prior to the participant joining the video conference.

3. The process of claim 1 , wherein the one or more objects forming the foreground comprises at least the participant.

4. The process of claim 1 , wherein the depth map is created using one or more of a structured light (SL) camera or a time-of-flight (TOF) camera.

5. The process of claim 1 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced using a machine learning algorithm.

6. The process of claim 1 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced by synthesizing the one or more objects from the background in the image that were occluded by the removed and replaced one or more objects from the middle ground.

7. A non-transitory machine-readable medium comprising instructions that when executed by a computer processor execute a process comprising:

receiving into the computer processor an image of a participant in a video conference;

receiving a depth map of the image;

identifying one or more objects forming a background in the image using the depth map;

identifying one or more objects forming a foreground in the image using the depth map;

identifying one or more objects forming a middle-ground in the image using the depth map;

removing the one or more objects from the middle-ground;

replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects; and

using the image with the removed and replaced one or more objects in a video stream associated with the video conference;

wherein the identifying the one or more objects forming a middle-ground in the image using the depth map, the removing the one or more objects from the middle-ground, the replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects, and the using the image with the removed and replaced one or more objects in a video stream associated with the video conference are executed a plurality of times during the video conference.

8. The non-transitory machine-readable medium of claim 7 , wherein the image is received into the computer processor prior to the participant joining the video conference.

9. The non-transitory machine-readable medium of claim 7 , wherein the one or more objects forming the foreground comprises at least the participant.

10. The non-transitory machine-readable medium of claim 7 , wherein the depth map is created using one or more of a structured light (SL) camera or a time-of-flight (TOF) camera.

11. The non-transitory machine-readable medium of claim 7 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced using a machine learning algorithm.

12. The non-transitory machine-readable medium of claim 7 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced by synthesizing the one or more objects from the background in the image that were occluded by the removed and replaced one or more objects from the middle ground.

13. A system comprising:

a computer processor; and

a computer memory coupled to the computer processor;

wherein one or more of the computer processor and the computer memory are operable for:

receiving into the computer processor an image of a participant in a video conference;

receiving a depth map of the image;

identifying one or more objects forming a background in the image using the depth map;

identifying one or more objects forming a foreground in the image using the depth map;

identifying one or more objects forming a middle-ground in the image using the depth map, removing the one or more objects from the middle-ground;

replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects; and

using the image with the removed and replaced one or more objects in a video stream associated with the video conference;

wherein the identifying the one or more objects forming a middle-ground in the image using the depth map, the removing the one or more objects from the middle-ground, the replacing the removed one or more objects from the middle-ground with the one or more objects from the background that are located behind the removed one or more objects, and the using the image with the removed and replaced one or more objects in a video stream associated with the video conference are executed a plurality of times during the video conference.

14. The system of claim 13 , wherein the image is received into the computer processor prior to the participant joining the video conference.

15. The system of claim 13 , wherein the one or more objects forming the foreground comprises at least the participant.

16. The system of claim 13 , wherein the depth map is created using one or more of a structured light (SL) camera or a time-of-flight (TOF) camera.

17. The system of claim 13 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced using a machine learning algorithm.

18. The system of claim 13 , wherein the removed and replaced one or more objects from the middle-ground are removed and replaced by synthesizing the one or more objects from the background in the image that were occluded by the removed and replaced one or more objects from the middle ground.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 3, 2022
From: LENOVO (UNITED STATES) INC.
To: LENOVO (SINGAPORE) PTE. LTD.
Reel/Frame 061880/0110 →
CORRECTIVE ASSIGNMENT TO CORRECT THE THE ASSIGNEE'S NAME PREVIOUSLY RECORDED AT REEL: 059299 FRAME: 0119. ASSIGNOR(S) HEREBY CONFIRMS THE ASSIGNMENT. Recorded May 17, 2022
From: NICHOLSON, JOHN WELDON; LOCKER, HOWARD J.; CROMER, DARYL C
To: LENOVO (UNITED STATES) INC.
Reel/Frame 060073/0961 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 17, 2022
From: NICHOLSON, JOHN WELDON; LOCKER, HOWARD J.; CROMER, DARYL C.
To: LENOVO (UNITED STATES), INC.
Reel/Frame 059299/0119 →
Continuity (1)
Related Publication 20230298143A1 · Sep 21, 2023
References Cited (14)
US 11671561B1 · Chang · 2023 [cited by examiner]
US 20080077953A1 · Fernandez · 2008 [cited by examiner]
US 20120050323A1 · Baron, Jr. · 2012 [cited by examiner]
US 20160182830A1 · Hu · 2016 [cited by examiner]
US 20170134656A1 · Burgess · 2017 [cited by examiner]
US 20170142371A1 · Barzuza · 2017 [cited by examiner]
US 20200077035A1 · Yao · 2020 [cited by examiner]
US 20210142497A1 · Pugh · 2021 [cited by examiner]
US 20210241422A1 · Burke, III · 2021 [cited by examiner]
US 20210385412A1 · Matula · 2021 [cited by examiner]
US 20220256116A1 · Chu · 2022 [cited by examiner]
US 20240029472A1 · Shen · 2024 [cited by examiner]
Liu, Guilin, “Image Inpainting for Irregular Holes Using Partial Convolutions”, Proceedings of theEuropean Conference on Computer Vision (ECCV), Sep. 2018., (Sep. 2018), 16 pgs. [cited by applicant]
Liu, Guilin, et al., “Image Inpainting for Irregular Holes Using Partial Convolutions”, Proceedings of the European Conference on Computer Vision (ECCV), Sep. 2018., (Sep. 2018), 16 pgs. [cited by applicant]