IP Library Granted Patent US 12,231,803
Granted Patent B2
US 12,231,803 · App. 17/877,790 · Granted Feb 18, 2025

Video conference background cleanup

Inventors: Shihwei Chang (Sammamish, WA); Robert Aaron Klegon (Chicago, IL); Cynthia Eshiuan Lee (Austin, TX); Nicholas Mueller (Fitchburg, WI); Shane Paul Springer (Manchester, MI)
Assignee: Zoom Communications, Inc.
H04N5/272G06T11/00G06V20/41G06V40/161G06T2200/24G06V20/48H04L65/403
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,231,803
App. No.
17/877,790
Granted
Feb 18, 2025
Kind
B2
Abstract

A computer stores a reference image representing a physical background within a field of view of a camera of a client device. The computer receives, via the camera and during a video conference to which the client device is connected, camera-generated visual data for output to at least one remote device connected to the video conference. The computer identifies, based on facial recognition applied to the camera-generated visual data, foreground imagery representing at least one person and background imagery representing content of the camera-generated visual data other than the foreground imagery. The computer identifies a difference between the background imagery and the reference image. The computer generates a composite image by replacing, within the background imagery of the camera-generated visual data, an item represented within the background imagery and within the identified difference with a co-located part of the reference image.

Claims (46)

1. A method, comprising:

receiving, during a video conference to which a client device is connected, camera-generated visual data for output to at least one remote device participating in the video conference;

identifying, using software-based image processing applied to the camera-generated visual data, foreground imagery representing a participant and background imagery representing content of the camera-generated visual data other than the foreground imagery;

identifying, within the background imagery, an extraneous item for removal, wherein the extraneous item comprises a portion of the background imagery;

determining whether to generate a composite image at the client device or at a server based on capabilities of the client device, wherein the capabilities comprise at least one of processing capabilities, memory capabilities, network access capabilities, software capabilities, and hardware capabilities, wherein the hardware capabilities are related to item recognition technology;

generating the composite image by removing the extraneous item from the camera-generated visual data and predicting, using replacement imagery prediction software, replacement imagery to replace the removed extraneous item; and

transmitting the composite image to the at least one remote device during the video conference.

2. The method of claim 1 , wherein the camera-generated visual data comprises at least one frame in a video generated by a camera.

3. The method of claim 1 , wherein, when the participant moves in front of the extraneous item, the composite image depicts the participant obscuring the replacement imagery.

4. The method of claim 1 , wherein generating the composite image comprises:

adding a virtual sticker to overlay a part of the background imagery; and

detecting, via a camera, that the participant has moved in front of a first part of the virtual sticker, wherein the composite image comprises the participant obscuring the part of the virtual sticker, wherein the composite image comprises a second part of the virtual sticker that is not obscured by the participant.

5. The method of claim 1 , wherein the composite image maintains at least a part of the background imagery from the camera-generated visual data, wherein the part of the background imagery is distinct from the replacement imagery.

6. The method of claim 1 , wherein the extraneous item is identified by receiving, via a graphical user interface, an input associated with drawing a border around the extraneous item.

7. The method of claim 1 , wherein the extraneous item is identified using extraneous item identification software, wherein the extraneous item identification software is trained based on previously identified extraneous items.

8. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

receiving, during a video conference to which a client device is connected, camera-generated visual data for output to at least one remote device participating in the video conference;

identifying, using software-based image processing applied to the camera-generated visual data, foreground imagery representing a participant and background imagery representing content of the camera-generated visual data other than the foreground imagery;

identifying, within the background imagery, an extraneous item for removal, wherein the extraneous item comprises a portion of the background imagery;

determining whether to generate a composite image at the client device or at a server based on capabilities of the client device, wherein the capabilities comprise at least one of processing capabilities, memory capabilities, network access capabilities, software capabilities, and hardware capabilities, wherein the hardware capabilities are related to item recognition technology;

generating the composite image by removing the extraneous item from the camera-generated visual data and predicting, using replacement imagery prediction software, replacement imagery to replace the removed extraneous item; and

transmitting the composite image to the at least one remote device during the video conference.

9. The non-transitory computer readable medium of claim 8 , wherein the camera-generated visual data comprises a frame in a video.

10. The non-transitory computer readable medium of claim 8 , wherein, when the participant moves in front of the extraneous item, the composite image depicts the participant in front of the replacement imagery.

11. The non-transitory computer readable medium of claim 8 , wherein generating the composite image comprises:

adding one or more virtual stickers to overlay a part of the background imagery; and

detecting, via a camera, that the participant has moved in front of a first part of the one or more virtual stickers, wherein the composite image comprises the participant obscuring the part of the one or more virtual stickers, wherein the composite image comprises a second part of the one or more virtual stickers that is not obscured by the participant.

12. The non-transitory computer readable medium of claim 8 , wherein the composite image maintains at least a part of the background imagery, wherein the part of the background imagery is distinct from the replacement imagery.

13. The non-transitory computer readable medium of claim 8 , wherein the extraneous item is identified by receiving, via a graphical user interface, a representation of a drawn border around the extraneous item.

14. The non-transitory computer readable medium of claim 8 , wherein the extraneous item is identified automatically using extraneous item identification software.

15. An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

receive, during a video conference to which a client device is connected, camera-generated visual data for output to at least one remote device participating in the video conference;

identify, using software-based image processing applied to the camera-generated visual data, foreground imagery representing a participant and background imagery representing content of the camera-generated visual data other than the foreground imagery;

identify, within the background imagery, an extraneous item for removal, wherein the extraneous item comprises a portion of the background imagery;

determine whether to generate a composite image at the client device or at a server based on capabilities of the client device, wherein the capabilities comprise at least one of processing capabilities, memory capabilities, network access capabilities, software capabilities, and hardware capabilities, wherein the hardware capabilities are related to item recognition technology;

generate the composite image by removing the extraneous item from the camera-generated visual data and predicting, using replacement imagery prediction software, replacement imagery to replace the removed extraneous item; and

transmit the composite image to the at least one remote device during the video conference.

16. The apparatus of claim 15 , wherein the camera-generated visual data comprises a video frame.

17. The apparatus of claim 15 , wherein, when the participant moves in front of the extraneous item, the composite image depicts the participant in front of the replacement imagery and forgoes depicting the extraneous item.

18. The apparatus of claim 15 , wherein generating the composite image comprises:

adding a sticker to overlay a part of the background imagery; and

detecting, via a camera, that the participant has moved in front of a first part of the sticker, wherein the composite image comprises the participant obscuring the first part of the sticker.

19. The method of claim 1 , wherein the method is performed by the server, wherein it is determined to generate the composite image at the client device, and wherein generating the composite image comprises sending an indication to the client device to generate the composite image.

20. The method of claim 1 , wherein the method is performed by the client device, wherein it is determined to generate the composite image at the server, and wherein generating the composite image comprises sending an indication to the server to generate the composite image.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069830/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 2, 2022
From: CHANG, SHIHWEI; KLEGON, ROBERT AARON; LEE, CYNTHIA ESHIUAN; MUELLER, NICHOLAS P.; SPRINGER, SHANE PAUL
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 060693/0425 →
Continuity (1)
Related Publication 20240040073A1 · Feb 1, 2024
References Cited (19)
US 7227567B1 · Beck · 2007 [cited by examiner]
US 7564476B1 · Coughlan · 2009 [cited by examiner]
US 11145334B2 · Chu · 2021 [cited by examiner]
US 11190735B1 · Trim · 2021 [cited by examiner]
US 11671561B1 · Chang · 2023 [cited by examiner]
US 20080030621A1 · Ciudad · 2008 [cited by examiner]
US 20080077953A1 · Fernandez et al. · 2008 [cited by applicant]
US 20120050323A1 · Baron, Jr. et al. · 2012 [cited by applicant]
US 20170142371A1 · Barzuza et al. · 2017 [cited by applicant]
US 20180114009A1 · Maresh et al. · 2018 [cited by applicant]
US 20210385412A1 · Matula · 2021 [cited by examiner]
US 20220141396A1 · Ruan · 2022 [cited by examiner]
US 20220350925A1 · Alexander · 2022 [cited by examiner]
US 20230244799A1 · Sharma · 2023 [cited by examiner]
US 20230298143A1 · Nicholson · 2023 [cited by examiner]
US 20230306610A1 · Nicholson · 2023 [cited by examiner]
Real or Virtual: A Video Conferencing Background Manipulation-Detection System, Ehsan Nowroozi, Yassine Mekdad, Mauro Conti, Simone Milani, A. Selcuk Uluagac, Berrin Yanikoglu, Apr. 27, 2022, 34 pages. [cited by applicant]
Background Buster: Peeking through Virtual Backgrounds in Online Video Calls, Mohd Sabra, Anindya Maiti, Murtuza Jadliwala, Jun. 27-30, 2022, 42 pages. [cited by applicant]
Pro IT Insights for Business, This Microsoft Teams update could bring the worst thing about high school to your work calls, Mike Moore, Jul. 11, 2022, 9 pages. [cited by applicant]