IP Library Granted Patent US 12694537
Granted Patent B1
US 12694537 · App. 18/315,930 · Granted Jul 28, 2026

Augmenting foreground of a conference participant

Inventor: Chi-chian Yu (San Ramon, CA)
Assignee: Zoom Communications, Inc.
G06T7/194G06T11/20G06V40/20G06T2210/12
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694537
App. No.
18/315,930
Granted
Jul 28, 2026
Kind
B1
Abstract

A foreground portion and a background portion are obtained from an image of a conference participant. Using a background replacement image, a replacement background portion corresponding to the background portion is obtained. An indication of an image portion that is included in the image and is outside the foreground portion is obtained. An output image that includes the replacement background portion, the foreground portion, and the image portion is generated such that the foreground portion and the image portion are visible in the output image. The output image is transmitted.

Claims (72)

1 . A method, comprising:

obtaining a foreground portion and a background portion from an image of a conference participant;

obtaining, using a background replacement image, a replacement background portion corresponding to the background portion;

obtaining, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein obtaining the indication of the image portion comprises:

determining that the gesture persists in images of the conference participant for a predetermined duration of time;

generating an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and

transmitting the output image.

2 . The method of claim 1 , further comprising:

receiving, from the conference participant, a bounding box surrounding another image portion; and

generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the output image.

3 . The method of claim 1 , further comprising:

obtaining an indication of another image portion verbally from the conference participant; and

generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the output image.

4 . The method of claim 1 , further comprising:

obtaining an indication of another image portion based on a pixel location received from a pointing device; and

generating another output image that includes the replacement background portion, the foreground portion, and the another image portion such that the foreground portion and the another image portion are visible in the another output image.

5 . The method of claim 1 , further comprising:

receiving an indication of the background replacement image to blur the background portion,

wherein generating the output image comprises:

unblurring the image portion based on the indication of the image portion.

6 . The method of claim 1 , further comprising:

presenting a preceding image on a display of the conference participant, wherein the preceding image precedes the image in an image stream of the conference participant; and

obtaining another indication of another object based on a markup of the preceding image by the conference participant.

7 . The method of claim 1 , further comprising:

applying a filter to the image portion to obtain a filtered image portion, wherein generating the output image comprises:

generating the output image by combining the background portion, the foreground portion, and the filtered image portion.

8 . A device, comprising:

a memory; and

a processor, the processor configured to execute instructions stored in the memory to:

obtain a foreground portion and a background portion from an image of a conference participant;

obtain, using a background replacement image, a replacement background portion corresponding to the background portion;

obtain, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein to obtain the indication of the image portion comprises to:

determine that the gesture persists in images of the conference participant for a predetermined duration of time;

generate an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and

transmit the output image.

9 . The device of claim 8 , wherein to obtain the indication of the image portion, the processor is configured to execute the instructions to:

obtain the indication from a preceding image of the conference participant, wherein the preceding image precedes the image in an image stream of the conference participant.

10 . The device of claim 8 , wherein the processor is further configured to execute instructions to:

obtain another indication of another image portion by determining that the another indication is identified in a predetermined number of images of the conference participant.

11 . The device of claim 8 , wherein the processor is configured to execute the instructions to:

obtain, using a machine learning model for object detection, a list of objects from the image of the conference participant;

obtain, from the conference participant, a selection of at least one object of the list of objects; and

identify another image portion based on the at least one object.

12 . The device of claim 8 , wherein the processor is further configured to execute the instructions to:

receive a request from the conference participant to stop displaying the image portion; and

responsive to the request, generate subsequent output images from respective subsequent images of the conference participant, wherein each subsequent output image consists of a foreground portion obtained from a corresponding subsequent image and the replacement background portion.

13 . The device of claim 8 , wherein the processor is configured to execute the instructions to:

identify, in another image captured by a camera, another gesture of the conference participant;

determine that the conference participant is holding an object based on the another gesture; and

configure the camera to focus on the object.

14 . The device of claim 8 , wherein the processor is further configured to execute the instructions to:

perform optical character processing on the image portion to extract text included in the image portion; and

transmit the text to participants of the conference.

15 . The device of claim 8 , wherein to generate the output image, the processor is configured to execute the instructions to:

apply a special effect to the image portion to obtain a modified image portion; and

generate the output image by combining the replacement background portion, the foreground portion, and the modified image portion.

16 . The device of claim 8 , wherein to generate the output image, the processor is configured to execute the instructions to:

apply an anti-glare filter to the image portion to obtain a filtered image portion; and

generate the output image by combining the replacement background portion, the foreground portion, and the filtered image portion.

17 . A non-transitory computer readable medium that stores instructions operable to cause one or more processors to perform operations including:

obtaining a foreground portion and a background portion from an image of a conference participant;

obtaining, using a background replacement image, a replacement background portion corresponding to the background portion;

obtaining, based on a gesture of the conference participant during a conference, an indication of an image portion that is included in the image and is outside the foreground portion, wherein obtaining the indication of the image portion comprises:

determining that the gesture persists in images of the conference participant for a predetermined duration of time;

generating an output image that includes the replacement background portion, the foreground portion, and the image portion such that the foreground portion and the image portion are visible in the output image; and

transmitting the output image.

18 . The non-transitory computer readable medium of claim 17 , wherein the output image further comprises a highlight of the image portion.

19 . The method of claim 1 , further comprising:

determining that the conference participant is holding an object; and

configuring a camera that captured the image of the conference participant to focus on the object.

20 . The non-transitory computer readable medium of claim 17 , wherein the operations further comprise:

configuring a camera to focus on an object in response to determining that the conference participant is holding the object.