IP Library Granted Patent US 12,361,620
Granted Patent B2
US 12,361,620 · App. 18/124,860 · Granted Jul 15, 2025

Visual asset display and controlled movement in a video communication session

Inventor: Richard Dean Legatski (Castle Rock, CO)
Assignee: Zoom Communications, Inc.
G06T13/00G06T3/40G06T7/55G06T2200/24G06T2207/10016G06T2207/20084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,620
App. No.
18/124,860
Granted
Jul 15, 2025
Kind
B2
Abstract

Methods, systems, and apparatus, including computer programs encoded on computer storage media related to visual asset display and controlled movement in a video communication session. The system provides for display, via a user interface, video of a meeting participant during a video communication session. A visual asset for display is selected. The visual asset is provided for display via the user interface. The visual asset is maneuvered along a path about the user interface such that the visual asset moves about the user interface during the video communication session.

Claims (54)

1. A computer-implemented method comprising:

providing for display, via a user interface, video of a meeting participant during a video communication session;

determining an estimated depth value of a background in the video using a machine learning network that is trained on background images, wherein each background image of the background images has an associated depth value;

selecting a first visual asset for display;

providing for display the first visual asset via the user interface; and

maneuvering the first visual asset along a path about the user interface such that the first visual asset moves about the user interface during the video communication session.

2. The computer-implemented method of claim 1 , further comprising:

receiving a control input to adjust the first visual asset for display; and

changing a display characteristic of the first visual asset according to the received control input.

3. The computer-implemented method of claim 1 , further comprising:

inputting the background images of the video into the machine learning network.

4. The computer-implemented method of claim 1 , further comprising:

based on the estimated depth value of the background, maneuvering the first visual asset in a display area depicting the meeting participant such that the first visual asset reduces in size to cause the first visual asset to appear to be moving into the background.

5. The computer-implemented method of claim 1 , wherein the first visual asset maneuvers over and behind a display area of the meeting participant.

6. The computer-implemented method of claim 1 , further comprising:

receiving a selection of a second visual asset; and

providing for display the first visual asset via the user interface.

7. The computer-implemented method of claim 1 , wherein the first visual asset is any one of an image, a video, an animation or animated graphic.

8. A non-transitory computer readable medium that stores executable program instructions that when executed by one or more computing devices configure the one or more computing devices to perform operations comprising:

providing for display, via a user interface, video of a meeting participant during a video communication session;

determining an estimated depth value of a background in the video using a machine learning network that is trained on background images, wherein each background image of the background images has an associated depth value;

selecting a first visual asset for display;

providing for display the first visual asset via the user interface; and

maneuvering the first visual asset along a path about the user interface such that the first visual asset moves about the user interface during the video communication session.

9. The non-transitory computer readable medium of claim 8 , further comprising the operations of:

receiving a control input to adjust the first visual asset for display; and

changing a display characteristic of the first visual asset according to the received control input.

10. The non-transitory computer readable medium of claim 8 , further comprising the operations of:

inputting the background images of the video into the machine learning network.

11. The non-transitory computer readable medium of claim 8 , further comprising the operations of:

based on the estimated depth value of the background, maneuvering the first visual asset in a display area depicting the meeting participant such that the first visual asset reduces in size to cause the first visual asset to appear to be moving into the background.

12. The non-transitory computer readable medium of claim 8 , wherein the first visual asset maneuvers over and behind a display area of the meeting participant.

13. The non-transitory computer readable medium of claim 8 , further comprising the operations of:

receiving a selection of a second visual asset; and

providing for display the first visual asset via the user interface.

14. The non-transitory computer readable medium of claim 8 , wherein the first visual asset is any one of an image, a video, an animation or animated graphic.

15. A system, comprising:

one or more processors configured to:

provide for display, via a user interface, video of a meeting participant during a video communication session;

determine an estimated depth value of a background in the video using a machine learning network that is trained on background images, wherein each background image of the background images has an associated depth value;

select a first visual asset for display;

provide for display the first visual asset via the user interface; and

maneuver the first visual asset along a path about the user interface such that the first visual asset moves about the user interface during the video communication session.

16. The system of claim 15 , wherein the one or more processors are further configured to:

receive a control input to adjust the first visual asset for display; and

change a display characteristic of the first visual asset according to the received control input.

17. The system of claim 15 , wherein the one or more processors are further configured to:

input the background images of the video into the machine learning network.

18. The system of claim 15 , wherein the one or more processors are further configured to:

based on the estimated depth value of the background, maneuvering the first visual asset in a display area depicting the meeting participant such that the first visual asset reduces in size to cause the first visual asset to appear to be moving into the background.

19. The system of claim 15 , wherein the first visual asset maneuvers over and behind a display area of the meeting participant.

20. The system of claim 15 , wherein the one or more processors are further configured to:

receive a selection of a second visual asset; and

provide for display the first visual asset via the user interface.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 22, 2023
From: LEGATSKI, RICHARD DEAN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 063062/0560 →
Continuity (1)
Related Publication 20240320890A1 · Sep 26, 2024
References Cited (3)
US 20210042950A1 · Wantland · 2021 [cited by examiner]
US 20240127518A1 · Shao · 2024 [cited by examiner]
CN 111225226A · 2020 [cited by examiner]