IP Library Patent Application 18766060
Patent Application
App. No. 18/766,060

Conference Gallery View

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
18/766,060
Abstract

A conference gallery view intelligence system determines at least two regions of interest within a conference room based on an input video stream received from a video capture device located within the conference room. An output video stream for rendering within conferencing software is produced for each of the at least two regions of interest. The output video stream for each of the at least two regions of interest is then transmitted to one or more client devices connected to the conferencing software.

Claims (44)

1 . A method, comprising:

determining a first region of interest for a first conference participant in a first input video stream;

determining a second region of interest for a second conference participant in the first input video stream; and

transmitting a first output video stream for the first region of interest and a second output video stream for the second region of interest for display in a video conference, wherein the first output video stream is rendered within a first view of a conferencing software user interface corresponding to the first conference participant and the second output video stream is rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

2 . The method of claim 1 , wherein the first input video stream is captured by a first video capture device, the method further comprising:

receiving a second input video stream from a second video capture device located in a same physical space as the first video capture device;

determining a third region of interest for a third conference participant in the second input video stream; and

transmitting a third output video stream for the third region of interest for display in the video conference, wherein the third output video stream is rendered within a third view of the conferencing software user interface corresponding to the third conference participant.

3 . The method of claim 2 , wherein a field of view of the first video capture device and a field of view of the second video capture device are partially overlapping within the physical space.

4 . The method of claim 1 , wherein the regions of interest are determined based on video, audio, and context.

5 . The method of claim 1 , wherein the first view is a primary view and the second view is a secondary view, and wherein the first output video stream is selected for the primary view based on a detected conversational context of the video conference.

6 . The method of claim 1 , wherein a total number of regions of interest within a physical space that includes the first conference participant and the second conference participant corresponds to a total number of faces located in the physical space.

7 . The method of claim 1 , wherein the regions of interest are determined at a first time during the video conference, the method further comprising:

determining a third region of interest based on changes in the first input video stream; and

producing a third output video stream for the third region of interest to change content rendered within the first view of the conferencing software user interface.

8 . The method of claim 7 , wherein the changes in the first input video stream correspond to conversational dynamics determined using a machine learning model.

9 . The method of claim 1 , wherein the conferencing software user interface includes a fixed number of views during the video conference.

10 . The method of claim 1 , the method further comprising:

transmitting a third output video stream that depicts the first conference participant and the second conference participant for display in the video conference, wherein the third output video stream is rendered within a third view of the conferencing software user interface corresponding to a physical space of the first conference participant and the second conference participant.

11 . The method of claim 10 , wherein the third output video stream is used in place of the first output video stream and the second output video stream in the conferencing software user interface when nobody is speaking in the video conference.

12 . An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

determine a first region of interest for a first conference participant in a first input video stream;

determine a second region of interest for a second conference participant in the first input video stream; and

transmit a first output video stream for the first region of interest and a second output video stream for the second region of interest for display in a video conference, wherein the first output video stream is rendered within a first view of a conferencing software user interface corresponding to the first conference participant and the second output video stream is rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

13 . The apparatus of claim 12 , wherein the regions of interest are determined at a first time during the video conference, and wherein the processor is further configured to execute the instructions to:

determine a change to the first region of interest based on changes within a physical space that includes the first conference participant and the second conference participant during the video conference; and

modify the first output video stream according to the change to the first region of interest to change content rendered within the first view of the conferencing software user interface.

14 . The apparatus of claim 12 , wherein the conferencing software user interface includes a fixed number of views during the video conference.

15 . The apparatus of claim 12 , wherein the first input video stream is captured by a first video capture device, and wherein a field of view of the first video capture device is adjustable to determine the regions of interest.

16 . A non-transitory computer readable storage device including program instructions that, when executed by a processor, cause the processor to perform operations, the operations comprising:

determining a first region of interest for a first conference participant in a first input video stream;

determining a second region of interest for a second conference participant in the first input video stream; and

transmitting a first output video stream for the first region of interest and a second output video stream for the second region of interest for display in a video conference, wherein the first output video stream is rendered within a first view of a conferencing software user interface corresponding to the first conference participant and the second output video stream is rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

17 . The non-transitory computer readable storage device of claim 16 , wherein other regions of interest are determined within a second input video stream received from a second video capture device located within a same physical space as a first video capture device that captured the first input video stream.

18 . The non-transitory computer readable storage device of claim 17 , wherein fields of view of the first video capture device and the second video capture device are at least partially overlapping.

19 . The non-transitory computer readable storage device of claim 16 , wherein the operations further comprise:

receiving a second input video stream from a second video capture device located in a same physical space as a first video capture device that captured the first input video stream;

determining a third region of interest for a third conference participant in the second input video stream; and

transmitting a third output video stream for the third region of interest for display in the video conference, wherein the third output video stream is rendered within a third view of the conferencing software user interface corresponding to the third conference participant.

20 . The non-transitory computer readable storage device of claim 16 , wherein the regions of interest are determined at a first time during the video conference, and wherein the operations further comprise:

determining a change to the first region of interest based on conversational dynamics determined using a machine learning model; and

modifying the first output video stream corresponding to the first region of interest according to the change.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 8, 2024
From: LEE, CYNTHIA ESHIUAN; SMITH, JEFFREY WILLIAM; YU, CHI-CHIAN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 067928/0694 →