IP Library Granted Patent US 12,068,872
Granted Patent B2
US 12,068,872 · App. 17/243,026 · Granted Aug 20, 2024

Conference gallery view intelligence system

Inventors: Cynthia Eshiuan Lee (Austin, TX); Jeffrey William Smith (Milpitas, CA); Chi-chian Yu (San Ramon, CA)
Assignee: Zoom Video Communications, Inc.
H04L12/1818G06N20/00G06V30/147H04L12/1822H04N19/167
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,068,872
App. No.
17/243,026
Granted
Aug 20, 2024
Kind
B2
Abstract

A conference gallery view intelligence system determines at least two regions of interest within a conference room based on an input video stream received from a video capture device located within the conference room. An output video stream for rendering within conferencing software is produced for each of the at least two regions of interest. The output video stream for each of the at least two regions of interest is then transmitted to one or more client devices connected to the conferencing software.

Claims (48)

1. A method, comprising:

determining, during a video conference, a first region of interest for a first conference participant within a conference room based on an input video stream received from a video capture device located within the conference room;

determining, during the video conference, a second region of interest for a second conference participant within the conference room based on the input video stream, wherein each of the regions of interest depicts a single conference participant;

producing a first output video stream for the first region of interest from the input video stream;

producing a second output video stream for the second region of interest from the input video stream; and

transmitting the first output video streams and the second output video stream to one or more client devices for display within a conferencing software user interface throughout the video conference, wherein the first output video stream is rendered within a first view of the conferencing software user interface corresponding to the first conference participant simultaneous with the second output video stream being rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

2. The method of claim 1 , wherein the video capture device is a first video capture device, the method further comprising:

receiving a second input video stream from a second video capture device located within the conference room,

wherein the regions of interest include one or more regions within a field of view of the second video capture device.

3. The method of claim 2 , wherein the field of view of the first video capture device and the field of view of the second video capture device are partially overlapping within the conference room.

4. The method of claim 1 , wherein the regions of interest are determined at a first time during the video conference, the method further comprising:

determining a third region of interest within a field of view of the video capture device based on changes within the conference room; and

producing a third output video stream for the third region of interest to change content rendered within the first view of the conferencing software user interface.

5. The method of claim 4 , wherein the changes correspond to conversational dynamics determined using a machine learning model.

6. The method of claim 4 , wherein the conferencing software user interface includes a fixed number of views during the video conference.

7. The method of claim 1 , wherein the regions of interest are determined based on video, audio, and context.

8. The method of claim 1 , wherein the first view is a primary view and the second view is a secondary view, and wherein the first output video stream is selected for the primary view based on a detected conversational context of the video conference.

9. The method of claim 1 , wherein a total number of regions of interest within the conference room corresponds to a total number of faces located in the conference room.

10. The method of claim 1 , wherein a default region of interest covering most of a field of view of the video capture device is used in place of the first region of interest and the second region of interest when a detected conversational context of the video conference is unclear.

11. The method of claim 1 , wherein a default region of interest covering most of a field of view of the video capture device is used in place of the first region of interest and the second region of interest when nobody is speaking in the video conference.

12. An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

determine, during a video conference, a first region of interest for a first conference participants within a conference room based on an input video stream received from a video capture device located within the conference room;

determine, during the video conference, a second region of interest for a second conference participant within the conference room based on the input video stream received from the video capture device located within the conference room, wherein each of the regions of interest depicts a single conference participant;

produce a first output video stream for the first region of interest from the input video stream;

produce a second output video stream for the second region of interest from the input video stream; and

transmit the first output video stream and the second output video stream to one or more client devices for display within a conferencing software user interface throughout the video conference, wherein the first output video stream is rendered within a first view of the conferencing software user interface corresponding to the first conference participant simultaneous with the second output video stream being rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

13. The apparatus of claim 12 , wherein the regions of interest are determined at a first time during the video conference, and wherein the processor is further configured to execute the instructions to:

determine a change to the first region of interest based on changes within the conference room during the video conference; and

modify the first output video stream according to the change to the first region of interest to change content rendered within the first view of the conferencing software user interface.

14. The apparatus of claim 13 , wherein the conferencing software user interface includes a fixed number of views during the video conference.

15. The apparatus of claim 12 , wherein a field of view of the video capture device is adjustable to determine the regions of interest within the conference room.

16. A non-transitory computer readable storage device including program instructions that, when executed by a processor, cause the processor to perform operations, the operations comprising:

determining, during a video conference, a first region of interest for a first conference participants within a conference room based on an input video stream received from a video capture device located within the conference room;

determining, during the video conference, a second region of interest for a second conference participant within the conference room based on the input video stream, wherein each of the regions of interest depicts a single conference participant;

producing a first output video stream for the first region of interest from the input video stream;

producing a second output video stream for the second region of interest from the input video stream; and

transmitting the first output video stream and the second output video stream to one or more client devices for display within a conferencing software user interface throughout the video conference, wherein the first output video stream is rendered within a first view of the conferencing software user interface corresponding to the first conference participant simultaneous with the second output video stream being rendered within a second view of the conferencing software user interface corresponding to the second conference participant.

17. The non-transitory computer readable storage device of claim 16 , wherein the operations further comprise:

determining a third region of interest within the conference room based on a second input video stream received from a second video capture device located within the conference room;

producing a third output video stream for the third region of interest from the second input video stream; and

transmitting the third output video stream for display within the conferencing software user interface.

18. The non-transitory computer readable storage device of claim 16 , wherein the regions of interest are determined at a first time during the video conference, and wherein the operations further comprise:

determining a change to the first region of interest based on conversational dynamics determined using a machine learning model; and

modifying the first output video stream corresponding to the first region of interest according to the change.

19. The non-transitory computer readable storage device of claim 16 , wherein other regions of interest are determined within a second input video stream received from a second video capture device located within the conference room.

20. The non-transitory computer readable storage device of claim 19 , wherein fields of view of the video capture device and the second video capture device are at least partially overlapping.

Assignments (3)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 8, 2022
From: LEE, CYNTHIA ESHIUAN; SMITH, JEFFREY WILLIAM; YU, CHI-CHIAN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 061695/0405 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2021
From: SMITH, JEFFREY WILLIAM
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 056073/0542 →
Continuity (1)
Related Publication 20220353096A1 · Nov 3, 2022