IP Library Granted Patent US 12,342,100
Granted Patent B2
US 12,342,100 · App. 18/175,698 · Granted Jun 24, 2025

Changing conference outputs based on conversational context

Inventors: Jeffrey William Smith (Layton, UT); Chi-chian Yu (San Ramon, CA)
Assignee: Zoom Communications, Inc.
H04N7/15G06T7/70G06V40/10G10L15/22G10L25/57H04N5/2624H04N5/268H04N7/142H04N23/90H04R1/326G06N20/00G06T2207/10016G06T2207/30196G10L2015/227G10L2015/228
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,342,100
App. No.
18/175,698
Granted
Jun 24, 2025
Kind
B2
Abstract

A conference gallery view intelligence system determines regions of interest for display within views of conferencing software based on input streams received from devices within a conference room during a conference. Conference participants are detected in the conference room based on an input video stream received from a video capture device. A direction of audio from the conference participants is determined based on an input audio stream received from a multi-directional audio capture device. A conversational context within the conference room is then determined based on the direction of the audio and locations of the one or more conference participants in the conference room. A region of interest to output within conferencing software is determined based on the conversational context, and the region of interest is output for display within a view of the conferencing software.

Claims (55)

1. A method, comprising:

determining, based on a direction of audio captured within a conference room and an input video stream depicting conference participants in the conference room, a conversational context within the conference room;

selecting a first region of interest of the input video stream for use in a primary view of conferencing software, wherein the first region of interest is selected based on a determination that one or more of the conference participants depicted in the first region of interest are actively participating in a conversation corresponding to the conversational context;

selecting a second region of interest of the input video stream for use in a secondary view of the conferencing software, wherein the second region of interest is selected based on a determination that one or more of the conference participants depicted in the second region of interest are not actively participating in the conversation;

outputting, based on the conversational context, the first region of interest within a primary view of the conferencing software and the second region of interest within a secondary view of the conferencing software; and

changing, based on a second conversational context within the conference room, the output within the secondary view of the conferencing software to a third region of interest of the input video stream, wherein the third region of interest of the input video stream depicts one or more of the conference participants who are actively participating in a second conversation corresponding to the second conversational context,

wherein the primary view of the conferencing software and the secondary view of the conferencing software are simultaneously rendered within a gallery view layout user interface of the conferencing software.

2. The method of claim 1 , wherein the conversational context corresponds to one or both of a context or a length of a conversation involving one or more of the conference participants.

3. The method of claim 1 , wherein the input video stream is captured using a video capture device located within the conference room and the direction of audio is determined using an input audio stream captured using a multi-directional audio capture device located within the conference room.

4. The method of claim 1 , comprising:

changing, based on a change in the conversational context, an output within the primary view of the conferencing software to a region of interest of the input video stream other than the first region of interest.

5. The method of claim 1 , wherein the first region of interest is determined using a machine learning model trained to determine regions of interest of video streams based on conversational dynamic processing using recordings of past video conferences.

6. The method of claim 1 , wherein changing the output within the secondary view of the conferencing software to the third region of interest of the input video stream comprises:

determining the second conversational context; and

selecting the third region of interest of the input video stream for use in the secondary view of the conferencing software based on a determination that the one or more of the conference participants depicted in the third region of interest are actively participating in the second conversation.

7. The method of claim 1 , wherein the primary view of the conferencing software is a largest view of a gallery layout of the conferencing software and the secondary view of the conferencing software is smaller than the primary view.

8. The method of claim 1 , wherein the secondary view of the conferencing software switches between regions of interest depicting one or more of the conference participants that are not actively participating in the conversation while the primary view depicts the first region of interest.

9. The method of claim 1 , comprising:

detecting the conference participants within the input video stream by:

using facial detection to identify one or more humans within the input video stream; and

segmenting the one or more humans from a background identified within the input video stream.

10. The method of claim 1 , comprising:

determining the direction of audio by:

processing an input audio stream to detect voice activity within the conference room; and

determining the direction of audio based on the voice activity.

11. A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

determining, based on a direction of audio captured within a conference room and an input video stream depicting conference participants in the conference room, a conversational context within the conference room;

selecting a first region of interest of the input video stream for use in a primary view of conferencing software, wherein the first region of interest is selected based on a determination that one or more of the conference participants depicted in the first region of interest are actively participating in a conversation corresponding to the conversational context;

selecting a second region of interest of the input video stream for use in a secondary view of the conferencing software, wherein the second region of interest is selected based on a determination that one or more of the conference participants depicted in the second region of interest are not actively participating in the conversation;

outputting, based on the conversational context, the first region of interest within a primary view of the conferencing software and the second region of interest within a secondary view of the conferencing software; and

changing, based on a second conversational context within the conference room, the output within the secondary view of the conferencing software to a third region of interest of the input video stream, wherein the third region of interest of the input video stream depicts one or more of the conference participants who are actively participating in a second conversation corresponding to the second conversational context,

wherein the primary view of the conferencing software and the secondary view of the conferencing software are simultaneously rendered within a gallery view layout user interface of the conferencing software.

12. The non-transitory computer readable medium of claim 11 , the operations comprising:

changing, based on a change in the conversational context, an output with the primary view of the conferencing software to a region of interest of the input video stream other than the first region of interest.

13. The non-transitory computer readable medium of claim 11 , wherein changing the output within the secondary view of the conferencing software to the third region of interest of the input video stream comprises:

determining the second conversational context within the conference room; and

selecting the third region of interest of the input video stream.

14. The non-transitory computer readable medium of claim 11 , the operations comprising:

detecting the conference participants within the input video stream; and

determining the direction of audio.

15. An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

determine, based on a direction of audio captured within a conference room and an input video stream depicting conference participants in the conference room, a conversational context within the conference room;

select a first region of interest of the input video stream for use in a primary view of conferencing software, wherein the first region of interest is selected based on a determination that one or more of the conference participants depicted in the first region of interest are actively participating in a conversation corresponding to the conversational context;

select a second region of interest of the input video stream for use in a secondary view of the conferencing software, wherein the second region of interest is selected based on a determination that one or more of the conference participants depicted in the second region of interest are not actively participating in the conversation;

output, based on the conversational context, the first region of interest within a primary view of the conferencing software and the second region of interest within a secondary view of the conferencing software; and

change, based on a second conversational context within the conference room, the output within the secondary view of the conferencing software to a third region of interest of the input video stream, wherein the third region of interest of the input video stream depicts one or more of the conference participants who are actively participating in a second conversation corresponding to the second conversational context,

wherein the primary view of the conferencing software and the secondary view of the conferencing software are simultaneously rendered within a gallery view layout user interface of the conferencing software.

16. The apparatus of claim 15 , wherein the conversational context is determined using a machine learning model trained to determine regions of interest within the input video stream.

17. The apparatus of claim 15 , wherein the processor is configured to execute the instructions to:

determine to output the first region of interest within the primary view of the conferencing software based on an evaluation of the conversational context and one or more other conversational contexts.

18. The apparatus of claim 15 , wherein the output within the secondary view of the conferencing software is further changed based on a change to the conversational context while the first region of interest remains output within the primary view of the conferencing software.

19. The apparatus of claim 15 , wherein the conversational context corresponds to a presentation led by a conference participant included in the one or more conference participants depicted in the first region of interest.

20. The apparatus of claim 15 , wherein the conversational context corresponds to a dialogue between two or more conference participants included in the one or more conference participants depicted in the first region of interest.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 28, 2023
From: SMITH, JEFFREY WILLIAM; YU, CHI-CHIAN
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 062823/0396 →
Continuity (2)
Continuation 17243004 · Apr 28, 2021
Related Publication 20230209014A1 · Jun 29, 2023
References Cited (78)
US 7668907B1 · Janakiraman et al. · 2010 [cited by applicant]
US 8150155B2 · El-Maleh et al. · 2012 [cited by applicant]
US 8842161B2 · Feng et al. · 2014 [cited by applicant]
US 9055189B2 · Su · 2015 [cited by applicant]
US 9100540B1 · Gates et al. · 2015 [cited by applicant]
US 9237307B1 · Vendrow · 2016 [cited by applicant]
US 9294726B2 · Decker et al. · 2016 [cited by applicant]
US 9591479B1 · Leavy et al. · 2017 [cited by applicant]
US 9706171B1 · Riley et al. · 2017 [cited by applicant]
US 9769424B2 · Michot · 2017 [cited by applicant]
US 9774823B1 · Gadnir et al. · 2017 [cited by applicant]
US 9858936B2 · Cartwright et al. · 2018 [cited by applicant]
US 9942518B1 · Tangeland et al. · 2018 [cited by applicant]
US 10057707B2 · Cartwright et al. · 2018 [cited by applicant]
US 10104338B2 · Goesnar · 2018 [cited by applicant]
US 10187579B1 · Wang et al. · 2019 [cited by applicant]
US 10440325B1 · Boxwell · 2019 [cited by examiner]
US 10516852B2 · Theien et al. · 2019 [cited by applicant]
US 10522151B2 · Cartwright et al. · 2019 [cited by applicant]
US 10574899B2 · Wang et al. · 2020 [cited by applicant]
US 10778941B1 · Childress, Jr. et al. · 2020 [cited by applicant]
US 10904485B1 · Childress, Jr. et al. · 2021 [cited by applicant]
US 10939045B2 · Wang et al. · 2021 [cited by applicant]
US 10991108B2 · Schnittman et al. · 2021 [cited by applicant]
US 11064256B1 · Voss et al. · 2021 [cited by applicant]
US 11076127B1 · Schaefer et al. · 2021 [cited by applicant]
US 11082661B1 · Pollefeys · 2021 [cited by applicant]
US 11165597B1 · Matsuguma et al. · 2021 [cited by applicant]
US 11212129B1 · Sipcic et al. · 2021 [cited by applicant]
US 11350029B1 · Ostap et al. · 2022 [cited by applicant]
US 20050081160A1 · Wee · 2005 [cited by examiner]
US 20090210491A1 · Thakkar et al. · 2009 [cited by applicant]
US 20100123770A1 · Friel et al. · 2010 [cited by applicant]
US 20100238262A1 · Kurtz et al. · 2010 [cited by applicant]
US 20110285808A1 · Feng et al. · 2011 [cited by applicant]
US 20120093365A1 · Aragane · 2012 [cited by examiner]
US 20120293606A1 · Watson et al. · 2012 [cited by applicant]
US 20130083153A1 · Lindbergh · 2013 [cited by applicant]
US 20130088565A1 · Buckler · 2013 [cited by applicant]
US 20130198629A1 · Tandon et al. · 2013 [cited by applicant]
US 20160057385A1 · Burenius · 2016 [cited by applicant]
US 20160073055A1 · Marsh · 2016 [cited by examiner]
US 20160277712A1 · Michot · 2016 [cited by applicant]
US 20170094222A1 · Tangeland et al. · 2017 [cited by applicant]
US 20170099461A1 · Nimri et al. · 2017 [cited by applicant]
US 20180063206A1 · Faulkner et al. · 2018 [cited by applicant]
US 20180063479A1 · Nimri et al. · 2018 [cited by applicant]
US 20180063480A1 · Luks et al. · 2018 [cited by applicant]
US 20180098026A1 · Gadnir et al. · 2018 [cited by applicant]
US 20180225852A1 · Um et al. · 2018 [cited by applicant]
US 20180232920A1 · Faulkner et al. · 2018 [cited by applicant]
US 20190215464A1 · Kumar et al. · 2019 [cited by applicant]
US 20190222892A1 · Faulkner · 2019 [cited by examiner]
US 20190341050A1 · Diamant et al. · 2019 [cited by applicant]
US 20200099890A1 · Tanaka et al. · 2020 [cited by applicant]
US 20200126513A1 · Lau et al. · 2020 [cited by applicant]
US 20200260049A1 · Erna · 2020 [cited by applicant]
US 20200267427A1 · Rogers · 2020 [cited by examiner]
US 20200275058A1 · Graham et al. · 2020 [cited by applicant]
US 20200403817A1 · Daredia et al. · 2020 [cited by applicant]
US 20210120208A1 · Morris et al. · 2021 [cited by applicant]
US 20210235040A1 · Childress, Jr. et al. · 2021 [cited by applicant]
US 20210405865A1 · Faulkner · 2021 [cited by applicant]
US 20210409893A1 · Sommer · 2021 [cited by applicant]
US 20220060660A1 · Chou · 2022 [cited by examiner]
US 20220086390A1 · Erna et al. · 2022 [cited by applicant]
US 20220353465A1 · Smith · 2022 [cited by applicant]
US 20220400216A1 · Wang et al. · 2022 [cited by applicant]
US 20230081717A1 · Hoang et al. · 2023 [cited by applicant]
GB 2594761B · 2022 [cited by applicant]
WO 2021243633A1 · 2021 [cited by applicant]
WO 2022078656A1 · 2022 [cited by applicant]
International Search Report and Written Opinion mailed on May 8, 2023 in corresponding PCT Application No. PCT/US2023/011368. [cited by applicant]
Paul Chippendale and Oswald Lanz: “Optimised Meeting Recording and Annotation Using Real-Time Video Analysis”, In: Andrej Popescu-Selis: “Machine Learning for Multimodal Interaction”, Proc. of 5th MLMI workshop, Sep. 10… [cited by applicant]
Rasmussen Troels Ammitsbol et al: “SceneCam: Improving Multti-camera Remote Collaboration using Augmented Reality”, 2019 IEEE International Symposium on Mixed and Augmented Reality Adjunct (!SMAR-ADJUNCT), IEEE, Oct. 10… [cited by applicant]
International Search Report and Written Opinion mailed on Jul. 4, 2022 in corresponding PCT Application No. PCT/US2022/024820. [cited by applicant]
International Search Report and Written Opinion mailed on Jul. 25, 2022 in corresponding PCT Application Numbber PCT/US2022/024822. [cited by applicant]
International Search Report and Written Opinion mailed on Dec. 22, 2022 in corresponding PCT Application No. PCT/US2022/042860. [cited by applicant]