IP Library Granted Patent US 12,483,436
Granted Patent B2
US 12,483,436 · App. 18/441,698 · Granted Nov 25, 2025

Recommendation based on video-based audience sentiment

Inventor: Vi Dinh Chau (Seattle, WA)
Assignee: Zoom Communications, Inc.
H04L12/1827G06V10/507G06V40/20H04L12/1831
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,483,436
App. No.
18/441,698
Granted
Nov 25, 2025
Kind
B2
Abstract

Video data from audience participants reacting to a speaker participation during a conference is obtained. The video data is processed to detect and recognize reactions based on a speaker presentation. Sentiment types are determined for the recognized reactions in view of a context of the speaker presentation. An engagement level is determined based on aggregated sentiment types for the audience participants. A real-time recommendation output is presented based on the engagement level. The real-time recommendation output provides suggestive actions for the speaker participant based on a positive or negative engagement level.

Claims (65)

1 . A method, comprising:

aggregating, by a server during multiple previous video conference sessions, previous conference session information including previous speaker participant behaviors, previous engagement levels, previous recommendations output by the server and effectiveness indications for the previous recommendations, wherein the effectiveness indications identify at least some of the previous recommendations as previously effective recommendations;

determining, by the server during a video conference, sentiment types based on reactions of one or more audience participants to speaker participant behaviors of a speaker participant;

determining, by the server during the video conference, a most frequently counted one of the sentiment types based on a count of the sentiment types for the one or more audience participants;

determining, by the server during the video conference, an engagement level based on the most frequently counted one of the sentiment types;

determining, by the server during the video conference using a machine learning system based in part on the previously effective recommendations from the previous conference session information, a real-time recommendation based on the engagement level and the speaker participant behaviors;

providing, by the server during the video conference, an outputting of the real-time recommendation at a device associated with the speaker participant to allow the speaker participant to change the speaker participant behaviors during the video conference;

maintaining, by the server during the video conference, engagement trends based on an engagement level before an output of the real-time recommendation, an engagement level after the output of the real-time recommendation, and associated timestamps; and

determining, by the server, an impact of the real-time recommendation on the engagement trends during the video conference.

2 . The method of claim 1 , wherein the sentiment types of the one or more audience participants are determined by:

aggregating, by the server, the sentiment types for the one or more audience participants which are perceptible and imperceptible to the speaker participant.

3 . The method of claim 1 , wherein determining the sentiment types of the one or more audience participants comprises:

determining, by a machine learning system, contexts of the reactions, wherein a respective context is used to determine a meaning for a respective one of the reactions.

4 . The method of claim 1 , further comprising:

determining performance characterizations of the speaker participant corresponding to the reactions during the video conference to determine which speaker participant behavior is effective when the engagement level is positive.

5 . The method of claim 1 , wherein a context indicates at least one of a setting or environment of the video conference and wherein a different meaning is assigned to the reactions based on the context.

6 . The method of claim 1 , further comprising:

obtaining, by the server, video data of the one or more audience participants during the video conference; and

determining, by the server, the reactions of the one or more audience participants based on the video data of the one or more audience participants.

7 . The method of claim 1 , further comprising:

obtaining, by the server, audio data of the one or more audience participants during the video conference; and

determining, by the server, the reactions of the one or more audience participants based on the audio data of the one or more audience participants.

8 . The method of claim 1 , further comprising:

obtaining, by the server, audio data of the speaker participant during the video conference;

generating, by the server using an automated speech recognition engine, a real-time transcription of the audio data of the speaker participant; and

determining, by the server, the speaker participant behaviors based on the real-time transcription of the audio data of the speaker participant.

9 . An apparatus, comprising:

a memory; and

a processor configured to execute instructions stored in the memory to:

aggregate, during multiple previous video conference sessions, previous conference session information including previous speaker participant behaviors, previous engagement levels, previous recommendations and effectiveness indications for the previous recommendations, wherein the effectiveness indications identify at least some of the previous recommendations as previously effective recommendations;

determine, during a video conference, sentiment types based on reactions of one or more audience participants to speaker participant behaviors of a speaker participant;

determine, during the video conference, a most frequently counted one of the sentiment types based on a count of the sentiment types for the one or more audience participants;

determine, during the video conference, an engagement level based on the most frequently counted one of the sentiment types;

determine, during the video conference using a machine learning system based in part on the previously effective recommendations from the previous conference session information, a real-time recommendation based on the engagement level and the speaker participant behaviors;

provide, during the video conference, an outputting of a real-time recommendation at a device associated with the speaker participant to allow the speaker participant to change the speaker participant behaviors during the video conference;

maintain, during the video conference, engagement trends based on an engagement level before an output of the real-time recommendation, an engagement level after the output of the real-time recommendation, and associated timestamps; and

determine an impact of the real-time recommendation on the engagement trends during the video conference.

10 . The apparatus of claim 9 , wherein the processor is configured to execute the instructions to:

determine, using the machine learning system, the reactions based on facial recognition and movement detection on video data of the one or more audience participants, and based on locations associated with the one or more audience participants.

11 . The apparatus of claim 9 , wherein the processor is further configured to execute the instructions stored in the memory to:

obtain video data and audio data of the one or more audience participants during the video conference; and

determine the reactions of the one or more audience participants based on the video data and the audio data of the one or more audience participants.

12 . The apparatus of claim 9 , wherein the processor is further configured to execute the instructions stored in the memory to:

obtain audio data of the speaker participant during the video conference;

generate, using an automated speech recognition engine, a real-time transcription of the audio data of the speaker participant; and

determine the speaker participant behaviors based on the real-time transcription of the audio data of the speaker participant.

13 . A non-transitory computer readable medium storing instructions operable to cause one or more processors to perform operations comprising:

aggregating, by a server during multiple previous video conference sessions, previous conference session information including previous speaker participant behaviors, previous engagement levels, previous recommendations output by the server and effectiveness indications for the previous recommendations, wherein the effectiveness indications identify at least some of the previous recommendations as previously effective recommendations;

determining, by the server during a video conference, sentiment types based on reactions of one or more audience participants to speaker participant behaviors of a speaker participant;

determining, by the server during the video conference, a most frequently counted one of the sentiment types based on a count of the sentiment types for the one or more audience participants;

determining, by the server during the video conference, an engagement level based on the most frequently counted one of the sentiment types;

determining, by the server during the video conference using a machine learning system based in part on the previously effective recommendations from the previous conference session information, a real-time recommendation based on the engagement level and the speaker participant behaviors;

providing, by the server during the video conference, an outputting of the real-time recommendation at a device associated with the speaker participant to allow the speaker participant to change the speaker participant behaviors during the video conference;

maintaining, by the server during the video conference, engagement trends based on an engagement level before an output of the real-time recommendation, an engagement level after the output of the real-time recommendation, and associated timestamps; and

determining, by the server, an impact of the real-time recommendation on the engagement trends during the video conference.

14 . The non-transitory computer readable medium of claim 13 , the operations further comprising:

generating a histogram with bins for different ones of the sentiment types; and

determining the most frequently counted one of the sentiment types based on most populated bin in the histogram.

15 . The non-transitory computer readable medium of claim 13 , further comprising:

obtaining, by the server, video data and audio data of the one or more audience participants during the video conference; and

determining, by the server, the reactions of the one or more audience participants based on the video data and the audio data of the one or more audience participants.

16 . The non-transitory computer readable medium of claim 13 , further comprising:

obtaining, by the server, audio data of the speaker participant during the video conference;

generating, by the server using an automated speech recognition engine, a real-time transcription of the audio data of the speaker participant; and

determining, by the server, the speaker participant behaviors based on the real-time transcription of the audio data of the speaker participant.

Assignments (2)
CHANGE OF NAME Recorded Jan 7, 2025
From: ZOOM VIDEO COMMUNICATIONS, INC.
To: ZOOM COMMUNICATIONS, INC.
Reel/Frame 069839/0593 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 15, 2024
From: CHAU, VI DINH
To: ZOOM VIDEO COMMUNICATIONS, INC.
Reel/Frame 066472/0975 →