IP Library Granted Patent US 11,122,240
Granted Patent B2
US 11,122,240 · App. 17/037,324 · Granted Sep 14, 2021

Enhanced video conference management

Inventors: Michael H Peters (Washington, DC); Alexander M. Stufflebeam (Indianapolis, IN)
Assignee: Michael H Peters
H04N7/152G06K9/00718
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,122,240
App. No.
17/037,324
Granted
Sep 14, 2021
Kind
B2
Abstract

Methods, systems, and apparatus, including computer-readable media storing executable instructions, for enhanced video conference management. In some implementations, a participant score for each participant in a set of multiple participants in the communication session. Each of the participant scores is based on at least one of facial image analysis or facial video analysis performed using image data or video data captured for the corresponding participant during the communication session. Each of the participant scores is indicative of an emotional or cognitive state of the corresponding participant. The participant scores are used to generate an aggregate representation of the emotional or cognitive states of the set of multiple participants. During the communication session, output data is comprising the aggregate representation of the emotional or cognitive states of the set of multiple participants is provided for display.

Claims (48)

1. A method performed by one or more computing devices, the method comprising:

during a communication session, obtaining, by the one or more computing devices, a participant score for each participant in a set of multiple participants in the communication session, wherein each of the participant scores is based on at least one of facial image analysis or facial video analysis performed using image data or video data captured for the corresponding participant during the communication session, wherein each of the participant scores is indicative of an emotional or cognitive state of the corresponding participant;

determining, by the one or more computing devices, amounts of speaking time in which different participants have spoken during the communication session;

selecting, by the one or more computing devices, an action to improve a level of engagement or emotion of the participants based on the participant scores and the amounts of speaking time, the selected action being selected from among set of multiple candidate actions;

using, by the one or more computing devices, the participant scores to generate an aggregate representation of the emotional or cognitive states of the set of multiple participants; and

providing, by the one or more computing devices, output data to devices of one or more of the participants during the communication session, the output data comprising the aggregate representation of the emotional or cognitive states of the set of multiple participants and a recommendation of the selected action to improve the level of engagement or emotion of the participants.

2. The method of claim 1 , wherein the one or more computing devices comprise a server system;

wherein at least some of the multiple participants participate in the communication session using respective endpoint devices, the endpoint devices each providing image data or video data over a communication network to the server system; and

wherein the server system obtains the participant scores by performing at least one of image facial image analysis or facial video analysis on the image data or video data received from the endpoint devices over the communication network.

3. The method of claim 1 , wherein at least some of the multiple participants participate in the communication session using respective endpoint devices; and

wherein obtaining the participant scores comprises receiving, over a communication network, participant scores that the respective endpoint devices determined by each end point device performing at least one of facial image analysis or facial video analysis using image data or video data captured by the endpoint device.

4. The method of claim 1 , wherein the one or more computing devices comprise an endpoint device for a person participating in the communication session; and

wherein obtaining the participant score, using the participant scores to generate the aggregate representation, and providing the output data for display are performed by the endpoint device, the aggregate representation being displayed at the endpoint device.

5. The method of claim 1 , wherein the aggregate representation characterizes or indicates the emotional or cognitive states or levels of engagement among the set of multiple participants using at least one of a score, a graph, a chart, a table, an animation, a symbol, or text.

6. The method of claim 1 , comprising, during the communication session, repeatedly (i) obtaining updated participant scores for the participants as additional image data or video data captured for the respective participants during the communication session, (ii) generating an updated aggregate representation of the emotional or cognitive states or levels of engagement of the set of multiple participants based on the updated participant scores, and (iii) providing updated output data indicative of the updated aggregate representation.

7. The method of claim 1 , wherein the participant scores are varied during the communication session based on captured image data or video data such that the aggregate representation provides a substantially real-time indicator of current emotion or engagement among the set of multiple participants.

8. The method of claim 1 , wherein the output data is provided for display by an endpoint device of a speaker or presenter for the communication session.

9. The method of claim 1 , wherein the output data is provided for display by an endpoint device of a teacher, and wherein the set of multiple participants is a set of students.

10. The method of claim 1 , wherein the set of participants comprises remote participants that are located remotely from each other during the communication session, each of the remote participants having image data or video data captured by a corresponding endpoint device.

11. The method of claim 1 , wherein the set of participants comprises local participants that are located in a same room as each other during the communication session.

12. The method of claim 1 , comprising:

tracking changes in emotion or engagement of the set of multiple participants over time during the communication session; and

providing during the communication session an indication of a change in the emotion or engagement of the set of multiple participants over time.

13. The method of claim 1 , comprising grouping the participants in the set of multiple participants into different groups based on the participant scores; and

wherein the aggregate representation comprises an indication of characteristics of the different groups.

14. The method of claim 1 , comprising determining, for each participant of a plurality of the multiple participants, a different action for the participant to improve a level of collaboration among the set of multiple participants during the communication session; and

providing different recommendations to devices of each of the plurality of participants, the different recommendations respectively indicating the different actions determined for the participants in the plurality of the multiple participants.

15. The method of claim 1 , comprising accessing information describing characteristics of the communication session or the set of multiple participants; and

using the characteristics of the communication session or the set of multiple participants to select the action.

16. The method of claim 15 , wherein the characteristics include a number of participants in the communication session, a duration of the communication session, or a communication session type for the communication session.

17. The method of claim 1 , wherein the action is selected based on a pattern or distribution of the participant scores determined for the communication session.

18. The method of claim 1 , comprising determining, based on the amounts of speaking time, that a particular participant has an excessive proportion of the speaking time;

wherein the action includes limiting speaking time of the particular participant; and

wherein the recommendation is provided to a device associated with the particular participant or to a participant having a role as host or moderator for the communication session.

19. A system comprising:

one or more computers; and

one or more computer-readable media storing instructions that are operable, when executed by one or more computing devices, to cause the one or more computing devices to perform operations comprising:

during a communication session, obtaining, by the one or more computing devices, a participant score for each participant in a set of multiple participants in the communication session, wherein each of the participant scores is based on at least one of facial image analysis or facial video analysis performed using image data or video data captured for the corresponding participant during the communication session, wherein each of the participant scores is indicative of an emotional or cognitive state of the corresponding participant;

determining, by the one or more computing devices, amounts of speaking time in which different participants have spoken during the communication session;

selecting, by the one or more computing devices, an action to improve a level of engagement or emotion of the participants based on the participant scores and the amounts of speaking time, the selected action being selected from among set of multiple candidate actions;

using, by the one or more computing devices, the participant scores to generate an aggregate representation of the emotional or cognitive states of the set of multiple participants; and

providing, by the one or more computing devices, output data to devices of one or more of the participants during the communication session, the output data comprising the aggregate representation of the emotional or cognitive states of the set of multiple participants and a recommendation of the selected action to improve the level of engagement or emotion of the participants.

20. One or more non-transitory computer-readable media storing instructions that are operable, when executed by one or more computing devices, to cause the one or more computing devices to perform operations comprising:

during a communication session, obtaining, by the one or more computing devices, a participant score for each participant in a set of multiple participants in the communication session, wherein each of the participant scores is based on at least one of facial image analysis or facial video analysis performed using image data or video data captured for the corresponding participant during the communication session, wherein each of the participant scores is indicative of an emotional or cognitive state of the corresponding participant;

determining, by the one or more computing devices, amounts of speaking time in which different participants have spoken during the communication session;

selecting, by the one or more computing devices, an action to improve a level of engagement or emotion of the participants based on the participant scores and the amounts of speaking time, the selected action being selected from among set of multiple candidate actions;

using, by the one or more computing devices, the participant scores to generate an aggregate representation of the emotional or cognitive states of the set of multiple participants; and

providing, by the one or more computing devices, output data to devices of one or more of the participants during the communication session, the output data comprising the aggregate representation of the emotional or cognitive states of the set of multiple participants and a recommendation of the selected action to improve the level of engagement or emotion of the participants.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2023
From: PETERS, MICHAEL H.
To: REELAY MEETINGS, INC.
Reel/Frame 063424/0284 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 24, 2021
From: STUFFLEBEAM, ALEXANDER M.
To: MICHAEL H PETERS
Reel/Frame 056651/0585 →
Continuity (7)
Continuation In Part 16993010 · Aug 13, 2020
Continuation 16516731 · Jul 19, 2019
Continuation 16128137 · Sep 11, 2018
Provisional Application 63072936 · Aug 31, 2020
Provisional Application 63075809 · Sep 8, 2020
Provisional Application 62556672 · Sep 11, 2017
Related Publication 20210176429A1 · Jun 10, 2021
Cited By (6)
US 12,333,258 US 12,387,291 US 12,494,293 US 12,548,330 US 12,641,194 US 12,718,380