IP Library Granted Patent US 11,290,686
Granted Patent B2
US 11,290,686 · App. 17/186,977 · Granted Mar 29, 2022

Architecture for scalable video conference management

Inventors: Michael H Peters (Washington, DC); Alexander M. Stufflebeam (Indianapolis, IN)
Assignee: Michael H Peters
H04N7/152G06K9/00718
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,290,686
App. No.
17/186,977
Granted
Mar 29, 2022
Kind
B2
Abstract

In some implementations, an endpoint device captures video data during a network-based communication session. The endpoint devices processes a stream of user state data indicating attributes of a user of the endpoint device at different times during the network-based communication session. The endpoint device transmits the stream of user state data over a communication network to a server system. The endpoint device receives, over the communication network, (i) content of the network-based communication session and (ii) additional content based on user state data generated by the respective endpoint devices each processing video data that the respective endpoint devices captured during the network-based communication session. The endpoint device presents a user interface providing the received content of the network-based communication session concurrent with the received additional content that is based on the user state data generated by the respective endpoint devices.

Claims (38)

1. A method of providing enhanced video communication over a network, the method comprising:

capturing, by a first endpoint device, video data during a network-based communication session using a camera of the first endpoint device to generate a video data stream, wherein the network-based communication session involves the first endpoint device and one or more other endpoint devices;

processing, by the first endpoint device, the captured video data to generate a stream of user state data indicating attributes of a user of the first endpoint device at different times during the network-based communication session, the processing including performing facial analysis on the captured video data to evaluate images of a face of the user of the first endpoint device;

transmitting, by the first endpoint device, the stream of user state data indicating the attributes over a communication network to a server system configured to aggregate and distribute, during the network-based communication session, user state data generated by the respective endpoint devices each processing video data that the respective endpoint devices captured during the network-based communication session;

receiving, by the first endpoint device over the communication network, (i) content of the network-based communication session and (ii) additional content based on user state data generated by the respective endpoint devices each processing video data that the respective endpoint devices captured during the network-based communication session; and

presenting, by the first endpoint device and during the network-based communication session, a user interface providing the received content of the network-based communication session concurrent with the received additional content based on the user state data generated by the respective endpoint devices.

2. The method of claim 1 , wherein processing the captured video data comprises using a trained machine learning model at the first endpoint device to process images of a face of the user of the first endpoint device, wherein the output of the trained machine learning model comprises scores for each of multiple emotions.

3. The method of claim 1 , wherein receiving the content of the network-based communication session comprises video data from by one or more of the other endpoint devices.

4. The method of claim 1 , wherein the determined user state data comprises cognitive or emotional attributes of the user of the first endpoint device.

5. The method of claim 1 , comprising applying a time-series transformation to a sequence of scores generated by the first endpoint device for attributes of the user of the first endpoint device at different times during the network-based communication session.

6. The method of claim 5 , wherein the time-series transformation includes determining at least one of a minimum, maximum, mean, mode, range, or variance for a sequence of scores generated by the first endpoint device.

7. The method of claim 5 , wherein the stream of user state data comprises a series of analysis results each determined based on different overlapping windows of the video data stream captured by the first endpoint device.

8. The method of claim 1 , further comprising transmitting the video data stream to the server system over the communication network.

9. The method of claim 1 , wherein the server system is a first server system, wherein the first endpoint device receives the additional content from the first server system, and wherein the endpoint device receives the content of the network-based communication session from a second server system that is different from the first server system.

10. The method of claim 1 , wherein the received additional content comprises a measure of participation, engagement, or collaboration for a group of multiple participants in the communication session, the measure being based on aggregated user state data for the multiple participants.

11. An endpoint device comprising:

one or more processors; and

one or more computer-readable media storing instructions that, when executed by the one or more processors, cause the endpoint device to perform operations comprising:

capturing, by the endpoint device, video data during a network-based communication session using a camera of the endpoint device to generate a video data stream, wherein the network-based communication session involves the endpoint device and one or more other endpoint devices;

processing, by the endpoint device, the captured video data to generate a stream of user state data indicating attributes of a user of the endpoint device at different times during the network-based communication session, the processing including performing facial analysis on the captured video data to evaluate images of a face of the user of the endpoint device;

transmitting, by the endpoint device, the stream of user state data indicating the attributes over a communication network to a server system configured to aggregate and distribute, during the network-based communication session, user state data generated by the respective endpoint devices each processing video data that the respective endpoint devices captured during the network-based communication session;

receiving, by the endpoint device over the communication network, (i) content of the network-based communication session and (ii) additional content based on user state data generated by the respective endpoint devices each processing video data that the respective endpoint devices captured during the network-based communication session; and

presenting, by the endpoint device and during the network-based communication session, a user interface providing the received content of the network-based communication session concurrent with the received additional content based on the user state data generated by the respective endpoint devices.

12. The endpoint device of claim 11 , wherein processing the captured video data comprises using a trained machine learning model at the endpoint device to process images of a face of the user of the endpoint device, wherein the output of the trained machine learning model comprises scores for each of multiple emotions.

13. The endpoint device of claim 11 , wherein receiving the content of the network-based communication session comprises video data from by one or more of the other endpoint devices.

14. The endpoint device of claim 11 , wherein the user state data comprises cognitive or emotional attributes of the user of the endpoint device.

15. The endpoint device of claim 11 , wherein the operations comprise applying a time-series transformation to a sequence of scores generated by the endpoint device for attributes of the user of the endpoint device at different times during the network-based communication session.

16. The endpoint device of claim 15 , wherein the time-series transformation includes determining at least one of a minimum, maximum, mean, mode, range, or variance for a sequence of scores generated by the endpoint device.

17. The endpoint device of claim 15 , wherein the stream of user state data comprises a series of analysis results each determined based on different overlapping windows of the video data stream captured by the endpoint device.

18. The endpoint device of claim 11 , wherein the server system is a first server system, wherein the endpoint device receives the additional content from the first server system, and wherein the operations comprise receiving the content of the network-based communication session from a second server system that is different from the first server system.

19. One or more non-transitory computer-readable media storing instructions that, when executed by one or more processors of an endpoint device, cause the endpoint device to perform operations comprising:

capturing, by the endpoint device, video data during a network-based communication session using a camera of the endpoint device to generate a video data stream, wherein the network-based communication session involves the endpoint device and one or more other endpoint devices;

processing, by the endpoint device, the video data captured by the endpoint device to generate a stream of user state data indicating attributes of a user of the endpoint device at different times during the network-based communication session, the processing including the endpoint device performing facial analysis on the video data by the endpoint device to evaluate images of a face of the user of the endpoint device;

transmitting, by the endpoint device, the stream of user state data indicating the attributes over a communication network to a server system configured to aggregate and distribute, during the network-based communication session, user state data generated by the respective endpoint devices each performing facial analysis on video data that the respective endpoint devices captured during the network-based communication session;

receiving, by the endpoint device over the communication network, (i) content of the network-based communication session and (ii) additional content based on user state data, transmitted to the server by the one or more other endpoint devices, that the respective endpoint devices each generated by performing facial analysis on the video data that the endpoint device captured of its user during the network-based communication session; and

presenting, by the endpoint device and during the network-based communication session, a user interface providing the received content of the network-based communication session concurrent with the received additional content that is based on the user state data generated by the respective endpoint devices by performing facial analysis on the video data that they respectively captured.

20. The one or more non-transitory computer-readable media of claim 19 , wherein the user state data generated and transmitted by the endpoint device comprises values determined by the endpoint device using facial analysis that indicate levels of emotions of the user of the endpoint device; and

wherein the user state data generated and transmitted by the one or more other endpoint devices comprises values, determined by the respective other endpoints using facial analysis, that indicate levels of emotions of the users of the respective other endpoints.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 24, 2023
From: PETERS, MICHAEL H.
To: REELAY MEETINGS, INC.
Reel/Frame 063424/0284 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 18, 2021
From: STUFFLEBEAM, ALEXANDER M.
To: MICHAEL H PETERS
Reel/Frame 057210/0884 →
Continuity (14)
Continuation In Part 17037324 · Sep 29, 2020
Continuation In Part 16993010 · Aug 13, 2020
Continuation 16516731 · Jul 19, 2019
Continuation 16128137 · Sep 11, 2018
Continuation 17186977 · Feb 26, 2021
Continuation In Part 16950888 · Nov 17, 2020
Continuation In Part 16993010 · Aug 13, 2020
Continuation 16516731 · Jul 19, 2019
Continuation 16128137 · Sep 11, 2018
Provisional Application 62556672 · Sep 11, 2017
Provisional Application 63072936 · Aug 31, 2020
Provisional Application 63075809 · Sep 8, 2020
Provisional Application 63088449 · Oct 6, 2020
Related Publication 20210185276A1 · Jun 17, 2021
Cited By (4)
US 12,494,293 US 12,548,330 US 12,556,654 US 12,641,194