IP Library › Granted Patent US 12,211,317
Granted Patent B2
US 12,211,317 · App. 17/357,497 · Granted Jan 28, 2025

Highlighting expressive participants in an online meeting

Inventors: Javier Hernandez Rivera (Sommerville, MA); Daniel J. McDuff (Cambridge, MA); Jin A. Suh (Seattle, WA); Kael R. Rowan (Box Elder, SD); Mary P. Czerwinski (Kirkland, WA); Prasanth Murali (Boston, MA); Mohammad Akram (Prague, CZ)
Assignee: Microsoft Technology Licensing, LLC
G06V40/20G06N3/08G06N7/01G06V20/46G06V40/174H04L12/1831
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,211,317
App. No.
17/357,497
Granted
Jan 28, 2025
Kind
B2
Abstract

The present disclosure relate to highlighting audience members with reactions to a presenter of an online meeting. Unlike physical, fact-to-face meeting that enables spontaneous interactions among the presenter and the audiences that are collocated with the presenter, presenting materials during an online meeting raises an issue of the present not being able to see real-time reactions or feedback by the audience members. The present disclosure addresses the issue by dynamically determining one or more audience members who indicate reactions during the online meeting or presentation and displaying faces of the one or more audience members under spotlight to the presenter. The presenter sees faces of the audience members with reactions during the online presentation and responds to the audience members and keep the audience engaged. The spotlight audience server analyzes video frames and determines types of reactions of the audience members.

Claims (62)

1. A computer-implemented method comprising:

receiving data of a video stream, the data associated with a plurality of audience members participating in an online presentation;

deriving, from the received data, a plurality of features-associated with a reaction made by one or more of the plurality of audience members, wherein the deriving a plurality of features comprises:

determining, via an image processing neural network, a head yaw rotation value associated with a head gesture in the video stream; and

inputting the head yaw rotation value into a yaw-detection model to determine the reaction comprises a head shake gesture;

generating at least one expressiveness score based on the reaction made by the one or more of the plurality of audience members, the expressiveness score being generated using a combination of the plurality of features, with one or more features of the plurality of features that are pre-defined as being less-preferred for their conveying of negative engagement by the one or more of the plurality of audience members being weighted less than one or more other features of the plurality of features that are pre-defined as being more-preferred for their conveying of positive engagement by the one or more of the plurality of audience members;

selecting, based on the at least one expressiveness score, a member of the plurality of audience members to highlight; and

transmitting information associated with the selected member of the plurality of audience members to cause an online meeting application to highlight a display of the selected member of the plurality of audience members to a presenter of the online meeting.

2. The computer-implemented method of claim 1 , wherein the plurality of features associated with the reaction made by the one or more of the plurality of audience members comprises one or more of smiles, downturned mouth, mouth open, and closed eyes.

3. The computer-implemented method of claim 1 , the method further comprising:

causing the online meeting application to periodically update, based on a predetermined time interval, the display of the selected member of the plurality of audience members.

4. The computer-implemented method of claim 1 , the method further comprising:

determining a first probabilistic score associated with a facial expression associated with a reaction made by an audience member of the plurality of audience members through classification;

determining a second probabilistic score associated with a furrowed brow by the audience member;

determining a third probabilistic score associated with the head shake gesture; and

generating, based at least on a combination of the first probabilistic score, the second probabilistic score, and the third probabilistic score, an expressiveness score associated with the audience member.

5. The computer-implemented method of claim 4 , wherein the received data associated with the audience member includes a video frame, and wherein determining the first probabilistic score associated with the facial expression uses the image processing neural network with one or regions of interests in the video frame as input.

6. The computer-implemented method of claim 5 , wherein determining the third probabilistic score associated with the head shake gesture uses the yaw-detection model, and wherein the yaw-detection model is a Hidden Markov Model (HMM).

7. The computer-implemented method of claim 6 , further comprising:

determining, by the image processing neural network, a Y-position of at least one facial landmark; and

determining, by the yaw detection model, using the Y-position and the head yaw rotation value, a head nod over time.

8. The computer-implemented method of claim 6 , wherein the second probabilistic score associated with a furrowed brow corresponds to a degree of confusion shown by the audience member.

9. The computer-implemented method of claim 1 , wherein highlighting the display of the selected member of the plurality of audience members comprises displaying live video data including placing a spotlight on the selected member.

10. A system for displaying reactive audience under spotlight in an online meeting, the system comprising:

a processor; and

a memory storing computer-executable instructions that when executed by the processor cause the system to:

receive data associated with a plurality of audience members participating in an online presentation;

derive, from the received data, a plurality of features associated with a reaction made by one or more of the plurality of audience members, wherein the deriving a plurality of features comprises:

determining, via an image processing neural network, a head yaw rotation value associated with a head gesture in the video stream; and

inputting the head yaw rotation value into a yaw-detecting model to determine the reaction comprises a head shake gesture;

generate at least one expressiveness score based on the reaction made by the one or more of the plurality of audience members, the expressiveness score being generated using a combination of the plurality of features, with one or more features of the plurality of features that are pre-defined as being less-preferred for their conveying of negative engagement by the one or more of the plurality of audience members being weighted less than one or more other features of the plurality of features that are pre-defined as being more-preferred for their conveying of positive engagement by the one or more of the plurality of audience members;

select, based on the at least one expressiveness score, a member of the plurality of audience members to highlight; and

transmit information associated with the selected member of the plurality of audience members to cause an online meeting application to highlight a display of the selected member of the plurality of audience members to a presenter of the online meeting.

11. The system of claim 10 , wherein the plurality of features associated with the reaction made by the one or more of the plurality of audience members comprises one or more of smiles, downturned mouth, and brow furrow.

12. The system of claim 10 , the computer-executable instructions when executed further cause the system to:

cause the online meeting application to periodically update, based on a predetermined time interval, the display of the selected member of the plurality of audience members.

13. The system of claim 10 , the computer-executable instructions when executed further cause the system to:

determine a first probabilistic score associated with a facial expression associated with a reaction made by an audience member of the plurality of audience members through classification;

determine a second probabilistic score associated with a furrowed brow by the audience member;

determine a third probabilistic score associated with the head shake gesture; and

generate, based at least on a combination of the first probabilistic score, the second probabilistic score; and the third probabilistic score, an expressiveness score associated with the audience member.

14. The system of claim 13 , wherein the received data associated with the audience member includes a video frame, and wherein determining the first probabilistic score associated with the facial expression uses the image processing neural network with one or regions of interests in the video frame as input.

15. The system of claim 14 , wherein determining the third probabilistic score associated with the head shake gesture uses the yaw-detection model, and wherein the yaw-detection model is a Hidden Markov Model (HMM).

16. A non-transitory computer readable storage media storing computer-executable instructions that when executed by a processor cause a computer system to:

receive data associated with a plurality of audience members participating in an online presentation;

derive, from the received data, a plurality of features associated with a reaction made by one or more of the plurality of audience members, wherein the deriving a plurality of features comprises:

determining, via an image processing neural network, a head yaw rotation value associated with a head gesture in the video stream; and

inputting the head yaw rotation value into a yaw-detecting model to determine the reaction comprises a head shake gesture;

generate at least one expressiveness score based on the reaction made by the one or more of the plurality of audience members, the expressiveness score being generated using a combination of the plurality of features, with one or more features of the plurality of features that are pre-defined as being less-preferred for their conveying of negative engagement by the one or more of the plurality of audience members being weighted less than one or more other features of the plurality of features that are pre-defined as being more-preferred for their conveying of positive engagement by the one or more of the plurality of audience members;

select, based on the at least one expressiveness score, a member of the plurality of audience members to highlight; and

transmit information associated with the selected member of the plurality of audience members to cause an online meeting application to highlight a display of the selected member of the plurality of audience members to a presenter of the online meeting.

17. The computer storage media of claim 16 , wherein the plurality of features associated with the reaction made by the one or more of the plurality of audience members comprises one or more of smiles, downturned mouth, and brow furrow.

18. The computer storage media of claim 16 , the computer-executable instructions when executed further cause the system to:

cause the online meeting application to periodically update, based on a predetermined time interval, the display of the selected member of the plurality of audience members.

19. The computer storage media of claim 16 , the computer-executable instructions when executed further cause the system to:

determine a first probabilistic score associated with a facial expression associated with a reaction made by an audience member of the plurality of audience members through classification;

determine a second probabilistic score associated with a furrowed brow by the audience member;

determine a third probabilistic score associated with the head shake gesture; and

generate, based at least on a combination of the first probabilistic score, the second probabilistic score, and the third probabilistic score, an expressiveness score associated with the audience member.

20. The computer storage media of claim 19 ,

wherein determining the first probabilistic score associated with the facial expression uses the image processing neural network with one or more regions of interests in the video frame as input, and

wherein determining the third probabilistic score associated with the head shake gesture uses the yaw-detection model, and wherein the yaw-detection model is a Hidden Markov Model (HMM).

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 2, 2021
From: RIVERA, JAVIER HERNANDEZ; MCDUFF, DANIEL J.; SUH, JIN A.; ROWAN, KAEL R.; CZERWINSKI, MARY P.; MURALI, PRASANTH; AKRAM, MOHAMMAD
To: MICROSOFT TECHNOLOGY LICENSING, LLC
Reel/Frame 056744/0367 →
Continuity (2)
Provisional Application 63185635 · May 7, 2021
Related Publication 20220358308A1 · Nov 10, 2022
References Cited (79)
US 9547908B1 · Kim · 2017 [cited by examiner]
US 10176365B1 · Ramanarayanan · 2019 [cited by examiner]
US 10311743B2 · Chen · 2019 [cited by examiner]
US 10713814B2 · De Villers-Sidani · 2020 [cited by examiner]
US 10943125B1 · Evans · 2021 [cited by examiner]
US 11277597B1 · Canberk · 2022 [cited by examiner]
US 11503998B1 · De Villers-Sidani · 2022 [cited by examiner]
US 11770197B2 · Zhang · 2023 [cited by examiner]
US 20110263946A1 · el Kaliouby · 2011 [cited by examiner]
US 20140201126A1 · Zadeh · 2014 [cited by examiner]
US 20170319123A1 · Voss · 2017 [cited by examiner]
US 20180189691A1 · Oehrle · 2018 [cited by examiner]
US 20190279393A1 · Ciuc · 2019 [cited by examiner]
US 20200184278A1 · Zadeh · 2020 [cited by examiner]
US 20200228359A1 · El Kaliouby et al. · 2020 [cited by applicant]
US 20200302187A1 · Wang · 2020 [cited by examiner]
US 20210076002A1 · Peters et al. · 2021 [cited by applicant]
US 20210185276A1 · Peters · 2021 [cited by examiner]
US 20220284220A1 · Wu · 2022 [cited by examiner]
WO 2015199871A1 · 2015 [cited by applicant]
Sun, et al., “How Presenters Perceive and React to Audience Flow Prediction In-situ: An Explorative Study of Live Online Lectures”, In Proceedings of the ACM on Human-Computer Interaction, vol. 3, Issue CSCW, Article No… [cited by applicant]
Tan, et al., “A real-time Head Nod and Shake Detector using HMMs”, In Journal of Expert Systems with Applications, vol. 25, Issue 3, Oct. 2003, pp. 461-466. [cited by applicant]
Teevan, et al., “Displaying Mobile Feedback During a Presentation”, In Proceedings of the 14th International Conference on Human-Computer Interaction with Mobile Devices and Services, Sep. 21, 2012, pp. 379-382. [cited by applicant]
Tsang, et al., “Boom Chameleon: Simultaneous Capture of 3D Viewpoint, Voice and Gesture Annotations on a Spatially-Aware Display”, In Proceedings of the 15th Annual ACM Symposium on User interface Software and Technolog… [cited by applicant]
Wegge, Jurgen, “Communication via Videoconference: Emotional and Cognitive Consequences of Affective Personality Dispositions, Seeing One's Own Picture, and Disturbing Events”, In Journal of Human-Computer Interaction, … [cited by applicant]
Zettl, Herbert, “Television Production Handbook”, In Publication of Cengage Learning, 2011. (No copy available). [cited by applicant]
Zolyomi, et al., “Managing Stress: The Needs of Autistic Adults in Video Calling”, In Proceedings of the ACM on Human-Computer Interaction, vol. 3, Issue CSCW, Article No. 134, Nov. 2019, pp. 1-29. [cited by applicant]
Zyto, et al., “Successful Classroom Deployment of a Social Document Annotation System”, In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, May 5, 2012, pp. 1883-1892. [cited by applicant]
“International Search Report and Written Opinion Issued in PCT Application No. PCT/US22/024719”, Mailed Date: Jul. 18, 2022, 11 Pages. [cited by applicant]
“Non-verbal Feedback and Meeting Reactions”, Retrieved from: https://web.archive.org/web/20210419004854/https://support.zoom.us/hc/en-us/articles/115001286183-Non-verbal-feedback-and-reactions-, Apr. 19, 2021, 5 Pages. [cited by applicant]
Akkil, et al., “GazeTorch: Enabling Gaze Awareness in Collaborative Physical Tasks”, In Proceedings of the CHI Conference Extended Abstracts on Human Factors in Computing Systems, May 7, 2016, pp. 1151-1158. [cited by applicant]
Amershi, et al., “Guidelines for Human-AI Interaction”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, May 4, 2019, pp. 1-13. [cited by applicant]
Ames, et al., “Making Love in the Network Closet: the Benefits and Work of Family Videochat”, In Proceedings of the ACM Conference on Computer Supported Cooperative Work, Feb. 6, 2010, pp. 145-154. [cited by applicant]
Barrett, et al., “Emotional Expressions Reconsidered: Challenges to Inferring Emotion From Human Facial Movements”, In Journal of Psychological Science in the Public Interest, vol. 20, Issue 1, Jul. 2019, pp. 1-68. [cited by applicant]
Barrett, Lisaf. , “The Theory of Constructed Emotion: An Active Inference Account of Interoception and Categorization”, In Journal of Social Cognitive and Affective Neuroscience, vol. 12, Issue 1, Oct. 19, 2016, pp. 1-2… [cited by applicant]
Barsoum, et al., “Training Deep Networks for Facial Expression Recognition with Crowd-Sourced Label Distribution”, In Proceedings of the 18th ACM International Conference on Multimodal Interaction, Oct. 31, 2016, pp. 27… [cited by applicant]
Bassett, et al., “The Effects of Positive and Negative Audience Responses on the Autonomic Arousal of Student Speakers”, In Journal of Southern Speech Communication Journal, vol. 38, Issue 3, Mar. 1, 1973, pp. 255-261. [cited by applicant]
Bellotti, et al., “Intelligibility and Accountability: Human Considerations in Context-Aware Systems”, In Journal of Human-Computer Interaction, vol. 16, Issue 2-4, Dec. 1, 2001, pp. 193-212. [cited by applicant]
Bickmore, et al., “Virtual Agents as Supporting Media for Scientific Presentations”, In Journal of Multimodal User Interfaces, vol. 15, Issue 2, Nov. 6, 2020, pp. 131-146. [cited by applicant]
Bishop, et al., “A Survey of Counseling Needs of Male and Female College Students”, In Journal of College Student Development, vol. 39, Issue 2, Jan. 1998, pp. 205-210. [cited by applicant]
Boehner, et al., “Affect: From Information to Interaction”, In Proceedings of the 4th Decennial Conference on Critical Computing: between Sense and Sensibility, Aug. 20, 2005, pp. 59-68. [cited by applicant]
Bohus, et al., “Rapid Development of Multimodal Interactive Systems: A Demonstration of Platform for Situated Intelligence”, In Proceedings of the 19th ACM International Conference on Multimodal Interaction, Nov. 13, 20… [cited by applicant]
Box, Harryc. , “Set Lighting Technician's Handbook: Film Lighting Equipment, Practice, and Electrical Distribution”, In Publication of Focal Press, 2nd Edition, Jan. 17, 1997. (No copy available). [cited by applicant]
Braun, et al., “Using Thematic Analysis in Psychology”, In Journal of Qualitative Research in Psychology, vol. 3, Issue 2, Jan. 1, 2006, pp. 77-101. [cited by applicant]
Brush, et al., “Supporting Interaction Outside of Class: Anchored Discussions vs. Discussion Boards”, In Proceedings of the Computer Support for Collaborative Learning, Jan. 2002, pp. 425-434. [cited by applicant]
Chamillard, A.T. , “Using a Student Response System in CS1 and CS2”, In Proceedings of the 42nd ACM Technical Symposium on Computer Science Education, Mar. 9, 2011, pp. 299-304. [cited by applicant]
Daly, et al., “Avoiding Communication: Shyness, Reticence, and Communication Apprehension”, In Publication of Hampton Press, 1997. (No Copy available). [cited by applicant]
Daly, et al., “Correlates and Consequences of Social-Communicative Anxiety”, In Publication of Sage Beverly Hills, 1984, pp. 21-71. [cited by applicant]
Donald, et al., “Fundamentals of TV Production”, In Publication of John Wiley & Sons, Jul. 1, 2000. [cited by applicant]
Ekman, Rosenberg, “What the Face Reveals: Basic and Applied Studies of Spontaneous Expression using the Facial Action Coding System (FACS)”, In Publication of Oxford University Press, 1997. (No Copy available). [cited by applicant]
Glassman, et al., “Mudslide: A Spatially Anchored Census of Student Confusion for Online Lecture Videos”, In Proceedings of the 33rd Annual ACM Conference on Human Factors in Computing Systems, Apr. 18, 2015, pp. 1555-1… [cited by applicant]
Haimson, et al., “What Makes Live Events Engaging on Facebook Live, Periscope, and Snapchat”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, May 6, 2017, pp. 48-60. [cited by applicant]
Hassib, et al., “A Design Space for Audience Sensing and Feedback Systems”, In Proceedings of Extended Abstracts of the CHI Conference on Human Factors in Computing Systems, Apr. 21, 2018, pp. 1-6. [cited by applicant]
Hassib, et al., “EngageMeter: A System for Implicit Audience Engagement Sensing Using Electroencephalography”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, May 6, 2017, pp. 5114-5119. [cited by applicant]
Hook, Kristina, “Affective Loop Experiences: Designing for Interactional Embodiment”, In Journal of Philosophical Transactions of the Royal Society B: Biological Sciences, vol. 364, Issue 1535, Dec. 12, 2009, pp. 3585-3… [cited by applicant]
Hsieh, et al., “Three Approaches to Qualitative Content Analysis”, In Journal of Qualitative Health Research, vol. 15, Issue 9, Nov. 2005, pp. 1277-1288. [cited by applicant]
Khan, et al., “Designing an Eyes-Reduced Document Skimming App for Situational Impairments”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, Apr. 25, 2020, pp. 1-14. [cited by applicant]
Khan, et al., “Spotlight: Directing Users' Attention on Large Displays”, In Proceedings of the SIGCHI Conference on Human Factors in Computing Systems, Apr. 2, 2005, pp. 791-798. [cited by applicant]
Kirschbaum, et al., “The ‘Trier Social Stress Test’—A Tool for Investigating Psychobiological Stress Responses in a Laboratory Setting”, In Journal of Neuropsychobiology, vol. 28, Issue 1-2, Feb. 1993, pp. 76-81. [cited by applicant]
Lugo-Fagundo, et al., “New Frontiers in Education: Facebook as a Vehicle for Medical Information Delivery”, In Journal of the American College of Radiology, vol. 13, issue 3, Mar. 1, 2016, pp. 316-319. [cited by applicant]
Macintyre, et al., “The Effects of Audience Interest, Responsiveness, and Evaluation on Public Speaking Anxiety and Related Variables”, In Journal of Communication Research Reports , vol. 14, Issue 2, Mar. 1, 1997, pp. … [cited by applicant]
Macintyre, et al., “The Effects of Speaker Personality on Anticipated Reactions to Public Speaking”, In Journal of Communication Research Reports, vol. 12, Issue 2, Sep. 1, 1995, pp. 125-133. [cited by applicant]
Mavadati, et al., “Extended DISFA Dataset: Investigating Posed and Spontaneous Facial Expressions”, In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition Workshops, Jun. 2016, pp. 1-8. [cited by applicant]
McCroskey, et al., “Oral Communication Apprehension”, In Publication of Springer, 1986, pp. 279-293. [cited by applicant]
McCroskey, et al., “Self-Report as an Approach to Measuring Communication Competence”, In Journal of Communication Research Reports, vol. 5, Issue 2, Dec. 1988. [cited by applicant]
McDuff, et al., “A Multimodal Emotion Sensing Platform for Building Emotion-Aware Applications”, In the Repository of arXiv:1903.12133, Mar. 28, 2019, 6 Pages. [cited by applicant]
Miller, et al., “Through the Looking Glass: The Effects of Feedback on Self-Awareness and Conversational Behaviour during Video Chat”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, May 6, 2… [cited by applicant]
Motley, Michaelt. , “Public Speaking Anxiety Qua Performance Anxiety: A Revised Model and an Alternative Therapy”, In Journal of Social Behavior and Personality, vol. 5, Issue 2, Jan. 1, 1990. (No copy available). [cited by applicant]
Murali, et al., “AffectiveSpotlight: Facilitating the Communication of Affective Responses from Audience Members during Online Presentations”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, … [cited by applicant]
Murali, et al., “Speaker Hand-Offs in Collaborative Human-Agent Oral Presentations”, In Proceedings of the 18th International Conference on Intelligent Virtual Agents, Nov. 5, 2018, pp. 153-158. [cited by applicant]
Parmar, et al., “Making It Personal: Addressing Individual Audience Members in Oral Presentations Using Augmented Reality”, In Proceedings of the ACM on Interactive, Mobile, Wearable and Ubiquitous Technologies, vol. 4,… [cited by applicant]
Pellacini, et al., “A User Interface for Interactive Cinematic Shadow Design”, In Journal of ACM Transactions on Graphics, vol. 21, Issue 3, Jul. 2002, pp. 563-566. [cited by applicant]
Picard, et al., “The Galvactivator: A Glove that Senses and Communicates Skin Conductivity”, In Proceedings of the 9th International Conference on HCI, Aug. 1, 2001, 6 Pages. [cited by applicant]
Radbourne, et al., “The Audience Experience: Measuring Quality in the Performing Art”, In International journal of arts management, vol. 11, Issue 3, Apr. 2009, pp. 16-29. [cited by applicant]
Rajcic, et al., “Mirror Ritual: An Affective Interface for Emotional Self-Reflection”, In Proceedings of the CHI Conference on Human Factors in Computing Systems, Apr. 25, 2020, pp. 1-13. [cited by applicant]
Reeves, et al., “The Effects of Screen Size and Message Content on Attention and Arousal”, In Journal of Media Psychology, vol. 1, Issue 1, Mar. 1, 1999, pp. 49-67. [cited by applicant]
Rivera-Pelayo, et al., “Live Interest Meter: Learning from Quantified Feedback in Mass Lectures”, In Proceedings of the Third International Conference on Learning Analytics and Knowledge, Apr. 8, 2013, pp. 23-27. [cited by applicant]
Solina, Franc, “15 Seconds of Fame”, In Journal of Leonardo, vol. 37, Issue 2, Apr. 1, 2004, pp. 105-110. [cited by applicant]
Straus, et al., “The Effects of Videoconference, Telephone, and Face-To-Face Media on Interviewer and Applicant Judgments in Employment Interviews”, In Journal of Management, vol. 27, Issue 3, Jun. 1, 2001, pp. 363-381. [cited by applicant]
Cited By (1)
US 12,573,198