IP Library › Granted Patent US 12,334,051
Granted Patent B2
US 12,334,051 · App. 17/811,867 · Granted Jun 17, 2025

Computationally reacting to a multiparty conversation

Inventors: Andrew Reece (San Francisco, CA); Peter Bull (San Francisco, CA); Gus Cooney (San Francisco, CA); Casey Fitzpatrick (San Francisco, CA); Gabriella Rosen Kellerman (San Francisco, CA); Ryan Sonnek (New Hope, MN)
Assignee: BetterUp, Inc.
G10L15/063G06N20/00G06V20/49G10L15/22
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,334,051
App. No.
17/811,867
Granted
Jun 17, 2025
Kind
B2
Abstract

Technology is provided for causing a computing system to extract conversation features from a multiparty conversation (e.g., between a coach and mentee), apply the conversation features to a machine learning system to generate conversation analysis indicators, and apply a mapping of conversation analysis indicators to actions and inferences to determine actions to take or inferences to make for the multiparty conversation. In various implementations, the actions and inferences can include determining scores for the multiparty conversation such as a score for progress toward a coaching goal, instant scores for various points throughout the conversation, conversation impact score, ownership scores, etc. These scores can be, e.g., surfaced in various user interfaces along with context and benchmark indicators, used to select resources for the coach or mentee, used to update coach/mentee matchings, used to provide real-time alerts to signify how the conversation is going, etc.

Claims (95)

1. A method for computationally reacting to conversations, the method comprising:

receiving one or more multiparty videos representing a conversation between at least a first user and a second user, each multiparty video including acoustic data and video data; and

generating conversation features based on the one or more multiparty videos by:

segmenting the one or more multiparty videos into multiple utterances;

identifying, for each the multiple utterances, data for multiple modalities; and

extracting conversation features, for each particular utterance of the multiple utterances, from each of the data for the multiple modalities associated with that particular utterance;

generating conversation analysis indicators, comprising one or more conversation scores for the conversation, by applying the conversation features to a plurality of machine learning systems, wherein generating the conversation analysis indicators comprises receiving the one or more conversation scores from the plurality of machine learning systems, wherein each machine learning system is individually trained using labeled data sets for a different conversation analysis indicator;

applying a mapping of the conversation analysis indicators to inferences or actions to determine inference or action results;

reacting to the conversation according to the inference or action results mapped to the conversation analysis indicators in the mapping;

transmitting a first alert notification to a user device associated with the first user if a progress score is above a first predetermined threshold value, the progress score being determined based on profile of the first user and the conversation analysis indicators; and

transmitting a second alert notification to a coaching device associated with the second user if the conversation analysis indicators are below a second predetermined threshold value.

2. The method of claim 1 , wherein the conversation analysis indicators comprise one or more conversation impact scores indicating a projected impact of the conversation on at least one of the first user and/or second user.

3. The method of claim 1 , wherein generating the conversation analysis indicators comprises:

wherein the mapping includes a mapping of progress scores being below a threshold to an action including transmitting an alert of a low-quality pairing between the first user and the second user.

4. The method of claim 1 , wherein the reacting to the conversation includes causing the first alert notification or the second alert notification to be transmitted with a control to initiate selection of a new pairing for the first user or the second user.

5. The method of claim 1 further comprising:

retrieving a first user profile for the first user, the first user profile including a progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

a conversation score benchmark indicating a comparison between the conversation analysis indicators and a reference score; and

the updated progress score in relation to a goal score.

6. The method of claim 1 further comprising:

retrieving a first user profile for the first user, the first user profile including a progress score;

updating the progress score based on the conversation analysis indicators; and generating data for a user interface including:

the updated progress score in relation to a goal score; and

a progress score chart indicating historical progress the first user has made toward having the progress score reaching the goal score.

7. The method of claim 1 further comprising:

retrieving a first user profile for the first user, the first user profile including a progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

the updated progress score in relation to a goal score; and

at least one sub-score from the conversation analysis indicators including one or more of an engagement score, an ownership score, an openness score, or any combination thereof.

8. A non-transitory computer-readable storage medium storing instructions that, when executed by a computing system, cause the computing system to perform operations comprising:

receiving one or more multiparty videos representing a conversation between at least a first user and a second user, each multiparty video including acoustic data and video data; and

generating conversation features based on the one or more multiparty videos by:

segmenting the one or more multiparty videos into multiple utterances;

identifying, for each the multiple utterances, data for multiple modalities; and

extracting conversation features, for each particular utterance of the multiple utterances, from each of the data for the multiple modalities associated with that particular utterance;

generating conversation analysis indicators, comprising one or more conversation scores for the conversation, by applying the conversation features to a plurality of machine learning systems, wherein generating the conversation analysis indicators comprises receiving the one or more conversation scores from the plurality of machine learning systems, wherein each machine learning system is individually trained using labeled data sets for a different conversation analysis indicator;

applying a mapping of the conversation analysis indicators to inferences or actions to determine inference or action results;

reacting to the conversation according to the inference or action results mapped to the conversation analysis indicators in the mapping;

transmitting a first alert notification to a user device associated with the first user if a progress score is above a first predetermined threshold value, the progress score being determined based on profile of the first user and the conversation analysis indicators; and

transmitting a second alert notification to a coaching device associated with the second user if the conversation analysis indicators are below a second predetermined threshold value.

9. The non-transitory computer-readable storage medium of claim 8 , wherein the conversation analysis indicators comprise one or more conversation impact scores indicating a projected impact of the conversation on at least one of the first user and/or second user.

10. The non-transitory computer-readable storage medium of claim 8 ,

wherein generating the conversation analysis indicators comprises:

wherein the mapping includes a mapping of progress scores being below a threshold to an action including transmitting an alert of a low-quality pairing between the first user and the second user.

11. The non-transitory computer-readable storage medium of claim 8 , wherein the reacting to the conversation includes causing the first alert notification or the second alert notification to be transmitted with a control to initiate selection of a new pairing for the first user or the second user.

12. The non-transitory computer-readable storage medium of claim 8 , further comprising:

retrieving a first user profile for the first user, the first user profile including the progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

a conversation score benchmark indicating a comparison between the conversation analysis indicators and a reference score; and

the updated progress score in relation to a goal score.

13. The non-transitory computer-readable storage medium of claim 8 , further comprising:

retrieving a first user profile for the first user, the first user profile including the progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

the updated progress score in relation to a goal score; and

a progress score chart indicating historical progress the first user has made toward having the progress score reaching the goal score.

14. The non-transitory computer-readable storage medium of claim 8 , further comprising:

retrieving a first user profile for the first user, the first user profile including the progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

the updated progress score in relation to a goal score; and

at least one sub-score from the conversation analysis indicators including one or more of an engagement score, an ownership score, an openness score, or any combination thereof.

15. A computing system comprising:

one or more processors; and

one or more memories storing instructions that, when executed by the one or more processors, cause the computing system to perform operations comprising:

receiving one or more multiparty videos representing a conversation between at least a first user and a second user, each multiparty video including acoustic data and video data; and

generating conversation features based on the one or more multiparty videos by:

segmenting the one or more multiparty videos into multiple utterances;

identifying, for each the multiple utterances, data for multiple modalities; and

extracting conversation features, for each particular utterance of the multiple utterances, from each of the data for the multiple modalities associated with that particular utterance;

generating conversation analysis indicators, comprising one or more conversation scores for the conversation, by applying the conversation features to a plurality of machine learning systems, wherein generating the conversation analysis indicators comprises receiving the one or more conversation scores from the plurality of machine learning systems, wherein each machine learning system is individually trained using labeled data sets for a different conversation analysis indicator;

applying a mapping of the conversation analysis indicators to inferences or actions to determine inference or action results;

reacting to the conversation according to the inference or action results mapped to the conversation analysis indicators in the mapping;

transmitting a first alert notification to a user device associated with the first user if a progress score is above a first predetermined threshold value, the progress score being determined based on profile of the first user and the conversation analysis indicators; and

transmitting a second alert notification to a coaching device associated with the second user if the conversation analysis indicators are below a second predetermined threshold value.

16. The system of claim 15 , wherein the conversation analysis indicators comprise one or more conversation impact scores indicating a projected impact of the conversation on at least one of the first user and/or second user.

17. The system of claim 15 , wherein generating the conversation analysis indicators comprises:

wherein the mapping includes a mapping of progress scores being below a threshold to an action including transmitting an alert of a low-quality pairing between the first user and the second user.

18. The system of claim 15 , wherein the reacting to the conversation includes causing the first alert notification or the second alert notification to be transmitted with a control to initiate selection of a new pairing for the first user or the second user.

19. The system of claim 15 , further comprising:

retrieving a first user profile for the first user, the first user profile including the progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

a conversation score benchmark indicating a comparison between the conversation analysis indicators and a reference score; and

the updated progress score in relation to a goal score.

20. The system of claim 15 , further comprising:

retrieving a first user profile for the first user, the first user profile including the progress score;

updating the progress score based on the conversation analysis indicators; and

generating data for a user interface including:

the updated progress score in relation to a goal score; and

a progress score chart indicating historical progress the first user has made toward having the progress score reaching the goal score.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 23, 2022
From: REECE, ANDREW; BULL, PETER; COONEY, GUS; FITZPATRICK, CASEY; KELLERMAN, GABRIELLA ROSEN; SONNEK, RYAN
To: BETTERUP, INC.
Reel/Frame 060872/0593 →
Continuity (2)
Continuation 16798248 · Feb 21, 2020
Related Publication 20220343899A1 · Oct 27, 2022
References Cited (98)
US 6611822B1 · Beams · 2003 [cited by examiner]
US 9392941B2 · Powch · 2016 [cited by examiner]
US 20020101505A1 · Gutta et al. · 2002 [cited by applicant]
US 20060031087A1 · Fox · 2006 [cited by examiner]
US 20060233347A1 · Tong · 2006 [cited by examiner]
US 20090210078A1 · Crowley · 2009 [cited by examiner]
US 20090271438A1 · Agapi et al. · 2009 [cited by applicant]
US 20090287738A1 · Colbran · 2009 [cited by examiner]
US 20110060591A1 · Chanvez et al. · 2011 [cited by applicant]
US 20110295392A1 · Cunnington et al. · 2011 [cited by applicant]
US 20160042226A1 · Cunico et al. · 2016 [cited by applicant]
US 20160073054A1 · Balasaygun et al. · 2016 [cited by applicant]
US 20160314620A1 · Reilly · 2016 [cited by examiner]
US 20180075376A1 · Dill · 2018 [cited by examiner]
US 20180190135A1 · Robichaux · 2018 [cited by examiner]
US 20180214075A1 · Falevsky et al. · 2018 [cited by applicant]
US 20190318641A1 · Segal · 2019 [cited by examiner]
US 20200106988A1 · Harpur et al. · 2020 [cited by applicant]
US 20200126437A1 · Fieldman · 2020 [cited by examiner]
US 20210076002A1 · Peters · 2021 [cited by examiner]
US 20210158235A1 · Sivasubramanian · 2021 [cited by examiner]
US 20210264900A1 · Reece et al. · 2021 [cited by applicant]
“Create Compatible Mentoring Matches In Minutes”, https://chronus.com/software/mentoring-software/mentor-matching, retrieved from: www.archive.org, archived on Aug. 29, 2018 (Year: 2018). [cited by examiner]
Poria et al. “Context-Dependent Sentiment Analysis in User-Generated Videos,” Proc. of 55th Annual Meeting of the Association for Computational Linguistics, pp. 873-883, VancouverCanada, Jul. 30-Aug. 4, 2017. [cited by applicant]
Poria eta. “MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversations,” Jun. 4, 2019, 10 pages. [cited by applicant]
Poria, Soujanya et al. “Multi-level Multiple Attentions for Contextual Multimodal Sentiment Analysis,” 6 pages. [retrieved Feb. 26, 2020]. [cited by applicant]
Poria, Soujanya. “A review of affective computing: From unimodal analysis to multimodal fusion,” Information Fusion 37 (2017) 98-125. [cited by applicant]
Ravanelli, M. et al. “The Pytorch-Kalki Speech Recognition Toolkit,” 2019. [accessed Mar. 2, 2020]. [cited by applicant]
Ravanelli, Mirco et al. “The Pytorch-Kaldi Speech Recognition Toolkit,” Feb. 15, 2019, 5 pages. [cited by applicant]
Roddy et al. “Detecting Conversational Gaze Aversion Using Unsupervised Learning,” 2017 25th European Signal Processing Conf., pp. 86-90. [cited by applicant]
Rogers, Shane L. et al. “Using dual eye tracking to uncover personal gaze patterns during social interaction,” Scientific Reports (2018) 8:4271. published Mar. 9, 2018, 9 pages. [cited by applicant]
Romero, Daniel M. et al. “Mimicry is Presidential: Linguistic Style Matching in Presidential Debates and Improved Poling Numbers,” 2014, 27 pages. [cited by applicant]
Sahay, Saurav et al. “Multimodal Relational Tensor Network for Sentiment and Emotion Classification,” Proceedings of Fist Grand Challenge on Human Multimodal Language, p. 20-27, Australia Jul. 20, 2018. [cited by applicant]
Satt et al. “Efficient Emotion Recognition from Speech Using Deep Learning on Spectrograms”, Interspeech Aug. 20-24, 2017, Stockholm Sweden, pp. 1089-1093. [cited by applicant]
Schrammel, Franziska et al. “Virtual friend or threat? The effects of facial expression and gaze interaction on psychophysiological responses and emotional experience,” Psychophysiology 46 (2009), 922-931. [cited by applicant]
Tripathi, Samarth, “Multi-Modal Emotion Recognition on IEMOCAP with Neural Networks,” Nov. 6, 2019, 5 pages. [cited by applicant]
Trustees of the University of Pennsylvania 1996. “Boston University Radio Speech Corpus” [accessed Feb. 26, 2020]. [cited by applicant]
Trustees of the University of Pennsylvania. “Emotional Prosody Speech and Transcripts” Jul. 23, 2002. [accessed Feb. 26, 2020]. [cited by applicant]
Tyiannak, pyAudioAnalysis, “A Python library for audio feature extraction, classification, segmentation and applications,” [accessed Mar. 2, 2020]. [cited by applicant]
University of Surrey, “Surrey Audio-Visual Expressed Emotion (SA VEE) Database” Apr. 2, 2015. [accessed Feb. 25, 2020]. [cited by applicant]
UT Dallas MSP Multimodal Signal Processing Laboratory. “MSP—Improv corpus: An emotional audiovisual database of spontaneous improvisations”. [accessed Feb. 25, 2020]. [cited by applicant]
Ververidis, Dimitrios et al. “A State of the Art Review on Emotional Speech Databases,” 2003, 11 pages. [cited by applicant]
Wu, Chung-Hsien et al. “Survey on audiovisual emotion recognition: databases, features, and data fusion strategies,” 2014, 18 pages. [cited by applicant]
Yang, Xitong et al. “Deep Multimodal Representation Learning from Temporal Data,” Apr. 11, 2017, 9 pages. [cited by applicant]
Zadeh, Amir et al. “Memory Fusion Network for Multi-view Sequential Learning,” Feb. 3, 2018, 9 pages. [cited by applicant]
Zadeh, Amir et al. “Tensor Fusion Network for Multimodal Sentiment Analysis,” Jul. 23, 2017, 12 pages. [cited by applicant]
Zhang et al. “Attention Based Fully Convolutional Network for Speech Emotion Recognition,” May 2, 2019, 5 pages. [cited by applicant]
Zhang, “Appearance-based Gaze Estimation in the Wild (MPIIGaze)”, Max-Pianck-Institut, 2015. [cited by applicant]
“Clmtrackr” [accessed Mar. 2, 2020]. [cited by applicant]
“CMU-MultimodaiSKD” [accessed Feb. 25, 2020]. [cited by applicant]
“Cohn-Kanade (CK and CK+) database Download Site” [accessed Feb. 25, 2020]. [cited by applicant]
“CREMA-D (Crowd-sourced Emotional Multimodal Actors Dataset)” [accessed Feb. 25, 2020]. [cited by applicant]
“DisVoice” [accessed Mar. 2, 2020]. [cited by applicant]
“Facial Expressions in the Wild Databases and Experimental Protocols” [accessed Feb. 25, 2020]. [cited by applicant]
“Ffmpy: Pythonic interface for FFmpeg/FFprobe command line,” [accessed Mar. 2, 2020]. [cited by applicant]
“Kaldi Speech Recognition Toolkit” [accessed Mar. 2, 2020]. [cited by applicant]
“Librosa: Python package for audio and music analysis” [accessed Mar. 2, 2020]. [cited by applicant]
“MELD: A Multimodal Multi-Party Dataset for Emotion Recognition in Conversation” . [accessed Feb. 25, 2020]. [cited by applicant]
“MMI Facial Expression Database” [accessed Feb. 25, 2020]. [cited by applicant]
“OpenCV on Wheels,” [accessed Mar. 2, 2020]. [cited by applicant]
“Real-world Affective Faces (RAF) Database” [accessed Feb. 25, 2020]. [cited by applicant]
“Scikit-video: video processing in python,” version 1.1.11, copyright 2015-2017 [accessed Mar. 2, 2020]. [cited by applicant]
“The Emotional Voices Database: Towards Controlling the Emotional Expressiveness in Voice Generation Systems” [accessed Feb. 26, 2020]. [cited by applicant]
“Tracking.js: A modern approach for Computer Vision on the web,” [accessed, <tracking.js>, [accessed Mar. 2, 2020]. [cited by applicant]
Amos, Brandon et al. “OpenFace: A general-purpose face recognition library with mobile applications,” Jun. 2016, Carnegie Mellon University, Pittsburgh PA, 20 pages. [cited by applicant]
U.S. Appl. No. 16/798,248, Examiner Interview Summary mailed Apr. 7, 2022, 1 pg. [cited by applicant]
U.S. Appl. No. 16/798,248, Non-Final Office Action mailed Nov. 12, 2021, 10 pgs. [cited by applicant]
U.S. Appl. No. 16/798,248, Notice of Allowance mailed Apr. 7, 2022, 14 pgs. [cited by applicant]
Arriaga et al. “Real-time Convolutional Neural Networks for Emotion and Gender Classification,” retrieved Feb. 26, 2010; 5 pages. [cited by applicant]
Astorfi, “SpeechPy Official Project Documentation” [accessed Mar. 2, 2020]. [cited by applicant]
AudEERING, “openSMILE” [accessed Mar. 2, 2020]. [cited by applicant]
Baltrusaitis, Tadas et al. “Multimodal Machine Learning: A Survey and Taxonomy,” Aug. 1, 2017, 20 pages. [cited by applicant]
Binetti, Nicola et al. “Pupil dilation as an index of preferred mutual gaze duration,” 2016 Royal Society Open Science 3: 160086; 11 pages. [cited by applicant]
Cambria, Erik et al. “SenticNet 5: Discovering Conceptual Primitives for Sentiment Analysis by Means of Context Embeddings,” The Thirty-Second AAAI Conference on Artificial Intelligence 2018, pp. 1795-1802. [cited by applicant]
Department of Linguistics, University of California, Santa Barbara. “Santa Barbara Corpus of Spoken American English” 2020. ˜ .[accessed Feb. 26, 2020]. [cited by applicant]
Eyben, Florian et al. “On-line emotion recognition in a 3-D activation-valence-time continuum using acoustic and linguistic cues,” J. Multimodal User Interfaces (2010) 3:7-19. [cited by applicant]
Google Inc. “AudioSet” [accessed Feb. 26, 2020]. [cited by applicant]
Gregory et al. “A Nonverbal Signal in Voices of Interview Partners Effectively Predicts Communication Accommodation and Social Status Perceptions,” J. Personality Social Psych 1996, V 70, No. 6, 1231-40. [cited by applicant]
Hannun, Awni et al. “Deep Speech: Scaling up end-to-end speech recognition,” Dec. 14, 2014, 12 pages. [retrieved Feb. 26, 2020]. [cited by applicant]
Hazarika et al. “ICON: Interactive Conversational Memory Network for Multimodal Emotion Detection,” Proc. of 2018 Conference on Empirical Methods in Natural Language Processing, p. 2594-2604, Oct. 31-Nov. 4, 2018, Belgi… [cited by applicant]
Hazarika, D. et al. “Conversational Memory Network for Emotion Recognition in Dyadic Dialogue Videos,” Proceedings of NAACL-HL T 2018, pp. 2122-2132, New Orleans LA Jun. 1-6, 2018. [cited by applicant]
Honma, M. et al. “Perceptual and not physical eye contact elicits pupillary dilation,” Jan. 2012, abstract. [cited by applicant]
Hossain, M. Shamim et al. “An emotion recognition system for mobile applications,” IEEE Access, Feb. 2017, 7 pages. [cited by applicant]
Hung and Puthran. “Emotion Detection through Speech,” Final Project Documentation CSYE7374, 9 pages. [retrieved Feb. 26, 2020]. [cited by applicant]
I3.FBK.EU. “DaFEx—a Database of Kinetic Facial Expressions” [accessed Feb. 25, 2020]. [cited by applicant]
Imperial College, London. “The SEMAINE database: annotated multimodal records of emotionally coloured conversations between a person and a limited agent” Jan. 2007. [accessed Feb. 26, 2020]. [cited by applicant]
Kahou et al. “Recurrent Neural Networks for Emotion Recognition in Video,” ICMI 2015, Nov. 9-13, 2015, Seattle, WA, 8 pages. [cited by applicant]
Livingstone SR, Russo FA, Apr. 5, 2018. “The Ryerson Audio-Visual Database of Emotional Speech and Song (RAVDESS)” [accessed Feb. 25, 2020]. [cited by applicant]
Mahoor PhD, Mohammad H. “AffectNet” 2020. [cited by applicant]
Majumder, N. et al. “DialogueRNN: An Attentive RNN for Emotion Detection in Conversations,” May 25, 2019, 8 pages. Association for the Advancement of Artificial Intelligence. [cited by applicant]
Majumder, N. et al. “Multimodal sentiment analysis using hierarchical fusion with context modeling,” Jul. 30, 2018, 10 pages. [retrieved Feb. 26, 2020]. [cited by applicant]
Majumder, N. et al. “Multimodal Sentiment Analysis using Hierarchical Fusion with Context Modeling,” Jun. 16, 2018, 28 pages. [cited by applicant]
Microsoft “FERPius” [accessed Feb. 25, 2020]. [cited by applicant]
Mike Boers, “PyAV: Pythonic bindings for FFmpeg's libraries,” [accessed Mar. 2, 2020]. [cited by applicant]
Mishra, Taniya et al. “Word Prominence Detection using Robust yet Simple Prosodic Features,” ISCA 13th Annual Conference, Portland OR, Sep. 9-13, 2012, 4 pages. [cited by applicant]
Mozilla, “Common Voice” [accessed Feb. 25, 2020]. [cited by applicant]
Ohio State University, “EmotioNet Challenge 2020 website” [accessed Feb. 25, 2020]. [cited by applicant]
Patwardhan, Amol. “Three-Dimensional, Kinematic, Human Behavioral Pattern-Based Features for Multimodal Emotion Recognition”, Multimodal Technologies and Interact. Sep. 11, 2017; 19 pages. [cited by applicant]