IP Library Granted Patent US 12,526,167
Granted Patent B2
US 12,526,167 · App. 18/392,849 · Granted Jan 13, 2026

Real-time tone feedback in video conferencing

Inventors: Vipul Raheja (San Francisco, CA); Dimitrios Alikaniotis (New York, NY)
H04L12/1831G06V10/764G06V40/176
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,526,167
App. No.
18/392,849
Granted
Jan 13, 2026
Kind
B2
Abstract

A computer-implemented process is programmed to programmatically receive, using a first computer system, electronic digital data representing input time-correlated speech data and video data, determine a first text sequence corresponding to the input time-correlated speech data, the first text sequence comprising unstructured natural language text, determining syntactic structure data associated with the first text sequence, inputting the time-correlated video data and the syntactic structure data associated with the first text sequence into one or more machine learning models, the machine learning models producing an output of one or more scores for at least a portion of the time-correlated video data and first text sequence, transforming the output of one or more scores to yield and output set of summary points and suggestions, and transmitting a graphical element of the output set of summary points and suggestions for display.

Claims (54)

1 . A computer-implemented method comprising:

using a computer system and under stored program control, receiving electronic digital data representing input time-correlated speech data and video data;

by the computer system, determining a first text sequence corresponding to the input time-correlated speech data;

by the computer system, determining a syntactic structure data associated with the first text sequence;

by the computer system, inputting the time-correlated video data and the syntactic structure data associated with the first text sequence into one or more machine-learning models, the machine-learning models having been trained to produce, and producing, an output set of summary points and suggestions; and

by the computer system, transmitting a graphical element of the output set of summary points and suggestions to a computing device, such that rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

2 . The computer-implemented method of claim 1 , further comprising:

by the computer system, presenting a plurality of functions of the output set of summary points and suggestions via a graphical user interface and receiving input via the graphical user interface specifying a plurality of selections of the plurality of functions;

by the computer system, updating the one or more machine-learning models using the plurality of selections, the first text sequence, and the syntactic structure data associated with the first text sequence;

by the computer system, applying the one or more machine-learning models to update the output set of summary points and suggestions; and

by the computer system, transmitting a graphical element of the updated output set of summary points and suggestions to a computing device, wherein rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

3 . The computer-implemented method of claim 1 , further comprising:

by the computer system, inputting the time-correlated video data and the syntactic structure data into one or more machine-learning models, the machine-learning models having been trained to determine a first expression and a second expression; and

by the computer system, transmitting a graphical element of the first and second expressions to a computing device, wherein rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

4 . The computer-implemented method of claim 1 , the machine-learning models comprising any one or more of expression determination systems and personality impression systems.

5 . The computer-implemented method of claim 4 , the one or more expression determination systems comprising a video-driven expression system to receive the time-correlated video data, the time-correlated video data having a plurality of frames that depict facial expressions from a video from whom the input time-correlated speech data and video data was obtained.

6 . The computer-implemented method of claim 4 , the one or more personality impression systems comprising a video-driven impression system to receive the time-correlated video data and an audio driven impression system to receive the time-correlated audio data, the time-correlated video data having a plurality of frames that depict facial expressions from a video from whom the input time-correlated speech data and video data was obtained.

7 . The computer-implemented method of claim 1 , the output set of summary points and suggestions comprising one or more of a classification of tone, speech, personality, and expression.

8 . The computer-implemented method of claim 1 , further comprising using a digital lexicon to associate the syntactic structure data for the first text sequence with a tone label.

9 . The computer-implemented method of claim 1 , further comprising, before the transmitting, ranking the output set of summary points and suggestions based on a ranking criterion.

10 . One or more non-transitory computer-readable media storing one or more sequences of instructions, execution of which causes a computer system to perform:

determining a first text sequence corresponding to the input time-correlated speech data;

determining a syntactic structure data associated with the first text sequence;

inputting the time-correlated video data and the syntactic structure data associated with the first text sequence into one or more machine-learning models, the machine-learning models having been trained to produce, and producing, an output set of summary points and suggestions; and

transmitting a graphical element of the output set of summary points and suggestions to a computing device, such that rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

11 . The one or more non-transitory computer-readable media of claim 10 , execution of the instructions further causing the computer system to perform:

presenting a plurality of functions of the output set of summary points and suggestions via a graphical user interface and receiving input via the graphical user interface specifying a plurality of selections of the plurality of functions;

updating the one or more machine-learning models using the plurality of selections, the first text sequence, and the syntactic structure data associated with the first text sequence;

applying the one or more machine-learning models to update the output set of summary points and suggestions; and

transmitting a graphical element of the updated output set of summary points and suggestions to a computing device, wherein rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

12 . The one or more non-transitory computer-readable media of claim 10 , execution of the instructions further causing the computer system to perform:

inputting the time-correlated video data and the syntactic structure data into one or more machine-learning models, the machine-learning models having been trained to determine a first expression and a second expression; and

transmitting a graphical element of the first and second expressions to a computing device, wherein rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

13 . The one or more non-transitory computer-readable media of claim 10 , the machine-learning models comprising any one or more of expression determination systems and personality impression systems.

14 . The one or more non-transitory computer-readable media of claim 13 , the one or more expression determination systems comprising a video-driven expression system to receive the time-correlated video data, the time-correlated video data having a plurality of frames that depict facial expressions from a video from whom the input time-correlated speech data and video data was obtained.

15 . The one or more non-transitory computer-readable media of claim 13 , the one or more personality impression systems comprising a video-driven impression system to receive the time-correlated video data and an audio driven impression system to receive the time-correlated audio data, the time-correlated video data having a plurality of frames that depict facial expressions from a video from whom the input time-correlated speech data and video data was obtained.

16 . The one or more non-transitory computer-readable media of claim 10 , the output set of summary points and suggestions comprising one or more of a classification of tone, speech, personality, and expression.

17 . The one or more non-transitory computer-readable media of claim 10 , execution of the instructions further causing the computer system to perform:

using a digital lexicon to associate the syntactic structure data for the first text sequence with a tone label.

18 . The one or more non-transitory computer-readable media of claim 10 , execution of the instructions further causing the computer system to perform:

before the transmitting, ranking the output set of summary points and suggestions based on a ranking criterion.

19 . A computer system comprising:

at least one processor;

at least one communication interface coupled to the at least one processor; and

at least one non-transitory program storage device coupled to the at least one processor and storing instructions, execution of which by the at least one processor causes the computer system to perform operations comprising:

determining a first text sequence corresponding to the input time-correlated speech data;

determining a syntactic structure data associated with the first text sequence;

inputting the time-correlated video data and the syntactic structure data associated with the first text sequence into one or more machine-learning models, the machine-learning models having been trained to produce, and producing, an output set of summary points and suggestions; and

transmitting a graphical element of the output set of summary points and suggestions to a computing device, such that rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

20 . The computer system of claim 19 , execution of the instructions further causing the computer system to perform:

presenting a plurality of functions of the output set of summary points and suggestions via a graphical user interface and receiving input via the graphical user interface specifying a plurality of selections of the plurality of functions;

updating the one or more machine-learning models using the plurality of selections, the first text sequence, and the syntactic structure data associated with the first text sequence;

applying the one or more machine-learning models to update the output set of summary points and suggestions; and

transmitting a graphical element of the updated output set of summary points and suggestions to a computing device, wherein rendering the graphical element using presentation functions of the computing device causes displaying the graphical element at the computing device.

Assignments (1)
CHANGE OF NAME Recorded Nov 21, 2025
From: GRAMMARLY, INC.
To: SUPERHUMAN PLATFORM INC.
Reel/Frame 073655/0136 →
Continuity (3)
Continuation 18180584 · Mar 8, 2023
Provisional Application 63321295 · Mar 18, 2022
Related Publication 20240205039A1 · Jun 20, 2024
References Cited (15)
US 7949517B2 · Eckert · 2011 [cited by examiner]
US 10014004B2 · Khaleghi · 2018 [cited by examiner]
US 11361151B1 · Guberman et al. · 2022 [cited by applicant]
US 20160005050A1 · Teman · 2016 [cited by applicant]
US 20200065612A1 · Xu · 2020 [cited by applicant]
US 20200175961A1 · Thomson · 2020 [cited by applicant]
US 20200366959A1 · Pau · 2020 [cited by applicant]
US 20210352380A1 · Duncan · 2021 [cited by applicant]
US 20210407520A1 · Neckermann · 2021 [cited by applicant]
US 20220093101A1 · Krishnan · 2022 [cited by applicant]
US 20230123574A1 · Guberman et al. · 2023 [cited by applicant]
AU 2023200677B2 · 2023 [cited by applicant]
AU 2023202100A1 · 2023 [cited by applicant]
AU 2023202256A1 · 2023 [cited by applicant]
JP 2023058747A · 2023 [cited by applicant]