IP Library › Granted Patent US 12,597,423
Granted Patent B2
US 12,597,423 · App. 17/850,454 · Granted Apr 7, 2026

Analysis of conversational attributes with real time feedback

Inventors: Dana Suskind (Chicago, IL); Arnoldo Muller-Molina (Naperville, IL); Snigdha Gupta (Mcdonald, PA); John List (Chicago, IL)
Assignee: The University of Chicago
G10L15/22G10L15/02G10L15/16G10L25/24G10L25/45G10L2015/225G10L17/02G10L17/26
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,597,423
App. No.
17/850,454
Granted
Apr 7, 2026
Kind
B2
Abstract

Devices and methods for providing real-time feedback of conversational attributes are provided. An audio signal comprising speech is received. The audio signal is divided into a plurality of sequential windows. A plurality of features is extracted from each sequential window of the audio signal. Each plurality of features is sequentially provided to a trained classifier and a speech attribute of the corresponding window of the audio signal is received therefrom. After receiving each speech attribute, and based upon that speech attribute and speech attributes of prior windows, a conversational attribute is generated. A user-perceivable output indicative of the conversational attribute is provided.

Claims (41)

1 . A method of providing real-time feedback of conversational attributes, the method comprising:

receiving an audio signal comprising speech;

dividing the audio signal into a plurality of sequential windows;

extracting a plurality of features from each sequential window of the audio signal;

sequentially providing each plurality of features to a trained artificial neural network, in particular a convolutional neural network, and receiving therefrom a speech attribute of the corresponding window of the audio signal, the speech attribute comprising a speaker type;

after receiving each speech attribute, and based upon that speech attribute and speech attributes of prior windows, generating a conversational attribute, wherein the conversational attribute comprises a conversational turn count (CTC) based on a change in speech attribute among that speech attribute and speech attributes of prior windows; and

providing a user-perceivable output indicative of the conversational attribute.

2 . The method of claim 1 , wherein each sequential window is about one second in length.

3 . The method of claim 1 , wherein each plurality of features comprise Mel-Frequency Cepstral Coefficients (MFCCs).

4 . The method of claim 1 , further comprising:

determining one or more environmental attributes during each of the plurality of sequential windows;

providing the one or more environmental attributes to the trained classifier with the plurality of features of the corresponding window.

5 . The method of claim 4 , wherein the environmental attributes comprises one or more of: temperature, positions, audio volume, vibration, and light level.

6 . The method of claim 1 , wherein each plurality of features comprise Mel-Frequency Cepstral Coefficients (MFCCs) and wherein providing each of the plurality of features to the trained artificial neural network comprises generating an image of the MFCCs.

7 . The method of claim 1 , wherein the speaker type is selected from: female, child, male, and non-human.

8 . The method of claim 1 , wherein the speech attribute comprises tone.

9 . The method of claim 1 , wherein the speech attribute comprises parentese.

10 . The method of claim 1 , wherein the user-perceivable output comprises a light, sounds, or vibration.

11 . The method of claim 1 , wherein providing the user-perceivable output comprises illuminating one of a plurality of lights according to whether the CTC exceeds a predetermined threshold.

12 . The method of claim 1 , wherein generating the conversational attribute comprises:

identifying child speech in one of the plurality of sequential windows; and

searching a predetermined number of prior windows for adult speech.

13 . The method of claim 1 , wherein the convolutional neural network includes exactly one convolutional layer.

14 . The method of claim 1 , wherein the convolutional neural network is quantized.

15 . A device for providing real-time feedback of conversational attributes, the device comprising:

a computing node configured to perform a method comprising:

receiving an audio signal comprising speech;

dividing the audio signal into a plurality of sequential windows;

extracting a plurality of features from each sequential window of the audio signal;

sequentially providing each plurality of features to a trained artificial neural network, in particular a convolutional neural network and receiving therefrom a speech attribute of the corresponding window of the audio signal, the speech attribute comprising a speaker type;

after receiving each speech attribute, and based upon that speech attribute and speech attributes of prior windows, generating a conversational attribute, wherein the conversational attribute comprises a conversational turn count (CTC) based on a change in speech attribute among that speech attribute and speech attributes of prior windows; and

providing a user-perceivable output indicative of the conversational attribute; and

a microphone configured to produce the audio signal;

a light, a speaker, or a vibration motor configured to produce the user-perceivable output.

16 . A computer program product comprising a computer readable storage medium having program instructions embodied therewith, the program instructions executable by a processor to cause the processor to perform a method comprising:

receiving an audio signal comprising speech;

dividing the audio signal into a plurality of sequential windows;

extracting a plurality of features from each sequential window of the audio signal;

sequentially providing each plurality of features to a trained artificial neural network in particular a convolutional neural network, and receiving therefrom a speech attribute of the corresponding window of the audio signal, the speech attribute comprising a speaker type;

after receiving each speech attribute, and based upon that speech attribute and speech attributes of prior windows, generating a conversational attribute, wherein the conversational attribute comprises a conversational turn count (CTC) based on a change in speech attribute among that speech attribute and speech attributes of prior windows; and

providing a user-perceivable output indicative of the conversational attribute.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 11, 2026
From: SUSKIND, DANA; MULLER-MOLINA, ARNOLDO; GUPTA, SNIGDHA; LIST, JOHN
To: THE UNIVERSITY OF CHICAGO
Reel/Frame 074042/0744 →
Continuity (1)
Related Publication 20230419961A1 · Dec 28, 2023
References Cited (15)
US 9548046B1 · Boggiano et al. · 2017 [cited by applicant]
US 9799348B2 · Paul et al. · 2017 [cited by applicant]
US 10134424B2 · Lacson et al. · 2018 [cited by applicant]
US 10789939B2 · Lacson et al. · 2020 [cited by applicant]
US 10959648B2 · Lacson et al. · 2021 [cited by applicant]
US 20030236663A1 · Dimitrova · 2003 [cited by examiner]
US 20150006168A1 · Mysore et al. · 2015 [cited by applicant]
US 20160042734A1 · Cetinturk · 2016 [cited by applicant]
US 20160203832A1 · Paul · 2016 [cited by examiner]
US 20170061978A1 · Wang et al. · 2017 [cited by applicant]
US 20180301158A1 · Zou et al. · 2018 [cited by applicant]
US 20190267027A1 · Boggiano et al. · 2019 [cited by applicant]
US 20210254800A1 · Bertken · 2021 [cited by examiner]
KR 102225308B1 · 2021 [cited by examiner]
International Search Report and Written Opinion for International Application No. PCT/US23/26278 dated Sep. 29, 2023. [cited by applicant]