IP Library › Granted Patent US 12,646,527
Granted Patent B2
US 12,646,527 · App. 17/831,407 · Granted Jun 2, 2026

System and method for identifying sentiment (emotions) in a speech audio input with haptic output

Inventors: Shannon Modine Brownlee (Carlsbad, CA); Chloe Jordan Duckworth (San Jose, CA)
G10L25/63G10L15/16G10L21/10G10L25/24G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,646,527
App. No.
17/831,407
Granted
Jun 2, 2026
Kind
B2
Abstract

In a system and method for enabling a user to identify the emotions of speakers to a conversation, spoken audio input is pre-processed using a one-dimensional Mel Spectrogram and/or a two-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix, reducing the two-dimensional matrix to a single dimension output, identifying at least one emotion in the audio input using a convolutional or recurrent neural network, and providing the user with haptic feedback corresponding to the at least one emotion in the audio input.

Claims (22)

1 . A system for enabling a user to tactilely feel the emotion in a verbal input, said system comprising:

a verbal input receiving device for receiving a spoken input signal;

a processor and memory configured with machine readable code to define a pre-processing stage and an emotion model in the form of a trained artificial neural network stage for extracting at least one emotion associated with the verbal input, and

one or more haptic feedback devices worn by one or more users, which are configured to emit a haptic feedback signal associated with said at least one emotion,

wherein the pre-processing stage is configured to generate a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken input signal, and to reduce the multi-dimensional matrix to a single dimensional output by taking the mean value for each time band as well as running mean normalization across all the data and feeding this into the artificial neural network stage, and wherein the haptic feedback signal comprises a vibration that is unique for each emotion or combination of emotions extracted by the neural network stage.

2 . A system of claim 1 , wherein the haptic feedback device comprises a wristband or other wearable item in communication with an output from the artificial neural network stage.

3 . A system of claim 1 , wherein the users are speakers taking part in live, face-to-face conversation, and the verbal input receiving device comprises a microphone.

4 . A system of claim 1 , further comprising representing at least one emotion in auditory form to a speaking participant.

5 . A system of claim 1 , wherein the artificial neural network stage is a recurrent neural network (RNN) that includes layers for performing one or more of:

reducing data overfitting, and transforming data into useful numbers.

6 . A system for enabling a user to tactilely feel the emotion in a verbal input, said system comprising:

a verbal input receiving device for receiving a spoken input signal;

a processor and memory configured with machine readable code to define a pre-processing stage, and an emotion model receiving the output of the pre-processing stage as its input, wherein the emotion model comprises a trained neural network stage for extracting at least one emotion associated with the verbal input,

wherein the pre-processing stage is configured to generate a multi-dimensional Mel Spectrogram or Mel-Frequency Cepstral Coefficient (MFCC) matrix from time bands defined in the spoken input signal, and take the mean value for each time band and also run mean normalization across all the data to reduce the multi-dimensional matrix to a single dimensional output; and

a haptic feedback device worn by a user, which is configured to emit a haptic feedback signal associated with said at least one emotion, wherein the haptic feedback signal comprises a vibration that is unique for each emotion or combination of emotions extracted by the neural network stage.

7 . A system of claim 6 wherein the neural network stage comprises a recurrent neural network (RNN) architecture that includes layers for

recurrently recognizing patterns; and layers to transform data into useful numbers.

8 . A method of analyzing spoken audio input obtained from one or more speaking participants for emotional content, comprising

capturing time bands of the spoken audio input,

pre-processing the time bands of the spoken audio input by transforming the time bands into a multi-dimensional Mel-Frequency Cepstral Coefficient (MFCC) matrix or Mel Spectrogram, and reducing the multi-dimensional matrix to a single-dimension output by taking the mean value for each time band as well as running mean normalization across all the data, the method further comprising,

feeding the single dimensional output into a neural network to identify at least one emotion in the audio input, and

converting the at least one emotion into haptic feedback, wherein the haptic feedback comprises a vibration that is unique for each emotion or combination of emotions identified by the neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 2, 2022
From: DUCKWORTH, CHLOE JORDAN; BROWNLEE, SHANNON MODINE
To: VALENCE VIBRATIONS, INC.
Reel/Frame 060091/0690 →
Continuity (2)
Provisional Application 63196264 · Jun 3, 2021
Related Publication 20230298616A1 · Sep 21, 2023
References Cited (14)
US 6691090B1 · Laurila · 2004 [cited by applicant]
US 11495215B1 · Wu · 2022 [cited by examiner]
US 11501794B1 · Kim · 2022 [cited by examiner]
US 11854538B1 · Rozgic · 2023 [cited by examiner]
US 20010056349A1 · St. John · 2001 [cited by applicant]
US 20020194002A1 · Petrushkin · 2002 [cited by applicant]
US 20160217807A1 · Gainsboro · 2016 [cited by applicant]
US 20180322863A1 · Bocklet · 2018 [cited by examiner]
US 20200035222A1 · Sypniewski · 2020 [cited by applicant]
US 20200075039A1 · Eleftheriou · 2020 [cited by examiner]
US 20220084543A1 · Sinha · 2022 [cited by applicant]
US 20220254332A1 · Cartwright · 2022 [cited by examiner]
Fernandes et al., “Speech Emotion Recognition using Mel Frequency Cepstral Coefficient and SVM Classifier,” 2018 International Conference on System Modeling & Advancement in Research Trends (SMART), Moradabad, India, 20… [cited by examiner]
S. K. Pandey, H. S. Shekhawat and S. R. M. Prasanna, “Deep Learning Techniques for Speech Emotion Recognition: A Review,” 2019 29th International Conference Radio elektronika (RADIOELEKTRONIKA), Pardubice, Czech Republi… [cited by examiner]