IP Library › Granted Patent US 12,119,022
Granted Patent B2
US 12,119,022 · App. 17/536,673 · Granted Oct 15, 2024

Cognitive assistant for real-time emotion detection from human speech

Inventors: Rishi Amit Sinha (San Jose, CA); Ria Sinha (San Jose, CA)
G10L25/63G10L21/0208G10L25/30H04L67/55
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,119,022
App. No.
17/536,673
Granted
Oct 15, 2024
Kind
B2
Abstract

Systems and methods used in a cognitive assistant for detecting human emotions from speech audio signals is described. The system obtains audio signals from an audio receiver and extracts human speech samples. Subsequently, it runs a machine learning based classifier to analyze the human speech signal and classify the emotion observed in it. The user is then notified, based on their preferences, with a summary of the emotion detected. Notifications can also be sent to other systems that have been configured to receive them. Optionally, the system may include the ability to store the speech sample and emotion classification detected for future analysis. The system's machine learning classifier is periodically re-trained based on labelled audio speech data and updated.

Claims (41)

1. A system comprising:

an audio receiver;

a processing system connected to the audio receiver;

a notification system connected to the processing system,

wherein the processing system is configured to

i) obtain audio signal from the audio receiver;

ii) process the audio signal to detect if human speech is present;

iii) responsive to the audio signal containing human speech, run a machine learning based classifier to analyze the speech audio signal and output an emotion detected in speech, wherein the detected emotion is one of calm, happy, sad, angry, fearful, surprise, and disgust;

v) send the emotion detected to a notification system;

vi) loop back to i),

wherein the machine learning classifier is periodically trained externally based on labelled audio sample data and updated in the system, and wherein the processing system is further configured to

receive feedback from a user that the detected emotion was incorrect or unknown, and

process the feedback for the labelled audio sample data.

2. The system of claim 1 , wherein the processing system has a filter and an amplifier to output an improved copy of the received audio signal or store it digitally.

3. The system of claim 1 , wherein the notification system is a mobile device push notification configured by the user.

4. The system of claim 1 , wherein the notification system responds to an API request from an external system.

5. The system of claim 1 , wherein notification preferences are configured by the user.

6. The system of claim 1 , where the system is running as an application on a mobile device, wherein the audio receiver is a microphone on the mobile device, the processing system is a processor on the mobile device and the notification system is a screen and vibration alerts.

7. The system of claim 1 , wherein the audio receiver is a separate device communicatively coupled to the processing system running on a computer.

8. A method comprising:

i) obtaining audio signal from an audio receiver;

ii) processing the audio signal to detect if human speech is present;

iii) responsive to the audio signal containing human speech, running a machine learning based classifier to analyze the speech audio signal and output an emotion detected in speech, wherein the detected emotion is one of calm, happy, sad, angry, fearful, surprise, and disgust;

v) sending the emotion detected to a notification system;

vi) looping back to i),

wherein the machine learning classifier is periodically trained externally based on labelled audio sample data and updated in the system, and wherein the method further comprises

receiving feedback from a user that the detected emotion was incorrect or unknown, and

processing the feedback for the labelled audio sample data.

9. The method of claim 8 , further comprising of a filter and an amplifier to output an improved copy of the received audio signal or store it digitally.

10. The method of claim 8 , wherein the notification method is a mobile device push notification configured by the user.

11. The method of claim 8 , wherein the notification method responds to an API request from an external system.

12. The method of claim 8 , wherein the notification method can be configured by the user.

13. A non-transitory computer-readable medium comprising instructions that, when executed, cause a processing system to perform steps of:

i) obtaining audio signal from an audio receiver;

ii) processing the audio signal to detect if human speech is present;

iii) responsive to the audio signal containing human speech, running a machine learning based classifier to analyze the speech audio signal and output an emotion detected in speech, wherein the detected emotion is one of calm, happy, sad, angry, fearful, surprise, and disgust;

v) sending the emotion detected to a notification system;

vi) looping back to i),

wherein the machine learning classifier is periodically trained externally based on labelled audio sample data and updated in the system, and wherein the steps further comprise

receiving feedback from a user that the detected emotion was incorrect or unknown, and

processing the feedback for the labelled audio sample data.

Continuity (2)
Continuation In Part 16747697 · Jan 21, 2020
Related Publication 20220084543A1 · Mar 17, 2022
Cited By (1)
US 12,657,455