IP Library Granted Patent US 11,133,025
Granted Patent B2
US 11,133,025 · App. 16/677,324 · Granted Sep 28, 2021

Method and system for speech emotion recognition

Inventors: Yatish Jayant Naik Raikar (Bangalore, IN); Varunkumar Tripathi (Karnataka, IN); Kiran Chittella (Bangalore, IN); Vinayak Kulkarni (Nargund, IN)
Assignee: SLING MEDIA PVT LTD
G10L25/63G10L15/02G10L15/22G10L15/26G10L2015/027
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,133,025
App. No.
16/677,324
Granted
Sep 28, 2021
Kind
B2
Abstract

A method for speech emotion recognition for enriching speech to text communications between users in speech chat sessions including: implementing a speech emotion recognition model to enable converting observed emotions in speech samples to enrich text with visual emotion content by: generating a data set of speech samples with labels of a plurality of emotion classes; extracting a set of acoustic features from each of the emotion classes; generating a machine learning (ML) model based on the acoustic features and data set; training the ML model from acoustic features from speech samples during speech chat sessions; predicting emotion content based on a trained ML model in the observed speech; generating enriched text based on predicted emotion content of the trained ML model; and presenting the enriched text in speech to text communications between users in the chat session for visual notice of an observed emotion in the speech sample.

Claims (78)

1. A method for speech emotion recognition for enriching speech to text communications between users in a speech chat session, the method comprising:

implementing a speech emotion recognition model to enable converting observed emotions in a speech sample to enrich text with visual emotion content in speech to text communications by:

generating a data set of speech samples with labels of a plurality of emotion classes; extracting a set of acoustic features from each of a plurality of emotion classes;

generating a machine learning (ML) model based on at least one of the set of acoustic features and data set;

training the ML model from acoustic features from speech samples during speech chat sessions;

predicting emotion content based on a trained ML model in the observed speech;

generating enriched text based on predicted emotion content of the trained ML model; and

presenting the enriched text in speech to text communications between users in the speech chat session for visual notice of an observed emotion in the speech sample.

2. The method of claim 1 , further comprising:

enriching text in the speech to text communications by changing a color of a select set of text of word in a phrase of a converted speech sample in the speech to text communications in the chat session for the visual notice of the observed emotion in the speech sample.

3. The method of claim 1 , further comprising:

enriching text in the speech to text communications by including an emoticon symbol and shading with one or more colors different parts of the emoticon symbol in speech to text communications of a converted speech sample in the speech to text communications in the chat session for the visual notice of the observed emotion in the speech sample by color shading in one or more of the different parts of the emoticon symbol.

4. The method of claim 3 , further comprising:

changing dynamically consistent with a change of the observed emotion in the speech sample, the color shading in the one or more of the different parts of the emoticon symbol wherein the change in color shading occurs gradually.

5. The method of claim 4 , further comprising:

changing dynamically consistent with a dramatic change of the observed emotion in the speech sample, the color shading in all the different parts of the emoticon symbol at once.

6. The method of claim 1 , further comprising:

implementing a speech emotion recognition model to enable converting observed emotions in speech samples to enrich text with visual emotion content in speech to text communications by:

displaying the visual emotion content with text by a color mapping with emoticons;

applying a threshold to compute duration of syllables of a phrase of words in the speech samples wherein each duration corresponds to a perceived excitement in the speech sample;

implementing word stretching on a computed word duration and on a ratio of the computed word duration to syllables of phrases of the word in the speech sample to gauge intensity of words in the speech sample at a word level; and

representing the perceived excitement in the speech samples by word stretching and by highlighting particular words.

7. The method of claim 6 , the highlighting of the particular words further comprising:

bolding the word in a phrase for the prominence for visual notice of the perceived excitement.

8. The method of claim 7 , further comprising:

stretching one or more letters in the word in the phrase for prominence for visual notice of the perceived excitement.

9. The method of claim 8 , further comprising:

computing the duration to syllable ratio based on a first, a second and a third threshold, comprising:

a first syllable count corresponding to a first duration to symbol ratio greater than a first threshold;

a second syllable count corresponding to a second duration to symbol ratio greater than a second threshold; and

a third syllable count corresponding to a third duration to symbol ratio greater than a third threshold.

10. A computer program product tangibly embodied in a computer-readable storage device and comprising instructions that when executed by a processor perform a method for speech emotion recognition for enriching speech to text communications between users in a speech chat session, the method comprising:

implementing a speech emotion recognition model to enable converting observed emotions in a speech sample to enrich text with visual emotion content in speech to text communications by:

generating a data set of speech samples with labels of a plurality of emotion classes; extracting a set of acoustic features from each emotion class;

generating a machine learning (ML) model based on at least one acoustic feature of the set of acoustic features and data set;

training the ML model from acoustic features from speech samples during speech chat sessions;

predicting emotion content based on a trained ML model in the observed speech; generating enriched text based on predicted emotion content of the trained ML model; and

presenting the enriched text in speech to text communications between users in the speech chat session for visual notice of an observed emotion in the speech sample.

11. The method of claim 10 , further comprising:

enriching text in the speech to text communications by changing a color of a select set of text of word in a phrase of a converted speech sample in the speech to text communications in the chat session for the visual notice of the observed emotion in the speech sample.

12. The method of claim 10 , further comprising:

enriching text in the speech to text communications by including an emoticon symbol and shading with one or more colors different parts of the emoticon symbol in speech to text communications of a converted speech sample in the speech to text communications in the chat session for the visual notice of the observed emotion in the speech sample by color shading in one or more of the different parts of the emoticon symbol.

13. The method of claim 12 , further comprising:

changing dynamically consistent with a change of the observed emotion in the speech sample, the color shading in the one or more of the different parts of the emoticon symbol wherein the change in color shading occurs gradually.

14. The method of claim 13 , further comprising:

changing dynamically consistent with a dramatic change of the observed emotion in the speech sample, the color shading in all the different parts of the emoticon symbol at once.

15. The method of claim 10 , further comprising:

implementing a speech emotion recognition model to enable converting observed emotions in speech samples to enrich text with visual emotion content in speech to text communications by:

displaying the visual emotion content with text by a color mapping with emoticons;

applying a threshold to compute duration of syllables of a phrase of words in the speech samples wherein each duration corresponds to a perceived excitement in the speech sample;

implementing word stretching on a computed word duration and on a ratio of the computed word duration to syllables of phrases of the word in the speech sample to gauge intensity of words in the speech sample at a word level; and

representing the perceived excitement in the speech samples by word stretching and by highlighting particular words.

16. The method of claim 15 , the highlighting of the particular words further comprising:

bolding the word in the phrase for the prominence for visual notice of the perceived excitement.

17. The method of claim 16 , further comprising:

stretching one or more letters in the word in the phrase for prominence for visual notice of the perceived excitement.

18. The method of claim 17 , further comprising:

computing the duration to syllable ratio based on a first, a second and a third threshold, comprising:

a first syllable count corresponding to a first duration to symbol ratio greater than a first threshold;

a second syllable count corresponding to a second duration to symbol ratio greater than a second threshold; and

a third syllable count corresponding to a third duration to symbol ratio greater than a third threshold.

19. A system comprising:

at least one processor; and

at least one computer-readable storage device comprising instructions that when executed causes performance of a method for processing speech samples for speech emotion recognition for enriching speech to text communications between users in speech chat, the system comprising:

a speech emotion recognition model implemented by the processor to enable converting observed emotions in a speech sample to enrich text with visual emotion content in speech to text communications, wherein the processor configured to:

generate a data set of speech samples with labels of a plurality of emotion classes; extract a set of acoustic features from each emotion class;

select a machine learning (ML) model based on at least one of the set of acoustic features and data set;

train the ML model from a particular acoustic feature from speech samples during speech chat sessions;

predict emotion content based on a trained ML model in the observed speech;

generate enriched text based on predicted emotion content of the trained ML model; and

present the enriched text in speech to text communications between users in the chat session for visual notice of an observed emotion in the speech sample.

20. The system of claim 19 , further comprising:

the processor further configured to:

implement a speech emotion recognition model to enable converting observed emotions in speech samples to enrich text with visual emotion content in speech to text communications by:

display the visual emotion content with text by a color mapping with emoticons;

apply a threshold to compute duration of syllables of a phrase of words in the speech samples wherein each duration corresponds to a perceived excitement in the speech sample;

implement word stretching on a computed word duration and on a ratio of the computed word duration to syllables of phrases of the word in the speech sample to gauge intensity of words in the speech sample at a word level; and

represent the perceived excitement in the speech samples by word stretching and by highlighting particular words.

Assignments (2)
CHANGE OF NAME Recorded Sep 1, 2022
From: SLING MEDIA PVT. LTD.
To: DISH NETWORK TECHNOLOGIES INDIA PRIVATE LIMITED
Reel/Frame 061365/0493 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 7, 2019
From: RAIKAR, YATISH JAYANT NAIK; TRIPATHI, VARUNKUMAR; CHITTELLA, KIRAN; KULKARNI, VINAYAK
To: SLING MEDIA PVT LTD
Reel/Frame 050951/0875 →
Continuity (1)
Related Publication 20210142820A1 · May 13, 2021
Cited By (1)
US 12,511,485