IP Library › Granted Patent US 12,725,630
Granted Patent B2
US 12,725,630 · App. 18/242,859 · Granted Sep 1, 2026

Method, system and computer-readable storage medium for cross-task unseen emotion class recognition

Inventors: Jeng-Lin Li (Hsinchu City, TW); Chi-Chun Lee (Hsinchu City, TW)
Assignee: NATIONAL TSING HUA UNIVERSITY
G10L25/63G10L15/02G10L15/063G10L15/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,725,630
App. No.
18/242,859
Granted
Sep 1, 2026
Kind
B2
Abstract

A method for unseen emotion class recognition comprises: receiving, with an emotion recognition model, a speech sample to be tested; calculating, with an encoder, a sample embedding to be tested of the speech sample to be tested; calculating a first distance metric between the sample embedding to be tested and a first registered emotion category representation, and a second distance metric between the sample embedding to be tested and a second registered emotion category representation, wherein the second registered emotion category is not included in a plurality of basic emotion categories; and determining an emotion category of the speech sample to be tested according to the first distance metric and the second distance metric.

Claims (70)

1 . A method for unseen emotion class recognition with a speech recognition device, comprising:

receiving, by an emotion recognition model, a speech sample to be tested;

calculating, by an encoder, a sample embedding to be tested of the speech sample to be tested;

calculating a first distance metric between the sample embedding to be tested and a first registered emotion category representation, and a second distance metric between the sample embedding to be tested and a second registered emotion category representation, wherein the second registered emotion category is not included in a plurality of basic emotion categories;

comparing the first distance metric and the second distance metric; and

determining an emotion category of the speech sample to be tested based on the comparison result,

wherein the method further comprises a training the emotion recognition model by means of:

receiving, by the emotion recognition model, a first training speech sample, wherein the first training speech sample indicates a first emotion category in the basic emotion categories, and the basic emotion categories include respectively a plurality of training speech samples;

calculating, by the encoder, a first embedding of the first training speech sample;

calculating respectively, by the encoder, a plurality of sample embeddings of the training speech samples of each of the basic emotion categories;

calculating a center-of-mass representation of each of the basic emotion categories, wherein the center-of-mass representations are respectively averages of the sample embeddings of different basic emotion categories;

calculating respectively a cosine similarity of the first embedding with the center-of-mass representation of each of the basic emotion categories, wherein an angular margin is added to the cosine similarity when calculating the cosine similarity with the first emotion category, and the angular margin is negative;

calculating a loss according to a loss function; and

adjusting a plurality of parameters of the emotion recognition model based on the calculated loss, wherein the loss function is associated with the cosine similarity.

2 . The method of claim 1 , wherein the encoder comprises an acoustic feature generator to fetch an acoustic feature from the speech sample to be tested.

3 . The method of claim 2 , wherein the encoder comprises a gated recurrent unit (GRU) model and a Transformer model to transform the acoustic feature into the sample embedding to be tested.

4 . The method of claim 1 , further comprising a register procedure which comprises:

receiving, with the emotion recognition model, a plurality of first registered speech samples and a plurality of second registered speech samples, wherein the first registered speech samples indicate the first registered emotion category and the second registered speech samples indicate the second registered emotion category;

calculating, with the encoder, a plurality of first registered sample embeddings of the first registered speech samples and a plurality of second registered sample embeddings of the second registered speech samples; and

calculating respectively averages of the first registered sample embeddings and the second registered sample embeddings, to generate respectively the first registered emotion category representation and the second registered emotion category representation.

5 . The method of claim 1 , wherein the loss function includes a cross-entropy loss.

6 . The method of claim 1 , wherein determining the emotion category of the speech sample to be tested comprises: determining the emotion category as the first registered emotion category according to the first distance metric being smaller than the second distance metric, or determining the emotion category as the second registered emotion category according to the second distance metric being smaller than the first distance metric.

7 . A speech recognition device for unseen emotion class recognition, comprising:

a memory having stored thereon a plurality of instructions; and

a processor coupled to the memory, wherein the processor is configured to, when executing the instructions:

receive, by an emotion recognition model, a speech sample to be tested;

calculate, by an encoder, a sample embedding to be tested of the speech sample to be tested;

calculate a first distance metric between the sample embedding to be tested and a first registered emotion category representation, and a second distance metric between the sample embedding to be tested and a second registered emotion category representation, wherein the second registered emotion category is not included in a plurality of basic emotion categories;

comparing the first distance metric and the second distance metric; and

determine an emotion category of the speech sample to be tested based on the comparison result,

wherein the processor is further configured to:

receive, by the emotion recognition model, a first training speech sample, wherein the first training speech sample indicates a first emotion category in the basic emotion categories, and the basic emotion categories include respectively a plurality of training speech samples;

calculate, by the encoder, a first embedding of the first training speech sample;

calculate respectively, by the encoder, a plurality of sample embeddings of the training speech samples of each of the basic emotion categories;

calculate a center-of-mass representation of each of the basic emotion categories, wherein the center-of-mass representations are respectively averages of the sample embeddings of different basic emotion categories;

calculate respectively a cosine similarity of the first embedding with the center-of-mass representation of each of the basic emotion categories, wherein an angular margin is added to the cosine similarity when calculating the cosine similarity with the first emotion category, and the angular margin is negative;

calculate a loss according to a loss function; and

adjust a plurality of parameters of the emotion recognition model based on the calculated loss, wherein the loss function is associated with the cosine similarity.

8 . The speech recognition device of claim 7 , wherein the encoder comprises an acoustic feature generator to fetch an acoustic feature from the speech sample to be tested.

9 . The speech recognition device of claim 8 , wherein the encoder comprises a gated recurrent unit (GRU) model and a Transformer model to transform the acoustic feature into the sample embedding to be tested.

10 . The speech recognition device of claim 7 , wherein the processor is further configured to:

receive, with the emotion recognition model, a plurality of first registered speech samples and a plurality of second registered speech samples, wherein the first registered speech samples indicate the first registered emotion category and the second registered speech samples indicate the second registered emotion category;

calculate, with the encoder, a plurality of first registered sample embeddings of the first registered speech samples and a plurality of second registered sample embeddings of the second registered speech samples; and

calculate respectively averages of the first registered sample embeddings and the second registered sample embeddings, to generate respectively the first registered emotion category representation and the second registered emotion category representation.

11 . The speech recognition device of claim 7 ,

wherein the loss function includes a cross-entropy loss.

12 . The speech recognition device of claim 7 , wherein the processor is configured to: determine the emotion category of the speech sample to be tested as the first registered emotion category according to the first distance metric being smaller than the second distance metric, or determine the emotion category as the second registered emotion category according to the second distance metric being smaller than the first distance metric.

13 . A computer-readable non-transitory storage medium for unseen emotion class recognition loaded with a computer-readable program capable of, after being read by a speech recognition device, configure the speech recognition device to:

receive, by an emotion recognition model, a speech sample to be tested;

calculate, by an encoder, a sample embedding to be tested of the speech sample to be tested;

calculate a first distance metric between the sample embedding to be tested and a first registered emotion category representation, and a second distance metric between the sample embedding to be tested and a second registered emotion category representation, wherein the second registered emotion category is not included in a plurality of basic emotion categories;

comparing the first distance metric and the second distance metric; and

determine an emotion category of the speech sample to be tested based on the comparison result;

wherein the computer-readable program is capable of, after being read by the speech recognition device, further configure the speech recognition device to:

receive, by the emotion recognition model, a first training speech sample, wherein the first training speech sample indicates a first emotion category in the basic emotion categories, and the basic emotion categories include respectively a plurality of training speech samples;

calculate, by the encoder, a first embedding of the first training speech sample;

calculate respectively, by the encoder, a plurality of sample embeddings of the training speech samples of each of the basic emotion categories;

calculate a center-of-mass representation of each of the basic emotion categories, wherein the center-of-mass representations are respectively averages of the sample embeddings of different basic emotion categories;

calculate respectively a cosine similarity of the first embedding with the center-of-mass representation of each of the basic emotion categories, wherein an angular margin is added to the cosine similarity when calculating the cosine similarity with the first emotion category, and the angular margin is negative;

calculating a loss according to a loss function; and

adjusting a plurality of parameters of the emotion recognition model based on the calculated loss, wherein the loss function is associated with the cosine similarity.

14 . The computer-readable non-transitory storage medium of claim 13 , wherein the encoder comprises an acoustic feature generator to fetch an acoustic feature from the speech sample to be tested.

15 . The computer-readable non-transitory storage medium of claim 14 , wherein the encoder comprises a gated recurrent unit (GRU) model and a Transformer model to transform the acoustic feature into the sample embedding to be tested.

16 . The computer-readable non-transitory storage medium of claim 13 , wherein the computer-readable program is capable of, after being read by the speech recognition device, further configure the speech recognition device to:

receive, with the emotion recognition model, a plurality of first registered speech samples and a plurality of second registered speech samples, wherein the first registered speech samples indicate the first registered emotion category and the second registered speech samples indicate the second registered emotion category;

calculate, with the encoder, a plurality of first registered sample embeddings of the first registered speech samples and a plurality of second registered sample embeddings of the second registered speech samples; and

calculate respectively averages of the first registered sample embeddings and the second registered sample embeddings, to generate respectively the first registered emotion category representation and the second registered emotion category representation.

17 . The computer-readable non-transitory storage medium of claim 13 ,

wherein the loss function includes a cross-entropy loss.

18 . The computer-readable non-transitory storage medium of claim 13 , wherein the computer-readable program is capable of, after being read by the speech recognition device, configure the speech recognition device to: determine the emotion category of the speech sample to be tested as the first registered emotion category according to the first distance metric being smaller than the second distance metric, or determine the emotion category as the second registered emotion category according to the second distance metric being smaller than the first distance metric.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2023
From: LI, JENG-LIN; LEE, CHI-CHUN
To: NATIONAL TSING HUA UNIVERSITY
Reel/Frame 064815/0548 →
Priority Claims (1)
TW 112119474 · May 25, 2023 · national
Continuity (1)
Related Publication 20240395280A1 · Nov 28, 2024
References Cited (82)
US 10276189B1 · Brochu · 2019 [cited by examiner]
US 10293260B1 · Evans · 2019 [cited by examiner]
US 10300394B1 · Evans · 2019 [cited by examiner]
US 10846601B1 · Howard · 2020 [cited by examiner]
US 10891629B1 · Barakat · 2021 [cited by examiner]
US 11064072B1 · Sawala · 2021 [cited by examiner]
US 11201966B1 · Restorff · 2021 [cited by examiner]
US 11461952B1 · Bosnak · 2022 [cited by examiner]
US 11704501B2 · Wu · 2023 [cited by examiner]
US 11842729B1 · Richter · 2023 [cited by examiner]
US 12205614B1 · Sharma · 2025 [cited by examiner]
US 20060281064A1 · Sato · 2006 [cited by examiner]
US 20070208569A1 · Subramanian · 2007 [cited by examiner]
US 20080040110A1 · Pereg · 2008 [cited by examiner]
US 20090234888A1 · Holmes · 2009 [cited by examiner]
US 20120105610A1 · Ahn · 2012 [cited by examiner]
US 20160171100A1 · Fujita · 2016 [cited by examiner]
US 20170257285A1 · Scholz · 2017 [cited by examiner]
US 20180032612A1 · Kariman · 2018 [cited by examiner]
US 20190253558A1 · Haukioja · 2019 [cited by examiner]
US 20190294868A1 · Martinez · 2019 [cited by examiner]
US 20200036810A1 · Howard · 2020 [cited by examiner]
US 20200210475A1 · Aguirre-Suarez · 2020 [cited by examiner]
US 20200250278A1 · Liu · 2020 [cited by examiner]
US 20200257975A1 · Chang · 2020 [cited by examiner]
US 20200394213A1 · Li · 2020 [cited by examiner]
US 20210258424A1 · Brown · 2021 [cited by examiner]
US 20210319780A1 · Aher · 2021 [cited by examiner]
US 20210398562A1 · Verbeke · 2021 [cited by examiner]
US 20220022790A1 · Kim · 2022 [cited by examiner]
US 20220032919A1 · Lee · 2022 [cited by examiner]
US 20220076693A1 · Bui et al. · 2022 [cited by applicant]
US 20220086393A1 · Peters · 2022 [cited by examiner]
US 20220101873A1 · Burmistrov · 2022 [cited by examiner]
US 20220132218A1 · Aher · 2022 [cited by examiner]
US 20220172711A1 · Hvelplund · 2022 [cited by examiner]
US 20220187847A1 · Cella et al. · 2022 [cited by applicant]
US 20220208180A1 · Eyben · 2022 [cited by examiner]
US 20220230623A1 · Byun · 2022 [cited by examiner]
US 20220292261A1 · Movshovitz-Attias et al. · 2022 [cited by applicant]
US 20220358935A1 · Hvelplund · 2022 [cited by examiner]
US 20220366197A1 · Mazza · 2022 [cited by examiner]
US 20230004738A1 · Wu · 2023 [cited by examiner]
US 20230007359A1 · Aher · 2023 [cited by examiner]
US 20230114150A1 · Lillelund · 2023 [cited by examiner]
US 20230289672A1 · Stumpf · 2023 [cited by examiner]
US 20230382407A1 · Donderici · 2023 [cited by examiner]
US 20230385441A1 · Donderici · 2023 [cited by examiner]
US 20230386138A1 · Donderici · 2023 [cited by examiner]
US 20230393553A1 · Donderici · 2023 [cited by examiner]
US 20230415751A1 · Donderici · 2023 [cited by examiner]
US 20230419950A1 · Khare · 2023 [cited by examiner]
US 20240010224A1 · Johanna · 2024 [cited by examiner]
US 20240011788A1 · Valle · 2024 [cited by examiner]
US 20240015248A1 · Johanna · 2024 [cited by examiner]
US 20240019864A1 · Elshenawy · 2024 [cited by examiner]
US 20240025451A1 · Grace · 2024 [cited by examiner]
US 20240034280A1 · Herse · 2024 [cited by examiner]
US 20240037612A1 · Rizk · 2024 [cited by examiner]
US 20240042953A1 · Chen · 2024 [cited by examiner]
US 20240078374A1 · Gokhale · 2024 [cited by examiner]
US 20240126981A1 · Shahinian · 2024 [cited by examiner]
US 20240246571A1 · Koduvayur · 2024 [cited by examiner]
US 20240253664A1 · Zhang · 2024 [cited by examiner]
US 20240296044A1 · Day · 2024 [cited by examiner]
US 20240317259A1 · Johanna · 2024 [cited by examiner]
US 20240326848A1 · Wachsman · 2024 [cited by examiner]
US 20240347037A1 · Lee · 2024 [cited by examiner]
US 20240359705A1 · Zhang · 2024 [cited by examiner]
US 20240391487A1 · Zheng · 2024 [cited by examiner]
US 20240394614A1 · McKnew · 2024 [cited by examiner]
US 20240420504A1 · Mendlovic · 2024 [cited by examiner]
CN 106782615A · 2017 [cited by applicant]
CN 108780653A · 2018 [cited by applicant]
CN 113535957A · 2021 [cited by applicant]
CN 113706541A · 2021 [cited by applicant]
CN 114298019A · 2022 [cited by applicant]
CN 114416991A · 2022 [cited by applicant]
CN 114420169A · 2022 [cited by applicant]
TW 324097B · 1998 [cited by applicant]
TW 200506657A · 2005 [cited by applicant]
WO WO2017075279A1 · 2017 [cited by applicant]