IP Library › Granted Patent US 12,548,556
Granted Patent B2
US 12,548,556 · App. 18/380,847 · Granted Feb 10, 2026

Method and system for fair speech emotion recognition

Inventors: Woan-Shiuan Chien (Hsinchu, TW); Chi-Chun Lee (Hsinchu, TW)
Assignee: National Tsing Hua University
G10L15/063G10L25/30G10L25/63
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,548,556
App. No.
18/380,847
Granted
Feb 10, 2026
Kind
B2
Abstract

A computer-implemented method for speech emotion recognition is provided. The computer-implemented method includes an emotion prediction corresponding to speech data that is generated based on the speech data and a speech emotion recognition model without bias. The speech emotion recognition model without bias is trained by training a fairness-constrained adversarial network based on a labeled training set with known bias and a loss function. The fairness-constrained adversarial network includes a domain classifier for bias classification and the speech emotion recognition model. The loss function used for training the fairness-constrained adversarial network is positively related to a Wasserstein distance (WD) loss.

Claims (47)

1 . A computer-implemented method for training a speech emotion recognition model, the method comprising:

providing a fairness-constrained adversarial network comprising a domain classifier and a speech emotion recognition model, the domain classifier configured to classify a bias in an input data, and the speech emotion recognition model configured to determine an emotion corresponding to the input data; and

training the fairness-constrained adversarial network, based on a labeled training set with known bias and a first loss function, to obtain the speech emotion recognition model without bias,

wherein;

the first loss function is positively related to a Wasserstein distance loss,

the labeled training set with known bias comprises a plurality of training data, and

each of the plurality of training data corresponds to an overall emotion label and a plurality of perspective emotion labels for a plurality of perspectives.

2 . The computer-implemented method of claim 1 , wherein the first loss function is negatively related to a first loss associated with the domain classifier and positively related to a second loss associated with the speech emotion recognition model.

3 . The computer-implemented method of claim 1 , wherein the domain classifier is configured to classify a binary bias.

4 . The computer-implemented method of claim 3 , wherein the binary bias comprises an evaluator gender bias.

5 . The computer-implemented method of claim 1 , further comprising:

generating a fairness embedding set corresponding to the labeled training set based on the speech emotion recognition model without bias; and

training a multi-perspective speech emotion recognition model based on the fairness embedding set and a second loss function,

wherein the multi-perspective speech emotion recognition model is configured to generate a plurality of perspective-specific emotion prediction results corresponding to the plurality of perspectives, and the second loss function is positively related to a metric learning loss.

6 . The computer-implemented method of claim 5 , wherein;

the metric learning loss comprises a triplet loss, and

an anchor, a positive sample, and a negative sample of the triplet loss are set based on the overall emotion label and the plurality of perspective emotion labels.

7 . The computer-implemented method of claim 5 , wherein the metric learning loss comprises at least one of a triplet loss, a quadruplet loss, an N-pair loss, a contrastive loss, or a center loss.

8 . The computer-implemented method of claim 5 , wherein the second loss function is positively related to a third loss associated with the multi-perspective speech emotion recognition model, and the third loss is associated with the plurality of perspective emotion labels.

9 . A computer-implemented method for speech emotion recognition, the method comprising:

receiving speech data; and

generating an emotion prediction result corresponding to the speech data based on the speech data and a speech emotion recognition model without bias, wherein:

the speech emotion recognition model without bias is trained by training a fairness-constrained adversarial network based on a labeled training set with known bias and a first loss function,

the first loss function is positively related to a Wasserstein distance loss,

the fairness-constrained adversarial network comprises a domain classifier for bias classification and the speech emotion recognition model, and

the labeled training set with known bias comprises a plurality of training data, and each of the plurality of training data corresponds to an overall emotion label and a plurality of perspective emotion labels for a plurality of perspectives.

10 . The computer-implemented method of claim 9 , wherein the first loss function is negatively related to a first loss associated with the domain classifier and positively related to a second loss associated with the speech emotion recognition model.

11 . The computer-implemented method of claim 9 , wherein the domain classifier is configured to classify a binary bias.

12 . The computer-implemented method of claim 11 , wherein the binary bias comprises an evaluator gender bias.

13 . The computer-implemented method of claim 9 , further comprising:

generating a fairness embedding corresponding to the speech data based on the speech data and the speech emotion recognition model without bias; and

generating at least one of a plurality of perspective-specific emotion prediction results corresponding to the speech data based on the fairness embedding and a multi-perspective speech emotion recognition model, wherein:

the multi-perspective speech emotion recognition model is trained based on a fairness embedding set corresponding to the labeled training set and a second loss function,

the fairness embedding set is generated based on the speech emotion recognition model without bias, and

the second loss function is positively related to a metric learning loss.

14 . The computer-implemented method of claim 13 , wherein;

the metric learning loss comprises a triplet loss, and

an anchor, a positive sample, and a negative sample of the triplet loss are set based on the overall emotion label and the plurality of perspective emotion labels.

15 . The computer-implemented method of claim 13 , wherein the metric learning loss comprises at least one of a triplet loss, a quadruplet loss, an N-pair loss, a contrastive loss, or a center loss.

16 . The computer-implemented method of claim 13 , wherein the second loss function is positively related to a third loss associated with the multi-perspective speech emotion recognition model, and the third loss is associated with the plurality of perspective emotion labels.

17 . A non-transitory computer-readable medium, comprising at least one instruction, when executed by a processor of an electronic device, causes the electronic device to:

receive speech data; and

generate an emotion prediction result corresponding to the speech data based on the speech data and a speech emotion recognition model without bias, wherein:

the speech emotion recognition model without bias is trained by training a fairness-constrained adversarial network based on a labeled training set with known bias and a first loss function,

the first loss function is positively related to a Wasserstein distance loss,

the fairness-constrained adversarial network comprises a domain classifier for bias classification and the speech emotion recognition model, and

the labeled training set with known bias comprises a plurality of training data, and each of the plurality of training data corresponds to an overall emotion label and a plurality of perspective emotion labels for a plurality of perspectives.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 17, 2023
From: CHIEN, WOAN-SHIUAN; LEE, CHI-CHUN
To: NATIONAL TSING HUA UNIVERSITY
Reel/Frame 065250/0638 →
Priority Claims (1)
TW 112120746 · Jun 2, 2023 · national
Continuity (1)
Related Publication 20240404511A1 · Dec 5, 2024
References Cited (17)
US 11227624B2 · Deshpande · 2022 [cited by examiner]
US 12229962B1 · Aravamudan · 2025 [cited by examiner]
US 12250201B1 · Rosenoer · 2025 [cited by examiner]
US 20200286506A1 · Deshpande · 2020 [cited by examiner]
US 20200335086A1 · Paraskevopoulos · 2020 [cited by examiner]
US 20220328065A1 · Li · 2022 [cited by examiner]
US 20220351068A1 · Mishraky et al. · 2022 [cited by applicant]
US 20220374637A1 · Liu · 2022 [cited by examiner]
US 20230069908A1 · Ando · 2023 [cited by examiner]
US 20240404511A1 · Chien · 2024 [cited by examiner]
US 20250022314A1 · Wei · 2025 [cited by examiner]
TW 202141343A · 2021 [cited by applicant]
TW 202213193A · 2022 [cited by applicant]
TW 202309876A · 2023 [cited by applicant]
WO WO2024186954A2 · 2024 [cited by examiner]
M. Abdelwahab and C. Busso, “Domain Adversarial for Acoustic Emotion Recognition,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, No. 12, pp. 2423-2435, Dec. 2018, doi: 10.1109/TASLP.2018.2… [cited by examiner]
M. Abdelwahab and C. Busso, “Domain Adversarial for Acoustic Emotion Recognition,” in IEEE/ACM Transactions on Audio, Speech, and Language Processing, vol. 26, No. 12, pp. 2423-2435, Dec. 2018, doi: 10.1109/TASLP.2018.2… [cited by examiner]