IP Library › Granted Patent US 12,387,048
Granted Patent B2
US 12,387,048 · App. 17/980,992 · Granted Aug 12, 2025

Apparatuses and methods for text classification

Inventors: Suleiman Ali Khan (Kista, SE); Simone Romano (Kista, SE); Mika Juuti (Kista, SE); Vladimir Poroshin (Kista, SE); Adrian Flanagan (Helsinki, FI); Kuan Eeik Tan (Helsinki, FI)
Assignee: Huawei Technologies Co., Ltd.
G06F40/30G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,387,048
App. No.
17/980,992
Granted
Aug 12, 2025
Kind
B2
Abstract

Apparatuses and methods are provided for classifying textual content using a text classifier for determining to which class the textual content belongs. After classification, the text classifier provides the classification result and a context relevant to the classification result to an explanation system. The explanation system predicts, from the classification result and the context relevant to the classification result, one or more reasons behind the classification result.

Claims (56)

1. A method, comprising:

receiving a text input to be classified;

predicting, using a text classifier, a class of the text input, to obtain a prediction result;

extracting a context relevant to the prediction result by identifying the context relevant to the prediction result by selecting words of the text input that are relevant to the prediction result using an interpretive multi-head attention module configured to perform;

identifying a multi-head attention importance of individual words of the text input as a normalized aggregate over attention head weights from words to the prediction result,

determining an adversarial importance of the individual words, by each respective word of the text input, by removing the respective word in the text input and computing a prediction probability of the text input when the respective word is removed,

applying a gradient-based filter defined as a normalized gradient for the individual words with respect to the prediction result, and

determining a weighted average of the multi-head attention importance, the adversarial importance, and the normalized gradient to generate the extracted context;

determining one or more reasons for the prediction result based on the extracted context, wherein the one or more reasons explain why the text input was classified into the predicted class; and

providing the prediction result and the determined one or more reasons as a classification result.

2. The method according to claim 1 , wherein determining the one or more reasons for the prediction result comprises:

determining, using a machine learning arrangement, the one or more reasons for the prediction result based on the identified context, using a reason classifier and a knowledge base configured to be used to predict a reason for classification.

3. The method according to claim 2 , wherein predicting the reason for classification comprises expanding the identified context using a knowledge base comprising semantical relationships of words.

4. The method according to claim 1 , further comprising computing a value representing a confidence of the prediction result and the determined one or more reasons.

5. The method according to claim 4 , further comprising:

comparing the computed value against a threshold; and

forwarding the text input, the prediction result, and the one or more reasons to a system operator when the computed value is lower than the threshold.

6. The method according to claim 1 , further comprising generating an explanation based on the one or more reasons.

7. The method according to claim 1 , wherein the text classifier is a language-representation based neural network.

8. An apparatus, comprising:

processing circuitry configured to:

receive a text input to be classified;

predict, using a text classifier, a class of the text input, to obtain a prediction result;

extract a context relevant to the prediction result by identifying the context relevant to the prediction result by selecting words of the text input that are relevant to the prediction result using an interpretive multi-head attention module, wherein the interpretive multi-head attention module is configured to:

identify a multi-head attention importance of individual words of the text input as a normalized aggregate over attention head weights from words to the prediction result,

determine an adversarial importance of the individual words, by each respective word of the text input, by removing the respective word in the text input and computing a prediction probability of the text input when the respective word is removed,

apply a gradient-based filter defined as a normalized gradient for the individual words with respect to the prediction result, and

determine a weighted average of the multi-head attention importance, the adversarial importance, and the normalized gradient to generate the extracted context;

determine one or more reasons for the prediction result based on the extracted context, wherein the one or more reasons explain why the text input was classified into the predicted class; and

provide the prediction result and the determined one or more reasons as a classification result.

9. The apparatus according to claim 8 , wherein the processing circuitry is configured to determine the one or more reasons for the prediction result by determining, using a machine learning arrangement, the one or more reasons for the prediction result based on the identified context and using a reason classifier and a knowledge base configured to be used to predict reasons for classification.

10. The apparatus according to claim 8 , wherein the processing circuitry is further configured to:

expand the identified context using a knowledge base comprising semantical relationships of words.

11. The apparatus according to claim 8 , wherein the processing circuitry is further configured to:

compute a value representing a confidence of the prediction result and the determined one or more reasons.

12. The apparatus according to claim 11 , wherein the processing circuitry is further configured to:

compare the computed value against a threshold; and

forward the text input, the prediction result, and the one or more reasons to a system operator when the computed value is lower than the threshold.

13. The apparatus according to claim 8 , wherein the processing circuitry is further configured to generate an explanation based on the one or more reasons.

14. An apparatus, comprising:

at least one processor; and

a non-transitory computer readable storage medium storing a program that is executable by the at least one processor, the program including instructions to:

receive a text input to be classified;

predict, using a text classifier, a class of the text input, to obtain a prediction result;

extract a context relevant to the prediction result by identifying the context relevant to the prediction result by selecting words of the text input that are relevant to the prediction result using an interpretive multi-head attention module, wherein the interpretive multi-head attention module is configured to:

identify a multi-head attention importance of individual words of the text input as a normalized aggregate over attention head weights from words to the prediction result,

determine an adversarial importance of the individual words, by each respective word of the text input, by removing the respective word in the text input and computing a prediction probability of the text input when the respective word is removed,

apply a gradient-based filter defined as a normalized gradient for the individual words with respect to the prediction result, and

determine a weighted average of the multi-head attention importance, the adversarial importance, and the normalized gradient to generate the extracted context;

determine one or more reasons for the prediction result based on the extracted context, wherein the one or more reasons explain why the text input was classified into the predicted class; and

provide the prediction result and the determined one or more reasons as a classification result.

15. The apparatus according to claim 14 , wherein the program includes instructions to determine the one or more reasons for the prediction result by determining, using a machine learning arrangement, the one or more reasons for the prediction result based on the identified context and using a reason classifier and a knowledge base configured to be used to predict reasons for classification.

16. The apparatus according to claim 14 , wherein the program further includes instructions to:

expand the identified context using a knowledge base comprising semantical relationships of words.

17. The apparatus according to claim 14 , wherein the program further includes instructions to:

compute a value representing a confidence of the prediction result and the determined one or more reasons.

Continuity (2)
Continuation PCTEP2020062449 · May 5, 2020
Related Publication 20230080261A1 · Mar 16, 2023
References Cited (26)
US 8868402B2 · Korolev et al. · 2014 [cited by applicant]
US 9836455B2 · Martens · 2017 [cited by examiner]
US 11182545B1 · Tsai · 2021 [cited by examiner]
US 11657222B1 · Anthony · 2023 [cited by examiner]
US 11748613B2 · Li · 2023 [cited by examiner]
US 20130138735A1 · Kanter et al. · 2013 [cited by applicant]
US 20180308487A1 · Goel · 2018 [cited by examiner]
US 20180365248A1 · Zheng · 2018 [cited by examiner]
US 20200042580A1 · Davis · 2020 [cited by examiner]
US 20200111023A1 · Pondicherry Murugappan · 2020 [cited by examiner]
US 20200311205A1 · Büttner · 2020 [cited by examiner]
US 20210201018A1 · Patel · 2021 [cited by examiner]
US 20210209142A1 · Tagra · 2021 [cited by examiner]
US 20230004604A1 · Li · 2023 [cited by examiner]
US 20230069587A1 · Lukyanenko · 2023 [cited by examiner]
US 20230139831A1 · Wang · 2023 [cited by examiner]
CN 108363753A · 2018 [cited by examiner]
CN 110222182A · 2019 [cited by examiner]
EP 3518142A1 · 2019 [cited by applicant]
WO 2017162919A1 · 2017 [cited by applicant]
WO WO2019050968A1 · 2019 [cited by examiner]
WO WO2019175571A1 · 2019 [cited by examiner]
Craig Bloem, “84 Percent of People Trust Online Reviews As Much As Friends. Here's How to Manage What They See,” Jul. 31, 2017, 12 pages, Inc.com. [cited by applicant]
Jacob Devlin et al., “BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding,” May 24, 2019, 16 pages, https://arxiv.org/abs/1810.04805. [cited by applicant]
Casey Newton, “Bodies in Seats,” Jun. 19, 2019, 21 pages, Vox Media, LLC. [cited by applicant]
Haizhou Du et al., “Hierarchical Gated Convolutional Networks with Multi-Head Attention for Text Classification,” The 2018 5th International Conference on Systems and Informatics (ICSAI 2018) Nov. 10, 2018, 6 pages, IEE… [cited by applicant]