IP Library › Granted Patent US 12,626,693
Granted Patent B2
US 12,626,693 · App. 18/077,826 · Granted May 12, 2026

Modeling attention to improve classification and provide inherent explainability

Inventors: Dalkandura Arachchige K.S.S. Gunaratna (San Jose, CA); Vijay Srinivasan (San Jose, CA); Hongxia Jin (San Jose, CA)
Assignee: SAMSUNG ELECTRONICS CO., LTD.
G10L15/1822G10L15/16G10L15/1815G10L15/22G06F40/30G10L15/18G10L15/183G10L2015/221G10L2015/223G10L2015/225
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,626,693
App. No.
18/077,826
Granted
May 12, 2026
Kind
B2
Abstract

In an artificial intelligence model (AI model), input data is processed to provide both classification of the input data and a visualization of the process of the AI model. This is done by performing intent and slot classification of the input data, generating weights and binary classifier logits, performing feature fusion and classification. A graphical explanation is then output as a visualization along with logits.

Claims (63)

1 . A method of visualizing a natural language understanding model, the method comprising:

parsing an utterance into a vector of tokens;

encoding the utterance with an encoder to obtain a vector of token embeddings;

applying an intent classifier, based on the vector of token embeddings, to obtain an estimated intent;

obtaining a vector of slot type weights for visualization, wherein the obtaining the vector of slot type weights uses an auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein obtaining the vector of slot type weights comprises, for each slot type, applying a binary classifier to the vector of multiple self-attentions to obtain the vector of slot type weights;

obtaining a vector of multiple self-attentions, wherein the obtaining the vector of multiple self-attentions uses the auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein the estimated intent includes a vector of intent logits, and wherein the obtaining the vector of multiple self-attentions comprises:

concatenating the vector of intent logits with the vector of token embeddings to obtain an expanded intent logits vector, and

obtaining the vector of multiple self-attentions based on the expanded intent logits vector;

visualizing the vector of slot type weights in a two column format, wherein the two column format comprises a first column and a second column;

performing a feature fusion based on the vector of slot type weights and based on the vector of token embeddings to obtain a vector of fused features; and

obtaining, based on the vector of fused features and using a slot classifier, a vector of classified slots corresponding to the utterance.

2 . The method of claim 1 , further comprising outputting the vector of classified slots for fulfillment by a voice-activated artificial intelligence-based personal assistant.

3 . The method of claim 1 , wherein the visualizing the vector of slot type weights comprises providing a visual presentation including the first column and the second column with bars connecting column entries from the first column to the second column, the first column and the second column both listing the vector of tokens, wherein a first bar corresponds to a correspondence between a first token in the first column with a second token in the second column.

4 . The method of claim 1 , wherein a training of the vector of slot type weights is based on an output of the binary classifier.

5 . The method of claim 4 , wherein the performing the feature fusion comprises:

computing a cross-attention vector based on the vector of slot type weights and based on the vector of token embeddings;

forming an intermediate vector as a sum of the cross-attention vector and the vector of token embeddings; and

forming the vector of fused features based on applying the intermediate vector to a linear layer and normalizing an output of the linear layer.

6 . The method of claim 1 , further comprising:

determining a vector of slot type attentions;

wherein the applying the slot classifier comprises operating on the vector of slot type specific attentions;

wherein the performing the feature fusion is further based on the vector of slot type specific attentions; and

wherein the method further comprises visualizing the vector of slot type specific attentions.

7 . A server for utterance recognition and model visualization, the server comprising:

one or more processors; and

one or more memories, the one or more memories storing a program, wherein execution of the program by the one or more processors is configured to cause the server to at least:

parse an utterance into a vector of tokens;

encode the utterance with an encoder to obtain a vector of token embeddings;

apply an intent classifier, based on the vector of token embeddings, to obtain an estimated intent;

obtain a vector of slot type weights for visualization, wherein the obtaining the vector of slot type weights uses an auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein execution of the program by the one or more processors is further configured to obtain the vector of slot type weights by, for each slot type, applying a binary classifier to the vector of multiple self-attentions to obtain the vector of slot type weights;

obtain a vector of multiple self-attentions, wherein the obtaining the vector of multiple self-attentions uses the auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein the estimated intent includes a vector of intent logits, and wherein execution of the program by the one or more processors is further configured to obtain the vector of multiple self-attentions by:

concatenating the vector of intent logits with the vector of token embeddings to obtain an expanded intent logits vector, and

obtaining the vector of multiple self-attentions based on the expanded intent logits vector;

visualize the vector of slot type weights in a two column format, wherein the two column format comprises a first column and a second column;

perform a feature fusion based on the vector of slot type weights and based on the vector of token embeddings to obtain a vector of fused features; and

obtain, based on the vector of fused features and using a slot classifier, a vector of classified slots corresponding to the utterance.

8 . The server of claim 7 , wherein execution of the program by the one or more processors is further configured to cause the server to output the vector of classified slots for fulfillment by a voice-activated artificial intelligence-based personal assistant.

9 . The server of claim 7 , wherein execution of the program by the one or more processors is further configured to provide information for a debugging engineer to alter the intent classifier and/or the slot classifier and/or model training data based on the vector of slot type weights visualized in the two column format.

10 . The server of claim 7 , wherein execution of the program by the one or more processors is further configured to visualize the vector of slot type weights by providing a visual presentation including the first column and the second column with bars connecting column entries from the first column to the second column, the first column and the second column both listing the vector of tokens, wherein a first bar corresponds to a correspondence between a first token in the first column with a second token in the second column, thereby permitting a person to recognize focus points on the utterance relevant to the classification by a natural language understanding model.

11 . The server of claim 7 , wherein a training of the vector of slot type weights is based on an output of the binary classifier.

12 . The server of claim 11 , wherein execution of the program by the one or more processors is further configured to perform the feature fusion by:

computing a cross-attention vector based on the vector of slot type weights and based on the vector of token embeddings;

forming an intermediate vector as a sum of the cross-attention vector and the vector of token embeddings; and

forming the vector of fused features based on a applying the intermediate vector to a linear layer and normalizing an output of the linear layer.

13 . The server of claim 7 , wherein execution of the program by the one or more processors is further configured to:

determine a vector of special slot type specific attentions;

wherein execution of the program by the one or more processors is further configured to apply the slot classifier by operating on the vector of special slot type specific attentions;

wherein execution of the program by the one or more processors is further configured to perform the feature fusion based on the vector of special slot type specific attentions; and

wherein execution of the program by the one or more processors is further configured to visualize the vector of special slot type specific attentions.

14 . A non-transitory computer readable medium configured to store a program for utterance recognition and model visualization, wherein execution of the program by one or more processors of a server is configured to cause the server to at least:

parse an utterance into a vector of tokens;

encode the utterance with an encoder to obtain a vector of token embeddings;

apply an intent classifier, based on the vector of token embeddings, to obtain an estimated intent;

obtain a vector of slot type weights for visualization, wherein the obtaining the vector of slot type weights uses an auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein execution of the program by the one or more processors is further configured to obtain the vector of slot type weights by, for each slot type, applying a binary classifier to the vector of multiple self-attentions to obtain the vector of slot type weights;

obtain a vector of multiple self-attentions, wherein the obtaining the vector of multiple self-attentions uses the auxiliary network and is based on the vector of token embeddings and based on the intent logits, wherein the estimated intent includes a vector of intent logits, and wherein execution of the program by the one or more processors is further configured to obtain the vector of multiple self-attentions by:

concatenating the vector of intent logits with the vector of token embeddings to obtain an expanded intent logits vector, and

obtaining the vector of multiple self-attentions based on the expanded intent logits vector;

visualize the vector of slot type weights in a two column format, wherein the two column format comprises a first column and a second column;

perform a feature fusion based on the vector of slot type weights and based on the vector of token embeddings to obtain a vector of fused features; and

obtain, based on the vector of fused features and using a slot classifier, a vector of classified slots corresponding to the utterance.

15 . The non-transitory computer readable medium of claim 14 , wherein execution of the program by the one or more processors of the server is configured to cause the server to output the vector of classified slots for fulfillment by a voice-activated artificial intelligence-based personal assistant.

16 . The non-transitory computer readable medium of claim 14 , wherein execution of the program by the one or more processors of the server is configured to provide information for a debugging engineer to alter the intent classifier and/or the slot classifier based on the vector of slot type weights visualized in the two column format.

17 . The non-transitory computer readable medium of claim 14 , wherein execution of the program by the one or more processors of the server is configured to visualize the vector of slot type weights by providing a visual presentation including the first column and the second column with bars connecting column entries from the first column to the second column, the first column and the second column both listing the vector of tokens, wherein a first bar corresponds to a correspondence between a first token in the first column with a second token in the second column, thereby permitting a person to recognize focus points on the utterance relevant to the classification by a natural language understanding model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 8, 2022
From: GUNARATNA, DALKANDURA ARACHCHIGE K.S.S.; SRINIVASAN, VIJAY; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 062031/0016 →
Continuity (2)
Provisional Application 63307592 · Feb 7, 2022
Related Publication 20230252982A1 · Aug 10, 2023
References Cited (26)
US 11281857B1 · Craft et al. · 2022 [cited by applicant]
US 11593560B2 · Zhou · 2023 [cited by examiner]
US 11593631B2 · Dalli · 2023 [cited by examiner]
US 20200043480A1 · Shen et al. · 2020 [cited by applicant]
US 20210183484A1 · Shaib · 2021 [cited by examiner]
US 20210217408A1 · Hakkani-Tur · 2021 [cited by examiner]
US 20220108688A1 · Wang · 2022 [cited by examiner]
US 20220171798A1 · Wang · 2022 [cited by examiner]
US 20220198254A1 · Dalli · 2022 [cited by examiner]
US 20220245326A1 · Stabler · 2022 [cited by examiner]
US 20230081598A1 · Rizk · 2023 [cited by examiner]
US 20230394317A1 · Zhai · 2023 [cited by examiner]
US 20240005654A1 · Ray · 2024 [cited by examiner]
CN 111309915A · 2020 [cited by applicant]
CN 112632961A · 2021 [cited by applicant]
EP 3522037A1 · 2019 [cited by examiner]
KR 102281581B1 · 2021 [cited by applicant]
Chen et al., “BERT for Joint Intent Classification and Slot Filling”, arXiv:1902.10909v1, Feb. 28, 2019, 7 total pages. [cited by applicant]
International Search Report and Written Opinion (PCT/ISA/210 and PCT/ISA/237) dated May 8, 2023, issued by International Searching Authority for International Application No. PCT/KR2023/001659. [cited by applicant]
Libo Qin et al., “GL-GIN: Fast and Accurate Non-Autoregressive Model for Joint Multiple Intent Detection and Slot Filling”, 11 pages, Jun. 3, 2021, arXiv:2106.01925v1 [cs.CL], XP081983541. [cited by applicant]
Momchil Hardalov et al., “Enriched Pre-trained Transformers for Joint Slot Filling and Intent Detection”, 12 pages, Oct. 5, 2021, arXiv:2004.14848v2 [cs.CL], XP091064531. [cited by applicant]
Sarah Wiegreffe et al., “Attention is not not Explanation”, 12 pages, Sep. 5, 2019, arXiv:1908.04626v2 [cs.CL], XP093215957. [cited by applicant]
Kalpa Gunaratna et al., “Explainable Slot Type Attentions to Improve Joint Intent Detection and Slot Filling”, 13 pages, Oct. 19, 2022, arXiv:2210.10227v1 [cs.LG], XP091347979. [cited by applicant]
Communication issued on Nov. 7, 2024 from the European Patent Office for European Patent Application No. 23750000.4. [cited by applicant]
A. Kumar et al., “MA-DST: Multi-Attention-Based Scalable Dialog State Tracking”, arXiv:2002.08898v1 [cs.CL], Association for the Advancement of Artificial Intelligence, www.aaai.org, Feb. 7, 2020, (9 total pages). [cited by applicant]
I. Shilin et al., “A Method for Dataset Creation for Dialogue State Classification in Voice Control Systems for the Internet of Things”, ResearchGate, https://www.researchgate.net/publication/328513309, Oct. 2018, (9 pa… [cited by applicant]