IP Library › Granted Patent US 12,211,486
Granted Patent B2
US 12,211,486 · App. 17/647,499 · Granted Jan 28, 2025

Apparatus and method for compositional spoken language understanding

Inventors: Avik Ray (Sunnyvale, CA); Yilin Shen (Santa Clara, CA); Hongxia Jin (San Jose, CA)
Assignee: Samsung Electronics Co., Ltd.
G10L15/063G10L15/14G10L15/22G10L2015/0633G10L2015/0638G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,211,486
App. No.
17/647,499
Granted
Jan 28, 2025
Kind
B2
Abstract

A method includes identifying multiple tokens contained in an input utterance. The method also includes generating slot labels for at least some of the tokens contained in the input utterance using a trained machine learning model. The method further includes determining at least one action to be performed in response to the input utterance based on at least one of the slot labels. The trained machine learning model is trained to use attention distributions generated such that (i) the attention distributions associated with tokens having dissimilar slot labels are forced to be different and (ii) the attention distribution associated with each token is forced to not focus primarily on that token itself.

Claims (269)

1. A method comprising:

identifying multiple tokens contained in an input utterance;

generating slot labels for at least some of the tokens contained in the input utterance using a trained machine learning model; and

determining at least one action to be performed in response to the input utterance based on at least one of the slot labels;

wherein the trained machine learning model is trained to use attention distributions generated such that (i) the attention distributions associated with tokens having dissimilar slot labels are forced to be different and (ii) the attention distribution associated with each token is forced to not focus primarily on that token itself; and

wherein the trained machine learning model is trained using an overall objective function that includes a slot-pair objective function and a non-degenerate objective function, the slot-pair objective function defining self-attention distributions for tokens, the non-degenerate objective function preventing the self-attention distributions for the tokens from converging to a degenerate solution based on, for each token, an average Kullback-Leibler distance between the token and its corresponding degenerate distribution.

2. The method of claim 1 , wherein the trained machine learning model is further trained by:

obtaining a training dataset comprising training utterances;

identifying different combinations of the training utterances, each combination having two or more training utterances with a common intent and disjoint sets of slot types;

concatenating the training utterances in each combination to generate at least one paired training sample for that combination;

adding the paired training samples to the training dataset in order to produce an augmented training dataset; and

training the machine learning model using the augmented training dataset.

3. The method of claim 1 , wherein:

the slot-pair objective function is based on a first divergence between different ones of the self-attention distributions for different tokens; and

the machine learning model is trained to increase the first divergence.

4. The method of claim 3 , wherein:

the non-degenerate objective function is based on a second divergence between the self-attention distribution for each token and its corresponding degenerate distribution; and

the machine learning model is trained to increase the second divergence.

5. The method of claim 4 , wherein the first divergence and the second divergence comprise Kullback-Leibler (KL) divergences.

6. The method of claim 1 , wherein the trained machine learning model is trained to reduce correlations between slot labels.

7. An electronic device comprising:

at least one processing device configured to:

identify multiple tokens contained in an input utterance;

generate slot labels for at least some of the tokens contained in the input utterance using a trained machine learning model; and

determine at least one action to be performed in response to the input utterance based on at least one of the slot labels;

wherein the trained machine learning model is trained to use attention distributions generated such that (i) the attention distributions associated with tokens having dissimilar slot labels are forced to be different and (ii) the attention distribution associated with each token is forced to not focus primarily on that token itself; and

wherein the trained machine learning model is trained using an overall objective function that includes a slot-pair objective function and a non-degenerate objective function, the slot-pair objective function defining self-attention distributions for tokens, the non-degenerate objective function preventing the self-attention distributions for the tokens from converging to a degenerate solution based on, for each token, an average Kullback-Leibler distance between the token and its corresponding degenerate distribution.

8. The electronic device of claim 7 , wherein the trained machine learning model is further trained by:

obtaining a training dataset comprising training utterances;

identifying different combinations of the training utterances, each combination having two or more training utterances with a common intent and disjoint sets of slot types;

concatenating the training utterances in each combination to generate at least one paired training sample for that combination;

adding the paired training samples to the training dataset in order to produce an augmented training dataset; and

training the machine learning model using the augmented training dataset.

9. The electronic device of claim 8 , wherein:

the slot-pair objective function is based on a first divergence between different ones of the self-attention distributions for different tokens; and

the machine learning model is trained to increase the first divergence.

10. The electronic device of claim 9 , wherein:

the non-degenerate objective function is based on a second divergence between the self-attention distribution for each token and its corresponding degenerate distribution; and

the machine learning model is trained to increase the second divergence.

11. The electronic device of claim 10 , wherein the first divergence and the second divergence comprise Kullback-Leibler (KL) divergences.

12. The electronic device of claim 7 , wherein the trained machine learning model is trained to reduce correlations between slot labels.

13. A non-transitory machine-readable medium containing instructions that when executed cause at least one processor of an electronic device to:

identify multiple tokens contained in an input utterance;

generate slot labels for at least some of the tokens contained in the input utterance using a trained machine learning model; and

determine at least one action to be performed in response to the input utterance based on at least one of the slot labels;

wherein the trained machine learning model is trained to use attention distributions generated such that (i) the attention distributions associated with tokens having dissimilar slot labels are forced to be different and (ii) the attention distribution associated with each token is forced to not focus primarily on that token itself; and

wherein the trained machine learning model is trained using an overall objective function that includes a slot-pair objective function and a non-degenerate objective function, the slot-pair objective function defining self-attention distributions for tokens, the non-degenerate objective function preventing the self-attention distributions for the tokens from converging to a degenerate solution based on, for each token, an average Kullback-Leibler distance between the token and its corresponding degenerate distribution.

14. The non-transitory machine-readable medium of claim 13 , wherein the trained machine learning model is further trained by:

obtaining a training dataset comprising training utterances;

identifying different combinations of the training utterances, each combination having two or more training utterances with a common intent and disjoint sets of slot types;

concatenating the training utterances in each combination to generate at least one paired training sample for that combination;

adding the paired training samples to the training dataset in order to produce an augmented training dataset; and

training the machine learning model using the augmented training dataset.

15. The non-transitory machine-readable medium of claim 13 , wherein:

the slot-pair objective function is based on a first divergence between different ones of the self-attention distributions for different tokens; and

the machine learning model is trained to increase the first divergence.

16. The non-transitory machine-readable medium of claim 15 , wherein:

the non-degenerate objective function is based on a second divergence between the self-attention distribution for each token and its corresponding degenerate distribution; and

the machine learning model is trained to increase the second divergence.

17. The non-transitory machine-readable medium of claim 13 , wherein the trained machine learning model is trained to reduce correlations between slot labels.

18. A method comprising:

obtaining a training dataset comprising training utterances;

identifying different combinations of the training utterances, each combination having two or more training utterances with a common intent and disjoint sets of slot types;

concatenating the training utterances in each combination to generate at least one paired training sample for that combination;

adding the paired training samples to the training dataset in order to produce an augmented training dataset; and

training a machine learning model using the augmented training dataset;

wherein training the machine learning model comprises using an overall objective function that includes a slot-pair objective function and a non-degenerate objective function, the slot-pair objective function defining self-attention distributions for tokens, the non-degenerate objective function preventing the self-attention distributions for the tokens from converging to a degenerate solution based on, for each token, an average Kullback-Leibler distance between the token and its corresponding degenerate distribution.

19. The method of claim 18 , wherein:

the slot-pair objective function is based on a first divergence between different ones of the self-attention distributions for different tokens; and

training the machine learning model comprises increasing the first divergence.

20. The method of claim 19 , wherein:

the non-degenerate objective function is based on a second divergence between the self-attention distribution for each token and its corresponding degenerate distribution; and

training the machine learning model comprises increasing the second divergence.

21. The method of claim 20 , wherein the first divergence and the second divergence comprise Kullback-Leibler (KL) divergences.

22. The method of claim 18 , further comprising:

dividing a dataset into the training dataset and a testing dataset, the testing dataset comprising testing utterances used to test the trained machine learning model.

23. The method of claim 22 , wherein dividing the dataset into the training dataset and the testing dataset comprises:

forming an initial training dataset and an initial testing dataset;

forming the training dataset used to train the machine learning model by removing any training utterances from the initial training dataset having a slot combination used in the initial testing dataset; and

forming the testing dataset used to test the machine learning model by:

removing any testing utterances from the initial testing dataset having an intent or a slot label not used in the training dataset; and

replacing any out-of-vocabulary slot values in the initial testing dataset with associated slot values used in the training dataset.

24. The method of claim 22 , wherein dividing the dataset into the training dataset and the testing dataset comprises:

forming an initial training dataset and an initial testing dataset;

forming the training dataset used to train the machine learning model by removing any training utterances from the initial training dataset having a number of slots above a threshold value; and

forming the testing dataset used to test the machine learning model by:

removing any testing utterances from the initial testing dataset having a slot combination used in the training dataset;

removing any testing utterances from the initial testing dataset having an intent or a slot label not used in the training dataset; and

replacing any out-of-vocabulary slot values in the initial testing dataset with associated slot values used in the training dataset.

25. The method of claim 1 , wherein the non-degenerate objective function is defined as:

ℒ

n

⁢

o

⁢

n

-

d

⁢

e

⁢

g

=

1

N

2

⁢

∑

h

⁢

∑

i

:

y

i

≠

0

⁢

K

⁢

L

⁡

(

P

i

h

,

1

i

)

where:

1 i represents a degenerate distribution with a value of “1” for an i th token and a value of “0” elsewhere;

N 2 represents a normalizing constant; and

non-deg represents a loss defined by the non-degenerate objective function.

26. The electronic device of claim 7 , wherein the non-degenerate objective function is defined as:

ℒ

n

⁢

o

⁢

n

-

d

⁢

e

⁢

g

=

1

N

2

⁢

∑

h

⁢

∑

i

:

y

i

≠

0

⁢

K

⁢

L

⁡

(

P

i

h

,

1

i

)

where:

1 i represents a degenerate distribution with a value of “1” for an i th token and a value of “0” elsewhere;

N 2 represents a normalizing constant; and

non-deg represents a loss defined by the non-degenerate objective function.

27. The non-transitory machine-readable medium of claim 13 , wherein the non-degenerate objective function is defined as:

ℒ

n

⁢

o

⁢

n

-

d

⁢

e

⁢

g

=

1

N

2

⁢

∑

h

⁢

∑

i

:

y

i

≠

0

⁢

K

⁢

L

⁡

(

P

i

h

,

1

i

)

where:

1 i represents a degenerate distribution with a value of “1” for an i th token and a value of “0” elsewhere;

N 2 represents a normalizing constant; and

non-deg represents a loss defined by the non-degenerate objective function.

28. The method of claim 18 , wherein the non-degenerate objective function is defined as:

ℒ

n

⁢

o

⁢

n

-

d

⁢

e

⁢

g

=

1

N

2

⁢

∑

h

⁢

∑

i

:

y

i

≠

0

⁢

K

⁢

L

⁡

(

P

i

h

,

1

i

)

where:

1 i represents a degenerate distribution with a value of “1” for an i th token and a value of “0” elsewhere;

N 2 represents a normalizing constant; and

non-deg represents a loss defined by the non-degenerate objective function.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 10, 2022
From: RAY, AVIK; SHEN, YILIN; JIN, HONGXIA
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 058595/0023 →
Continuity (2)
Provisional Application 63190695 · May 19, 2021
Related Publication 20220375457A1 · Nov 24, 2022
References Cited (11)
US 9978374B2 · Heigold et al. · 2018 [cited by applicant]
US 10515625B1 · Metallinou et al. · 2019 [cited by applicant]
US 10600406B1 · Shapiro et al. · 2020 [cited by applicant]
US 11562735B1 · Gupta · 2023 [cited by examiner]
US 20190259380A1 · Biyani et al. · 2019 [cited by applicant]
US 20190371307A1 · Zhao · 2019 [cited by examiner]
US 20200410986A1 · Shen et al. · 2020 [cited by applicant]
US 20220084510A1 · Peng · 2022 [cited by examiner]
Jixuan Wang, Kai Wei, Martin Radfar, Weiwei Zhang, Clement Chung, “Encoding Syntactic Knowledge in Transformer Encoder for Intent Detection and Slot Filling”, arXiv″2012.11689, Dec. 21, 2020 (Year: 2020). [cited by examiner]
Chen et al., “Compositional Generalization via Neural-Symbolic Stack Machines”, 34th Conference on Neural Information Processing Systems, 2020, 12 pages. [cited by applicant]
Hudson et al., “Learning by Abstraction: The Neural State Machine”, 33rd Conference on Neural Information Processing Systems, 2019, 14 pages. [cited by applicant]