IP Library › Granted Patent US 11,983,626
Granted Patent B2
US 11,983,626 · App. 17/109,490 · Granted May 14, 2024

Method and apparatus for improving quality of attention-based sequence-to-sequence model

Inventor: Min-Joong Lee (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G06N3/08G06N3/02G06N3/04G06N3/045G10L15/16
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,983,626
App. No.
17/109,490
Granted
May 14, 2024
Kind
B2
Abstract

A method and apparatus for improving the quality of an attention-based sequence-to-sequence model. The method includes determining an output sequence corresponding to an input sequence based on an attention-based sequence-to-sequence model, selecting at least one target attention head from among a plurality of attention heads, detecting at least one error output token among output tokens constituting the output sequence based on the target attention head, and correcting the output sequence based on the error output token.

Claims (55)

1. A method of improving the quality of an attention-based sequence-to-sequence model, the method comprising:

determining an output sequence corresponding to an input sequence based on an attention-based sequence-to-sequence model;

selecting at least one target attention head from among a plurality of attention heads each configured to generate a respective attention weight matrix;

detecting at least one error output token among output tokens constituting the output sequence based on the target attention head; and

correcting the output sequence based on the at least one error output token.

2. The method of claim 1 , wherein the selecting comprises selecting, as the target attention head, an attention head generating a predetermined attention weight matrix trained to be a target attention weight matrix corresponding to the target attention head.

3. The method of claim 2 , wherein the predetermined attention weight matrix is trained based on a guide weight matrix having a predetermined shape.

4. The method of claim 3 , wherein the guide weight matrix is determined based on any one or any combination of an output sequence length, an input frame length, a start shift, an end shift, and a diffusion ratio.

5. The method of claim 2 , wherein the predetermined attention weight matrix is trained to have a different distribution of attention weights for each step.

6. The method of claim 2 , wherein the predetermined attention weight matrix is trained to determine an attention weight of a current step based on a cumulative sum of attention weights of previous steps.

7. The method of claim 1 , wherein the selecting comprises selecting, as the target attention head, an attention head generating an attention weight matrix most suitable for a predetermined purpose.

8. The method of claim 1 , wherein the selecting comprises selecting the target attention head based on a guide weight matrix having a predetermined shape according to a predetermined purpose.

9. The method of claim 1 , wherein the selecting comprises selecting the target attention head by performing monotonic regression analysis on attention weight matrices generated by the plurality of attention heads, in response to the attention-based sequence-to-sequence model having monotonic properties.

10. The method of claim 1 , wherein the selecting comprises selecting the target attention head based on entropy of attention weight matrices generated by the plurality of attention heads.

11. The method of claim 10 , wherein the selecting of the target attention head based on the entropy comprises selecting, as the target attention head, an attention head generating an attention weight matrix having the largest entropy from among the attention weight matrices.

12. The method of claim 10 , wherein the selecting of the target attention head based on the entropy comprises selecting the target attention head based on a Kullback-Leibler divergence.

13. The method of claim 1 , wherein the selecting comprises selecting, as the target attention head, an attention head generating an attention weight matrix having a largest distance between distributions of rows therein.

14. The method of claim 1 , wherein the detecting comprises:

detecting at least one error attention weight in which differences between attention weights of the target attention head and a guide weight matrix are greater than or equal to a threshold value, among attention weights between the input sequence and the output sequence of the target attention head; and

determining an output token corresponding to the at least one error attention weight to be the at least one error output token.

15. The method of claim 1 , wherein the detecting comprises:

detecting at least one error attention weight in which a similarity to an attention weight of a previous step is greater than or equal to a threshold value, among attention weights between the input sequence and the output sequence of the target attention head; and

determining an output token corresponding to the at least one error attention weight to be the at least one error output token.

16. The method of claim 1 , wherein the correcting comprises excluding the at least one error output token from the output sequence.

17. The method of claim 1 , wherein the correcting comprises determining a next input token among other output token candidates other than the at least one error output token.

18. The method of claim 17 , further comprising determining an input token of a step in which the at least one error output token is output, to be the next input token.

19. The method of claim 1 , wherein the number of attention heads corresponds to a product of the number of attention layers and the number of decoder layers in the attention-based sequence-to-sequence model.

20. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

21. An electronic device of an attention-based sequence-to-sequence model, the electronic device comprising:

a processor configured to determine an output sequence corresponding to an input sequence based on an attention-based sequence-to-sequence model, select at least one target attention head from among a plurality of attention heads each configured to generate a respective attention weight matrix, detect at least one error output token among output tokens constituting the output sequence based on the target attention head, and correct the output sequence based on the at least one error output token.

22. The electronic device of claim 21 , wherein the processor is configured to select, as the target attention head, an attention head generating a predetermined attention weight matrix trained to be a target attention weight matrix corresponding to the target attention head.

23. The electronic device of claim 22 , wherein the predetermined attention weight matrix is trained based on a guide weight matrix having a predetermined shape.

24. The electronic device of claim 22 , wherein the predetermined attention weight matrix is trained to have a different distribution of attention weights for each step.

25. The electronic device of claim 21 , wherein the processor is configured to select, as the target attention head, an attention head generating an attention weight matrix most suitable for a predetermined purpose.

26. The electronic device of claim 21 , wherein the processor is configured to select the target attention head based on a guide weight matrix having a predetermined shape according to a predetermined purpose.

27. The electronic device of claim 21 , wherein the processor is configured to select the target attention head by performing monotonic regression analysis on attention weight matrices generated by the plurality of attention heads, in response to the attention-based sequence-to-sequence model having monotonic properties.

28. The electronic device of claim 21 , wherein the processor is configured to select the target attention head based on entropy of attention weight matrices generated by the plurality of attention heads.

29. The electronic device of claim 28 , wherein the processor is configured to select, as the target attention head, an attention head generating an attention weight matrix having the largest entropy from among the attention weight matrices.

30. The electronic device of claim 21 , wherein the processor is configured to select, as the target attention head, an attention head generating an attention weight matrix having a largest distance between distributions of rows therein.

31. The electronic device of claim 21 , wherein the processor is configured to detect at least one error attention weight in which differences between attention weights of the target attention head and a guide weight matrix are greater than or equal to a threshold value, among attention weights between the input sequence and the output sequence of the target attention head, and determine an output token corresponding to the at least one error attention weight to be the at least one error output token.

32. The electronic device of claim 21 , wherein the processor is configured to detect at least one error attention weight in which a similarity to an attention weight of a previous step is greater than or equal to a threshold value, among attention weights between the input sequence and the output sequence of the target attention head, and determine an output token corresponding to the at least one error attention weight to be the at least one error output token.

33. The electronic device of claim 21 , wherein the processor is configured to exclude the at least one error output token from the output sequence.

34. The electronic device of claim 21 , wherein the processor is configured to determine a next input token among other output token candidates other than the at least one error output token.

35. The electronic device of claim 34 , wherein the processor is configured to determine an input token of a step in which the at least one error output token is output, to be the next input token.

36. An electronic device comprising:

an encoder and a decoder configured to input an input sequence and to output an output sequence based on the input sequence; and

one or more processors configured to:

select, from a plurality of attention heads each configured to generate a respective attention weight matrix and included in an attention-based sequence-to-sequence model of the decoder, a target attention head;

detect an error output token included in the output sequence based on the target attention head; and

correct the output sequence based on the error output token and output a corrected output sequence.

37. The electronic device of claim 36 , wherein the encoder and the decoder are included in an artificial neural network.

38. The electronic device of claim 36 , wherein the one or more processors are configured to select the target attention head from among a plurality of attention weight matrices stored in the attention-based sequence-to-sequence model.

39. The electronic device of claim 38 , wherein the attention-based sequence-to-sequence model stores the plurality of attention weight matrices with respect to all of the attention heads by inputting example inputs of correct results into the encoder and the decoder, and

the one or more processors are configured to select, as the target attention head, an attention head that generates an attention weight matrix best suited for an operation of the sequence-to-sequence model.

40. The electronic device of claim 36 , wherein the one or more processors are configured to train a specific attention head, from among the plurality of attention heads, to generate an attention weight matrix of a desired shape and to select the specific attention head as the target attention head.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 2, 2020
From: LEE, MIN-JOONG
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 054516/0757 →
Priority Claims (1)
KR 10-2020-0062450 · May 25, 2020 · national
Continuity (1)
Related Publication 20210366501A1 · Nov 25, 2021