IP Library › Granted Patent US 12,062,365
Granted Patent B2
US 12,062,365 · App. 17/514,338 · Granted Aug 13, 2024

Apparatus and method for training dialogue summary model

Inventors: Hyun Jae Lee (Seoul, KR); Hyun Jin Choi (Seoul, KR); Jae Woong Yun (Seoul, KR); Ju Dong Kim (Seoul, KR); Bong Kyu Hwang (Seoul, KR); Seong Ho Joe (Seoul, KR); Young June Gwon (Seoul, KR)
Assignee: SAMSUNG SDS CO., LTD.
G10L15/18G10L15/063G10L15/22G10L2015/223
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,062,365
App. No.
17/514,338
Granted
Aug 13, 2024
Kind
B2
Abstract

An apparatus for training a dialogue summary model according to an embodiment includes a parameter transferer configured to transfer one or more learning parameter values of a pre-trained natural language processing model to a sequence-to-sequence-based dialogue summary model, and a model trainer configured to train the dialogue summary model by using the transferred learning parameter values as initial values for learning parameters of each of an encoder and a decoder in the dialogue summary model.

Claims (49)

1. An apparatus for training a dialogue summary model, the apparatus comprising at least one processor to implement:

a parameter transferer configured to transfer one or more learning parameter values of a natural language processing model, which is pre-trained, to a sequence-to-sequence-based dialogue summary model; and

a model trainer configured to train the sequence-to-sequence-based dialogue summary model by using the transferred learning parameter values as initial values for learning parameters of each of an encoder and a decoder in the sequence-to-sequence-based dialogue summary model,

wherein the natural language processing model is trained by selecting, as samples, dialogue text-summary text pairs of a preset batch size, and by self-supervised learning based on a masking speaker task using the selected samples, and

wherein the masking speaker task is performed by:

comparing a probability value of a variable α, which is randomly determined for each sample, with a first preset threshold, the variable α being a rational number equal to or greater than 0 and equal to or less than 1; and

for each sample, determining an execution target of the masking speaker task between a dialogue text and a summary text of a dialogue text-summary text pair, based on a result of the comparing.

2. The apparatus of claim 1 , wherein each of the encoder and the decoder includes at least a portion of a network structure of the natural language processing model.

3. The apparatus of claim 1 , wherein the model trainer is further configured to update the learning parameters of the sequence-to-sequence-based dialogue summary model based on a loss between an inferred summary text output by the sequence-to-sequence-based dialogue summary model receiving an input dialogue text and a correct answer summary text corresponding to the input dialogue text.

4. The apparatus of claim 1 , wherein the model trainer is further configured to share at least some of learning parameter values of the encoder and at least some of learning parameter values of the decoder.

5. The apparatus of claim 1 , wherein the masking speaker task is performed by:

comparing a probability value of a variable β, which is randomly determined for each utterance of a sample with a second preset threshold, the variable β being a rational number equal to or greater than 0 and equal to or less than 1; and

determining whether to perform the masking speaker task for a corresponding utterance based on a result of the comparing the probability value of the variable β with the second preset threshold.

6. The apparatus of claim 1 , wherein a begin of sentence (BOS) token is inserted at a beginning of each utterance constituting a first dialogue text and an end of sentence (EOS) token at an end of each utterance constituting the first dialogue text, and

wherein the first dialogue text and a summary text, corresponding to the first dialogue text, are combined to form a dialogue text-summary text pair by using the inserted BOS token and EOS token as a delimiter.

7. The apparatus of claim 1 , wherein the natural language processing model is further trained by self-supervised learning based on at least one of a switching speaker task, a switching utterance task, or an inserting utterance task.

8. The apparatus of claim 7 , wherein the masking speaker task is performed by replacing at least a portion of one or more speaker tokens included in one of the dialogue text and the summary text of the dialogue text-summary text pair, with a mask token; and

the natural language processing model is pre-trained based on a loss between a predicted result of the natural language processing model on a speaker token corresponding to a position of the mask token and the replaced speaker token.

9. The apparatus of claim 7 , wherein the switching speaker task is performed by switching at least a portion of one or more speaker tokens included in one of the dialogue text and the summary text of the dialogue text-summary text pair; and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not switching has been performed for a speaker token of each utterance included in the dialogue text and a correct answer indicating whether or not there is actual switching.

10. The apparatus of claim 7 , wherein the switching utterance task is performed by switching at least a portion of one or more utterances included in one of the dialogue text and the summary text of the dialogue text-summary text pair; and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not switching has been performed for each of the portion of one or more utterances included in the dialogue text and a correct answer indicating whether or not there is actual switching.

11. The apparatus of claim 7 , wherein the inserting utterance task is performed by selecting one or more utterances from at least one dialogue text and at least one summary text included in the same batch and inserting the selected one or more utterances into a dialogue text to be input to the natural language processing model, and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not each utterance included in the dialogue text, which has been input to the natural language processing model, is inserted from another dialogue text and a correct answer indicating whether or not there is actual inserting.

12. The apparatus of claim 7 , wherein the switching speaker task is performed by:

comparing a probability value of a variable γ, which is randomly determined for each utterance of a sample with a third preset threshold, the variable γ being a rational number equal to or greater than 0 and equal to or less than 1; and

determining whether to perform the switching speaker task for a corresponding utterance based on a result of the comparing the probability value of the variable γ with the third preset threshold.

13. The apparatus of claim 7 , wherein the switching utterance task is performed by:

comparing a probability value of a variable x, which is randomly determined for each utterance of a sample with a fourth preset threshold, the variable x being a rational number equal to or greater than 0 and equal to or less than 1; and

determining whether to perform the switching speaker task for a corresponding utterance based on a result of the comparing the probability value of the variable x with the fourth preset threshold.

14. A method for training a dialogue summary model, comprising:

transferring one or more learning parameter values of a natural language processing model, which is pre-trained, to a sequence-to-sequence-based dialogue summary model; and

training the sequence-to-sequence-based dialogue summary model by using the transferred learning parameter values as initial values for learning parameters of each of an encoder and a decoder in the sequence-to-sequence-based dialogue summary model,

wherein the natural language processing model is trained by selecting, as samples, dialogue text-summary text pairs of a preset batch size, and by self-supervised learning based on a masking speaker task by using the selected samples, and

wherein the masking speaker task is performed by:

comparing a probability value of a variable α, which is randomly determined for each sample, with a first preset threshold, the variable α being a rational number equal to or greater than 0 and equal to or less than 1; and

for each sample, determining an execution target of the masking speaker task between a dialogue text and a summary text of a dialogue text-summary text pair, based on a result of the comparing.

15. The method of claim 14 , wherein the training comprises updating the learning parameters of the sequence-to-sequence-based dialogue summary model based on a loss between an inferred summary text output by the sequence-to-sequence-based dialogue summary model receiving an input dialogue text and a correct answer summary text corresponding to the input dialogue text.

16. The method of claim 14 , wherein the training comprises sharing at least some of learning parameter values of the encoder and at least some of learning parameter values of the decoder.

17. The method of claim 14 , wherein the natural language processing model is further trained by self-supervised learning based on at least one of a switching speaker task, a switching utterance task, and an inserting utterance task.

18. The method of claim 17 , wherein the masking speaker task is performed by replacing at least a portion of one or more speaker tokens included in one of the dialogue text and the summary text of the dialogue text-summary text pair with a mask token; and

the natural language processing model is pre-trained based on a loss between a predicted result of the natural language processing model on a speaker token corresponding to a position of the mask token and the replaced speaker token.

19. The method of claim 17 , wherein the switching speaker task is performed by switching at least a portion of one or more speaker tokens included in one of the dialogue text and the summary text of the dialogue text-summary text pair; and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not switching has been performed for a speaker token of each utterance included in the dialogue text and a correct answer indicating whether or not there is actual switching.

20. The method of claim 17 , wherein the switching utterance task is performed by switching at least a portion of one or more utterances included in one of the dialogue text and the summary text of the dialogue text-summary text pair; and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not switching has been performed for each of the portion of one or more utterances included in the dialogue text and a correct answer indicating whether or not there is actual switching.

21. The method of claim 17 , wherein the inserting utterance task is performed by selecting one or more utterances from at least one dialogue text and at least one summary text included in the same batch and inserting the one or more utterances into a dialogue text to be input to the natural language processing model; and

the natural language processing model is pre-trained based on a loss between a prediction result of the natural language processing model on whether or not each utterance included in the dialogue text, which has been input to the natural language processing model, is inserted from another dialogue text and a correct answer indicating whether or not there is actual inserting.

22. The method of claim 14 , wherein each of the encoder and the decoder includes at least a portion of a network structure of the natural language processing model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 29, 2021
From: LEE, HYUN JAE; CHOI, HYUN JIN; YUN, JAE WOONG; KIM, JU DONG; HWANG, BONG KYU; JOE, SEONG HO; GWON, YOUNG JUNE
To: SAMSUNG SDS CO., LTD.
Reel/Frame 057961/0914 →
Priority Claims (2)
KR 10-2021-0037976 · Mar 24, 2021 · national
KR 10-2021-0069043 · May 28, 2021 · national
Continuity (1)
Related Publication 20220310075A1 · Sep 29, 2022