IP Library › Granted Patent US 12,093,297
Granted Patent B2
US 12,093,297 · App. 17/577,561 · Granted Sep 17, 2024

Summary generation model training method and apparatus, device and storage medium

Inventors: Wenhao Wu (Beijing, CN); Wei Li (Beijing, CN); Xinyan Xiao (Beijing, CN); Jiachen Liu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F16/345G06F40/51G06F40/56
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,093,297
App. No.
17/577,561
Granted
Sep 17, 2024
Kind
B2
Abstract

The present disclosure provides a summary generation model training method and apparatus, a device and a storage medium, and relates to the field of computer technologies, and in particular, to the field of artificial intelligence such as natural language processing and deep learning. The summary generation model training method includes: acquiring a document representation corresponding to a document sample; constructing, based on the document representation, a summary representation corresponding to the document representation, the summary representation including a positive summary representation and a negative summary representation; and constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation, and training a summary generation model based on the total contrastive loss function. The present disclosure may improve accuracy of the summary generation model.

Claims (66)

1. A computer-implemented summary generation model training method, wherein the summary generation model is configured to process a document to obtain a summary corresponding to the document, the method comprising:

acquiring a document representation corresponding to a document sample, the document representation describing data in a vector form;

constructing, based on the document representation, a summary representation corresponding to the document representation, the summary representation comprising a positive summary representation and a negative summary representation; and

constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation, and training a summary generation model based on the total contrastive loss function,

wherein the summary generation model comprises: an encoder and a decoder, and the step of acquiring a document representation corresponding to a document sample comprises:

processing the document sample by using the encoder, to obtain an encoding representation;

processing the encoding representation by using the decoder, to obtain a decoding representation; and

taking the encoding representation and the decoding representation as the document representation,

wherein the document representation comprises the encoding representation and the decoding representation, and the step of constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation comprises:

constructing a first contrastive loss function based on the encoding representation, the positive summary representation and the negative summary representation;

constructing a second contrastive loss function based on the decoding representation, the positive summary representation and the negative summary representation; and

constructing the total contrastive loss function based on the first contrastive loss function and the second contrastive loss function,

wherein the positive summary representation refers to a representation of the positive summary sample corresponding to the document sample, the negative summary representation refers to a representation of the negative summary sample corresponding to the document sample, and the positive summary sample and the negative summary sample are constructed based on a generation text corresponding to the decoding representation.

2. The method according to claim 1 , wherein the step of constructing a positive summary sample comprises:

performing loopback translation on the generation text to obtain a loopback translation result, and taking the loopback translation result as the positive summary sample.

3. The method according to claim 1 , wherein the step of constructing a negative summary sample comprises at least one of the following:

performing entity replacement on the generation text to obtain an entity replacement result, and taking the entity replacement result as the negative summary sample;

performing pronoun replacement on the generation text to obtain a pronoun replacement result, and taking the pronoun replacement result as the negative summary sample;

performing emotion replacement on the generation text to obtain an emotion replacement result, and taking the emotion replacement result as the negative summary sample;

acquiring a similar text of the generation text, and taking the similar text as the negative summary sample; and

performing virtual adversarial training on the generation text to obtain a virtual adversarial result, and taking the virtual adversarial result as the negative summary sample.

4. An electronic device, comprising:

at least one processor; and

a memory communicatively connected with the at least one processor;

wherein the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to perform a summary generation model training method, wherein the summary generation model is configured to process a document to obtain a summary corresponding to the document, and the summary generation model training method comprises:

acquiring a document representation corresponding to a document sample, the document representation describing data in a vector form;

constructing, based on the document representation, a summary representation corresponding to the document representation, the summary representation comprising a positive summary representation and a negative summary representation; and

constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation, and training a summary generation model based on the total contrastive loss function,

wherein the summary generation model comprises: an encoder and a decoder, and the step of acquiring a document representation corresponding to a document sample comprises:

processing the document sample by using the encoder, to obtain an encoding representation;

processing the encoding representation by using the decoder, to obtain a decoding representation; and

taking the encoding representation and the decoding representation as the document representation,

wherein the document representation comprises the encoding representation and the decoding representation, and the step of constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation comprises:

constructing a first contrastive loss function based on the encoding representation, the positive summary representation and the negative summary representation;

constructing a second contrastive loss function based on the decoding representation, the positive summary representation and the negative summary representation; and

constructing the total contrastive loss function based on the first contrastive loss function and the second contrastive loss function,

wherein the positive summary representation refers to a representation of the positive summary sample corresponding to the document sample, the negative summary representation refers to a representation of the negative summary sample corresponding to the document sample, and the positive summary sample and the negative summary sample are constructed based on a generation text corresponding to the decoding representation.

5. The electronic device according to claim 4 , wherein the step of constructing a positive summary sample comprises:

performing loopback translation on the generation text to obtain a loopback translation result, and taking the loopback translation result as the positive summary sample.

6. The electronic device according to claim 4 , wherein the step of constructing a negative summary sample comprises at least one of the following:

performing entity replacement on the generation text to obtain an entity replacement result, and taking the entity replacement result as the negative summary sample;

performing pronoun replacement on the generation text to obtain a pronoun replacement result, and taking the pronoun replacement result as the negative summary sample;

performing emotion replacement on the generation text to obtain an emotion replacement result, and taking the emotion replacement result as the negative summary sample;

acquiring a similar text of the generation text, and taking the similar text as the negative summary sample; and

performing virtual adversarial training on the generation text to obtain a virtual adversarial result, and taking the virtual adversarial result as the negative summary sample.

7. A non-transitory computer readable storage medium with computer instructions stored thereon, wherein the computer instructions are used for causing a computer to perform a summary generation model training method, wherein the summary generation model is configured to process a document to obtain a summary corresponding to the document, and the summary generation model training method comprises:

acquiring a document representation corresponding to a document sample, the document representation describing data in a vector form;

constructing, based on the document representation, a summary representation corresponding to the document representation, the summary representation comprising a positive summary representation and a negative summary representation; and

constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation, and training a summary generation model based on the total contrastive loss function,

wherein the summary generation model comprises: an encoder and a decoder, and the step of acquiring a document representation corresponding to a document sample comprises:

processing the document sample by using the encoder, to obtain an encoding representation;

processing the encoding representation by using the decoder, to obtain a decoding representation; and

taking the encoding representation and the decoding representation as the document representation,

wherein the document representation comprises the encoding representation and the decoding representation, and the step of constructing a total contrastive loss function based on the document representation, the positive summary representation and the negative summary representation comprises:

constructing a first contrastive loss function based on the encoding representation, the positive summary representation and the negative summary representation;

constructing a second contrastive loss function based on the decoding representation, the positive summary representation and the negative summary representation; and

constructing the total contrastive loss function based on the first contrastive loss function and the second contrastive loss function,

wherein the positive summary representation refers to a representation of the positive summary sample corresponding to the document sample, the negative summary representation refers to a representation of the negative summary sample corresponding to the document sample, and the positive summary sample and the negative summary sample are constructed based on a generation text corresponding to the decoding representation.

8. The non-transitory computer readable storage medium according to claim 7 , wherein the step of constructing a positive summary sample comprises:

performing loopback translation on the generation text to obtain a loopback translation result, and taking the loopback translation result as the positive summary sample.

9. The non-transitory computer readable storage medium according to claim 7 , wherein the step of constructing a negative summary sample comprises at least one of the following:

performing entity replacement on the generation text to obtain an entity replacement result, and taking the entity replacement result as the negative summary sample;

performing pronoun replacement on the generation text to obtain a pronoun replacement result, and taking the pronoun replacement result as the negative summary sample;

performing emotion replacement on the generation text to obtain an emotion replacement result, and taking the emotion replacement result as the negative summary sample;

acquiring a similar text of the generation text, and taking the similar text as the negative summary sample; and

performing virtual adversarial training on the generation text to obtain a virtual adversarial result, and taking the virtual adversarial result as the negative summary sample.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 18, 2022
From: WU, WENHAO; LI, WEI; XIAO, XINYAN; LIU, JIACHEN
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 058676/0986 →
Priority Claims (1)
CN 202110734020.1 · Jun 30, 2021 · national
Continuity (1)
Related Publication 20230004589A1 · Jan 5, 2023
Cited By (1)
US 12,455,913