IP Library › Granted Patent US 12,314,677
Granted Patent B2
US 12,314,677 · App. 17/889,218 · Granted May 27, 2025

Method for pre-training model, device, and storage medium

Inventors: Junyuan Shang (Beijing, CN); Shuohuan Wang (Beijing, CN); Siyu Ding (Beijing, CN); Yanbin Zhao (Beijing, CN); Chao Pang (Beijing, CN); Yu Sun (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F40/40G06F40/289
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,314,677
App. No.
17/889,218
Granted
May 27, 2025
Kind
B2
Abstract

A method and apparatus for pre-training a model, a device, a storage medium, and a program product. An embodiment of the method includes: acquiring a sample natural language text; generating N types of prompt words based on the sample natural language text, where N is a positive integer; generating sample input data based on the sample natural language text and the N types of prompt words; and training an initial language model based on the sample input data, to obtain a pre-trained language model.

Claims (92)

1. A method for pre-training a model, the method comprising:

acquiring a sample natural language text;

generating N types of prompt words based on the sample natural language text, wherein N is a positive integer, and the N types comprise at least one of a task type, a topic type, a key phrase type, or a sentiment type;

generating sample input data based on the sample natural language text and the N types of prompt words; and

training an initial language model based on the sample input data, to obtain a pre-trained language model,

wherein the generating sample input data based on the sample natural language text and the N types of prompt words, comprises:

generating random sampling probabilities of the N types of prompt words respectively;

selecting, from the N types of prompt words, a prompt word whose random sampling probability is greater than a preset probability threshold;

intercepting a sample prefix text fragment from the sample natural language text; and

splicing the selected prompt word with the sample prefix text fragment to generate the sample input data.

2. The method according to claim 1 , wherein the types of prompt words comprise the task type; and

the generating N types of prompt words based on the sample natural language text, comprises:

determining a target task type of the sample natural language text;

acquiring a vocabulary of consecutive prompt words associated with the target task type, wherein one task type is associated with one vocabulary of consecutive prompt words; and

acquiring consecutive prompt words of a random length from the vocabulary of consecutive prompt words associated with the target task type, as prompt words of the task type of the sample natural language text.

3. The method according to claim 1 , wherein the types of prompt words comprise the topic type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained topic classification model, to obtain a prompt word of the topic type of the sample natural language text.

4. The method according to claim 1 , wherein the types of prompt words comprise the key phrase type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained key phrase extraction model, to obtain a prompt word of the key phrase type of the sample natural language text.

5. The method according to claim 1 , wherein the types of prompt words comprise the sentiment type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained sentiment analysis model, to obtain a prompt word of the sentiment type of the sample natural language text.

6. The method according to claim 1 , wherein the types of prompt words comprise a generated length type; and

the generating N types of prompt words based on the sample natural language text, comprises:

using a length of the sample natural language text as a prompt word of the generated length type of the sample natural language text.

7. The method according to claim 1 , wherein the N types comprise at least two of: the task type, the topic type, the key phrase type, or the sentiment type.

8. The method according to claim 1 , wherein the training the initial language model based on the sample input data to obtain the pre-trained language model, comprises:

inputting the sample input data into the initial language model, to obtain sample pseudo-natural language text; and

adjusting parameters of the initial language model based on a difference between the sample pseudo-natural language text and the sample natural language text, to obtain the pre-trained language model.

9. A method for generating text by using a pre-trained language model obtained by training using the method according to claim 1 , the method comprising:

acquiring a prefix text fragment and at least one type of prompt word;

splicing the prefix text fragment with the at least one type of prompt word to generate input data; and

inputting the input data into a pre-trained language model to generate pseudo-natural language text.

10. An electronic device, comprising:

at least one processor; and

a memory communicatively connected to the at least one processor; wherein,

the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform operations, the operations comprising:

acquiring a sample natural language text;

generating N types of prompt words based on the sample natural language text, wherein N is a positive integer, and the N types comprise at least one of a task type, a topic type, a key phrase type, or a sentiment type;

generating sample input data based on the sample natural language text and the N types of prompt words; and

training an initial language model based on the sample input data, to obtain a pre-trained language model,

wherein the generating sample input data based on the sample natural language text and the N types of prompt words, comprises:

generating random sampling probabilities of the N types of prompt words respectively;

selecting, from the N types of prompt words, a prompt word whose random sampling probability is greater than a preset probability threshold;

intercepting a sample prefix text fragment from the sample natural language text; and

splicing the selected prompt word with the sample prefix text fragment to generate the sample input data.

11. The electronic device according to claim 10 , wherein the types of prompt words comprise the task type; and

the generating N types of prompt words based on the sample natural language text, comprises:

determining a target task type of the sample natural language text;

acquiring a vocabulary of consecutive prompt words associated with the target task type, wherein one task type is associated with one vocabulary of consecutive prompt words; and

acquiring consecutive prompt words of a random length from the vocabulary of consecutive prompt words associated with the target task type, as prompt words of the task type of the sample natural language text.

12. The electronic device according to claim 10 , wherein the types of prompt words comprise the topic type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained topic classification model, to obtain a prompt word of the topic type of the sample natural language text.

13. The electronic device according to claim 10 , wherein the types of prompt words comprise the key phrase type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained key phrase extraction model, to obtain a prompt word of the key phrase type of the sample natural language text.

14. The electronic device according to claim 10 , wherein the types of prompt words comprise the sentiment type; and

the generating N types of prompt words based on the sample natural language text, comprises:

inputting the sample natural language text into a pre-trained sentiment analysis model, to obtain a prompt word of the sentiment type of the sample natural language text.

15. The electronic device according to claim 10 , wherein the types of prompt words comprise a generated length type; and

the generating N types of prompt words based on the sample natural language text, comprises:

using a length of the sample natural language text as a prompt word of the generated length type of the sample natural language text.

16. The electronic device according to claim 10 , wherein the generating sample input data based on the sample natural language text and the N types of prompt words, comprises:

generating random sampling probabilities of the N types of prompt words respectively;

selecting, from the N types of prompt words, a prompt word whose random sampling probability is greater than a preset probability threshold;

intercepting a sample prefix text fragment from the sample natural language text; and

splicing the selected prompt word with the sample prefix text fragment to generate the sample input data.

17. The electronic device according to claim 10 , wherein the training the initial language model based on the sample input data to obtain the pre-trained language model, comprises:

inputting the sample input data into the initial language model, to obtain sample pseudo-natural language text; and

adjusting parameters of the initial language model based on a difference between the sample pseudo-natural language text and the sample natural language text, to obtain the pre-trained language model.

18. An electronic device for generating text, comprising:

at least one processor; and

a memory communicatively connected to the at least one processor; wherein,

the memory stores instructions executable by the at least one processor, and the instructions, when executed by the at least one processor, cause the at least one processor to perform the method according to claim 9 .

19. A non-transitory computer readable storage medium, storing computer instructions thereon, wherein, the computer instructions, when executed by a computer, cause the computer to perform operations, the operations comprising:

acquiring a sample natural language text;

generating N types of prompt words based on the sample natural language text, wherein N is a positive integer;

generating sample input data based on the sample natural language text and the N types of prompt words, and the N types comprise at least one of a task type, a topic type, a key phrase type, or a sentiment type; and

training an initial language model based on the sample input data, to obtain a pre-trained language model,

wherein the generating sample input data based on the sample natural language text and the N types of prompt words, comprises:

generating random sampling probabilities of the N types of prompt words respectively;

selecting, from the N types of prompt words, a prompt word whose random sampling probability is greater than a preset probability threshold;

intercepting a sample prefix text fragment from the sample natural language text; and

splicing the selected prompt word with the sample prefix text fragment to generate the sample input data.

20. The computer readable storage medium according to claim 19 , wherein the types of prompt words comprise the task type; and

the generating N types of prompt words based on the sample natural language text, comprises:

determining a target task type of the sample natural language text;

acquiring a vocabulary of consecutive prompt words associated with the target task type, wherein one task type is associated with one vocabulary of consecutive prompt words; and

acquiring consecutive prompt words of a random length from the vocabulary of consecutive prompt words associated with the target task type, as prompt words of the task type of the sample natural language text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 13, 2022
From: SHANG, JUNYUAN; WANG, SHUOHUAN; DING, SIYU; ZHAO, YANBIN; PANG, CHAO; SUN, YU
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061411/0795 →
Priority Claims (1)
CN 202111260446.4 · Oct 28, 2021 · national
Continuity (1)
Related Publication 20230040095A1 · Feb 9, 2023
References Cited (15)
US 20180239815A1 · Yi · 2018 [cited by examiner]
CN 110263158A · 2019 [cited by applicant]
CN 112183091A · 2021 [cited by applicant]
CN 113127624A · 2021 [cited by applicant]
CN 113468877A · 2021 [cited by applicant]
CN 113962315A · 2022 [cited by applicant]
JP 2003263441A · 2003 [cited by applicant]
JP 2015170241A · 2015 [cited by applicant]
JP 2016091078A · 2016 [cited by applicant]
WO WO2018126213A1 · 2018 [cited by applicant]
Aghajanyan, A., Okhonko, D., Lewis, M., Joshi, M., Xu, H., Ghosh, G., & Zettlemoyer, L. (2021). Htlm: Hyper-text pre-training and prompting of language models. arXiv preprint arXiv:2107.06955. (Year: 2021). [cited by examiner]
Fan et al., “Controllable Abstractive Summarization”, Facebook AI Research, May 18, 2018, 10 pages. [cited by applicant]
He et al., “CTRLsum: Towards Generic Controllable Text Summarization”, Dec. 8, 2020, 35 pages. [cited by applicant]
Liu et al., “Pre-train, Prompt, and Predict: A Systematic Survey of Prompting Methods in Natural Language Processing”, Jul. 28, 2021, 46 pages. [cited by applicant]
Wang et al., “ERNIE 3.0 Titan: Exploring Larger-scale Knowledge Enhanced Pre-training for Language Understanding and Generation”, Dec. 23, 2023, 28 pages. [cited by applicant]
Cited By (1)
US 12,737,631