IP Library › Granted Patent US 11,423,222
Granted Patent B2
US 11,423,222 · App. 17/243,097 · Granted Aug 23, 2022

Method and apparatus for text error correction, electronic device and storage medium

Inventors: Ruiqing Zhang (Beijing, CN); Chuanqiang Zhang (Beijing, CN); Zhongjun He (Beijing, CN); Zhi Li (Beijing, CN); Hua Wu (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F40/232G06F40/166G06F40/279G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,423,222
App. No.
17/243,097
Granted
Aug 23, 2022
Kind
B2
Abstract

A method for text error correction includes: obtaining a text to be corrected; obtaining a pinyin sequence of the text to be corrected; and inputting the text to be corrected and the pinyin sequence to a text error correction model, to obtain a corrected text.

Claims (70)

1. A method for text error correction, comprising:

obtaining a text to be corrected;

obtaining a pinyin sequence of the text to be corrected; and

inputting the text to be corrected and the pinyin sequence to a text error correction model, to obtain a corrected text;

wherein, inputting the text to be corrected and the pinyin sequence to the text error correction model, to obtain the corrected text, comprises:

detecting a wrong word in the text to be corrected by the text error correction model, to determine the wrong in the text to be corrected;

obtaining a pinyin in the pinyin sequence corresponding to the wrong word by the text error corrected model, and replacing the wrong with the pinyin, to obtain a pinyin-text to be corrected;

correcting the pinyin-text to be corrected by the text error correction model, to obtain the corrected text; and

obtaining the pinyin-text to be corrected by the formula:

X wp =W w *O det +X p *(1-O det ),

where X wp is the pinyin-text to be corrected, W w is the text to be corrected, X p is the pinyin sequence, O det is an error detection labeling sequence of the text to be corrected, and the error detection labeling sequence corresponds to the text to be corrected one by one.

2. The method of claim 1 , wherein, training the text error correction model by:

obtaining a sample text and a sample pinyin sequence corresponding to the sample text;

obtaining a target text of the sample text;

inputting the sample text and the sample pinyin sequence to the text error correction model, to generate a predicted sample correction text; and

generating a loss value according to the predicted sample correction text and the target text, and training the text error correction model according to the loss value.

3. The method of claim 1 , wherein, training the text error correction model by:

obtaining a sample text and a sample pinyin sequence corresponding to the sample obtaining a target pinyin text and a target text of the sample text;

inputting the sample text and the sample pinyin sequence to the model error correction model, to generate a predicted sample pinyin text and a predicted sample correction text;

generating a first loss value according to the predicted sample pinyin text and the target pinyin text, and generating a second loss value according to the predicted sample correction text and the target text; and

training the text error correction model according to the first loss value and the second loss value.

4. The method of claim 3 , wherein, the sample text comprises one or more of a masked sample text, a sample text with confusing words and a pinyin sample text with confusing words.

5. An electronic device, comprising:

at least one processor; and

a memory communicatively coupled to the at least one processor;

wherein the at least one processor is configured to:

obtain a text to be corrected;

obtain a pinyin sequence of the text to be corrected; and

input the text to be corrected and the pinyin sequence to a text error correction model, to obtain a corrected text;

wherein the at least one processor is further configured to:

detect a wrong word in the text to be corrected by the text error correction model, to determine the wrong word in the text to be corrected;

obtain a pinyin in the pinyin sequence corresponding to the wrong word by the text error correction model, and replacing the wrong word with the pinyin, to obtain a pinyin-text to be corrected; and

correct the pinyin-text to be corrected by the text error correction model, to obtain the corrected text;

wherein the pinyin-text to be corrected is obtained by the formula:

X wp =W w *O det +X p *(1-O det ),

where X wp is the pinyin-text to be corrected, W w is the text to be corrected, X p is the pinyin sequence, O det is an error detection labeling sequence of the text to be corrected, and the error detection labeling sequence corresponds to the text to be corrected one by one.

6. The electronic device of claim 5 , wherein the at least one processor is further configured to:

obtain a sample text and a sample pinyin sequence corresponding to the sample text;

obtain a target text of the sample text; input the sample text and the sample pinyin sequence to the text error correction model, to generate a predicted sample correction text; and

generate a loss value according to the predicted sample correction text and the target text, and train the text error correction model according to the loss value.

7. The electronic device of claim 5 , wherein the at least one processor is further configured to:

obtain a sample text and a sample pinyin sequence corresponding to the sample text;

obtain a target pinyin text and a target sample of the sample text;

input the sample text and the sample pinyin sequence to the text error correction model, to generate a predicted sample pinyin text and a predicted sample correction text;

generate a first loss value according to the predicted sample pinyin text and the target pinyin text, and generate a second loss value according to the predicted sample correction text and the target sample text; and

train the text error correction model according to the first loss value and the second loss value.

8. The electronic device of claim 7 , wherein the sample text comprises one or more of a masked sample text, a sample text with confusing words and a pinyin sample text with confusing words.

9. A non-transitory computer-readable storage medium having computer instructions stored thereon, wherein the computer instructions are configured to cause a computer to execute a method for text error correction, the method comprising:

obtaining a text to be corrected;

obtaining a pinyin sequence of the text to be corrected; and

inputting the text to be corrected and the pinyin sequence to a text error correction model, to obtain a corrected text;

wherein, inputting the text to be corrected and the pinyin sequence to the text error correction model, to obtain the corrected text, comprises:

detecting a wrong word in the text to be corrected by the text error correction model, to determine the wrong word in the text to be corrected;

obtaining a pinyin in the pinyin sequence corresponding to the wrong word by the text error correction model, and replacing the wrong word with the pinyin, to obtain a pinyin-text to be corrected;

correcting the pinyin-text to be corrected by the text error correction model, to obtain the corrected text; and

obtaining the pinyin-text to be corrected by the formula:

X wp =W w *O det +X p *(1−O det ),

where X wp is the pinyin-text to be corrected, W w is the text to be corrected, X p is the pinyin sequence, O det is an error detection labeling sequence of the text to be corrected, and the error detection labeling sequence corresponds to the text to be corrected one by one.

10. The storage medium of claim 9 , wherein, training the text error correction model by:

obtaining a sample text and a sample pinyin sequence corresponding to the sample text;

obtaining a target text of the sample text;

inputting the sample text and the sample pinyin sequence to the text error correction model, to generate a predicted sample correction text; and

generating a loss value according to the predicted sample correction text and the target text, and training the text error correction model according to the loss value.

11. The storage medium of claim 9 , wherein, training the text error correction model by:

obtaining a sample text and a sample pinyin sequence corresponding to the sample text;

obtaining a target pinyin text and a target text of the sample text;

inputting the sample text and the sample pinyin sequence to the model error correction model, to generate a predicted sample pinyin text and a predicted sample correction text;

generating a first loss value according to the predicted sample pinyin text and the target pinyin text, and generating a second loss value according to the predicted sample correction text and the target text; and

training the text error correction model according to the first loss value and the second loss value.

12. The storage medium of claim 11 , wherein, the sample text comprises one or more of a masked sample text, a sample text with confusing words and a pinyin sample text with confusing words.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2021
From: ZHANG, RUIQING; ZHANG, CHUANQIANG; HE, ZHONGJUN; LI, ZHI; WU, HUA
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 056073/0194 →
Priority Claims (1)
CN 202011442447.6 · Dec 11, 2020 · national
Continuity (1)
Related Publication 20210248309A1 · Aug 12, 2021