IP Library › Granted Patent US 11,392,779
Granted Patent B2
US 11,392,779 · App. 16/891,705 · Granted Jul 19, 2022

Bilingual corpora screening method and apparatus, and storage medium

Inventors: Jingwei Li (Beijing, CN); Yuhui Sun (Beijing, CN); Xiang Li (Beijing, CN)
Assignee: Beijing Xiaomi Mobile Software Co., Ltd.
G06F40/58G06F40/263G06F40/44G06F40/51G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,392,779
App. No.
16/891,705
Granted
Jul 19, 2022
Kind
B2
Abstract

A bilingual corpora screening method includes: acquiring multiple pairs of bilingual corpora, wherein each pair of the bilingual corpora comprises a source corpus and a target corpus; training a machine translation model based on the multiple pairs of bilingual corpora; obtaining a first feature of each pair of bilingual corpora based on the trained machine translation model; training a language model based on the multiple pairs of bilingual corpora; obtaining feature vectors of each pair of bilingual corpora and determining a second feature of each pair of bilingual corpora based on the trained language model; determining a quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora; and screening each pair of bilingual corpora according to the quality value of each pair of bilingual corpora.

Claims (74)

1. A bilingual corpora screening method, comprising:

acquiring multiple pairs of bilingual corpora, wherein each pair of the bilingual corpora comprises a source corpus and a target corpus;

training a machine translation model based on the multiple pairs of bilingual corpora;

obtaining a first feature of each pair of bilingual corpora based on the trained machine translation model;

training a language model with a noise-reduction autoencoder based on the multiple pairs of bilingual corpora, wherein the training the language model comprises: obtaining a corpus from the multiple pairs of bilingual corpora; adding the corpus with noise to obtain a noise-added corpus; inputting the noise-added corpus to an encoder to obtain a feature vector; and inputting the feature vector to a decoder for decoding to obtain a newly generated corpus;

obtaining feature vectors of each pair of bilingual corpora and determining a second feature of each pair of bilingual corpora based on the trained language model;

determining a quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora; and

screening each pair of bilingual corpora according to the quality value of each pair of bilingual corpora.

2. The method according to claim 1 , wherein the machine translation model comprises a first translation model and a second translation model, and each first feature comprises a first probability feature and a second probability feature; and

obtaining the first feature of each pair of bilingual corpora based on the trained machine translation model comprises:

inputting the source corpus in each pair of the bilingual corpora to a trained first translation model, and determining the first probability feature of the bilingual corpora based on a result output by the trained first translation model, wherein the first probability feature is a probability that the trained first translation model predicts the source corpus as the target corpus corresponding to the source corpus in the bilingual corpora, and

inputting the target corpus in each pair of the bilingual corpora to a trained second translation model, and determining the second probability feature of the bilingual corpora based on a result output by the trained second translation model, wherein the second probability feature is a probability that the trained second translation model predicts the target corpus as the source corpus corresponding to the target corpus in the bilingual corpora.

3. The method according to claim 1 , wherein the language model comprises a first language model and a second language model, and each of the feature vectors comprises a first feature vector and a second feature vector; and

obtaining the feature vectors of each pair of bilingual corpora and determining the second feature of each pair of bilingual corpora based on the trained language model, comprises:

for each pair of the bilingual corpora, obtaining the first feature vector corresponding to the source corpus by inputting the source corpus in the pair of bilingual corpora to a trained first language model,

obtaining the second feature vector corresponding to the target corpus by inputting the target corpus in the pair of bilingual corpora to a trained second language model, and

determining a semantic similarity between the source corpus and the target corpus in the pair of bilingual corpora as the second feature of the pair of bilingual corpora based on the first feature vector and the second feature vector.

4. The method according to claim 3 , wherein the first language model comprises a first encoder obtained by training the source corpora in each pair of bilingual corpora, and the second language model comprises a second encoder obtained by training the target corpora in each pair of bilingual corpora,

wherein each of the first encoder and the second encoder is a noise-reduction autoencoder.

5. The method according to claim 4 , wherein a model parameter with which the first encoder encodes the source corpora is the same as a model parameter with which the second encoder encodes the target corpora.

6. The method according to claim 3 , wherein the semantic similarity is one of a Manhattan distance, a Euclidean distance, or a cosine similarity.

7. The method according to claim 1 , wherein determining the quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora comprises:

performing a weighted calculation on the first feature and the second feature of each pair of bilingual corpora to obtain the quality value of each pair of bilingual corpora.

8. The method according to claim 1 , wherein screening each pair of bilingual corpora according to the quality value of each pair of bilingual corpora comprises:

ranking each pair of bilingual corpora according to the quality value of each pair of bilingual corpora; and

screening each pair of bilingual corpora according to a ranking result.

9. A bilingual corpora screening apparatus, comprising:

a processor; and

a memory storing instructions executable by the processor,

wherein the processor is configured to:

acquire multiple pairs of bilingual corpora, wherein each pair of the bilingual corpora comprises a source corpus and a target corpus;

train a machine translation model with a noise-reduction autoencoder based on the multiple pairs of bilingual corpora, wherein training the language model comprises: obtaining a corpus from the multiple pairs of bilingual corpora; adding the corpus with noise to obtain a noise-added corpus; inputting the noise-added corpus to an encoder to obtain a feature vector; and inputting the feature vector to a decoder for decoding to obtain a newly generated corpus;

obtain a first feature of each pair of bilingual corpora based on the trained machine translation model;

train a language model based on the multiple pairs of bilingual corpora;

obtain feature vectors of each pair of bilingual corpora and determine a second feature of each pair of bilingual corpora based on the trained language model;

determine a quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora; and

screen each pair of bilingual corpora according to the quality value of each pair of bilingual corpora.

10. The apparatus according to claim 9 , wherein the machine translation model comprises a first translation model and a second translation model, and each first feature comprises a first probability feature and a second probability feature; and

the processor is further configured to:

input the source corpus in each pair of the bilingual corpora to a trained first translation model, and determine the first probability feature of the bilingual corpora based on a result output by the trained first translation model, wherein the first probability feature is a probability that the trained first translation model predicts the source corpus as the target corpus corresponding to the source corpus in the bilingual corpora, and

input the target corpus in each pair of the bilingual corpora to a trained second translation model, and determine the second probability feature of the bilingual corpora based on a result output by the trained second translation model, wherein the second probability feature is a probability that the trained second translation model predicts the target corpus as the source corpus corresponding to the target corpus in the bilingual corpora.

11. The apparatus according to claim 9 , wherein the language model comprises a first language model and a second language model, and each of the feature vectors comprises a first feature vector and a second feature vector; and

the processor is further configured to:

for each pair of the bilingual corpora, obtain the first feature vector corresponding to the source corpus by inputting the source corpus in the pair of bilingual corpora to a trained first language model,

obtain the second feature vector corresponding to the target corpus by inputting the target corpus in the pair of bilingual corpora to a trained second language model, and

determine a semantic similarity between the source corpus and the target corpus in the pair of bilingual corpora as the second feature of the pair of bilingual corpora based on the first feature vector and the second feature vector.

12. The apparatus according to claim 11 , wherein the first language model comprises a first encoder obtained by training source corpora in each pair of bilingual corpora, and the second language model comprises a second encoder obtained by training target corpora in each pair of bilingual corpora,

wherein each of the first encoder and the second encoder is a noise-reduction autoencoder.

13. The apparatus according to claim 12 , wherein a model parameter with which the first encoder encodes the source corpora is the same as a model parameter with which the second encoder encodes the target corpora.

14. The apparatus according to claim 11 , wherein the semantic similarity is one of a Manhattan distance, a Euclidean distance, or a cosine similarity.

15. The apparatus according to claim 9 , wherein the processor is further configured to:

perform a weighted calculation on the first feature and the second feature of each pair of bilingual corpora to obtain the quality value of each pair of bilingual corpora.

16. The apparatus according to claim 9 , wherein the processor is further configured to:

rank each pair of bilingual corpora according to the quality value of each pair of bilingual corpora; and

screen each pair of bilingual corpora according to a ranking result.

17. A non-transitory computer-readable storage medium having stored therein instructions that, when executed by a processor of a device, cause the device to perform a bilingual corpora screening method, the method comprising:

acquiring multiple pairs of bilingual corpora, wherein each pair of the bilingual corpora comprises a source corpus and a target corpus;

training a machine translation model based on the multiple pairs of bilingual corpora;

obtaining a first feature of each pair of bilingual corpora based on the trained machine translation model;

training a language model with a noise-reduction autoencoder based on the multiple pairs of bilingual corpora, wherein the training the language model comprises: obtaining a corpus from the multiple pairs of bilingual corpora; adding the corpus with noise to obtain a noise-added corpus; inputting the noise-added corpus to an encoder to obtain a feature vector; and inputting the feature vector to a decoder for decoding to obtain a newly generated corpus;

obtaining feature vectors of each pair of bilingual corpora and determining a second feature of each pair of bilingual corpora based on the trained language model;

determining a quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora; and

screening each pair of bilingual corpora according to the quality value of each pair of bilingual corpora.

18. The non-transitory computer-readable storage medium according to claim 17 , wherein the machine translation model comprises a first translation model and a second translation model, and each first feature comprises a first probability feature and a second probability feature; and

obtaining the first feature of each pair of bilingual corpora based on the trained machine translation model comprises:

inputting the source corpus in each pair of the bilingual corpora to a trained first translation model, and determining the first probability feature of the bilingual corpora based on a result output by the trained first translation model, wherein the first probability feature is a probability that the trained first translation model predicts the source corpus as the target corpus corresponding to the source corpus in the bilingual corpora, and

inputting the target corpus in each pair of the bilingual corpora to a trained second translation model, and determining the second probability feature of the bilingual corpora based on a result output by the trained second translation model, wherein the second probability feature is a probability that the trained second translation model predicts the target corpus as the source corpus corresponding to the target corpus in the bilingual corpora.

19. The non-transitory computer-readable storage medium according to claim 17 , wherein the language model comprises a first language model and a second language model, and each of the feature vectors comprises a first feature vector and a second feature vector; and

obtaining the feature vectors of each pair of bilingual corpora and determining the second feature of each pair of bilingual corpora based on the trained language model, comprises:

for each pair of the bilingual corpora, obtaining the first feature vector corresponding to the source corpus by inputting the source corpus in the pair of bilingual corpora to a trained first language model,

obtaining the second feature vector corresponding to the target corpus by inputting the target corpus in the pair of bilingual corpora to a trained second language model, and

determining a semantic similarity between the source corpus and the target corpus in the pair of bilingual corpora as the second feature of the pair of bilingual corpora based on the first feature vector and the second feature vector.

20. The non-transitory computer-readable storage medium according to claim 17 , wherein determining the quality value of each pair of bilingual corpora according to the first feature and the second feature of each pair of bilingual corpora comprises:

performing a weighted calculation on the first feature and the second feature of each pair of bilingual corpora to obtain the quality value of each pair of bilingual corpora.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 3, 2020
From: LI, JINGWEI; SUN, YUHUI; LI, XIANG
To: BEIJING XIAOMI MOBILE SOFTWARE CO., LTD.
Reel/Frame 052827/0229 →
Priority Claims (1)
CN 201911269664.7 · Dec 11, 2019 · national
Continuity (1)
Related Publication 20210182503A1 · Jun 17, 2021