IP Library › Granted Patent US 12,572,828
Granted Patent B2
US 12,572,828 · App. 17/493,365 · Granted Mar 10, 2026

Method for industry text increment and electronic device

Inventors: Zhou Fang (Beijing, CN); Yabing Shi (Beijing, CN); Ye Jiang (Beijing, CN); Chunguang Chai (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
G06N5/04G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,572,828
App. No.
17/493,365
Granted
Mar 10, 2026
Kind
B2
Abstract

A method for an industry text increment, as well as an electronic device and a computer readable storage medium for the same are provided. The method may include: acquiring an original industry text in a target industry field, an order of magnitude of a number of the original industry text being smaller than a preset first order of magnitude; and performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts is greater than a preset second order of magnitude, wherein the preset second order of magnitude is not smaller than the preset first order of magnitude.

Claims (72)

1 . A method for an industry text increment, comprising:

acquiring an original industry text in a target industry field, an order of magnitude of a number of the original industry text being smaller than a preset first order of magnitude, wherein an industry text refers to a text content used to describe a specific object in a corresponding industry field; and

performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts being greater than a preset second order of magnitude, wherein the preset second order of magnitude is not smaller than the preset first order of magnitude, wherein the performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts is greater than a preset second order of magnitude, comprises:

performing a first sample incremental processing on the original industry text by using the distant supervision method, to obtain a first added industry text;

performing a second sample incremental processing on the original industry text and the first added industry text respectively by adopting a subject-object replacement method, to obtain a second added industry text, the subject-object replacement method comprising replacing an original subject and an original object with a new subject and a new object while maintaining a subject-object relationship provided by a predicate of a subject-predicate-object triple set; and

removing a text having a content error, a text having a logic error and a duplicate text from the first added industry text and the second added industry text, to obtain the increased industry texts, the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein if the order of magnitude of the number of the increased industry texts after the text having the content error, the text having the logic error and the duplicate text are removed, is not greater than the preset second order of magnitude, the sample incremental processing is performed on the increased industry texts again, until the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein the method further comprises:

training a language model based on the increased industry texts, and obtaining a trained language model;

extracting a subject-predicate-object triple set from an actual industry text by using the trained language model;

constructing a knowledge graph of a target industry field according to the extracted subject-predicate-object triple set, wherein the subject-predicate-object triple set comprises a plurality of subject-predicate-object triples; and

performing a query by using the knowledge graph constructed according to the subject-predicate-object triple set.

2 . The method according to claim 1 , wherein the performing a sample incremental processing on the original industry text by using a distant supervision method comprises:

extracting an initial subject-predicate-object triple set from the original industry text of the target industry field;

determining, in an another industry text of a non-target industry field and a public corpus, a text having a subject and a predicate of the initial subject-predicate-object triple set as a target text; and

using the target text as an added industry text of the original industry text distantly supervised.

3 . The method according to claim 1 , wherein the extracting a subject-predicate-object triple set from an actual industry text by using the trained language model comprises:

inputting a to-be-processed industry text into the trained language model to obtain an outputted text vector containing a context feature;

extracting a first result from the text vector by using a preset multi-pointer model, the multi-pointer model representing a corresponding relationship between the text vector and start and end positions of a relationship pair having a multi-layer nested relationship and existing in the text vector;

predicting a second result from the text vector by using a preset prediction sub-model, the prediction sub-model being used to predict at least one of a number of predicate categories, a number of subject-predicate-object triple sets or an entity type, contained in the to-be-processed industry text, according to a labeled label category; and

weighting the first result and the second result based on a preset model weighting coefficient, and extracting the subject-predicate-object triple set from an integrated result after the weighting.

4 . The method according to claim 1 , further comprising:

determining, in response to receiving a knowledge query request, an actual industry field to which the knowledge query request belongs according to the knowledge query request; and

invoking the knowledge graph of the actual industry field to perform a query, and feeding back target knowledge corresponding to the knowledge query request.

5 . An electronic device, comprising:

at least one processor; and

a storage device, communicated with the at least one processor,

wherein the storage device stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor, to enable the at least one processor to perform operations comprising:

acquiring an original industry text in a target industry field, an order of magnitude of a number of the original industry text being smaller than a preset first order of magnitude, wherein an industry text refers to a text content used to describe a specific object in a corresponding industry field; and

performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts being greater than a preset second order of magnitude, wherein the preset second order of magnitude is not smaller than the preset first order of magnitude, wherein the performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts is greater than a preset second order of magnitude, comprises:

performing a first sample incremental processing on the original industry text by using the distant supervision method, to obtain a first added industry text;

performing a second sample incremental processing on the original industry text and the first added industry text respectively by adopting a subject-object replacement method, to obtain a second added industry text, the subject-object replacement method comprising replacing an original subject and an original object with a new subject and a new object while maintaining a subject-object relationship provided by a predicate of a subject-predicate-object triple set; and

removing a text having a content error, a text having a logic error and a duplicate text from the first added industry text and the second added industry text, to obtain the increased industry texts, the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein if the order of magnitude of the number of the increased industry texts after the text having the content error, the text having the logic error and the duplicate text are removed, is not greater than the preset second order of magnitude, the sample incremental processing is performed on the increased industry texts again, until the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein the method further comprises:

training a language model based on the increased industry texts, and obtaining a trained language model;

extracting a subject-predicate-object triple set from an actual industry text by using the trained language model;

constructing a knowledge graph of a target industry field according to the extracted subject-predicate-object triple set, wherein the subject-predicate-object triple set comprises a plurality of subject-predicate-object triples; and

performing a query by using the knowledge graph constructed according to the subject-predicate-object triple set.

6 . The electronic device according to claim 5 , wherein the performing a sample incremental processing on the original industry text by using a distant supervision method comprises:

extracting an initial subject-predicate-object triple set from the original industry text of the target industry field;

determining, in another industry text of a non-target industry field and a public corpus, a text having a subject and a predicate of the initial subject-predicate-object triple set as a target text; and

using the target text as an added industry text of the original industry text distantly supervised.

7 . The electronic device according to claim 5 , wherein the extracting a subject-predicate-object triple set from an actual industry text by using the trained language model comprises:

inputting a to-be-processed industry text into the trained language model to obtain an outputted text vector containing a context feature;

extracting a first result from the text vector by using a preset multi-pointer model, the multi-pointer model representing a corresponding relationship between the text vector and start and end positions of a relationship pair having a multi-layer nested relationship and existing in the text vector;

predicting a second result from the text vector by using a preset prediction sub-model, the prediction sub-model being used to predict at least one of a number of predicate categories, a number of subject-predicate-object triple sets or an entity type, contained in the to-be-processed industry text, according to a labeled label category; and

weighting the first result and the second result based on a preset model weighting coefficient, and extracting the subject-predicate-object triple set from an integrated result after the weighting.

8 . The electronic device according to claim 5 , further comprising:

determining, in response to receiving a knowledge query request, an actual industry field to which the knowledge query request belongs according to the knowledge query request; and

invoking the knowledge graph of the actual industry field to perform a query, and feeding back target knowledge corresponding to the knowledge query request.

9 . A non-transitory computer readable storage medium, storing computer instructions, wherein the computer instructions, when executed by a computer, cause the computer to perform operations comprising:

acquiring an original industry text in a target industry field, an order of magnitude of a number of the original industry text being smaller than a preset first order of magnitude, wherein an industry text refers to a text content used to describe a specific object in a corresponding industry field; and

performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts being greater than a preset second order of magnitude, wherein the preset second order of magnitude is not smaller than the preset first order of magnitude, wherein the performing a sample incremental processing on the original industry text by using a distant supervision method, to obtain increased industry texts, an order of magnitude of a number of the increased industry texts is greater than a preset second order of magnitude, comprises:

performing a first sample incremental processing on the original industry text by using the distant supervision method, to obtain a first added industry text;

performing a second sample incremental processing on the original industry text and the first added industry text respectively by adopting a subject-object replacement method, to obtain a second added industry text, the subject-object replacement method comprising replacing an original subject and an original object with a new subject and a new object while maintaining a subject-object relationship provided by a predicate of a subject-predicate-object triple set; and

removing a text having a content error, a text having a logic error and a duplicate text from the first added industry text and the second added industry text, to obtain the increased industry texts, the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein if the order of magnitude of the number of the increased industry texts after the text having the content error, the text having the logic error and the duplicate text are removed, is not greater than the preset second order of magnitude, the sample incremental processing is performed on the increased industry texts again, until the order of magnitude of the number of the increased industry texts is greater than the preset second order of magnitude,

wherein the method further comprises:

training a language model based on the increased industry texts, and obtaining a trained language model;

extracting a subject-predicate-object triple set from an actual industry text by using the trained language model;

constructing a knowledge graph of a target industry field according to the extracted subject-predicate-object triple set, wherein the subject-predicate-object triple set comprises a plurality of subject-predicate-object triples; and

performing a query by using the knowledge graph constructed according to the subject-predicate-object triple set.

10 . The storage medium according to claim 9 , wherein the performing a sample incremental processing on the original industry text by using a distant supervision method comprises:

extracting an initial subject-predicate-object triple set from the original industry text of the target industry field;

determining, in another industry text of a non-target industry field and a public corpus, a text having a subject and a predicate of the initial subject-predicate-object triple set as a target text; and

using the target text as an added industry text of the original industry text distantly supervised.

11 . The storage medium according to claim 10 , wherein the extracting a subject-predicate-object triple set from an actual industry text by using the trained language model comprises:

inputting a to-be-processed industry text into the trained language model to obtain an outputted text vector containing a context feature;

extracting a first result from the text vector by using a preset multi-pointer model, the multi-pointer model representing a corresponding relationship between the text vector and start and end positions of a relationship pair having a multi-layer nested relationship and existing in the text vector;

predicting a second result from the text vector by using a preset prediction sub-model, the prediction sub-model being used to predict at least one of a number of predicate categories, a number of subject-predicate-object triple sets or an entity type, contained in the to-be-processed industry text, according to a labeled label category; and

weighting the first result and the second result based on a preset model weighting coefficient, and extracting the subject-predicate-object triple set from an integrated result after the weighting.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2025
From: SHI, YABING; CHAI, CHUNGUANG
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 072420/0729 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 30, 2025
From: BAIDU.COM TIMES TECHNOLOGY (BEIJING) CO., LTD.
To: BEIJING BAIDU NETCOM SCIENCE AND TECHNOLOGY CO., LTD.
Reel/Frame 072420/0910 →
EMPLOYMENT AGREEMENT Recorded Sep 30, 2025
From: FANG, ZHOU
To: BAIDU.COM TIMES TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 072989/0763 →
EMPLOYMENT AGREEMENT Recorded Sep 30, 2025
From: JIANG, YE
To: BAIDU.COM TIMES TECHNOLOGY (BEIJING) CO., LTD.
Reel/Frame 072989/0781 →
Priority Claims (1)
CN 202110189733.4 · Feb 19, 2021 · national
Continuity (1)
Related Publication 20220027766A1 · Jan 27, 2022
References Cited (23)
US 11734937B1 · Pushkin · 2023 [cited by examiner]
US 20160132357A1 · Kuraishi et al. · 2016 [cited by applicant]
US 20200372395A1 · Mahmud et al. · 2020 [cited by applicant]
US 20220198149A1 · Wu · 2022 [cited by examiner]
US 20220245362A1 · Nizar · 2022 [cited by examiner]
US 20230153526A1 · Wang · 2023 [cited by examiner]
CN 107145503A · 2017 [cited by applicant]
CN 108984683A · 2018 [cited by applicant]
CN 109086660A · 2018 [cited by applicant]
CN 109101583A · 2018 [cited by applicant]
CN 109284396A · 2019 [cited by applicant]
CN 111241813A · 2020 [cited by applicant]
CN 111339407A · 2020 [cited by applicant]
CN 111597795A · 2020 [cited by applicant]
CN 111651614A · 2020 [cited by applicant]
CN 111831788A · 2020 [cited by applicant]
KR 1020200094627A · 2020 [cited by applicant]
KR 1020200096133A · 2020 [cited by applicant]
Q. Wang, Z. Mao, B. Wang and L. Guo, “Knowledge Graph Embedding: A Survey of Approaches and Applications,” Dec. 1, 2017, in IEEE Transactions on Knowledge and Data Engineering, vol. 29, No. 12, pp. 2724-2743 (Year: 2017… [cited by examiner]
Chinese Office Action for Chinese Application No. 202110189733.4, dated May 11, 2022, 7 pages. [cited by applicant]
Extended European Search Report for European Application No. 21196648.6, dated Mar. 10, 2022, 28 pages. [cited by applicant]
Mintz et al., “Distant supervision for relation extraction without labeled data,” Proceedings of the 47 [cited by applicant]
Shleifer, “Low Resource Text Classification with ULMFit and Backtranslation,” Cornell University Library, 2018, 9 pages. [cited by applicant]
Cited By (1)
US 12,718,025