IP Library › Granted Patent US 12,333,251
Granted Patent B2
US 12,333,251 · App. 17/954,900 · Granted Jun 17, 2025

Extracting triplets from text with relationship prediction matrix, entity prediction matrix, and alignment matrix

Inventors: Jiandong Sun (Beijing, CN); Yabing Shi (Beijing, CN); Ye Jiang (Beijing, CN); Chunguang Chai (Beijing, CN)
Assignee: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
G06F40/289G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,333,251
App. No.
17/954,900
Granted
Jun 17, 2025
Kind
B2
Abstract

Disclosed are an information extraction method, an electronic device and a readable storage medium, which relate to the field of artificial intelligence technologies, and particularly to the field of knowledge graph technologies. The information extraction method includes: acquiring to-be-processed text to obtain a semantic vector of each token in the to-be-processed text; generating a relationship prediction matrix, an entity prediction matrix and an alignment matrix according to each token in the to-be-processed text and the semantic vector of each token; and extracting a target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix, and taking the target triplet as an information extraction result of the to-be-processed text.

Claims (92)

1. An information extraction method, comprising:

acquiring to-be-processed text to obtain a semantic vector of each token in the to-be-processed text;

generating a relationship prediction matrix, an entity prediction matrix and an alignment matrix according to each token in the to-be-processed text and the semantic vector of each token; and

extracting a target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix, and taking the target triplet as an information extraction result of the to-be-processed text,

wherein the extracting the target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix comprises:

determining a subject start token and an object start token corresponding to a same relationship type according to the relationship prediction matrix;

determining an entity start token and an entity end token corresponding to a same entity type according to the entity prediction matrix;

determining an entity and an object corresponding to the same relationship type in the to-be-processed text according to the subject start token and the object start token corresponding to the same relationship type, as well as the entity start token and the entity end token corresponding to the same entity type;

combining each relationship type and the entity and the object corresponding to the relationship type to obtain at least one candidate triplet; and

selecting a triplet meeting a preset requirement from the at least one candidate triplet as the target triplet according to the alignment matrix.

2. The method according to claim 1 , wherein the generating the relationship prediction matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

acquiring at least one relationship type, each relationship type comprising a relationship subject type and a relationship object type;

taking the at least one relationship type as a row in the relationship prediction matrix, and taking each token in the to-be-processed text as a column in the relationship prediction matrix; and

obtaining values of different elements in the relationship prediction matrix according to the semantic vector of the token of each column and the relationship type of each row.

3. The method according to claim 2 , wherein the obtaining values of different elements in the relationship prediction matrix according to the semantic vector of the token of each column and the relationship type of each row comprises:

for each element in the relationship prediction matrix, determining a token and a relationship type corresponding to the element;

performing calculation according to the semantic vector of the determined token and the determined relationship type to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a first preset threshold, setting the value of the element to 1.

4. The method according to claim 1 , wherein the generating the entity prediction matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

acquiring at least one entity type, each entity type comprising an entity start type and an entity end type;

taking the at least one entity type as a row in the entity prediction matrix, and taking each token in the to-be-processed text as a column in the entity prediction matrix; and

obtaining values of different elements in the entity prediction matrix according to the semantic vector of the token of each column and the entity type of each row.

5. The method according to claim 4 , wherein the obtaining values of different elements in the entity prediction matrix according to the semantic vector of the token of each column and the entity type of each row comprises:

for each element in the entity prediction matrix, determining a token and an entity type corresponding to the element;

performing calculation according to the semantic vector of the determined token and the determined entity type to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a second preset threshold, setting the value of the element to 1.

6. The method according to claim 1 , wherein the generating the alignment matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

taking each token in the to-be-processed text both as a row and a column in the alignment matrix; and

obtaining values of different elements in the alignment matrix according to the semantic vector of the token of each column and the semantic vector of the token of each row.

7. The method according to claim 6 , wherein the obtaining values of different elements in the alignment matrix according to the semantic vector of the token of each column and the semantic vector of the token of each row comprises:

for each element in the alignment matrix, determining a row token and a column token corresponding to the element;

performing calculation according to a semantic vector of the determined row token and a semantic vector of the determined column token to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a third preset threshold, setting the value of the element to 1.

8. The method according to claim 1 , wherein the determining the subject start token and the object start token corresponding to the same relationship type according to the relationship prediction matrix comprises:

taking an element with the value of 1 in the relationship prediction matrix as a target element; and

taking a token of a column where the target element is located as a start token of a relationship subject type or a start token of a relationship object type of a row where the target element is located.

9. The method according to claim 1 , wherein the determining the entity start token and the entity end token corresponding to the same entity type according to the entity prediction matrix comprises:

taking an element with the value of 1 in the entity prediction matrix as a target element; and

taking a token of a column where the target element is located as a start token of an entity start type or an end token of an entity end type of a row where the target element is located.

10. The method according to claim 1 , wherein the selecting a triplet meeting a preset requirement from the at least one candidate triplet as the target triplet according to the alignment matrix comprises:

for each candidate triplet, under a condition that a subject end token and an object end token in the candidate triplet are determined to have element values of 1 in the alignment matrix, taking the candidate triplet as the target triplet.

11. An electronic device, comprising:

at least one processor; and

a memory connected with the at least one processor communicatively;

wherein the memory stores instructions executable by the at least one processor to enable the at least one processor to perform an information extraction method comprising:

acquiring to-be-processed text to obtain a semantic vector of each token in the to-be-processed text;

generating a relationship prediction matrix, an entity prediction matrix and an alignment matrix according to each token in the to-be-processed text and the semantic vector of each token; and

extracting a target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix, and taking the target triplet as an information extraction result of the to-be-processed text,

wherein the extracting the target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix comprises:

determining a subject start token and an object start token corresponding to a same relationship type according to the relationship prediction matrix;

determining an entity start token and an entity end token corresponding to a same entity type according to the entity prediction matrix;

determining an entity and an object corresponding to the same relationship type in the to-be-processed text according to the subject start token and the object start token corresponding to the same relationship type, as well as the entity start token and the entity end token corresponding to the same entity type;

combining each relationship type and the entity and the object corresponding to the relationship type to obtain at least one candidate triplet; and

selecting a triplet meeting a preset requirement from the at least one candidate triplet as the target triplet according to the alignment matrix.

12. The electronic device according to claim 11 , wherein the generating the relationship prediction matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

acquiring at least one relationship type, each relationship type comprising a relationship subject type and a relationship object type;

taking the at least one relationship type as a row in the relationship prediction matrix, and taking each token in the to-be-processed text as a column in the relationship prediction matrix; and

obtaining values of different elements in the relationship prediction matrix according to the semantic vector of the token of each column and the relationship type of each row, comprising:

for each element in the relationship prediction matrix, determining a token and a relationship type corresponding to the element;

performing calculation according to the semantic vector of the determined token and the determined relationship type to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a first preset threshold, setting the value of the element to 1.

13. The electronic device according to claim 11 , wherein the generating the entity prediction matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

acquiring at least one entity type, each entity type comprising an entity start type and an entity end type;

taking the at least one entity type as a row in the entity prediction matrix, and taking each token in the to-be-processed text as a column in the entity prediction matrix; and

obtaining values of different elements in the entity prediction matrix according to the semantic vector of the token of each column and the entity type of each row comprising:

for each element in the entity prediction matrix, determining a token and an entity type corresponding to the element;

performing calculation according to the semantic vector of the determined token and the determined entity type to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a second preset threshold, setting the value of the element to 1.

14. The electronic device according to claim 11 , wherein the generating the alignment matrix according to each token in the to-be-processed text and the semantic vector of each token comprises:

taking each token in the to-be-processed text both as a row and a column in the alignment matrix; and

obtaining values of different elements in the alignment matrix according to the semantic vector of the token of each column and the semantic vector of the token of each row, comprising:

for each element in the alignment matrix, determining a row token and a column token corresponding to the element;

performing calculation according to a semantic vector of the determined row token and a semantic vector of the determined column token to obtain a calculation result of the element; and

under a condition that the calculation result is determined to exceed a third preset threshold, setting the value of the element to 1.

15. The electronic device according to claim 11 , wherein the determining the subject start token and the object start token corresponding to the same relationship type according to the relationship prediction matrix comprises:

taking an element with the value of 1 in the relationship prediction matrix as a target element; and

taking a token of a column where the target element is located as a start token of a relationship subject type or a start token of a relationship object type of a row where the target element is located.

16. The electronic device according to claim 11 , wherein the determining the entity start token and the entity end token corresponding to the same entity type according to the entity prediction matrix comprises:

taking an element with the value of 1 in the entity prediction matrix as a target element; and

taking a token of a column where the target element is located as a start token of an entity start type or an end token of an entity end type of a row where the target element is located.

17. The electronic device according to claim 11 , wherein the selecting a triplet meeting a preset requirement from the at least one candidate triplet as the target triplet according to the alignment matrix comprises:

for each candidate triplet, under a condition that a subject end token and an object end token in the candidate triplet are determined to have element values of 1 in the alignment matrix, taking the candidate triplet as the target triplet.

18. A non-transitory computer readable storage medium with computer instructions stored thereon, wherein the computer instructions are used for causing a computer to perform an information extraction method comprising:

acquiring to-be-processed text to obtain a semantic vector of each token in the to-be-processed text;

generating a relationship prediction matrix, an entity prediction matrix and an alignment matrix according to each token in the to-be-processed text and the semantic vector of each token; and

extracting a target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix, and taking the target triplet as an information extraction result of the to-be-processed text,

wherein the extracting the target triplet in the to-be-processed text using the relationship prediction matrix, the entity prediction matrix and the alignment matrix comprises:

determining a subject start token and an object start token corresponding to a same relationship type according to the relationship prediction matrix;

determining an entity start token and an entity end token corresponding to a same entity type according to the entity prediction matrix;

determining an entity and an object corresponding to the same relationship type in the to-be-processed text according to the subject start token and the object start token corresponding to the same relationship type, as well as the entity start token and the entity end token corresponding to the same entity type;

combining each relationship type and the entity and the object corresponding to the relationship type to obtain at least one candidate triplet; and

selecting a triplet meeting a preset requirement from the at least one candidate triplet as the target triplet according to the alignment matrix.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 29, 2022
From: SUN, JIANDONG; SHI, YABING; JIANG, YE; CHAI, CHUNGUANG
To: BEIJING BAIDU NETCOM SCIENCE TECHNOLOGY CO., LTD.
Reel/Frame 061250/0041 →
Priority Claims (1)
CN 202111300797.3 · Nov 4, 2021 · national
Continuity (1)
Related Publication 20230133717A1 · May 4, 2023
References Cited (28)
US 10664660B2 · Li · 2020 [cited by examiner]
US 10977282B2 · Nitta · 2021 [cited by examiner]
US 11151179B2 · Li · 2021 [cited by examiner]
US 20070282814A1 · Gupta · 2007 [cited by examiner]
US 20120323558A1 · Nolan · 2012 [cited by examiner]
US 20190122111A1 · Min · 2019 [cited by examiner]
US 20200073933A1 · Zhao · 2020 [cited by examiner]
US 20210216819A1 · He · 2021 [cited by examiner]
US 20210241050A1 · Gunaratna · 2021 [cited by examiner]
US 20220027766A1 · Fang · 2022 [cited by examiner]
US 20220309254A1 · Kotnis · 2022 [cited by examiner]
US 20230016403A1 · Wang · 2023 [cited by examiner]
US 20230087667A1 · Dash · 2023 [cited by examiner]
US 20230103728A1 · Liu · 2023 [cited by examiner]
CN 110310721A · 2019 [cited by applicant]
CN 111008276A · 2020 [cited by applicant]
CN 111444305A · 2020 [cited by applicant]
CN 113568969A1 · 2021 [cited by applicant]
CN 113590784A · 2021 [cited by applicant]
WO 2021135910A1 · 2021 [cited by applicant]
WO WO2021147726A1 · 2021 [cited by examiner]
WO 2021212682A1 · 2021 [cited by applicant]
Lin et al., “Relation Extraction Based on Label Constraints”, 2020 IEEE 6th International Conference on Computer and Communications (ICCC), Dec. 11-14, 2020, pp. 2166 to 2170. (Year: 2020). [cited by examiner]
Sun et al., Cross-lingual Entity Alignment via Joint Attribute-Preserving Embedding, arXiv:1708.05045v2 [cs. CL] Sep. 26, 2017, 16 pages. [cited by applicant]
Extended European Search Report of European patent application No. 22200126.5 dated Mar. 24, 2023, 5 pages. [cited by applicant]
Zheng, et al., PRGC: Potential Relation and Global Correspondence Based Joint Relational Triple Extraction, arXiv:2106.09895v1 [cs.CL] Jun. 18, 2021, 11 pages. [cited by applicant]
Pang et al., A Deep Neural Network Model for Joint Entity and Relation Extraction, IEEE Access, vol. 7, 179143 to 79150, Oct. 23, 2019. [cited by applicant]
Sun et al., Chinese Entity Relation Extraction Algorithms Based on COAE2016 Datasets, Journal of Shandong University (Natural Science), vol. 52, No. 9, Sep. 2017, 8 pages. [cited by applicant]