IP Library Granted Patent US 11,221,856
Granted Patent B2
US 11,221,856 · App. 15/993,688 · Granted Jan 11, 2022

Joint bootstrapping machine for text analysis

Inventor: Pankaj Gupta (Munich, DE)
Assignee: SIEMENS AKTIENGESELLSCHAFT
G06F9/4401G06F16/3344
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,221,856
App. No.
15/993,688
Granted
Jan 11, 2022
Kind
B2
Abstract

Present invention concerns a method of relation extraction from a text corpus, the method comprising extracting instances from the text corpus based on seeds, wherein the seeds include at least one set of template seeds and at least one set of entity seeds. The invention also pertains to related devices and methods.

Claims (24)

1. A method of relation extraction, the method comprising:

providing a text corpus, the text corpus including a plurality of entities represented by at least one of a word, term, and phrase, at least one entity pair including a set of entities of the plurality of entities, and a plurality of templates representing the context of the at least one entity pair,

extracting, by a processor of a computing system, instances from the text corpus based on seeds, wherein the seeds comprise at least one set of template seeds and at least one set of entity seeds, wherein instances are extracted iteratively using at least a first hop, a second hop, and a third hop, wherein the second hop uses instances extracted from the first hop as second seeds for the second hop, wherein the second hop generates a set of extractors based on the second seeds, wherein the third hop uses the set of extractors to generate candidate instances, and wherein the third hop includes performing a confidence computation for each candidate instance.

2. The method according to claim 1 , wherein instances are extracted based on a similarity metric.

3. The method according to claim 1 , wherein output instances are extracted from the candidate instances based on reliability of a respective extractor of the set of extractors.

4. The method according to claim 3 , wherein the output instances are fed back to augment the seeds.

5. The method according to claim 1 , wherein instances are extracted from the text corpus based on reliability determined for a cluster of instances.

6. The method according to claim 1 , wherein the seeds comprise positive seeds that fulfil a relation and negative seeds that do not fulfil the relation.

7. The method according to claim 1 , wherein the template comprises a first vector representing a context before a first entity of a respective entity pair, a second vector representing a context between the first entity of the respective entity pair and a second entity of the respective entity pair, and a third vector representing a context after the second entity of the respective entity pair.

8. The method according to claim 1 , wherein the template is typed.

9. A text analysis system for relation extraction from a text corpus, the system being adapted for extracting instances from the text corpus based on seeds, wherein the seeds comprise at least one set of template seeds and at least one set of entity seeds, wherein the text corpus includes a plurality of entities represented by at least one of a word, term, and phrase, at least one entity pair including a set of entities of the plurality of entities, and a plurality of templates representing the context of the at least one entity pair, and wherein the text analysis system is adapted to iteratively extract instances using at least a first hop, a second hop, and a third hop, wherein the second hop uses instances extracted from the first hop as second seeds, wherein the second hop generates a set of extractors based on the second seeds, wherein the third hop uses the set of extractors to generate candidate instances, and wherein the third hop includes performing a confidence computation for each candidate instance.

10. A system according to claim 9 , the system being adapted for extracting instances based on a similarity metric.

11. The system according to claim 9 , wherein output instances are extracted from the candidate instances based on reliability of a respective extractor of the set of extractors.

12. The system according to claim 9 , the system being adapted for extracting instances from the text corpus based on reliability determined for a cluster of instances.

13. The system according to claim 9 , wherein the output instances are fed back to augment the seeds.

14. The system according to claim 9 , wherein the seeds comprise positive seeds that fulfil a relation and negative seeds that do not fulfil the relation.

15. The system according to claim 9 , wherein the seeds comprise positive seeds that fulfil a relation and negative seeds that do not fulfil the relation.

16. The system according to claim 9 , wherein the template comprises a first vector representing a context before a first entity of a respective entity pair, a second vector representing a context between the first entity of the respective entity pair and a second entity of the respective entity pair, and a third vector representing a context after the second entity of the respective entity pair.

17. A computer program comprising instructions causing a processor of a computer system to perform and/or control a method of relation extraction from a text corpus, the method comprising:

extracting, by the processor, instances from the text corpus based on seeds, wherein the seeds comprise at least one set of template seeds and at least one set of entity seeds, wherein the text corpus includes a plurality of entities represented by at least one of a word, term, and phrase, at least one entity pair including a set of entities of the plurality of entities, and a plurality of templates representing the context of the at least one entity pair, wherein instances are extracted iteratively using at least a first hop, a second hop, and a third hop, wherein the second hop uses instances extracted from the first hop as second seeds, wherein the second hop generates a set of extractors based on the second seeds, wherein the third hop uses the set of extractors to generate candidate instances, and wherein the third hop includes performing a confidence computation for each candidate instance.

18. A storage medium storing a computer program according to claim 17 .

19. The computer program according to claim 17 , wherein the seeds comprise positive seeds that fulfil a relation and negative seeds that do not fulfil the relation.

20. The computer program according to claim 17 , wherein the template comprises a first vector representing a context before a first entity of a respective entity pair, a second vector representing a context between the first entity of the respective entity pair and a second entity of the respective entity pair, and a third vector representing a context after the second entity of the respective entity pair.

21. The computer program according to claim 17 , wherein the template is typed.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 11, 2026
From: SIEMENS AKTIENGESELLSCHAFT
To: DRIMCO GMBH
Reel/Frame 073761/0214 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 25, 2018
From: GUPTA, PANKAJ
To: SIEMENS AKTIENGESELLSCHAFT
Reel/Frame 046455/0441 →
Continuity (1)
Related Publication 20190370007A1 · Dec 5, 2019