IP Library Granted Patent US 12,645,876
Granted Patent B2
US 12,645,876 · App. 18/304,382 · Granted Jun 2, 2026

Auto-correcting framework for open information extraction systems

Inventors: Kiril Gashteovski (Heidelberg, DE); Ammar Shaker (Heidelberg, DE)
Assignee: NEC CORPORATION
G06F40/253
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,876
App. No.
18/304,382
Granted
Jun 2, 2026
Kind
B2
Abstract

A method for auto-correcting information extracted by an open information extraction system includes inputting an input triple extracted from an input sentence to a grammaticality model trained to output whether the input triple is grammatical or ungrammatical. The input triple is corrected based on the grammaticality model outputting that the input triple is ungrammatical. Embodiments of the present invention can be used in a variety of machine learning and artificial intelligence applications, including but not limited to several anticipated use cases in material informatics, data security, data extraction, and medical/healthcare, for example, for supporting decision making or optimizing and calibrating of extracted chemical compounds from libraries of textbooks and publications.

Claims (35)

1 . A method for auto-correcting information extracted by an open information extraction system, the method comprising:

inputting an input triple extracted from an input sentence to a grammaticality model trained to output whether the input triple is grammatical or ungrammatical; and

correcting the input triple based on the grammaticality model outputting that the input triple is ungrammatical.

2 . The method according to claim 1 , further comprising identifying which slot or slots of the input triple are ungrammatical.

3 . The method according to claim 2 , wherein the identifying which slot or slots of the input triple are ungrammatical is performed by computing a grammaticality score for each slot of the input triple and comparing the grammaticality scores of the slots to an overall grammaticality score computed for the input triple.

4 . The method according to claim 3 , further comprising generating a set of new triples by sampling from the input sentence and replacing the slot or slots of the input triple identified to be ungrammatical with candidate words or phrases given a fixed relation with slot or slots of the input triple that were not identified to be ungrammatical.

5 . The method according to claim 4 , wherein the input triple and the set of new triples are input to a similarity score model trained to determine similarity scores between the input triple and each of the new triples.

6 . The method according to claim 5 , wherein the input triple is corrected using one of the new triples determined by the similarity score model to be most similar to the input triple.

7 . The method according to claim 5 , wherein the similarity score model is trained by:

classifying triples extracted by the open information extraction system as being positive or negative; and

determining a shortest string similarity between each of the negative triples and the positive triples to determine one of the positive triples as being a correct triple for the respective negative triple, wherein the correct triples are used in training data for the similarity score model, the training data being in a form of an anchor, a positive and a random positive, wherein the anchor is the respective negative triple, the positive is the correct triple determined for the respective negative triple, and the random positive is a different one of the positive triples.

8 . The method according to claim 7 , wherein the positive triples are determined from annotated sentences used as input for the training, and wherein the negative triples are determined using an evaluation protocol and comparing the triples extracted by the open information extraction system to the positive triples.

9 . The method according to claim 8 , wherein the evaluation protocol indicates which slot or slots is/are erroneous for each of the negative triples, and wherein the shortest string similarity is determined between each of the erroneous slots and a same slot of the positive triples that have other slots that correspond to the respective negative triple.

10 . The method according to claim 9 , wherein the shortest string similarity is determined using a Levenshtein distance.

11 . The method according to claim 1 , wherein the grammaticality model is trained by:

computing abstract representations of unannotated sentences;

computing n-gram grammaticality scores of the abstract representations;

classifying triples extracted by the open information extraction system as being positive or negative;

computing abstract phrases for each of the positive and negative triples; and

using the n-gram grammaticality scores and the abstract phrases for each of the positive and negative triples as training data to a machine learning model.

12 . The method according to claim 11 , wherein the n-gram grammaticality scores are determined using perplexity measurements.

13 . The method according to claim 11 , wherein the machine learning model is a neural network classifier based on a long short-term memory or transformer-based neural model.

14 . A system for auto-correcting information extracted by an open information extraction system comprising one or more hardware processors which, alone or in combination, are configured to provide for execution of the following steps:

inputting an input triple extracted from an input sentence to a grammaticality model trained to output whether the input triple is grammatical or ungrammatical; and

correcting the input triple based on the grammaticality model outputting that the input triple is ungrammatical.

15 . A tangible, non-transitory computer-readable medium containing instructions which, upon being executed by one or more processors, provide for execution of the following steps:

inputting an input triple extracted from an input sentence to a grammaticality model trained to output whether the input triple is grammatical or ungrammatical; and

correcting the input triple based on the grammaticality model outputting that the input triple is ungrammatical.

16 . The method according to claim 1 , wherein the grammaticality model was trained based on n-gram grammaticality scores computed for unannotated sentences from medical reports, and wherein the input sentence is taken from one of the medical reports or a different medical report that is being processed by the open information extraction system.

17 . A method for training a grammaticality model for an open information extraction system, the method comprising:

computing abstract representations of unannotated sentences;

computing n-gram grammaticality scores of the abstract representations;

classifying triples extracted by the open information extraction system as being positive or negative;

computing abstract phrases for each of the positive and negative triples; and

using the n-gram grammaticality scores and the abstract phrases for each of the positive and negative triples as training data to a machine learning model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 28, 2026
From: NEC LABORATORIES EUROPE GMBH
To: NEC CORPORATION
Reel/Frame 074495/0362 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 21, 2023
From: GASHTEOVSKI, KIRIL; SHAKER, AMMAR
To: NEC LABORATORIES EUROPE GMBH
Reel/Frame 063397/0593 →
Continuity (2)
Provisional Application 63444276 · Feb 9, 2023
Related Publication 20240273288A1 · Aug 15, 2024
References Cited (29)
US 8954359B2 · Huang et al. · 2015 [cited by applicant]
US 9442917B2 · Dou et al. · 2016 [cited by applicant]
US 10482384B1 · Stoilos et al. · 2019 [cited by applicant]
US 10978056B1 · Challa · 2021 [cited by examiner]
US 11494420B2 · Chen et al. · 2022 [cited by applicant]
US 20190354544A1 · Hertz · 2019 [cited by examiner]
US 20200073933A1 · Zhao et al. · 2020 [cited by applicant]
US 20200380211A1 · Fan et al. · 2020 [cited by applicant]
US 20210081612A1 · Saito et al. · 2021 [cited by applicant]
US 20220156582A1 · Sengupta et al. · 2022 [cited by applicant]
US 20220300544A1 · Potter · 2022 [cited by examiner]
US 20220309254A1 · Kotnis et al. · 2022 [cited by applicant]
US 20220383867A1 · Faulkner · 2022 [cited by examiner]
JP 2018156332A · 2018 [cited by applicant]
JP 2019212289A · 2019 [cited by applicant]
WO WO2020177142A1 · 2020 [cited by applicant]
Abedini, Farhad, Mohammad Reza Keyvanpour, and Mohammad Bagher Menhaj. “Correction Tower: A general embedding method of the error recognition for the knowledge graph correction.” International Journal of Pattern Recogni… [cited by examiner]
Melo, André, and Heiko Paulheim. “An approach to correction of erroneous links in knowledge graphs.” CEUR Workshop Proceedings. vol. 2065. RWTH Aachen, 2017. (Year: 2017). [cited by examiner]
Caldeira, João, et al. “Profiling Software Developers with Process Mining and N-Gram Language Models.” arXiv preprint arXiv: 2101.06733 (2021). (Year: 2021). [cited by examiner]
Yao, Liang, Chengsheng Mao, and Yuan Luo. “KG-BERT: BERT for knowledge graph completion.” arXiv preprint arXiv:1909.03193 (2019). (Year: 2019). [cited by examiner]
Abedini, Farhad et al.; “Correction Tower: A General Embedding Method of the Error Recognition for the Knowledge Graph Correction”; [cited by applicant]
Pereira Gomes Costa, Hermani et al.; “Automatic Extraction and Validation of Lexical Ontologies from text”; MSC Thesis in Informatics Engineering; Oct. 7, 2011; pp. 1-124; LAP Lambert Academic Publishing; Saarbruecken, … [cited by applicant]
Del Corro, Luciano et al.; “ClausIE: Clause-Based Open Information Extraction”; [cited by applicant]
Devereux, Barry et al.; “Large-Scale Acquisition of Feature-Based Conceptual Representations from Textual Corpora”; [cited by applicant]
Fensel, Dieter et al.; “ [cited by applicant]
Gashteovski, Kiril et al.; “OPIEC: An Open Information Extraction Corpus”; [cited by applicant]
Gashteovski, Kiril et al.; “BenchIE: A Framework for Multi-Faceted Fact-Based Open Information Extraction Evaluation”; [cited by applicant]
Jiang, Zhengbao et al.; “Improving Open Information Extraction via Iterative Rank-Aware Learning”; [cited by applicant]
Nakamachi, Akifumi et al.; “Text Simplification with Reinforcement Learning using Supervised Rewards on Grammaticality, Meaning Preservation, and Simplicity”; [cited by applicant]