IP Library › Granted Patent US 12,632,724
Granted Patent B2
US 12,632,724 · App. 17/480,270 · Granted May 19, 2026

Canonicalization of data within open knowledge graphs

Inventors: Sarthak Dash (Jersey City, NJ); Gaetano Rossiello (Brooklyn, NY); Nandana Mihindukulasooriya (Cambridge, MA); Sugato Bagchi (White Plains, NY); Alfio Massimiliano Gliozzo (Brooklyn, NY)
Assignee: International Business Machines Corporation
G06N3/08G06N5/022G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,632,724
App. No.
17/480,270
Granted
May 19, 2026
Kind
B2
Abstract

Embodiments of the present invention provide computer-implemented methods, computer program products and computer systems. Embodiments of the present invention can, in response to receiving information, learn entity representations and cluster assignments of respective entity representations in a joint manner for both entities and relations of respective entities.

Claims (59)

1 . A computer-implemented method comprising:

in response to receiving information, learning entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:

training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:

building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;

training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and

training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:

calculating a first score during the encoder section training and a second score during the decoder section training; and

calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.

2 . The computer-implemented method of claim 1 , wherein received information includes OpenIE triples obtained from a corpus of text.

3 . The computer-implemented method of claim 1 , wherein learning entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprises:

canonicalizing entity and relation mentions embedded in the received information.

4 . The computer-implemented method of claim 3 , wherein canonicalizing entity and relation mentions embedded in the received information comprises:

clustering entity mentions using a first variational auto-encoder for respective entities in the received information;

clustering relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and

leveraging structural knowledge present within an open knowledge base of mentions.

5 . The computer-implemented method of claim 3 , further comprising:

calculating an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.

6 . The computer-implemented method of claim 5 , further comprising:

encoding a constraint loss that enforces certain mention pairs to be within proximity of each other in an embedding space.

7 . A computer program product comprising:

one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising:

program instructions to, in response to receiving information, learn entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:

training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:

building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;

training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and

training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:

calculating a first score during the encoder section training and a second score during the decoder section training; and

calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.

8 . The computer program product of claim 7 , wherein received information includes OpenIE triples obtained from a corpus of text.

9 . The computer program product of claim 7 , wherein the program instructions stored on the one or more computer readable storage media further comprise learn entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprise:

program instructions to canonicalize entity and relation mentions embedded in the received information.

10 . The computer program product of claim 9 , wherein the program instructions to canonicalize entity and relation mentions embedded in the received information comprise:

program instructions to cluster entity mentions using a first variational auto-encoder for respective entities in the received information;

program instructions to cluster relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and

program instructions to leverage structural knowledge present within an open knowledge base of mentions.

11 . The computer program product of claim 9 , program instructions stored on the one or more computer readable storage media further comprise:

program instructions to calculate an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.

12 . The computer program product of claim 11 , wherein the program instructions stored on the one or more computer readable storage media further comprise:

program instructions to encode a constraint loss that enforces certain mention pairs to be within proximity of each other in an embedding space.

13 . A computer system comprising:

one or more computer processors;

one or more computer readable storage media; and

program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising:

program instructions to, in response to receiving information, learn entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:

training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:

building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;

training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and

training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:

calculating a first score during the encoder section training and a second score during the decoder section training; and

calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.

14 . The computer system of claim 13 , wherein received information includes OpenIE triples obtained from a corpus of text.

15 . The computer system of claim 13 , wherein the program instructions stored on the one or more computer readable storage media further comprise learn entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprise:

program instructions to canonicalize entity and relation mentions embedded in the received information.

16 . The computer system of claim 15 , wherein the program instructions to canonicalize entity and relation mentions embedded in the received information comprise:

program instructions to cluster entity mentions using a first variational auto-encoder for respective entities in the received information;

program instructions to cluster relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and

program instructions to leverage structural knowledge present within an open knowledge base of mentions.

17 . The computer system of claim 15 , program instructions stored on the one or more computer readable storage media further comprise:

program instructions to calculate an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 21, 2021
From: DASH, SARTHAK; ROSSIELLO, GAETANO; MIHINDUKULASOORIYA, NANDANA; BAGCHI, SUGATO; GLIOZZO, ALFIO MASSIMILIANO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057541/0949 →
Continuity (1)
Related Publication 20230087667A1 · Mar 23, 2023
References Cited (27)
US 10127289B2 · Manning · 2018 [cited by applicant]
US 11830476B1 · Karanasou · 2023 [cited by examiner]
US 20190294732A1 · Srinivasan · 2019 [cited by applicant]
US 20210224610A1 · Jha · 2021 [cited by examiner]
US 20210357585A1 · Surdeanu · 2021 [cited by examiner]
US 20220300831A1 · Friede · 2022 [cited by examiner]
Luis Galarraga, et al. “Canonicalizing Open Knowledge Bases”, Nov. 3, 2014 (Year: 2014). [cited by examiner]
Quan Wang, Bin Wang, Li Guo, “Knowledge Base Completion Using Embeddings and Rules”, 2015 (Year: 2015). [cited by examiner]
Yan Liang, Xin Liu, Jianwen Zhang, Yangqiu Song, “Relation Discovery with Out-of-Relation Knowledge Base as Supervision”, Apr. 19, 2019 (Year: 2019). [cited by examiner]
Chaoyang Xu, Yuanfei Dai, Renjie Lin, Shiping Wang, “Deep clustering by maximizing mutual information in variational auto-encoder”, Jul. 16, 2020 (Year: 2020). [cited by examiner]
Zheng-Xin Yong, Tiago Timponi Torrent, “Semi-supervised Deep Embedded Clustering with Anomaly Detection for Semantic Frame Induction”, May 2020 (Year: 2020). [cited by examiner]
Pushpendre Rastogi, Adam Poliak, Vince Lyzinski, Benjamin Van Durme, “Neural variational entity set expansion for automatically populated knowledge graphs” Oct. 25, 2018 (Year: 2018). [cited by examiner]
Bovi et al., “Knowledge Base Unification via Sense Embeddings and Disambiguation”, Proceedings of the 2015 Conference on Empirical Methods in Natural Language Processing, 2015, 11 pages, <https://www.aclweb.org/antholog… [cited by applicant]
Broscheit et al., “Can We Predict New Facts with Open Knowledge Graph Embeddings? A Benchmark for Open Link Prediction”, Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Jul. 5-10… [cited by applicant]
Dash et al., “Joint Entity and Relation Canonicalization in Open Knowledge Graphs using Variational Autoencoders”, Grace Period Disclosure, arXiv:2012.04780v1 [cs.CL] Dec. 8, 2020, 11 pages. [cited by applicant]
Galarraga et al., “Canonicalizing Open Knowledge Bases”, Copyright 2014 ACM 978-1-4503-2598-1/14/11, 10 pages, <http://luisgalarraga.de/docs/km1298-galarraga.pdf>. [cited by applicant]
Gupta et al., “CaRe: Open Knowledge Graph Embeddings”, Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processingand the 9th International Joint Conference on Natural Language Processing, 201… [cited by applicant]
Jiang et al., “Canonicalizing Open Knowledge Bases with Multi-Layered Meta-Graph Neural Network”, arXiv:2006.09610v1 [cs.CL] Jun. 17, 2020, 7 pages, <https://arxiv.org/pdf/2006.09610.pdf>. [cited by applicant]
Jiang et al., “Variational Deep Embedding: An Unsupervised and Generative Approach to Clustering”, arXiv:1611.05148v3 [cs.CV] Jun. 28, 2017, 22 pages, <https://arxiv.org/pdf/1611.05148.pdf>. [cited by applicant]
Krishnamurthy et al., “Which Noun Phrases Denote Which Concepts?”, Proceedings of the 49th Annual Meeting of the Association for Computational Linguistics: Human Language Technologies, 2011, 11 pages, <https://www.aclwe… [cited by applicant]
Lin et al., “Canonicalization of Open Knowledge Bases with Side Information from the Source Text”, 2019 IEEE 35th International Conference on Data Engineering (ICDE), 2019, DOI 10.1109/ICDE.2019.00089, 12 pages. [cited by applicant]
Lin et al., “KBPearl: A Knowledge Base Population System Supported by Joint Entity and Relation Linking”, Proceedings of the VLDB Endowment 13, No. 7, 2020, 15 pages, <https://doi.org/10.14778/3384345.3384352>. [cited by applicant]
Pal et al., “Co-Clustering Triples from Open Information Extraction”, Proceedings of the 7th ACM IKDD CoDS and 25th COMAD, 2020, 5 pages, <https://doi.org/10.1145/3371158.3371183>. [cited by applicant]
Vashishth et al., “Cesi: Canonicalizing Open Knowledge Bases using Embeddings and Side Information”, Proceedings of the 2018 World Wide Web Conference. 2018, ACM ISBN 978-1-4503-5639-8/18/04, 10 pages, <http://malllabii… [cited by applicant]
Wu et al., “MULCE: Multi-level Canonicalization with Embeddings of Open Knowledge Bases”, International Conference on Web Information Systems Engineering, 2020, 13 pages. [cited by applicant]
Yates et al., “Unsupervised Methods for Determining Object and Relation Synonyms on the Web”, Journal of Artificial Intelligence Research 34 (2009): 255-296, Submitted Oct. 2008; published Mar. 2009, <https://www.aaai.o… [cited by applicant]
Zhu et al., “Iterative Entity Alignment via Joint Knowledge Embeddings”, IJCAI. vol. 17. 2017, Aug. 2017, 8 pages, <https://www.researchgate.net/profile/Hao_Zhu31/publication/318830326_Iterative_Entity_Alignment_via_Joi… [cited by applicant]