Canonicalization of data within open knowledge graphs
Embodiments of the present invention provide computer-implemented methods, computer program products and computer systems. Embodiments of the present invention can, in response to receiving information, learn entity representations and cluster assignments of respective entity representations in a joint manner for both entities and relations of respective entities.
1 . A computer-implemented method comprising:
in response to receiving information, learning entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:
training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:
building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;
training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and
training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:
calculating a first score during the encoder section training and a second score during the decoder section training; and
calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.
2 . The computer-implemented method of claim 1 , wherein received information includes OpenIE triples obtained from a corpus of text.
3 . The computer-implemented method of claim 1 , wherein learning entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprises:
canonicalizing entity and relation mentions embedded in the received information.
4 . The computer-implemented method of claim 3 , wherein canonicalizing entity and relation mentions embedded in the received information comprises:
clustering entity mentions using a first variational auto-encoder for respective entities in the received information;
clustering relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and
leveraging structural knowledge present within an open knowledge base of mentions.
5 . The computer-implemented method of claim 3 , further comprising:
calculating an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.
6 . The computer-implemented method of claim 5 , further comprising:
encoding a constraint loss that enforces certain mention pairs to be within proximity of each other in an embedding space.
7 . A computer program product comprising:
one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions comprising:
program instructions to, in response to receiving information, learn entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:
training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:
building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;
training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and
training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:
calculating a first score during the encoder section training and a second score during the decoder section training; and
calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.
8 . The computer program product of claim 7 , wherein received information includes OpenIE triples obtained from a corpus of text.
9 . The computer program product of claim 7 , wherein the program instructions stored on the one or more computer readable storage media further comprise learn entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprise:
program instructions to canonicalize entity and relation mentions embedded in the received information.
10 . The computer program product of claim 9 , wherein the program instructions to canonicalize entity and relation mentions embedded in the received information comprise:
program instructions to cluster entity mentions using a first variational auto-encoder for respective entities in the received information;
program instructions to cluster relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and
program instructions to leverage structural knowledge present within an open knowledge base of mentions.
11 . The computer program product of claim 9 , program instructions stored on the one or more computer readable storage media further comprise:
program instructions to calculate an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.
12 . The computer program product of claim 11 , wherein the program instructions stored on the one or more computer readable storage media further comprise:
program instructions to encode a constraint loss that enforces certain mention pairs to be within proximity of each other in an embedding space.
13 . A computer system comprising:
one or more computer processors;
one or more computer readable storage media; and
program instructions stored on the one or more computer readable storage media for execution by at least one of the one or more computer processors, the program instructions comprising:
program instructions to, in response to receiving information, learn entity representations, and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities, wherein the learning comprises:
training a resulting neural network architecture using a variational autoencoder (VAE) loss function and a Knowledge Base Completion (KBC) loss function, the training comprising:
building a hierarchical agglomerative clustering model with complete linkage criterion using pretrained GloVe embeddings;
training an encoder section while keeping a decoder section fixed by using clustering results as a source of weak supervision using a constraint-based loss; and
training the decoder sections only while keeping the encoder sections fixed using the constraint-based loss, wherein the constraint-based loss comprises:
calculating a first score during the encoder section training and a second score during the decoder section training; and
calculating a weighted mean squared error between the first score and the second score that is L2 norm of a different of input embeddings.
14 . The computer system of claim 13 , wherein received information includes OpenIE triples obtained from a corpus of text.
15 . The computer system of claim 13 , wherein the program instructions stored on the one or more computer readable storage media further comprise learn entity representations and cluster assignments of respective entity representations in a joint manner for entities and relations of respective entities comprise:
program instructions to canonicalize entity and relation mentions embedded in the received information.
16 . The computer system of claim 15 , wherein the program instructions to canonicalize entity and relation mentions embedded in the received information comprise:
program instructions to cluster entity mentions using a first variational auto-encoder for respective entities in the received information;
program instructions to cluster relation mentions using a second variational auto-encoder for respective relation mentions in the received information, wherein the first variational auto-encoder and the second variational auto-encoder use a Mixture-of-Gaussians algorithm within its latent space; and
program instructions to leverage structural knowledge present within an open knowledge base of mentions.
17 . The computer system of claim 15 , program instructions stored on the one or more computer readable storage media further comprise:
program instructions to calculate an average of unit-normalized GloVe embeddings of tokens for multiple token terms based on the entity mentions and the relation mentions using the multiple token terms.