IP Library Granted Patent US 12,566,974
Granted Patent B2
US 12,566,974 · App. 17/495,214 · Granted Mar 3, 2026

Method, system, and computer program product for knowledge graph based embedding, explainability, and/or multi-task learning

Inventors: Adit Krishnan (San Francisco, CA); Mahashweta Das (Campbell, CA); Mangesh Bendre (Sunnyvale, CA); Azita Nouri (Highland Park, NJ); Fei Wang (Fremont, CA); Hao Yang (San Jose, CA)
Assignee: Visa International Service Association
G06N5/02G06F9/3851G06F18/214G06N5/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,566,974
App. No.
17/495,214
Filed
Oct 6, 2021
Granted
Mar 3, 2026
Kind
B2
Art Unit
2129
USPC
706/45
Abstract

Methods, systems, and computer program products for knowledge graph based embedding, explainability, and/or multi-task learning may connect task-specific inductive models with knowledge graph completion and enrichment processes.

Claims (166)

1 . A computer-implemented method comprising:

reading, with at least one processor, in parallel, with a plurality of threads, graph data associated with a plurality of edges and a plurality of nodes for the plurality of edges of a graph;

for each edge of the plurality of edges, with a thread of the plurality of threads that read the graph data associated with that edge, one of: (i) discarding that edge, (ii) sampling that edge to generate one or more samples and providing the one or more samples to a random queue of a plurality of queues, and (iii) oversampling that edge to generate the one or more samples and providing the one or more samples to the random queue of the plurality of queues, based on frequencies of nodes for that edge, wherein the plurality of queues corresponds to a plurality of groups of threads; and

training, with at least one other processor, in parallel, with the plurality of groups of threads, a plurality of embeddings of a plurality of samples provided to the plurality of queues,

wherein edges for lower frequent nodes having lower degrees in the graph than other nodes having higher degrees in the graph are collected in a shared memory that is accessible by each thread of the plurality of threads, and wherein, for each lower frequent node, each edge for that lower frequent node is represented in the training for a number of times proportional to a degree of that lower frequent node.

2 . The computer-implemented method of claim 1 , wherein at least one central processing unit (CPU) executes the plurality of threads, and wherein a different graphics processing unit (GPU) of a plurality of GPUs executes each group of threads of the plurality of groups of threads.

3 . The computer-implemented method of claim 1 , further comprising:

converting, with the at least one processor, using a hash table, the graph to the graph data associated with the plurality of edges of the graph, wherein the graph data associated with the plurality of edges of the graph includes frequencies of the plurality of edges and frequencies of nodes for the plurality of edges.

4 . The computer-implemented method of claim 1 , wherein the plurality of nodes includes a plurality of different types of nodes, and wherein for each edge of the plurality of edges, the thread of the plurality of threads that read the graph data associated with that edge determines to perform the one of: (i), (ii), and (iii) based on the frequencies of the nodes for that edge only with respect to frequencies of other nodes of a same type of node of the plurality of different types of nodes.

5 . The computer-implemented method of claim 1 , wherein providing the one or more samples to the random queue of the plurality of queues further includes:

for each of the one or more samples, determining, with the at least one processor, a distance between that sample and a negative sample including at least one of a different node than that sample, a same node having a different node type than the same node of that sample, and a same edge having a different edge type than the same edge of that sample; and

in response to determining that the distance satisfies a threshold distance, providing, with the at least one processor, the negative sample to the random queue of the plurality of queues.

6 . The computer-implemented method of claim 1 , wherein training, with the at least one other processor, in parallel, with the plurality of groups of threads, the plurality of embeddings includes:

for each queue of the plurality of queues, generating, with a group of threads of the plurality of groups of threads corresponding to that queue, node embeddings for two nodes and an edge embedding for an edge connecting the two nodes, based on samples provided to that queue, using an objective function that depends on embeddings of two nodes and an embedding of an edge connecting the two nodes; and

storing, in a shared memory, with each group of threads of the plurality of groups of threads, the embeddings generated by that group of threads.

7 . The computer-implemented method of claim 6 , further comprising:

for each queue of the plurality of queues:

determining, with the at least one processor, that (i) the queue is at full capacity and (ii) the group of threads corresponding to that queue is ready for a next batch of training samples; and

in response to determining that (i) the queue is at full capacity and (ii) the group of threads corresponding to that queue is ready for the next batch of training samples, copying, with the at least one processor, the samples provided to that queue from that queue to a memory of the group of threads corresponding to that queue.

8 . The computer-implemented method of claim 6 ,

wherein the objective function is defined according to the following Equation:

L

=

R

r

R

(

e

1

,

r

,

e

2

)

R

r

σ

(

r

(

e

1

e

2

)

)

where entity types are represented as E 1 , E 2 . . . E |ε| where ε={E 1 , E 2 . . . E |ε| } is a set of all entity types, R={R 1 , R 2 , . . . R |ε| } denotes a set of relations where each relation R r :E 1 r →E 2 r is a collection of links between two entity types E 1 r , E 2 r ∈ε, each edge is denoted as (e 1 , r, e 2 ) where e 1 ∈E 1 r , e 2 r ∈E 2 r denotes head and tail entities, respectively, and r is a relation type of the head and tail entities, {right arrow over (e)} 1 , {right arrow over (e)} 2 and r respectively denote an embedding representation of the two nodes and the connecting edge of the two nodes, σ(x)=1/(1+e −x ), and ⊗ is an outer product of the embeddings {right arrow over (e)} 1 , {right arrow over (e)} 2 .

9 . A system comprising:

at least one central processing unit (CPU) programmed and/or configured to:

read, in parallel, with a plurality of threads, graph data associated with a plurality of edges and a plurality of nodes for the plurality of edges of a graph; and

for each edge of the plurality of edges, with a thread of the plurality of threads that read the graph data associated with that edge, one of: (i) discard that edge, (ii) sample that edge to generate one or more samples and provide the one or more samples to a random queue of a plurality of queues, and (iii) oversample that edge to generate the one or more samples and provide the one or more samples to the random queue of the plurality of queues, based on frequencies of nodes for that edge, wherein the plurality of queues corresponds to a plurality of groups of threads; and

a plurality of graphics processing units (GPUs), wherein a different GPU of the plurality of GPUs executes each group of threads of the plurality of groups of threads, and wherein the plurality of GPUs are programmed and/or configured to:

train, in parallel, with the plurality of groups of threads, a plurality of embeddings of a plurality of samples provided to the plurality of queues,

wherein edges for lower frequent nodes having lower degrees in the graph than other nodes having higher degrees in the graph are collected in a shared memory that is accessible by each thread of the plurality of threads, and wherein, for each lower frequent node, each edge for that lower frequent node is represented in the training for a number of times proportional to a degree of that lower frequent node.

10 . The system of claim 9 , wherein the at least one CPU is further programmed and/or configured to:

convert, using a hash table, the graph to the graph data associated with the plurality of edges of the graph, wherein the graph data associated with the plurality of edges of the graph includes frequencies of the plurality of edges and frequencies of nodes for the plurality of edges.

11 . The system of claim 9 , wherein the plurality of nodes includes a plurality of different types of nodes, and wherein for each edge of the plurality of edges, the thread of the plurality of threads that read the graph data associated with that edge determines to perform the one of: (i), (ii), and (iii) based on the frequencies of the nodes for that edge only with respect to frequencies of other nodes of a same type of node of the plurality of different types of nodes.

12 . The system of claim 9 , wherein the at least one CPU is further programmed and/or configured to provide the one or more samples to the random queue of the plurality of queues by:

for each of the one or more samples, determine a distance between that sample and a negative sample including at least one of a different node than that sample, a same node having a different node type than the same node of that sample, and a same edge having a different edge type than the same edge of that sample; and

in response to determining that the distance satisfies a threshold distance, provide the negative sample to the random queue of the plurality of queues.

13 . The system of claim 9 , wherein the plurality of GPUs are programmed and/or configured to train, in parallel, with the plurality of groups of threads, the plurality of embeddings by:

for each queue of the plurality of queues, generating, with a group of threads of the plurality of groups of threads corresponding to that queue, node embeddings for two nodes and an edge embedding for an edge connecting the two nodes, based on samples provided to that queue, using an objective function that depends on embeddings of two nodes and an embedding of an edge connecting the two nodes; and

storing, in a shared memory, with each group of threads of the plurality of groups of threads, the embeddings generated by that group of threads.

14 . The system of claim 13 , wherein the at least one CPU is further programmed and/or configured to:

for each queue of the plurality of queues:

determine, that (i) the queue is at full capacity and (ii) the GPU of the group of threads corresponding to that queue is ready for a next batch of training samples; and

in response to determining that (i) the queue is at full capacity and (ii) the group of threads corresponding to that queue is ready for the next batch of training samples, copy the samples provided to that queue from that queue to a memory of the GPU of the group of threads corresponding to that queue.

15 . The system of claim 13 , wherein the objective function is defined according to the following Equation:

L

=

R

r

R

(

e

1

,

r

,

e

2

)

R

r

σ

(

r

(

e

1

e

2

)

)

where entity types are represented as E 1 , E 2 . . . E |ε| where ε={E 1 , E 2 . . . E |ε| } is a set of all entity types, R={R 1 , R 2 , . . . R |R| } denotes a set of relations where each relation R r :E 1 r →E 2 r is a collection of links between two entity types E 1 r , E 2 r ∈ε, each edge is denoted as (e 1 , r, e 2 ) where e 1 ∈E 1 r , e 2 ∈E 2 r denotes head and tail entities, respectively, and r is a relation type of the head and tail entities, {right arrow over (e)} 1 , {right arrow over (e)} 2 and r respectively denote an embedding representation of the two nodes and the connecting edge of the two nodes, σ(x)=1/(1+e −x ), and ⊗ is an outer product of the embeddings {right arrow over (e)} 1 , {right arrow over (e)} 2 .

16 . A computer program product comprising at least one non-transitory computer-readable medium including program instructions that, when executed by at least one processor, cause the at least one processor to:

read, in parallel, with a plurality of threads, graph data associated with a plurality of edges and a plurality of nodes for the plurality of edges of a graph;

for each edge of the plurality of edges, with a thread of the plurality of threads that read the graph data associated with that edge, one of: (i) discard that edge, (ii) sample that edge to generate one or more samples and provide the one or more samples to a random queue of a plurality of queues, and (iii) oversample that edge to generate the one or more samples and provide the one or more samples to the random queue of the plurality of queues, based on frequencies of nodes for that edge, wherein the plurality of queues corresponds to a plurality of groups of threads; and

train, in parallel, with the plurality of groups of threads, a plurality of embeddings of a plurality of samples provided to the plurality of queues,

wherein edges for lower frequent nodes having lower degrees in the graph than other nodes having higher degrees in the graph are collected in a shared memory that is accessible by each thread of the plurality of threads, and wherein, for each lower frequent node, each edge for that lower frequent node is represented in the training for a number of times proportional to a degree of that lower frequent node.

17 . The computer program product of claim 16 , wherein the at least one processor includes at least one central processing unit (CPU) and a plurality of graphics processing units (GPUs), wherein the at least one CPU executes the plurality of threads, and wherein a different GPU of the plurality of GPUs executes each group of threads of the plurality of groups of threads.

18 . The computer program product of claim 17 , wherein the plurality of groups of threads, train, in parallel, the plurality of embeddings by:

for each queue of the plurality of queues, generating, with a group of threads of the plurality of groups of threads corresponding to that queue, node embeddings for two nodes and an edge embedding for an edge connecting the two nodes, based on samples provided to that queue, using an objective function that depends on embeddings of two nodes and an embedding of an edge connecting the two nodes; and

storing, in a shared memory, with each group of threads of the plurality of groups of threads, the embeddings generated by that group of threads.

19 . The computer program product of claim 17 ,

wherein the program instructions, when executed by the at least one CPU, cause the at least one CPU to:

for each queue of the plurality of queues:

determine, that (i) the queue is at full capacity and (ii) the group of threads corresponding to that queue is ready for a next batch of training samples; and

in response to determining that (i) the queue is at full capacity and (ii) the group of threads corresponding to that queue is ready for the next batch of training samples, copy the samples provided to that queue from that queue to a memory of the group of threads corresponding to that queue.

20 . The computer program product of claim 17 ,

wherein an objective function is defined according to the following Equation:

L

=

R

r

R

(

e

1

,

r

,

e

2

)

R

r

σ

(

r

(

e

1

e

2

)

)

where entity types are represented as E 1 , E 2 . . . E |ε| where ε={E 1 , E 2 . . . E |ε| } is a set of all entity types, R={R 1 , R 2 , . . . R |R| } denotes a set of relations where each relation R r :E 1 r →E 2 r is a collection of links between two entity types E 1 r , E 2 r ∈ε, each edge is denoted as (e 1 , r, e 2 ) where e 1 ∈E 1 r , e 2 r ∈E 2 r denotes head and tail entities, respectively, and r is a relation type of the head and tail entities, {right arrow over (e)} 1 , {right arrow over (e)} 2 and r respectively denote an embedding representation of the two nodes and the connecting edge of the two nodes, σ(x)=1/(1+e −x ), and ⊗ is an outer product of the embeddings {right arrow over (e)} 1 , {right arrow over (e)} 2 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 16, 2022
From: KRISHNAN, ADIT; DAS, MAHASHWETA; BENDRE, MANGESH; NOURI, AZITA; WANG, FEI; YANG, HAO
To: VISA INTERNATIONAL SERVICE ASSOCIATION
Reel/Frame 058668/0001 →
Continuity (3)
Provisional Application 63092717 · Oct 16, 2020
Provisional Application 63089841 · Oct 9, 2020
Related Publication 20220114456A1 · Apr 14, 2022
References Cited (66)
US 11455512B1 · Al-Rfou' · 2022 [cited by examiner]
US 20150186464A1 · Seputis · 2015 [cited by examiner]
US 20200167426A1 · Scheideler · 2020 [cited by examiner]
US 20200210228A1 · Wu · 2020 [cited by examiner]
US 20210064959A1 · Gui · 2021 [cited by examiner]
US 20220012190A1 · Cheng · 2022 [cited by examiner]
Voudigari et al., “Rank Degree: An Efficient Algorithm for Graph Sampling,” 2016 IEEE/ACM International Conference on Advances in Social Networks Analysis and Mining (ASONAM) (Year: 2016). [cited by examiner]
Zhu et al., “Enhancing Stratified Graph Sampling Algorithms based on Approximate Degree Distribution,” 2018, arXiv:1801.04624 (Year: 2018). [cited by examiner]
Ying et al., “Graph Convolutional Neural Networks for Web-Scale Recommender Systems,” 2018, arXiv:1806.01973v1 (Year: 2018). [cited by examiner]
Zhang et al., “Degree-biased random walk for large-scale network embedding,” 2019, Future Generation Computer Systems, 100, 198-209 (Year: 2019). [cited by examiner]
Liu et al., “Principled Multilayer Network Embedding,” 2017, arXiv:1709.03551v3 (Year: 2017). [cited by examiner]
Hubler et al., “Metropolis Algorithms for Representative Subgraph Sampling,” 2008, Eighth IEEE International Conference on Data Mining (Year: 2008). [cited by examiner]
Bhatia et al., “An Efficient Algorithm for Sampling of a Single Large Graph,” 2018, Tenth International Conference on Contemporary Computing (Year: 2018). [cited by examiner]
Al et al., “Learning Heterogeneous Knowledge Base Embeddings for Explainable Recommendation”, Algorithms, 2018, vol. 11, No. 9, pp. 1-17. [cited by applicant]
Bordes et al., “Translating Embeddings for Modeling Multi-relational Data”, Advances in Neural Information Processing Systems (NIPS), 2018, pp. 1-9. [cited by applicant]
Cao et al., “Unifying Knowledge Graph Learning and Recommendation: Towards a Better Understanding of User Preferences”, International World Wide Web Conference (WWW), 2019, pp. 1-11. [cited by applicant]
Casale et al., “Gaussian Process Prior Variational Autoencoders”, Advances in Neural Information Processing Systems (NIPS), 2018, pp. 1-18. [cited by applicant]
Chandrahas et al., “Towards Understanding the Geometry of Knowledge Graph Embeddings”, ACL, 2018, pp. 122-131. [cited by applicant]
Chen et al., “Embedding Uncertain Knowledge Graphs”, Thirty-Third AAAI Conference on Artificial Intelligence, 2019, pp. 3363-3370. [cited by applicant]
Cheng et al., “Knowledge Graph-based Event Embedding Framework for Financial Quantitative Investments”, International Conference on Research and Development in Information Retrieval (SIGIR), 2020, pp. 2221-2230. [cited by applicant]
Daume, III, et al., “Domain Adaptation for Statistical Classifiers”, Journal of Artificial Intelligence Research 26, 2006, pp. 101-126. [cited by applicant]
Ernst et al., “KnowLife: a versatile approach for constructing a large knowledge graph for biomedical sciences”, BMC Bioinformatics, 2015, vol. 16, No. 157, pp. 1-13. [cited by applicant]
Foerster et al., “Counterfactual Multi-Agent Policy Gradients”, Thirty-Second AAAI Conference on Artificial Intelligence, 2018, pp. 2974-2982. [cited by applicant]
Guo et al., “Knowledge Graph Embedding with Iterative Guidance from Soft Rules”, Thirty-Second AAAI Conference on Artificial Intelligence, 2017, pp. 4816-4823. [cited by applicant]
He et al., “Translation-based Recommendation”, Eleventh ACM Conference on Recommender Systems, ACM, 2017, pp. 161-169. [cited by applicant]
Huang et al., “Knowledge Graph Embedding Based Question Answering”, International Conference on Web Search and Data Mining (WSDM), 2019, pp. 1-9. [cited by applicant]
Inokuchi et al., “Complete Mining of Frequent Patterns from Graphs: Mining Graph Data”, Journal of Machine earning vol. 50, No. 3, 2003, pp. 321-354. [cited by applicant]
Jagerman et al., “To Model or to Intervene: A Comparison of Counterfactual and Online Learning to Rank from User Interactions”, 42nd International ACM SIGIR Conference on Research and Development in Information Retrieva… [cited by applicant]
Ji et al., “Knowledge Graph Completion with Adaptive Sparse Transfer Matrix”, Thirtieth AAAI Conference on Artificial Intelligence, 2016, pp. 985-991. [cited by applicant]
Ji et al., “Knowledge Graph Embedding via Dynamic Mapping Matrix”, Association for Computational Linguistics and International Joint Conference on Natural Language Processing (ACL-IJCNLP), 2015, pp. 687-696. [cited by applicant]
Jia et al., “Locally Adaptive Translation for Knowledge Graph Embedding”, Thirtieth AAAI Conference on Artificial Intelligence, 2016, pp. 992-998. [cited by applicant]
Joachims et al., “Counterfactual Evaluation and Learning for Search, Recommendation and Ad Placement”, 39th International ACM SIGIR Conference on Research and Development in Information Retrieval, 2016, pp. 1199-1201. [cited by applicant]
Johansson et al., “Learning Representations for Counterfactual Inference”, 33rd International Conference on Machine Learning (ICML), 2016, pp. 1-10. [cited by applicant]
“Knowledge Graphs for Financial Services”, Deloitte, White Report, 2020, pp. 1-16. [cited by applicant]
Krishnan et al., “Insights from the Long-Tail: Learning Latent Representations of Online User Behavior in the Presence of Skew and Sparsity”, International Conference on Information and Knowledge Management (CIKM), 2018… [cited by applicant]
Krishnan et al., “Transfer Learning via Contextual Invariants for One-to-Many Cross-Domain Recommendation”, Association for Computing Machinery (ACM), 2020, pp. 1-10. [cited by applicant]
Kusner et al., “Counterfactual Fairness”, 31st Conference on Neural Information Processing Systems (NIPS), 2017, pp. 1-18. [cited by applicant]
Lerer et al., “PyTorch-BigGraph: A Large-scale Graph Embedding System”, 2nd SysML Conference, 2019, pp. 1-12. [cited by applicant]
Li et al., “Unit Selection Based on Counterfactual Logic”, International Joint Conferences on Artificial Intelligence (IJCAI), 2019, pp. 1-96. [cited by applicant]
Li, “Zipf's Law Everywhere”, Glottometrics 5, 2002, pp. 14-21. [cited by applicant]
Lin et al., “Learning Entity and Relation Embeddings for Knowledge Graph Completion”, Twenty-Ninth AAAI Conference on Artificial Intelligence, 2015, pp. 2181-2187. [cited by applicant]
Loughlin, “Why Knowledge Graph for Financial Services? Real World Use Cases”, Cambridge Semantics Blog, 2019, pp. 1-10. [cited by applicant]
Mansour et al., “Domain Adaptation: Learning Bounds and Algorithms”, 2009, pp. 1-16. [cited by applicant]
Mikolov et al., “Distributed Representations of Words and Phrases and their Compositionality”, NIPS, 2013, pp. 1-9. [cited by applicant]
Morgan et al., “Counterfactuals and Causal Inference”, Cambridge University Press, 2015, pp. 1-319. [cited by applicant]
Mothilal et al., “Explaining Machine Learning Classifiers through Diverse Counterfactual Explanations”, 2020 Conference on Fairness, Accountability, and Transparency, 2020, pp. 607-617. [cited by applicant]
Narita et al., “Efficient Counterfactual Learning from Bandit Feedback”, Thirty-Third AAAI Conference on Artificial Intelligence, 2019, pp. 4634-4641. [cited by applicant]
Nguyen et al., “A Capsule Network-based Embedding Model for Knowledge Graph Completion and Search Personalization”, North American Chapter of the Association for Computational Linguistics: Human Language Technologies (N… [cited by applicant]
Nickel et al., “An Analysis of Tensor Models for Learning on Structured Data”, European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2013, pp. 272-287. [cited by applicant]
Nickel et al., “Tensor Factorization for Multi-relational Learning”, European Conference on Machine Learning and Principles and Practice of Knowledge Discovery in Databases (ECML PKDD), 2018, pp. 617-621. [cited by applicant]
Ouyang et al., “Query Associations Over Big Financial Knowledge Graph”, Big SDM, 2019, pp. 199-211. [cited by applicant]
Pasricha et al., “Translation-based Factorization Machines for Sequential Recommendation”, International Conference on Recommender Systems (RecSys), 2018, pp. 1-9. [cited by applicant]
Pearl, “Causal Diagrams for Empirical Research”, Biometrika, 1994, vol. 82, No. 4, pp. 1-36. [cited by applicant]
Perozzi et al., “DeepWalk: Online Learning of Social Representations”, 20th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 2014, pp. 701-710. [cited by applicant]
Ren et al., “Query2box: Reasoning over Knowledge Graphs in Vector Space using Box Embeddings”, ICLR 2020, 2020, pp. 1-17. [cited by applicant]
Rubin, “Estimating Causal Effects of Treatments in Randomized And Nonrandomized Studies”, Journal of Educational Psychology, 1974, vol. 66, No. 5, pp. 688-701. [cited by applicant]
Sun et al., “Recurrent Knowledge Graph Embedding For Effective Recommendation”, International Conference on Recommender Systems (RecSys), 2018, pp. 297-305. [cited by applicant]
Sun et al., “Rotate: Knowledge Graph Embedding by Relational Rotation in Complex Space”, ICLR 2019, 2019, pp. 1-18. [cited by applicant]
Tang et al., “LINE: Large-scale Information Network Embedding”, International World Wide Web Conference Committee (IW3C2), 2015, pp. 1067-1077. [cited by applicant]
Trouillon et al., “Complex Embeddings for Simple Link Prediction”, 33rd International Conference on Machine Learning (ICML), 2016, pp. 1-10. [cited by applicant]
Wang et al., “Explainable Reasoning over Knowledge Graphs for Recommendation”, Thirty-Third AAAI Conference on Artificial Intelligence, 2019, pp. 5329-5336. [cited by applicant]
Wang et al., “KGAT: Knowledge Graph Attention Network for Recommendation”, International Conference on Knowledge Discovery Data Mining (SIGKDD), 2019, pp. 1-9. [cited by applicant]
Wang et al., “Knowledge Graph Embedding by Translating on Hyperplanes”, Twenty-Eighth AAAI Conference on Artificial Intelligence, 2014, pp. 1112-1119. [cited by applicant]
Wang et al., “XLore: A Largescale English-Chinese Bilingual Knowledge Graph”, International SemanticWeb Conference (ISWC), 2013, pp. 1-4. [cited by applicant]
Yang et al., “Embedding Entities and Relations for Learning and Inference in Knowledge Bases”, ICLR 2015, 2014, pp. 1-12. [cited by applicant]
Zalite, “Case study: The Practicalities of building an enterprise knowledge graph”, Data Management Summit in New York, Deutsche Bank, 2019, pp. 1-10. [cited by applicant]