IP Library › Granted Patent US 12,541,706
Granted Patent B2
US 12,541,706 · App. 17/150,343 · Granted Feb 3, 2026

System and method for out-of-sample representation learning

Inventors: Marjan Albooyeh (Montreal, CA); Seyed Mehran Kazemi (Montreal, CA)
Assignee: ROYAL BANK OF CANADA
G06N20/00G06F16/9024G06F7/58G06N3/045G06N5/022
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,541,706
App. No.
17/150,343
Granted
Feb 3, 2026
Kind
B2
Abstract

Disclosed are systems, methods, and devices for out-of-sample representation learning using knowledge graphs. An embedding data structure reflective of a knowledge graph embedding model is received. A training data set including a plurality of training data entries, each of the training data entries reflective of a head entity, a tail entity, and a relation therebetween, wherein at least one of the head entities or the tail entities includes an out-of-sample entity, is received. A plurality of knowledge graph embedding model processors is provided. A random number is generated and compared to at least one criterion. A knowledge graph embedding model processor is selected from among the plurality of knowledge graph embedding model processors based at least in part on the comparing. The embedding data structure is processed with the selected knowledge graph embedding model processor.

Claims (67)

1 . A computer-implemented method for out-of-sample representation learning using non-attributed knowledge graphs to make predictions about out-of-sample entities at inference time, said method comprising:

receiving an embedding data structure reflective of a non-attributed knowledge graph embedding model;

receiving a training data set including a plurality of training data entries, each of said training data entries reflective of a head entity, a tail entity, and a relation therebetween, wherein at least one of said head entities or said tail entities includes an out-of-sample entity;

providing a plurality of knowledge graph embedding model processors;

for a given training data entry of said plurality of training data entries, said given training data entry reflective of a given head entity, a given tail entity, and a given relation:

with stochastic frequency ψ/2, selecting a first knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the first knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, a tail embedding for said given tail entity; and

generate a head embedding for said given head entity using an aggregation function with no learnable parameters, wherein the aggregation function provides the head embedding at inference time for said given head entity based on embeddings from said knowledge graph embedding model data structure;

with stochastic frequency ψ/2, selecting a second knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the second knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, the head embedding for said given head entity; and

generate the tail embedding for said given tail entity using the aggregation function with no learnable parameters, wherein the aggregation function provides the tail embedding at inference time for said given tail entity based on embeddings from said knowledge graph embedding model data structure,

wherein ψ encourages learning embeddings for predictions of out-of-sample entities; and

retrieving, from said knowledge graph embedding model data structure, a relation embedding for said given relation; and

upon processing said head embedding, said tail embedding, and said relation embedding, generating a score reflective of a degree of belief that said given relation holds between said given head entity and said given tail entity.

2 . The computer-implemented method of claim 1 , wherein the plurality of knowledge graph embedding model processors includes at least three knowledge graph embedding model processors.

3 . The computer-implemented method of claim 2 , further comprising:

with stochastic frequency 1-ψ, selecting a third knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the third knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, the head embedding for said given head entity and the tail embedding for said given tail entity.

4 . The computer-implemented method of claim 1 , further comprising generating a loss according to said score.

5 . The computer-implemented method of claim 4 , further comprising updating said knowledge graph embedding model data structure based on said loss.

6 . The computer-implemented method of claim 5 , wherein said updating includes computing a loss gradient.

7 . The computer-implemented method of claim 5 , wherein said updating includes applying gradient descent.

8 . The computer-implemented method of claim 1 , further comprising: repeating said generating for a plurality of data training data entries of said training data set.

9 . The computer-implemented method of claim 8 , wherein said repeating comprises processing said data training data entries in batches.

10 . The computer-implemented method of claim 1 , wherein said aggregation function provides an average of the embeddings from said knowledge graph embedding model data structure.

11 . The computer-implemented method of claim 1 , wherein said aggregation function provides a solution to a least squares problem to maximize said score using said head embedding, said tail embedding, and said relation embedding.

12 . A computer-implemented system for out-of-sample representation learning using non-attributed knowledge graphs to make predictions about out-of-sample entities at inference time, said system comprising:

at least one processor;

memory in communication with said at least one processor, and

software code stored in said memory, which when executed by said at least one processor causes said system to:

receive an embedding data structure reflective of a non-attributed knowledge graph embedding model;

receive a training data set including a plurality of training data entries, each of said training data entries reflective of a head entity, a tail entity, and a relation therebetween, wherein at least one of said head entities or said tail entities includes an out-of-sample entity;

provide a plurality of knowledge graph embedding model processors;

for a given training data entry of said plurality of training data entries, said given training data entry reflective of a given head entity, a given tail entity, and a given relation:

with stochastic frequency ψ/2, select a first knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the first knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, a tail embedding for said given tail entity; and

generate a head embedding for said given head entity using an aggregation function with no learnable parameters, wherein the aggregation function provides the head embedding at inference time for said given head entity based on embeddings from said knowledge graph embedding model data structure;

with stochastic frequency ψ/2, select a second knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the second knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, the head embedding for said given head entity; and

generate the tail embedding for said given tail entity using the aggregation function with no learnable parameters, wherein the aggregation function provides the tail embedding at inference time for said given tail entity based on embeddings from said knowledge graph embedding model data structure,

wherein ψ encourages learning embeddings for predictions of out-of-sample entities; and

retrieve, from said knowledge graph embedding model data structure, a relation embedding for said given relation; and

upon processing said head embedding, said tail embedding, and said relation embedding, generate a score reflective of a degree of belief that said given relation holds between said given head entity and said given tail entity.

13 . The computer-implemented system of claim 12 , wherein the plurality of knowledge graph embedding model processors includes at least three knowledge graph embedding model processors.

14 . The computer-implemented system of claim 13 , wherein said software code stored in said memory, when executed by said at least one processor further causes said system to:

with stochastic frequency 1-ψ, select a third knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the third knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, the head embedding for said given head entity and the tail embedding for said given tail entity.

15 . The computer-implemented system of claim 12 , wherein said software code stored in said memory, when executed by said at least one processor further causes said system to:

calculate a loss according to said score.

16 . The computer-implemented system of claim 15 , wherein said software code stored in said memory, when executed by said at least one processor further causes said system to:

update said knowledge graph embedding model data structure based on said loss.

17 . The computer-implemented system of claim 12 , wherein said aggregation function provides an average of the embeddings from said knowledge graph embedding model data structure.

18 . The computer-implemented system of claim 12 , wherein said aggregation function provides a solution to a least squares problem to maximize said score using said head embedding, said tail embedding, and said relation embedding.

19 . A non-transitory computer-readable medium having stored thereon machine interpretable instructions which, when executed by a processor, cause the processor to perform a computer implemented method for out-of-sample representation learning using non-attributed knowledge graphs to make predictions about out-of-sample entities at inference time, said method comprising:

receiving an embedding data structure reflective of a non-attributed knowledge graph embedding model;

receiving a training data set including a plurality of training data entries, each of said training data entries reflective of a head entity, a tail entity, and a relation therebetween, wherein at least one of said head entities or said tail entities includes an out-of-sample entity;

providing a plurality of knowledge graph embedding model processors;

for a given training data entry of said plurality of training data entries, said given training data entry reflective of a given head entity, a given tail entity, and a given relation:

with stochastic frequency ψ/2, selecting a first knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the first knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, a tail embedding for said given tail entity; and

generate a head embedding for said given head entity using an aggregation function with no learnable parameters, wherein the aggregation function provides the head embedding at inference time for said given head entity based on embeddings from said knowledge graph embedding model data structure;

with stochastic frequency ψ/2, selecting a second knowledge graph embedding model processor from among the plurality of knowledge graph embedding model processors to process the knowledge graph embedding model data structure, wherein the second knowledge graph embedding model processor is configured to:

retrieve, from said knowledge graph embedding model data structure, the head embedding for said given head entity; and

generate the tail embedding for said given tail entity using the aggregation function with no learnable parameters, wherein the aggregation function provides the tail embedding at inference time for said given tail entity based on embeddings from said knowledge graph embedding model data structure,

wherein ψ encourages learning embeddings for predictions of out-of-sample entities; and

retrieving, from said knowledge graph embedding model data structure, a relation embedding for said given relation; and

upon processing said head embedding, said tail embedding, and said relation embedding, generating a score reflective of a degree of belief that said given relation holds between said given head entity and said given tail entity.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2021
From: ALBOOYEH, MARJAN; KAZEMI, SEYED MEHRAN
To: ROYAL BANK OF CANADA
Reel/Frame 055365/0071 →
Continuity (2)
Provisional Application 62963591 · Jan 21, 2020
Related Publication 20210224690A1 · Jul 22, 2021
References Cited (41)
US 11164078B2 · Jin · 2021 [cited by examiner]
US 20070255543A1 · Chatfield · 2007 [cited by examiner]
US 20170039188A1 · Allen · 2017 [cited by examiner]
WO WO2011059537A1 · 2011 [cited by examiner]
Grégoire Montavon, Geneviève B. Orr, Klaus-Robert Müller, Neural Networks: Tricks of the Trade. Springer, Sep. 2012, Second edition. (Year: 2012). [cited by examiner]
Takuo Hamaguchi, Hidekazu Oiwa, Masashi Shimbo, Yuji Matsumoto; Knowledge Transfer for Out-of-Knowledge-Base Entities: A Graph Neural Network Approach,https://arxiv.org/abs/1706.05674, Jun. 2017 (Year: 2017). [cited by examiner]
James Bergstra, Yoshua Bengio, “Random Search for Hyper-Parameter Optimization”, JMLR, https://www.jmlr.org/papers/volume13/bergstra12a/bergstra12a.pdf, Feb. 2012, (Year: 2012). [cited by examiner]
Bollacker et al.; Freebase: A Collaboratively Created Graph Database For Structuring Human Knowledge; SIGMOD '08: Proceedings of the ACM SIGMOD International Conference on Management of Data, Jun. 2008. [cited by applicant]
Bordes et al.; Translating Embeddings for Modeling Multi-relational Data; NIPS'13: Proceedings of the 26th International Conference on Neural Information Processing Systems—vol. 2, Dec. 2013. [cited by applicant]
Chen et al; FastGCN: Fast Learning with Graph Convolutional Networks via Importance Sampling; International Conference on Learning Representations (ICLR 2018)—arXiv:1801.10247v1. [cited by applicant]
Defferrard et al.; Convolutional Neural Networks on Graphs with Fast Localized Spectral Filtering; NIPS'16: Proceedings of the 30th International Conference on Neural Information Processing Systems, Dec. 2016. [cited by applicant]
Dettmers et al.; Convolutional 2D Knowledge Graph Embeddings; Proceedings of the 32th AAAI Conference on Artificial Intelligence, 2018, pp. 1811-1818. [cited by applicant]
Duchi et al.; Adaptive Subgradient Methods for Online Learning and Stochastic Optimization; Journal of Machine Learning Research, vol. 12, pp. 2121-2159, Jul. 2011. [cited by applicant]
Hamilton et al.; Inductive Representation Learning on Large Graphs; NIPS'17: Proceedings of the 31st International Conference on Neural Information Processing Systems, Dec. 2017. [cited by applicant]
Hamilton et al.; Representation Learning on Graphs: Methods and Applications; IEEE Data Engineering Bulletin, vol. 40, pp. 52-74, 2017. [cited by applicant]
Hammond et al.; Wavelets on graphs via spectral graph theory; Applied and Computational Harmonic Analysis, vol. 30, Issue 2, Mar. 2011, pp. 129-150. [cited by applicant]
Ji et al.; Knowledge Graph Embedding via Dynamic Mapping Matrix; Proceedings of the 53rd Annual Meeting of the Association for Computational Linguistics and the 7th International Joint Conference on Natural Language Pro… [cited by applicant]
Kazemi et al.; SimpIE Embedding for Link Prediction in Knowledge Graphs; NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, Dec. 2018. [cited by applicant]
Kipf et al.; Semi-Supervised Classification with Graph Convolutional Networks; International Conference on Learning Representations (ICLR 2017)—arXiv:1609.02907v4. [cited by applicant]
Ma et al.; DepthLGP: Learning Embeddings of Out-of-Sample Nodes in Dynamic Networks; Proceedings of the 32th AAAI Conference on Artificial Intelligence, 2018, pp. 370-377. [cited by applicant]
Miller, George A.; WordNet: A Lexical Database for English; Communications of the ACM, Nov. 1995, vol. 38, No. 11, pp. 39-41. [cited by applicant]
Zhang et al.; Inductive Matrix Completion Based on Graph Neural Networks; International Conference on Learning Representations (ICLR 2020)—arXiv:1904.12058v3. [cited by applicant]
Nguyen et al.; STransE: a novel embedding model of entities and relationships in knowledge bases; Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human … [cited by applicant]
Nickel et al.; A Review of Relational Machine Learning for Knowledge Graphs; Proceedings of the IEEE, vol. 104, No. 1, pp. 11-33, Jan. 2016—arXiv:1503.00759v3. [cited by applicant]
Paszke et al.; Automatic Differentiation in PyTorch; 31st Conference on Neural Information Processing Systems (NIPS 2017). [cited by applicant]
Schlichtkrull et al.; Modeling Relational Data with Graph Convolutional Networks; ESWC 2018—arXiv:1703.06103v4. [cited by applicant]
Sun et al.; RotatE: Knowledge Graph Embedding by Relational Rotation in Complex Space; International Conference on Learning Representations (ICLR 2019)—arXiv:1902.10197v1. [cited by applicant]
Teru et al.; Inductive Relation Prediction on Knowledge Graphs; arXiv:1911.06962v1. [cited by applicant]
Toutanova et al.; Observed versus latent features for knowledge base and text inference; Proceedings of the 3rd Workshop on Continuous Vector Space Models and their Compositionality (CVSC), pp. 57-66, 2015. [cited by applicant]
Trouillon et al.; Complex Embeddings for Simple Link Prediction; Proceedings of the 33rd International Conference on Machine Learning (ICML 2016), vol. 48, 2016. [cited by applicant]
Veličković et al.; Graph Attention Networks; International Conference on Learning Representations (ICLR 2018)—arXiv:1710.10903v3. [cited by applicant]
Xie et al.; Representation Learning of Knowledge Graphs with Entity Descriptions; Proceedings of the Thirtieth AAAI Conference on Artificial Intelligence (AAAI-16), Feb. 2016, pp. 2659-2665. [cited by applicant]
Yang et al.; Embedding Entities and Relations for Learning and Inference in Knowledge Bases; International Conference on Learning Representations (ICLR 2015)—arXiv:1412.6575v4. [cited by applicant]
Yang et al.; Revisiting Semi-Supervised Learning with Graph Embeddings; Proceedings of The 33rd International Conference on Machine Learning (ICML 2016), vol. 48, 2016. [cited by applicant]
Zhang et al.; Quaternion Knowledge Graph Embedding; 33rd Conference on Neural Information Processing Systems (NeurIPS 2019)—arXiv:1904.10281v3. [cited by applicant]
Zhao et al.; Zero-Shot Embedding for Unseen Entities in Knowledge Graph; IEICE Transactions on Information and Systems, vol. E100.D, No. 7, pp. 1440-1447, Jul. 2017. [cited by applicant]
Fatemi et al.; Knowledge Hypergraphs: Prediction Beyond Binary Relations; Proceedings of the Twenty-Ninth International Joint Conference on Artificial Intelligence (IJCAI-20), pp. 2191-2197. [cited by applicant]
Goel et al.; Diachronic Embedding for Temporal Knowledge Graph Completion; arXiv:1907.03143v1. [cited by applicant]
Hajimoradlou et al.; Stay Positive: Knowledge Graph Embedding Without Negative Sampling; Graph Representation Learning and Beyond Workshop (ICML 2020). [cited by applicant]
Hamaguchi et al.; Knowledge Transfer for Out-of-Knowledge-Base Entities: A Graph Neural Network Approach; Proceedings of the Twenty-Sixth International Joint Conference on Artificial Intelligence (IJCAI-17), pp. 1802-18… [cited by applicant]
Kingma et al.; Adam: A Method for Stochastic Optimization; International Conference on Learning Representations (ICLR 2015)—arXiv:1412.6980v9. [cited by applicant]