IP Library Granted Patent US 12,688,366
Granted Patent B2
US 12,688,366 · App. 17/357,157 · Granted Jul 21, 2026

Apparatus and method for filling a knowledge graph by way of strategic data splits

Inventors: Hanna Wecker (Munich, DE); Annemarie Friedrich (Weil der Stadt, DE); Heike Adel-Vu (Stuttgart, DE)
Assignee: ROBERT BOSCH GMBH
G06F40/30G06F18/23213G06F18/285G06F40/205G06N5/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,688,366
App. No.
17/357,157
Filed
Jun 24, 2021
Granted
Jul 21, 2026
Kind
B2
Art Unit
2144
USPC
706/45
Abstract

A method for filling a knowledge graph. A first and second subset of data points are determined. A data point to which a label is assigned is associated with a cluster from among a set of clusters, depending on whether a distribution of labels from data points that are already associated with the cluster satisfies a condition. Data points that are associated with the cluster are associated with the first or second subset. Models for classification are trained depending on data points from the first subset. For at least one of the models, a value of a quality factor is determined depending on data points from the second subset. A model for classification is selected from the models depending on the value. A classification that defines a relationship, node, or type of node in the knowledge graph for the sentence is determined using the selected model.

Claims (39)

1 . A method for filling a knowledge graph, comprising the following steps:

determining a first subset of a set of data points and a second subset of the set of data points;

associating a data point to which a label is assigned with a cluster from among a set of clusters, the data point being associated with the cluster in response to a distribution of labels from data points that are already associated with the cluster satisfies a condition;

associating data points that are associated with the cluster either with the first subset or with the second subset;

training a plurality of models for classification depending on data points from the first subset;

determining, for at least one of the trained models from among the plurality of trained models, a value of a quality factor depending on data points from the second subset;

selecting a trained model for classification from the plurality of trained models depending on the value; and

determining a classification that defines a relationship or a node or a type of node in the knowledge graph for a sentence using the trained model, selected depending on the value, for classification for data.

2 . The method as recited in claim 1 , wherein a number of data points, which are associated with the cluster and with which the label is associated, is determined, the data point being associated with the cluster if the number satisfies a condition, and otherwise not being associated.

3 . The method as recited in claim 2 , wherein a series of data points with which the label is associated is determined, the data points being associated with the cluster in accordance with the series as long as the number of data points satisfies the condition.

4 . The method as recited in claim 3 , wherein the series encompasses the data point, the series being determined depending on a difference between a distance of the data point from a first cluster center of a first cluster from among the set of clusters, and a distance of the data point from a second cluster center of a second cluster from among the set of clusters.

5 . The method as recited in claim 4 , wherein the data point is placed in the series either before another data point if the difference is greater than a reference, or otherwise after the other data point in the series.

6 . The method as recited in claim 4 , wherein the first cluster center is that cluster center, from among the set of cluster centers, which is located closest to the data point, the second cluster center being that cluster center, from among the set of cluster centers, which is located farthest from the data point.

7 . The method as recited in claim 6 , further comprising:

determining a third cluster center depending on data points that are associated with the first cluster;

determining a fourth cluster center depending on data points that are associated with the second cluster;

determining a first distance of a first data point of the first cluster from the third cluster center, and a second distance of the first data point from the fourth cluster center;

determining a third distance of a second data point of the second cluster from the fourth cluster center, and a fourth distance of the second data point from the third cluster;

determining a first difference between the first distance and the second distance;

determining a second difference between the third distance and the fourth distance; and

based on the labels of the first and second data points being the same, and based on the first difference satisfying a first condition and the second difference satisfying a second condition, associating the first data point with the second cluster, and the second data point with the first cluster.

8 . The method as recited in claim 1 , wherein a plurality of cluster centers is furnished, data points from among the set of data points being associated with one of the cluster centers from among the plurality of cluster centers.

9 . The method as recited in claim 8 , wherein a number of subsets for the data points is predefined, the number of subsets defining a number of cluster centers.

10 . An apparatus configured to fill a knowledge graph, the apparatus configured to:

determine a first subset of a set of data points and a second subset of the set of data points;

associate a data point to which a label is assigned with a cluster from among a set of clusters, the data point being associated with the cluster in response to a distribution of labels from data points that are already associated with the cluster satisfies a condition;

associate data points that are associated with the cluster either with the first subset or with the second subset;

train a plurality of models for classification depending on data points from the first subset;

determine, for at least one of the trained models from among the plurality of trained models, a value of a quality factor depending on data points from the second subset;

select a trained model for classification from the plurality of trained models depending on the value; and

determine a classification that defines a relationship or a node or a type of node in the knowledge graph for a sentence using the trained model, selected depending on the value, for classification for data.

11 . A non-transitory computer-readable medium on which is stored a computer program including computer-readable instructions for filling a knowledge graph, the instructions, when executed by a computer, causing the computer to perform the following steps:

determining a first subset of a set of data points and a second subset of the set of data points;

associating a data point to which a label is assigned with a cluster from among a set of clusters, the data point being associated with the cluster in response to a distribution of labels from data points that are already associated with the cluster satisfies a condition;

associating data points that are associated with the cluster either with the first subset or with the second subset;

training a plurality of models for classification depending on data points from the first subset;

determining, for at least one of the trained models from among the plurality of trained models, a value of a quality factor depending on data points from the second subset;

selecting a trained model for classification from the plurality of trained models depending on the value; and

determining a classification that defines a relationship or a node or a type of node in the knowledge graph for a sentence using the trained model, selected depending on the value, for classification for data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2022
From: WECKER, HANNA; FRIEDRICH, ANNEMARIE; ADEL-VU, HEIKE
To: ROBERT BOSCH GMBH
Reel/Frame 058733/0701 →
Priority Claims (1)
DE 102020208041.0 · Jun 29, 2020 · national
Continuity (1)
Related Publication 20210406702A1 · Dec 30, 2021
References Cited (21)
US 10719301B1 · Dasgupta · 2020 [cited by examiner]
US 11379501B2 · Yadav · 2022 [cited by examiner]
US 20090319457A1 · Cheng et al. · 2009 [cited by applicant]
US 20130097103A1 · Chari et al. · 2013 [cited by applicant]
US 20170293608A1 · Seow · 2017 [cited by examiner]
US 20210097140A1 · Chatterjee · 2021 [cited by examiner]
US 20210248193A1 · Cho · 2021 [cited by examiner]
CN 106991296A · 2017 [cited by applicant]
CN 110413856A · 2019 [cited by applicant]
JP 2019185665A · 2019 [cited by applicant]
Alessandro Lulli, etc., “Scalable k-NN based text clustering”, published via 2015 IEEE International Conference on Big Data, Santa Clara, CA, USA, Oct. 29-Nov. 1, 2015, retrieved Dec. 5, 2024. (Year: 2015). [cited by examiner]
Liang Yao, etc., “Graph Convolutional Networks for Text Classification”, published via arXiv as of Nov. 13, 2018, retrieved Dec. 5, 2024. (Year: 2018). [cited by examiner]
Smail Sellah, etc., “A Document Clustering Approach for Automatic Building of Ontologies”, published via 2018 Fifth International Conference on Social Networks Analysis, Management and Security, Valencia, Spain, Oct. 15… [cited by examiner]
Derek Greene, etc., “TextLuas: Tracking and Visualizing Document and Term Clusters in Dynamic Text Data”, published via arXiv as of Nov. 3, 2014, retrieved Dec. 5, 2024. (Year: 2014). [cited by examiner]
Lirong Qiu, etc., “Chinese-Uyghur-English Semantic Search Based on the Knowledge Graphs”, 2017 IEEE International Conference on Computational Science and Engineering (CSE) and IEEE International Conference on Embedded a… [cited by examiner]
Xiaoli Wang, etc., “An Improved K_means Algorithm for Document Clustering Based on Knowledge Graphs”, published via 2018 11th International Congress on Image and Signal Processing, BioMedical Engineering and Informatics… [cited by examiner]
John Violos, etc., “Text Classification Using the N-Gram Graph Representation Model Over High Frequency Data Streams”, published via Frontiers in Applied Mathematics and Statistics, vol. 4, article 41, Sep. 2018, retrie… [cited by examiner]
Mikolov et al., “Efficient Estimation of Word Representations in Vector Space,” ICLR Workshop, 2013, pp. 1-12. <https://arxiv.org/pdf/1301.3781.pdf> Downloaded Jun. 24, 2021. [cited by applicant]
Arthur et al., “K-Means++: the Advantages of Careful Seeding,” Proceedings of the Eighteenth Annual Acm-Siam Symposium on Discrete Algorithms, 2007, pp. 1-11. <http://ilpubs.stanford.edu:8090/778/1/2006-13.pdf> Download… [cited by applicant]
Wecker et al., “Clusterdatasplit: Exploring Challenging Clustering-Based Data Splits for Model Performance Evaluation,” Proceedings of the First Workshop on Evaluation and Comparison of NLP Systems (EVAL4NLP ), 2020, pp… [cited by applicant]
Vanderplas, “Hyperparameters and Model Validation,” Python Data Science Handbook, 2016, pp. 1-19. [cited by applicant]