IP Library › Granted Patent US 12,455,945
Granted Patent B2
US 12,455,945 · App. 17/654,417 · Granted Oct 28, 2025

Device and in particular a computer-implemented method for classifying data sets

Inventors: Lukas Lange (Pforzheim, DE); Jannik Stroetgen (Karlsruhe, DE)
Assignee: ROBERT BOSCH GMBH
G06F18/2413G06F18/214G06F18/2431
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,455,945
App. No.
17/654,417
Granted
Oct 28, 2025
Kind
B2
Abstract

A device and method for classifying data sets are provided. A model for solving a task, and training data sets are predefined. For each of the training data sets, a trained model for solving the task is determined by pretraining the model on the training data set and training the model on a reference training data set. A trained reference model for solving the task is determined by training the model on the reference training data set without pretraining with the plurality of training data sets. The trained models are classified as suitable or unsuitable for the pretraining as a function of a deviation of their particular quality from a reference quality. In the plurality of training data sets, nearest neighbors of a data set are determined. Each data set is classified as suitable or unsuitable for the pretraining.

Claims (61)

1. A computer-implemented method for classifying data sets, the method comprising:

predefining a model for solving a task;

predefining a plurality of training data sets;

defining, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determining a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determining, for each respective trained model, a respective quality of solving the task;

determining, for the trained reference model, a reference quality of solving the task;

classifying each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determining, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classifying the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classifying the nearest neighbors of the data set as suitable for the pretraining, wherein for training of the model, the model is pretrained with at least one of the training data sets, which is classified as suitable for the pretraining, and the model for solving the task is subsequently trained with the data set.

2. A computer-implemented method for classifying data sets, the method comprising:

predefining a model for solving a task;

predefining a plurality of training data sets;

defining, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determining a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determining, for each respective trained model, a respective quality of solving the task;

determining, for the trained reference model, a reference quality of solving the task;

classifying each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determining, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classifying the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classifying the nearest neighbors of the data set as suitable for the pretraining, wherein for training of the model, the model is pretrained for solving the task with the data set when the data set is classified, by a classifier, as suitable for the pretraining, and otherwise the model for solving the task is not pretrained with the data set.

3. A computer-implemented method for classifying data sets, the method comprising:

predefining a model for solving a task;

predefining a plurality of training data sets;

defining, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determining a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determining, for each respective trained model, a respective quality of solving the task;

determining, for the trained reference model, a reference quality of solving the task;

classifying each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determining, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classifying the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classifying the nearest neighbors of the data set as suitable for the pretraining, wherein for each training data set of the plurality of training data sets, at least one distance from the data set is determined, either a predefined number of training data sets from the plurality of training data sets being determined as nearest neighbors, whose distance is less than the other of the training data sets from the plurality of training data sets, or the training data sets from the plurality of training data sets being determined as nearest neighbors, whose distance is less than a predefined distance.

4. The method as recited in claim 3 , wherein for each training data set of the plurality of training data sets, a plurality of distances from the data set is determined using various distance measures, the training data sets from the plurality of training data sets being determined as nearest neighbors, for which at least one distance from the plurality of distances is less than the predefined distance.

5. A computer-implemented device for classifying data sets, the computer-implemented device having a processor that executes computer-readable instructions configured to:

predefine a model for solving a task;

predefine a plurality of training data sets;

define, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determine a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determine, for each respective trained model, a respective quality of solving the task;

determine, for the trained reference model, a reference quality of solving the task;

classify each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determine, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classify the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classify the nearest neighbors of the data set as suitable for the pretraining, wherein for training of the model, the model is pretrained with at least one of the training data sets, which is classified as suitable for the pretraining, and the model for solving the task is subsequently trained with the data set.

6. A computer-implemented device for classifying data sets, the computer-implemented device having a processor that executes computer-readable instructions configured to:

predefine a model for solving a task;

predefine a plurality of training data sets;

define, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determine a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determine, for each respective trained model, a respective quality of solving the task;

determine, for the trained reference model, a reference quality of solving the task;

classify each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determine, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classify the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classify the nearest neighbors of the data set as suitable for the pretraining, wherein the computer-readable instructions executed by the processor of the computer-implemented device are configured to classify the data set, and configured to determine the model for solving a task, which is trained or pretrained with the data set, when the data set is classified as suitable for the pretraining, and otherwise to determine the model without training or pretraining with the data set.

7. A non-transitory computer-readable medium on which is stored a computer program including computer-readable instructions for classifying data sets, the instructions, when executed by a computer, causing the computer to perform the following steps:

predefining a model for solving a task;

predefining a plurality of training data sets;

defining, for each training data set from the plurality of training data sets, a respective trained model for solving the task by pretraining the model on the training data set and training the model on a reference training data set;

determining a trained reference model for solving the task by training the model on the reference training data set without pretraining with the plurality of training data sets;

determining, for each respective trained model, a respective quality of solving the task;

determining, for the trained reference model, a reference quality of solving the task;

classifying each respective trained model as suitable or unsuitable for the pretraining as a function of a deviation of the respective quality from the reference quality;

determining, in the plurality of training data sets, nearest neighbors of a data set of the plurality of training data sets; and

either classifying the data set as suitable or unsuitable for the pretraining as a function of how the trained models, which have been trained with the nearest neighbors, are classified, or classifying the nearest neighbors of the data set as suitable for the pretraining, wherein for training of the model, the model is pretrained with at least one of the training data sets, which is classified as suitable for the pretraining, and the model for solving the task is subsequently trained with the data set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 8, 2022
From: LANGE, LUKAS; STROETGEN, JANNIK
To: ROBERT BOSCH GMBH
Reel/Frame 060742/0930 →
Priority Claims (1)
DE 10 2021 202 564.1 · Mar 16, 2021 · national
Continuity (1)
Related Publication 20220300750A1 · Sep 22, 2022
References Cited (12)
US 11210605B1 · Gupta · 2021 [cited by examiner]
US 20180018587A1 · Kobayashi · 2018 [cited by examiner]
US 20180060738A1 · Achin · 2018 [cited by examiner]
US 20180157971A1 · Fusi · 2018 [cited by examiner]
US 20190377984A1 · Ghanta · 2019 [cited by examiner]
US 20200380378A1 · Moharrer · 2020 [cited by examiner]
US 20210390458A1 · Blumstein · 2021 [cited by examiner]
US 20220044078A1 · Sathe · 2022 [cited by examiner]
US 20220114473A1 · Awasthy · 2022 [cited by examiner]
US 20230229965A1 · Ishikawa · 2023 [cited by examiner]
Vu et al., “Exploring and Predicting Transferability Across NLP Tasks,” Cornell University, 2020, pp. 1-45. <https://arxiv.org/pdf/2005.00770.pdf> Downloaded Mar. 9, 2022. [cited by applicant]
Zamir, et al.: “Taskonomy: Disentangling Task Transfer Learning”, Proceedings of the 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, (2018), pp. 3712-3722; 001: https://doi.org/10.1109/CVPR.2018.003… [cited by applicant]