IP Library › Granted Patent US 12,567,238
Granted Patent B2
US 12,567,238 · App. 17/646,958 · Granted Mar 3, 2026

Generating a data structure for specifying visual data sets

Inventors: Christoph Gladisch (Renningen, DE); Christian Heinzemann (Ludwigsburg, DE); Martin Herrmann (Korntal, DE); Matthias Woehrle (Bietigheim-Bissingen, DE); Nadja Schalm (Renningen, DE)
Assignee: ROBERT BOSCH GMBH
G06V10/771G06V10/58G06V10/774G06V10/776G06V10/7796G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,567,238
App. No.
17/646,958
Granted
Mar 3, 2026
Kind
B2
Abstract

Facilitating the description or configuration of a computer vision model by generating a data structure comprising a plurality of language entities defining a semantic mapping of visual parameters to a visual parameter space based on a sensitivity analysis of the computer vision model.

Claims (67)

1 . A computer-implemented method for generating a data structure including a plurality of language entities defining a semantic mapping of visual parameters to a visual parameter space, the method comprising the following steps:

obtaining a computer vision model configured to perform a computer vision function of characterizing elements of observed scenes;

obtaining a first visual parameter set including a plurality of initial visual parameters, wherein an item of visual data provided based on an extent of the at least one initial visual parameter is capable of affecting a classification or regression performance of the computer vision model;

providing a visual data set including a subset of items of visual data compliant with the first visual parameter set, and a corresponding subset of items of groundtruth data;

applying the subset of items of visual data to the computer vision model to obtain a plurality of performance scores characterizing performance of the computer vision model when applied to the subset of items of visual data of the visual data set, using the corresponding groundtruth data;

performing a sensitivity analysis of the plurality of performance scores over a domain of the first visual parameter set;

generating a second visual parameter set including at least one updated visual parameter, wherein the second visual parameter set includes at least one initial visual parameter modified based on the outcome of the sensitivity analysis to provide the at least one updated visual parameter; and

generating a data structure including at least one language entity based on the visual parameters of the second visual parameter set, thus providing a semantic mapping to the visual parameters of the second visual parameter set.

2 . The computer-implemented method according to claim 1 , wherein the obtaining of the plurality of performance scores includes:

generating, using the computer vision model, a plurality of predictions of elements of observed scenes in the subset of items of visual data, wherein the plurality of predictions include at least one prediction of a classification label and/or at least one regression value of at least one item in the subset of visual data;

comparing the plurality of predictions of elements in the subset of items of visual data with the corresponding subset of groundtruth data, to obtain the plurality of performance scores.

3 . The computer-implemented method according to claim 2 , wherein the performance score comprises, or is based on, any one or combination, of: a list of a confusion matrix, a precision score, a recall score, an F1 score, a union intersection score, a mean average score.

4 . The computer-implemented method according to claim 2 , wherein the computer vision model is a neural network, or a neural-network-like model.

5 . The computer-implemented method according to claim 1 , wherein the performing of the sensitivity analysis includes:

computing a plurality of variances of respective performance scores of the plurality of performance scores with respect to the initial visual parameters of the first visual parameter set and/or with respect to one or more combinations of visual parameters of the first visual parameter set.

6 . The computer-implemented method according to claim 5 , wherein the performing of the sensitivity analysis further includes ranking the initial visual parameters of the first visual parameter set and/or the one or more combinations of visual parameters based on the computed plurality of variances of performance scores.

7 . The computer-implemented method according to claim 5 , further comprising:

identifying at least one range of the initial visual parameter of the first visual parameter set using the plurality of performance scores and/or the plurality of variances of performance scores, wherein the generating of the second visual parameter set includes modifying the range of the at least one initial visual parameter by enlarging or shrinking a scope of the at least one initial visual parameter range on its domain to thus yield a modified visual parameter range.

8 . The computer-implemented method according to claim 5 , further comprising:

identifying at least one combination of visual parameters including at least two initial visual parameter sets or at least two initial visual parameter ranges or at least one initial visual parameter set and one initial visual parameter range from the first visual parameter set using the plurality of performance scores and/or the plurality of variances of performance scores, and wherein the generating of the second visual parameter set includes concatenating the at least one combination of initial visual parameters, thus defining a further language entity.

9 . The computer-implemented method according to claim 8 , wherein the identifying of the at least one combination of visual parameters is automated according to at least one predetermined criterion based a plurality of variances of performance scores.

10 . The computer-implemented method according to claim 9 , wherein the at least one combination of visual parameters is identified, when the corresponding variance of performance scores exceeds a predetermined threshold value.

11 . The computer-implemented method according to claim 1 , further comprising:

identifying, based on an identification condition, at least one initial visual parameter set of the first visual parameter set using the plurality of performance scores and/or the plurality of variance of performance scores, and wherein generating the second visual parameter set includes modifying the at least one initial visual parameter set by dividing the at least one initial visual parameter set into at least a first and a second visual parameter subset, thus defining two further language entities;, and/or

concatenating at least a third and a fourth visual parameter set of the first visual parameter set into a combined visual parameter subset.

12 . The computer-implemented method according to claim 1 , wherein the domain of the first visual parameter set includes a subset, in a finite-dimensional vector space, of numerical representations that visual parameters are allowed to lie in, or a multi-dimensional interval of continuous or discrete visual parameters, or a set of numerical representations of visual parameters in the finite-dimensional vector space.

13 . The computer-implemented method according to claim 1 , wherein the providing of the semantic mapping from visual parameters of the second visual parameter set to items of visual data and corresponding items of groundtruth data includes:

sampling the at least one initial visual parameter included in the first visual parameter set to obtain a set of sampled initial visual parameter values, wherein the sampling of the at least one initial visual parameter range is performed using a sampling method including combinatorial testing and/or Latin hypercube sampling; and

obtaining a visual data set by one or a combination of:

generating, using a synthetic visual data generator, a synthetic visual data set including synthetic visual data and groundtruth data synthesized according to the samples of the second visual parameter set; and/or

sampling items of visual data from a database including specimen images associated with corresponding items of groundtruth data according to the samples of the second visual parameter set; and/or

specifying experimental requirements according to the samples of the second visual parameter set, and performing live experiments to obtain the visual data set and to gather groundtruth data.

14 . The computer-implemented method according to claim 13 , further comprising outputting the set of visual data and corresponding items of groundtruth data as a training data set.

15 . The computer-implemented method according to claim 1 , wherein the at least one data structure includes at least one language entity based on the visual parameters of the second visual parameter set is received via an input interface of a computing device, and the language entity is displayed to a user via an output interface of the computing device.

16 . A computer-implemented method for training a computer vision model, comprising:

obtaining a further computer vision model configured to perform a computer vision function of characterising elements of observed scenes;

obtaining a set of training data by:

generating a data structure, including:

obtaining a computer vision model configured to perform a computer vision function of characterizing elements of observed scenes;

obtaining a first visual parameter set including a plurality of initial visual parameters, wherein an item of visual data provided based on an extent of the at least one initial visual parameter is capable of affecting a classification or regression performance of the computer vision model;

providing a visual data set including a subset of items of visual data compliant with the first visual parameter set, and a corresponding subset of items of groundtruth data;

applying the subset of items of visual data to the computer vision model to obtain a plurality of performance scores characterizing performance of the computer vision model when applied to the subset of items of visual data of the visual data set, using the corresponding groundtruth data;

performing a sensitivity analysis of the plurality of performance scores over a domain of the first visual parameter set;

generating a second visual parameter set including at least one updated visual parameter, wherein the second visual parameter set includes at least one initial visual parameter modified based on the outcome of the sensitivity analysis to provide the at least one updated visual parameter;

generating the data structure including at least one language entity based on the visual parameters of the second visual parameter set, thus providing a semantic mapping to the visual parameters of the second visual parameter set; and

outputting the set of visual data and corresponding items of groundtruth data as the training data set; and

training the computer vision model using the set of training data.

17 . An apparatus for generating a data structure comprising a plurality of language entities defining a semantic mapping of visual parameters to a visual parameter space, comprising:

an input interface;

a processor;

a memory; and

an output interface;

wherein the input interface is configured to obtain a computer vision model configured to perform a computer vision function of characterizing elements of observed scenes, and to obtain a first visual parameter set including a plurality of initial visual parameters, wherein an item of visual data provided based on an extent of the at least one initial visual parameter is capable of affecting a classification or regression performance of the computer vision model, and

wherein the processor is configured to:

provide a visual data set including a subset of items of visual data compliant with the first visual parameter set, and a corresponding subset of items of groundtruth data,

apply the subset of items of visual data to the computer vision model to obtain a plurality of performance scores characterizing the performance of the computer vision model when applied to the subset of items of visual data of the visual data set, using the corresponding groundtruth data,

perform a sensitivity analysis of the plurality of performance scores over a domain of the first visual parameter set,

generate a second visual parameter set including at least one updated visual parameter, wherein the second visual parameter set includes at least one initial visual parameter modified based on the outcome of the sensitivity analysis to provide the at least one updated visual parameter, and

generate a data structure comprising at least one language entity based on the visual parameters of the second visual parameter set, thus providing a semantic mapping to visual parameters of the second visual parameter set.

18 . A non-transitory computer readable medium on which is stored a computer program for generating a data structure including a plurality of language entities defining a semantic mapping of visual parameters to a visual parameter space, the computer program, when executed by a processor, causing the processor to perform the following steps:

obtaining a computer vision model configured to perform a computer vision function of characterizing elements of observed scenes;

obtaining a first visual parameter set including a plurality of initial visual parameters, wherein an item of visual data provided based on an extent of the at least one initial visual parameter is capable of affecting a classification or regression performance of the computer vision model;

providing a visual data set including a subset of items of visual data compliant with the first visual parameter set, and a corresponding subset of items of groundtruth data;

applying the subset of items of visual data to the computer vision model to obtain a plurality of performance scores characterizing performance of the computer vision model when applied to the subset of items of visual data of the visual data set, using the corresponding groundtruth data;

performing a sensitivity analysis of the plurality of performance scores over a domain of the first visual parameter set;

generating a second visual parameter set including at least one updated visual parameter, wherein the second visual parameter set includes at least one initial visual parameter modified based on the outcome of the sensitivity analysis to provide the at least one updated visual parameter; and

generating a data structure including at least one language entity based on the visual parameters of the second visual parameter set, thus providing a semantic mapping to the visual parameters of the second visual parameter set.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 12, 2022
From: GLADISCH, CHRISTOPH; HEINZEMANN, CHRISTIAN; HERRMANN, MARTIN; WOEHRLE, MATTHIAS; SCHALM, NADJA
To: ROBERT BOSCH GMBH
Reel/Frame 060484/0046 →
Priority Claims (1)
DE 10 2021 200 347.8 · Jan 15, 2021 · national
Continuity (1)
Related Publication 20220230072A1 · Jul 21, 2022
References Cited (15)
US 11716300B2 · Ravine · 2023 [cited by examiner]
US 11908178B2 · Gladisch · 2024 [cited by examiner]
US 12192600B2 · Mazaheri · 2025 [cited by examiner]
US 12198275B2 · Bautista Martin · 2025 [cited by examiner]
US 12223696B2 · Gladisch · 2025 [cited by examiner]
US 12283119B2 · Vineet · 2025 [cited by examiner]
US 20060262959A1 · Tuzel · 2006 [cited by examiner]
US 20220222926A1 · Gladisch · 2022 [cited by examiner]
US 20220230418A1 · Gladisch · 2022 [cited by examiner]
US 20220237897A1 · Heinzemann · 2022 [cited by examiner]
US 20220414928A1 · Venkataraman · 2022 [cited by examiner]
Bargoti and Underwood: “Utiising Metadata to Aid Image Classification in Orchards”, IEEE International Conference on Intelligent Robots and Systems (IROS), Workshop on Alternative Sensing for Robot Perception (WASROP), … [cited by applicant]
Engelbrecht, et al.: “Determining the Significance of Input Parameters Using Sensitivity Analysis”, International Workshop on Artificial Neural Networks, Springer, Berlin, Heidelberg, (1995), pp. 382-388. [cited by applicant]
Teodoro, et al.: “Algorithm sensitivity analysis and parameter tuning for tissue image segmentation pipelines”, Bioimage informatics, Bioinformatics 33(7), (2017), pp. 1064-1072. [cited by applicant]
Dosovitskiy, Alexey et al. “Carla: An Open Urban Driving Simulator” 1st Conference on Robot Learning (CoRL 2017) Mountain View, United States. Nov. 10, 2017. arXiv:1711.03938v1. Retreived from the Internet on Dec. 29, 2… [cited by applicant]