IP Library Granted Patent US 12,243,283
Granted Patent B2
US 12,243,283 · App. 17/894,358 · Granted Mar 4, 2025

Device and method for determining a semantic segmentation and/or an instance segmentation of an image

Inventors: Chaithanya Kumar Mummadi (Pittsburgh, PA); Jan Hendrik Metzen (Boeblingen, DE); Robin Hutmacher (Renningen, DE)
Assignee: ROBERT BOSCH GMBH
G06V10/26G06V10/764G06V10/7715G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,243,283
App. No.
17/894,358
Granted
Mar 4, 2025
Kind
B2
Abstract

A computer-implemented method for determining an output signal characterizing a semantic segmentation and/or an instance segmentation of an image. The method includes: determining a first intermediate output signal from a machine learning system, wherein the first intermediate output signal characterizes a semantic segmentation and/or an instance segmentation of the image; adapting parameters of the machine learning system based on a loss function, wherein the loss function characterizes an entropy or a cross-entropy of the first intermediate output signal; determining the output signal from the machine learning system based on the image and the adapted parameters.

Claims (40)

1. A computer-implemented method for determining an output signal characterizing a semantic segmentation and/or an instance segmentation of an image, the method comprising the following steps:

determining a first intermediate output signal from a machine learning system, the first intermediate output signal characterizing a semantic segmentation and/or an instance segmentation of the image;

adapting parameters of the machine learning system based on a loss function, wherein the loss function characterizes an entropy or a cross-entropy of the first intermediate output signal; and

determining the output signal from the machine learning system based on the image and the adapted parameters;

wherein the loss function characterizes a mean entropy of classifications obtained for pixels of the image;

wherein a second intermediate output signal is determined based on the first intermediate output signal by using a transformation function,

wherein the second intermediate output signal characterizes a semantic segmentation and/or an instance segmentation of the image and the cross-entropy is determined based on the first intermediate output signal and the second intermediate output signal; and

wherein the transformation function characterizes an edge-preserving smoothing filter.

2. The method according to claim 1 , wherein the loss function further characterizes a likelihood of at least a part of the first intermediate output signal, wherein the likelihood is determined based on a density model of the at least part of the image.

3. The method according to claim 1 , wherein the loss function further characterizes a log-likelihood, of at least a part of the first intermediate output signal, wherein the log-likelihood is determined based on a density model of the at least part of the image.

4. The method according to claim 2 , wherein the likelihood characterizes an average likelihood of a plurality of patches, wherein the plurality of patches is determined based on the first intermediate output signal.

5. The method according to claim 4 , wherein the likelihood of a patch of the patches is determined by determining a feature representation of the patch using a feature extractor and providing a likelihood of the feature representation as likelihood of the patch, wherein the likelihood of the feature representation is determined by the density model.

6. The method according to claim 2 , wherein the density model is characterized by a mixture model, or a Gaussian mixture model, or a normal distribution.

7. The method according to claim 1 , wherein for determining the output signal, the machine learning system includes a normalization transformation and the loss further characterizes a Kullback-Leibler divergence between an output of the normalization transformation and a predefined probability distribution.

8. The method according to claim 7 , wherein the predefined probability distribution is characterized by a standard normal distribution.

9. The method according to claim 7 , wherein the machine learning system includes a neural network for determining the output signal, wherein the neural network includes a normalization layer and wherein the loss further characterizes a Kullback-Leibler divergence between an output of the normalization layer and a predefined probability distribution.

10. A machine learning system configured to determine an output signal characterizing a semantic segmentation and/or an instance segmentation of an image, the machine learning system configured to:

determine a first intermediate output signal from the machine learning system, the first intermediate output signal characterizing a semantic segmentation and/or an instance segmentation of the image;

adapt parameters of the machine learning system based on a loss function, wherein the loss function characterizes an entropy or a cross-entropy of the first intermediate output signal; and

determine the output signal from the machine learning system based on the image and the adapted parameters;

wherein the loss function characterizes a mean entropy of classifications obtained for pixels of the image;

wherein a second intermediate output signal is determined based on the first intermediate output signal by using a transformation function,

wherein the second intermediate output signal characterizes a semantic segmentation and/or an instance segmentation of the image and the cross-entropy is determined based on the first intermediate output signal and the second intermediate output signal; and

wherein the transformation function characterizes an edge-preserving smoothing filter.

11. A control system configured to determine a control signal, wherein the control signal is configured to control an actuator and/or a display, and wherein the control signal is determined based on an output signal determined by:

determining a first intermediate output signal from a machine learning system, the first intermediate output signal characterizing a semantic segmentation and/or an instance segmentation of the image;

adapting parameters of the machine learning system based on a loss function, wherein the loss function characterizes an entropy or a cross-entropy of the first intermediate output signal; and

determining the output signal from the machine learning system based on the image and the adapted parameters;

wherein the loss function characterizes a mean entropy of classifications obtained for pixels of the image;

wherein a second intermediate output signal is determined based on the first intermediate output signal by using a transformation function,

wherein the second intermediate output signal characterizes a semantic segmentation and/or an instance segmentation of the image and the cross-entropy is determined based on the first intermediate output signal and the second intermediate output signal; and

wherein the transformation function characterizes an edge-preserving smoothing filter.

12. A non-transitory machine-readable storage medium on which is stored a computer program for determining an output signal characterizing a semantic segmentation and/or an instance segmentation of an image, the computer program, when executed by a computer, causing the computer to perform the following steps:

determining a first intermediate output signal from a machine learning system, the first intermediate output signal characterizing a semantic segmentation and/or an instance segmentation of the image;

adapting parameters of the machine learning system based on a loss function, wherein the loss function characterizes an entropy or a cross-entropy of the first intermediate output signal; and

determining the output signal from the machine learning system based on the image and the adapted parameters;

wherein the loss function characterizes a mean entropy of classifications obtained for pixels of the image;

wherein a second intermediate output signal is determined based on the first intermediate output signal by using a transformation function,

wherein the second intermediate output signal characterizes a semantic segmentation and/or an instance segmentation of the image and the cross-entropy is determined based on the first intermediate output signal and the second intermediate output signal; and

wherein the transformation function characterizes an edge-preserving smoothing filter.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 11, 2022
From: MUMMADI, CHAITHANYA KUMAR; METZEN, JAN HENDRIK; HUTMACHER, ROBIN
To: ROBERT BOSCH GMBH
Reel/Frame 061739/0647 →
Priority Claims (1)
EP 21198202 · Sep 22, 2021 · regional
Continuity (1)
Related Publication 20230101810A1 · Mar 30, 2023
References Cited (31)
US 8879813B1 · Solanki · 2014 [cited by examiner]
US 10467500B1 · Bao · 2019 [cited by examiner]
US 10685446B2 · Fleishman · 2020 [cited by examiner]
US 10783622B2 · Wang · 2020 [cited by examiner]
US 11074504B2 · Chen · 2021 [cited by examiner]
US 11163989B2 · Sun · 2021 [cited by examiner]
US 11958529B2 · Mandlekar · 2024 [cited by examiner]
US 12051261B2 · Rejeb Sfar · 2024 [cited by examiner]
US 12094124B2 · Guo · 2024 [cited by examiner]
US 12106828B2 · Kostem · 2024 [cited by examiner]
US 20080292194A1 · Schmidt · 2008 [cited by examiner]
US 20170262735A1 · Ros Sanchez · 2017 [cited by examiner]
US 20200027002A1 · Hickson · 2020 [cited by examiner]
US 20210035304A1 · Jie · 2021 [cited by examiner]
US 20210241034A1 · Laradji · 2021 [cited by examiner]
US 20210319315A1 · Hutmacher · 2021 [cited by examiner]
US 20210407090A1 · Li · 2021 [cited by examiner]
US 20220092368A1 · Hao · 2022 [cited by examiner]
US 20220375211A1 · Tolstikhin · 2022 [cited by examiner]
US 20230107917A1 · Pabbaraju · 2023 [cited by examiner]
US 20230186622A1 · Laszlo · 2023 [cited by examiner]
EP 3869387A1 · 2021 [cited by examiner]
EP 3879461A1 · 2021 [cited by examiner]
WO WO2020049087A1 · 2020 [cited by examiner]
WO WO2020237215A1 · 2020 [cited by examiner]
WO WO2022023646A1 · 2022 [cited by examiner]
WO WO2022037170A1 · 2022 [cited by examiner]
Mummadi et al., “Test-Time Adaptation To Distribution Shift By Confidence Maximization and Input Transformation,” Cornell University, 2021, pp. 1-16. [cited by applicant]
Wittich, “Deep Domain Adaptation by Weighted Entropy Minimization for the Classification of Aerial Images,” ISPRS Annals of the Photogrammetry, Remote Sensing and Spatial Information Sciences, vol. V—Feb. 2020, XXIV ISP… [cited by applicant]
Vu et al., “Advent: Adversarial Entropy Minimization for Domain Adaptation in Semantic Segmentation,” IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2019, pp. 2512-2521. [cited by applicant]
Wang et al., “Tent: Fully Test-Time Adaptation By Entropy Minimization,” Cornell University, 2021, pp. 1-15. [cited by applicant]