IP Library › Granted Patent US 12,511,889
Granted Patent B2
US 12,511,889 · App. 18/034,897 · Granted Dec 30, 2025

Neural network models for semantic image segmentation

Inventors: Mohsen Ghafoorian (Diemen, NL); Klaus Michael Hofmann (Amsterdam, NL); Erik Stammes (Barcelona, ES)
Assignee: TomTom Global Content B.V.
G06V10/82G06V10/72
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,511,889
App. No.
18/034,897
Filed
May 2, 2023
Granted
Dec 30, 2025
Kind
B2
Art Unit
2673
USPC
382/155
Abstract

The invention relates to a method of training a model for use in semantic image segmentation. The model includes a first classifier neural network and a second classifier neural network. The method includes training the first network by inputting a first training image containing a target object to the first network, using the first network to identify and erase pixels of the first training image that are discriminative for the target object, inputting the first training image to the second network, using the second network to determine a likelihood of the first modified training image containing the target object and updating weights of the first network using a first loss function that is a monotonically-increasing function of the determined likelihood.

Claims (56)

1 . A method for training a model for use in semantic image segmentation, wherein the model comprises:

a first classifier neural network; and

a second classifier neural network,

wherein the method comprises training the first classifier neural network by:

inputting a first training image containing a target object to the first classifier neural network;

using the first classifier neural network to identify pixels of the first training image that are discriminative for the target object;

generating a first modified training image by attenuating or erasing the identified pixels of the first training image;

inputting the first modified training image to the second classifier neural network;

using the second classifier neural network to determine a likelihood of the first modified training image containing the target object; and

updating weights of the first classifier neural network using a first loss function that is a monotonically-increasing function of the determined likelihood, and wherein the method comprises training the second classifier neural network by:

inputting a second training image, containing the target object, to the first classifier neural network;

using the first classifier neural network to identify pixels of the second training image that are discriminative for the target object;

generating a second modified training image by attenuating or erasing the identified pixels of the second training image;

inputting a set of training images to the second classifier neural network, the set comprising the second modified training image and one or more training images that do not contain the target object; and

for each training image of the set, using the second classifier neural network to determine a likelihood of the respective training image containing the target object, and updating weights of the second classifier neural network using a second loss function, different from the first loss function, wherein the second loss function is a monotonically-decreasing function of the determined likelihood when the training image contains the target object, and is a monotonically-increasing function of the determined likelihood when the training image does not contain the target object.

2 . The method of claim 1 , wherein the first classifier neural network and the second classifier neural network are independent convolutional neural networks having independent sets of weights.

3 . The method of claim 1 , comprising alternately training the first classifier neural network and the second classifier neural network for a plurality of training cycles.

4 . The method of claim 1 , comprising:

training the first classifier neural network with the second classifier neural network operating in an inference mode;

updating the weights of the first classifier neural network without changing any weights of the second classifier neural network;

training the second classifier neural network with the first classifier neural network operating in an inference mode; and

updating the weights of the second classifier neural network without changing any weights of the first classifier neural network.

5 . The method of claim 1 , comprising training the first classifier neural network and the second classifier neural network on a common plurality of training images.

6 . The method of claim 1 , wherein the pixels of the first training image identified by the first classifier neural network are pixels that are relatively more discriminative, for the first classifier neural network, than any other pixels of the image.

7 . The method of claim 1 , comprising identifying the pixels of the first training image by applying a hard or soft threshold to attention-map data obtained using the first classifier neural network.

8 . The method of claim 1 , comprising training the first classifier neural network to classify the target object.

9 . The method of claim 1 , wherein the first classifier neural network and the second classifier neural network are image-level classifier networks, and wherein the model is trained using training data that comprises training images and associated image-level labels.

10 . The method of claim 1 , comprising training the first classifier neural network to favour identifying smaller sets of discriminative pixels by the first loss function additionally being a monotonically-increasing function of the number of pixels identified by the first classifier neural network, or of the sum of values in all or part of an attention map for the first training image generated using the first classifier neural network.

11 . A non-transitory computer readable storage medium storing instructions that, when executed on a computer processing system, cause the computer processing system to perform semantic image segmentation using the first classifier neural network of a model trained by the method of claim 1 .

12 . A computer processing system configured to implement the first classifier neural network of a model trained by the method of claim 1 , for performing semantic image segmentation.

13 . The computer processing system of claim 12 , wherein the computer processing system is a computer processing system for a vehicle, and comprises an input for receiving image data from a camera, and an output for outputting segmentation data to an autonomous driving system for the vehicle.

14 . A computer processing system for training a model for use in semantic image segmentation,

wherein the model comprises:

a first classifier neural network; and

a second classifier neural network,

wherein the computer processing system is configured to train the first classifier neural network by:

inputting a first training image containing a target object to the first classifier neural network;

using the first classifier neural network to identify pixels of the first training image that are discriminative for the target object;

generating a first modified training image by attenuating or erasing the identified pixels of the first training image;

inputting the first modified training image to the second classifier neural network;

using the second classifier neural network to determine a likelihood of the first modified training image containing the target object; and

updating weights of the first classifier neural network using a first loss function that is a monotonically-increasing function of the determined likelihood, and wherein the computer processing system is configured to train the second classifier neural network by:

inputting a second training image, containing the target object, to the first classifier neural network;

using the first classifier neural network to identify pixels of the second training image that are discriminative for the target object;

generating a second modified training image by attenuating or erasing the identified pixels of the second training image;

inputting a set of training images to the second classifier neural network, the set comprising the second modified training image and one or more training images that do not contain the target object; and

for each training image of the set, using the second classifier neural network to determine a likelihood of the respective training image containing the target object, and updating weights of the second classifier neural network using a second loss function, different from the first loss function, wherein the second loss function is a monotonically-decreasing function of the determined likelihood when the training image contains the target object, and is a monotonically-increasing function of the determined likelihood when the training image does not contain the target object.

15 . The computer processing system of claim 14 , configured to train the first classifier neural network and the second classifier neural network alternately for a plurality of training cycles.

16 . The computer processing system of claim 14 , configured to:

train the first classifier neural network with the second classifier neural network operating in an inference mode;

update the weights of the first classifier neural network without changing any weights of the second classifier neural network;

train the second classifier neural network with the first classifier neural network operating in an inference mode; and

update the weights of the second classifier neural network without changing any weights of the first classifier neural network.

17 . The computer processing system of claim 14 , wherein the pixels of the first training image identified by the first classifier neural network are pixels that are relatively more discriminative, for the first classifier neural network, than any other pixels of the image.

18 . The computer processing system of claim 14 , configured to identify the pixels of the first training image by applying a hard or soft threshold to attention-map data obtained using the first classifier neural network.

19 . The computer processing system of claim 14 , configured to train the first classifier neural network to favour identifying smaller sets of discriminative pixels by the first loss function additionally being a monotonically-increasing function of the number of pixels identified by the first classifier neural network, or of the sum of values in all or part of an attention map for the first training image generated using the first classifier neural network.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 12, 2025
From: GHAFOORIAN, MOHSEN; HOFMANN, KLAUS MICHAEL; STAMMES, ERIK
To: TOMTOM GLOBAL CONTENT B.V.
Reel/Frame 072876/0162 →
Priority Claims (1)
GB 2017369 · Nov 2, 2020 · national
Continuity (1)
Related Publication 20230419648A1 · Dec 28, 2023
References Cited (15)
US 10713794B1 · He et al. · 2020 [cited by applicant]
US 20200027002A1 · Hickson et al. · 2020 [cited by applicant]
US 20200143204A1 · Nakano et al. · 2020 [cited by applicant]
CN 109191392A · 2019 [cited by applicant]
CN 109543502A · 2019 [cited by applicant]
CN 110111340A · 2019 [cited by applicant]
CN 110245665A · 2019 [cited by applicant]
CN 110458221A · 2019 [cited by applicant]
CN 110880001A · 2020 [cited by applicant]
Wei et al, “Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach”, 2017, Proceedings of the IEEE conference on computer vision and pattern recognition, pp. 1568-1576 (… [cited by examiner]
Zhang et al, “Automatic Image Labelling at Pixel Level”, 2020, IEEE Transaction on Image Processing, arXiv: 2007.07415v2(13 Pages) (Year: 2020). [cited by examiner]
International Search Report dated Dec. 23, 2021 for International patent application No. PCT/EP2021/073856. [cited by applicant]
GB Search Report dated Apr. 22, 2021 for GB application No. GB2017369.6. [cited by applicant]
Wei, Y et al., Object Region Mining with Adversarial Erasing: A Simple Classification to Semantic Segmentation Approach, arXiv:1703.08448, https://arxiv.org/abs/1703.08448, 2018. [cited by applicant]
Zhang, X et al., Adversarial Complementary Learning for Weakly Supervised Object Localization, arXiv:1804.06962, https://arxiv.org/abs/1804.06962, 2018. [cited by applicant]