IP Library Granted Patent US 12,518,521
Granted Patent B2
US 12,518,521 · App. 18/298,613 · Granted Jan 6, 2026

Method for training a convolutional neural network

Inventors: Tamas Kapelner (Hildesheim, DE); Thomas Wenzel (Hamburg, DE)
Assignee: ROBERT BOSCH GMBH
G06V10/82G06N3/0464G06N3/08G06T3/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,518,521
App. No.
18/298,613
Granted
Jan 6, 2026
Kind
B2
Abstract

A method for training a convolutional neural network. For each of a multiplicity of training input images, the method includes processing of the training input image by the convolutional network; processing of a scaled version of the training input image by the convolutional network; determining a pair of convolutional layers of the convolutional network so that a convolutional layer of the pair generates a first feature map for the training input image which has the same size as a second feature map which is generated by the other convolutional layer of the pair for the scaled version of the training input image; and calculating a loss between the first feature map and the second feature map; and training the convolutional neural network to reduce an overall loss which includes the calculated losses.

Claims (26)

1 . A method for training a convolutional neural network, the method comprising:

for each of a multitude of training input images:

processing the training input image by the convolutional network,

processing a scaled version of the training input image by the convolutional network,

determining a pair of convolutional layers of the convolutional network so that a convolutional layer of the pair generates a first feature map for the training input image which has the same size as a second feature map which is generated by the other convolutional layer of the pair for the scaled version of the training input image, and

calculating a loss between the first feature map and the second feature map; and

training the convolutional neural network to reduce an overall loss which includes the calculated losses.

2 . The method as recited in claim 1 , wherein the scaled version of the training input image is generated by scaling the training input image using a downsampling factor between successive convolutional layers of the convolutional network, or a power of the downsampling factor, or a reciprocal value of the downsampling factor, or a power of the reciprocal value of the downsampling factor.

3 . The method as recited in claim 1 , wherein the overall loss furthermore has a training loss for training the convolutional network for a predefined task, and the method further comprises weighting the calculated losses in the overall loss with regard to the training loss.

4 . The method as recited in claim 1 , further comprising:

determining multiple pairs of convolutional layers of the convolutional network for each of the multiplicity of training input images so that, for each pair of the multiple pairs, a convolutional layer of the pair generates a first feature map for the training input image, which has the same size as a second feature map which is generated by the other convolutional layer of the pair for a respective scaled version of the training input image, and a loss between the first feature map and the second feature map is calculated, and the overall loss includes the losses calculated for the pairs.

5 . The method as recited in claim 1 , wherein the neural network is trained using a training dataset of training input images, the multiplicity of training input images is selected from the training dataset, and scaled versions of the training input images are generated for the training input images so that the convolutional network has a pair of convolutional layers for each training input image of the multiplicity of training input images and each scaled version of the training input image, so that a convolutional layer of the pair generates a feature map for the training input image which has the same size as a feature map which is generated by the other convolutional layer of the pair for the scaled version of the training input image.

6 . A training device for a convolutional neural network, the training device configured to:

for each of a multitude of training input images:

process the training input image by the convolutional network,

process a scaled version of the training input image by the convolutional network,

determine a pair of convolutional layers of the convolutional network so that a convolutional layer of the pair generates a first feature map for the training input image which has the same size as a second feature map which is generated by the other convolutional layer of the pair for the scaled version of the training input image, and

calculate a loss between the first feature map and the second feature map; and

train the convolutional neural network to reduce an overall loss which includes the calculated losses.

7 . A non-transitory computer-readable medium on which is stored commands which, when executed by a processor, cause the processor to perform the following steps:

for each of a multitude of training input images:

processing the training input image by a convolutional network,

processing a scaled version of the training input image by the convolutional network,

determining a pair of convolutional layers of the convolutional network so that a convolutional layer of the pair generates a first feature map for the training input image which has the same size as a second feature map which is generated by the other convolutional layer of the pair for the scaled version of the training input image, and

calculating a loss between the first feature map and the second feature map; and

training the convolutional neural network to reduce an overall loss which includes the calculated losses.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 24, 2023
From: KAPELNER, TAMAS; WENZEL, THOMAS
To: ROBERT BOSCH GMBH
Reel/Frame 063749/0294 →
Priority Claims (1)
DE 10 2022 204 722.2 · May 13, 2022 · national
Continuity (1)
Related Publication 20230368334A1 · Nov 16, 2023
References Cited (6)
US 20230021551A1 · Lu · 2023 [cited by examiner]
Liu, Yanfei, Yanfei Zhong, and Qianqing Qin. “Scene classification based on multiscale convolutional neural network.” IEEE Transactions on Geoscience and Remote Sensing 56.12 (2018): 7109-7121. (Year: 2018). [cited by examiner]
Kim, Yonghyun, Bong-Nam Kang, and Daijin Kim. “SAN: Learning Relationship Between Convolutional Features for Multi-scale Object Detection.” European Conference on Computer Vision. Cham: Springer International Publishing… [cited by examiner]
Wu, Jialian, et al. “Self-mimic learning for small-scale pedestrian detection.” Proceedings of the 28th ACM International Conference on Multimedia. 2020. (Year: 2020). [cited by examiner]
Takimoglu, “What is Data Augmentation? Techniques & Examples In 2022,” AI Multiple, The Way Back Machine, 2022, pp. 1-11. [cited by applicant]
Li et al., “Delta: Deep Learning Transfer Using Feature Map With Attention for Convolutional Networks,” ICLR 2019 Conference Blind Submission, 2019, pp. 1-13. [cited by applicant]