IP Library › Granted Patent US 12,462,549
Granted Patent B2
US 12,462,549 · App. 18/262,984 · Granted Nov 4, 2025

Method for determining the encoder architecture of a neural network

Inventors: David Ivan (Balatonszemes, HU); Regina Deak-Meszlenyi (Budapest, HU); Csaba Nemes (Dunakeszi, HU)
Assignee: CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
G06V10/82G06N3/0455
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,462,549
App. No.
18/262,984
Granted
Nov 4, 2025
Kind
B2
Abstract

A method for determining an encoder architecture of a convolutional neural network configured to process image processing tasks. For each image processing task), characteristic scale distribution is calculated based on training data. Encoder architecture candidates are generated, each including a shared encoder layer providing computational operations for image processing tasks and branches which span over encoder layers providing at least partly different computational operations for the image processing tasks. Each branch is associated with a certain image processing task. Receptive encoder layer field sizes and assessment measures are calculated, each assessment measure referring to a combination of a certain encoder architecture and a certain image processing task, and including information regarding matching quality of characteristic scale distribution associated with the assessment measure to the receptive field sizes of the encoder layers. The assessment measures are compared and a comparison result established. An encoder architecture is selected based on the comparison result.

Claims (28)

1 . A method for determining an architecture of an encoder of a convolutional neural network, the neural network being configured to process multiple different image processing tasks, the method comprising:

for each image processing task, calculating a characteristic scale distribution based on training data, the characteristic scale distribution indicating a size distribution of objects to be detected by the respective image processing task;

generating multiple encoder architecture candidates, each encoder architecture of the encoder architecture candidates comprising at least one shared encoder layer which provides computational operations for multiple image processing tasks and multiple branches which span over one or more encoder layers which provide at least partly different computational operations for the image processing tasks, wherein each branch is associated with a certain image processing task;

calculating receptive field sizes of the encoder layers of the multiple encoder architectures;

calculating multiple assessment measures, each assessment measure referring to a combination of a certain encoder architecture of the multiple encoder architectures and a certain image processing task, each assessment measure including information regarding the quality of matching of characteristic scale distribution of the image processing task associated with the assessment measure to the receptive field sizes of the encoder layers of the encoder architecture associated with the assessment measure;

comparing the calculated assessment measures and establishing a comparison result; and

selecting an encoder architecture based on the comparison result.

2 . The method according to claim 1 , further comprising determining for the characteristic scale distribution of each image processing task, one or more percentiles, values of the one or more percentiles depending on an interface of a task-specific decoder.

3 . The method according to claim 2 , wherein a number of determined percentiles is chosen according to a number of input connections required by the task-specific decoder.

4 . The method according to claim 2 , wherein the values of the percentiles are distributed across a percentile range defined by a minimum and a maximum percentile value.

5 . The method according to claim 4 , wherein the values of the percentiles are equally distributed across the percentile range.

6 . The method according to claim 1 , wherein at least some of the encoder layers of a task-specific branch of the encoder comprise a feature map with a resolution which matches to a feature resolution of an input connection of a task-specific decoder.

7 . The method according to claim 1 , wherein generating multiple encoder architecture candidates comprises selecting certain building blocks, which are adapted for an encoder of a neural network, out of a set of building blocks and connecting the selected building blocks in order to obtain an encoder architecture.

8 . The method according to claim 1 , wherein the encoder architectures of the encoder architecture candidates are determined by defining a runtime limit and the encoder architectures are created such that, at a given hardware, the encoder architectures provide computational results in a runtime range of 90% to 100% of the runtime limit.

9 . The method according to claim 1 , wherein calculating multiple assessment measures comprises comparing the receptive field size of an encoder layer which matches with a resolution of a task-specific decoder with a determined percentile value and determining a distance between the receptive field size and the percentile value.

10 . The method according to claim 1 , wherein calculating multiple assessment measures comprises comparing multiple receptive field sizes of encoder layers which match with a resolution of a task-specific decoder with multiple determined percentile values and determining a sum of distances between the receptive field sizes and the percentile values.

11 . The method according to claim 1 , wherein calculating multiple assessment measures uses a loss function with a least absolute deviations distance measure, the absolute deviations distance measure comprising an L1 distance measure or a cosine distance measure.

12 . The method according to claim 1 , wherein calculating multiple assessment measures uses a loss function with a penalty term, the penalty term being adapted to increase the assessment measure with a decreasing number of layers shared between multiple image processing tasks of an encoder architecture.

13 . The method according to claim 1 , wherein comparing the calculated assessment measures and establishing a comparison result comprises determining an encoder architecture for each image processing task which has a lowest assessment measure.

14 . The method to claim 1 , wherein selecting an encoder architecture based on the comparison result comprises selecting, for each image processing task, the encoder architecture which has a lowest assessment measure.

15 . A computer program product for determining the architecture of an encoder of a convolutional neural network, the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions being executable by a processor to cause the processor to execute a method comprising:

for each image processing task, calculating a characteristic scale distribution based on training data, the characteristic scale distribution indicating a size distribution of objects to be detected by the respective image processing task;

generating multiple encoder architecture candidates, each encoder architecture of the encoder architecture candidates comprising at least one shared encoder layer which provides computational operations for multiple image processing tasks and multiple branches which span over one or more encoder layers which provide at least partly different computational operations for the image processing tasks, wherein each branch is associated with a certain image processing task;

calculating receptive field sizes of the encoder layers of the multiple encoder architectures;

calculating multiple assessment measures, each assessment measure referring to a combination of a certain encoder architecture of the multiple encoder architectures and a certain image processing task, each assessment measure including information regarding the quality of matching of characteristic scale distribution of the image processing task associated with the assessment measure to the receptive field sizes of the encoder layers of the encoder architecture associated with the assessment measure;

comparing the calculated assessment measures and establishing a comparison result; and

selecting an encoder architecture based on the comparison result.

16 . The computer program product according to claim 15 , wherein the method executed by the processor further comprises determining, for the characteristic scale distribution of each image processing task, one or more percentiles, values of the one or more percentiles depending on an interface of a task-specific decoder.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2025
From: IVAN, DAVID; DEAK-MESZLENYI, REGINA
To: CONTINENTAL AUTONOMOUS MOBILITY GERMANY GMBH
Reel/Frame 069838/0701 →
Priority Claims (1)
EP 21153462 · Jan 26, 2021 · regional
Continuity (1)
Related Publication 20240127585A1 · Apr 18, 2024
References Cited (37)
US 11189049B1 · Chakravarty · 2021 [cited by examiner]
US 20180247201A1 · Liu · 2018 [cited by examiner]
US 20190354818A1 · Reisswig · 2019 [cited by examiner]
US 20190370648A1 · Zoph et al. · 2019 [cited by applicant]
US 20200193269A1 · Park · 2020 [cited by examiner]
US 20210056691A1 · Gernand · 2021 [cited by examiner]
US 20210271870A1 · Ni · 2021 [cited by examiner]
US 20210357674A1 · Ogawa · 2021 [cited by examiner]
US 20210358177A1 · Park · 2021 [cited by examiner]
CN 109858372A · 2019 [cited by applicant]
CN 110705695A · 2020 [cited by applicant]
CN 111209383A · 2020 [cited by applicant]
CN 111667728A · 2020 [cited by applicant]
CN 111819580A · 2020 [cited by applicant]
JP 2021111388A · 2021 [cited by applicant]
Eric Crawford et al. , “Spatially Invariant Unsupervised Object Detection with Convolutional Neural Networks, ”Jul. 17, 2019, AAAI Technical Track: Machine Learning, vol. 33 No. 01,AAAI-19, IAAI-19, EAAI-20, pp. 3412-34… [cited by examiner]
Zhongling Huang et al., “Transfer Learning with Deep Convolutional Neural Network for SAR Target Classification with Limited Labeled Data,” Aug. 31, 2017, Remote Sens. 2017, 9(9), 907, pp. 1-17. [cited by examiner]
Wenjie Luo et al., “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks, ”Nov. 13, 2024,30th Conference on Neural Information Processing Systems (NIPS 2016), Barcelona, Spain, pp. 1-7. [cited by examiner]
Simon Vandenhende et al., “Branched Multi-Task Networks: Deciding What Layers to Share,” Aug. 13, 2020, arXiv:1904.02920, pp. 1-14. [cited by examiner]
Trevor Standley et al., “Which Tasks Should Be Learned Together in Multi-task Learning?,” Jul. 13, 2020, ICML'20: Proceedings of the 37th International Conference on Machine Learning, Article No. 846, pp. 1-8. [cited by examiner]
Asifullah Khan et al., A survey of the recent architectures of deep convolutional neural networks,Apr. 21, 2020, Artificial Intelligence Review (2020) 53, pp. 5455-5490. [cited by examiner]
Syed Shakib Sarwar et al., “Incremental Learning in Deep Convolutional Neural Networks Using Partial Network Sharing, ”Jan. 8, 2020, IEEE Access, vol. 8, 2020, pp. 4615-4623. [cited by examiner]
Timothy Hospedales et al., “Meta-Learning in Neural Networks: A Survey,” Aug. 4, 2022, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 44, No. 9, Sep. 2022, pp. 5149-5160. [cited by examiner]
Liang-Chieh Chen et al., “Searching for Efficient Multi-Scale Architectures for Dense Image Prediction, ”Dec. 3, 2018, NIPS'18: Proceedings of the 32nd International Conference on Neural Information Processing Systems, … [cited by examiner]
Notice of Reasons for Refusal drafted Mar. 11, 2024 for the counterpart Japanese Patent Application No. 2023-537026 and machine translation of same. [cited by applicant]
European Examination Report dated Apr. 23, 2024 for the priority European Patent Application No. 21 153 462.3. [cited by applicant]
European Search Report dated Jul. 9, 2021 for the counterpart European Patent Application No. 21153462.3. [cited by applicant]
The International Search Report and the Written Opinion of the International Searching Authority mailed on May 27, 2022 for the priority PCT Application No. PCT/EP2022/051427. [cited by applicant]
Thomas Elsken et al: “Neural Architecture Search: A Survey”, CORR (ARXIV), vol. 1808.05377, No. v3, Apr. 26, 2019 (Apr. 26, 2019), pp. 1-21, XP055710476, Journal of Machine Learning Research, .vol. 20, pp. 1-21, Mar. 19… [cited by applicant]
Timothy Hospedales et al: “Meta-Learning in Neural Networks: A Survey”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Nov. 7, 2020 (Nov. 7, 2020), XP081797264. [cited by applicant]
Joaquin Vanschoren: “Meta-Learning: A Survey”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Oct. 8, 2018 (Oct. 8, 2018), XP081058788. [cited by applicant]
Liang-Chieh Chen et al.: “Searching for Efficient Multi-Scale Architectures for Dense Image Prediction”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Sep. 12, 2018 (Sep. 1… [cited by applicant]
Nikolas Adaloglou: “Understanding the receptive field of deep convolutional networks,” AI Sumner, AI, Jul. 2, 2020, XP055920333, https//theaisummer.com/receptive-field. [cited by applicant]
Wenjie Luo et al : “Understanding the Effective Receptive Field in Deep Convolutional Neural Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jan. 16, 2017 (Jan. 16… [cited by applicant]
Simon Vandenhende et al: “Branched Multi-Task Networks: Deciding what Layers to Share”, arXiv:1904.02920 [cs.CV], Nov. 2, 2019. [cited by applicant]
Trevor Standley et al: “Which Tasks Should Be Learned Together in Multi-task Learning?”, arXiv:1905.07553 [cs.CV], May 21, 2019. [cited by applicant]
Office Action (The First Office Action) issued Jul. 28, 2025, by the State Intellectual Property Office of People's Republic of China in corresponding Chinese Patent Application No. 202280008670.0 and an English transla… [cited by applicant]