IP Library › Granted Patent US 12,657,900
Granted Patent B2
US 12,657,900 · App. 18/467,017 · Granted Jun 16, 2026

Multi-task transfer learning using weight divergence constraints

Inventors: Simon Ekman Von Huth (Nacka, SE); Mohammadreza Malek-Mohammadi (Solna, SE); Saeed Dabbaghchian (Bromma, SE)
Assignee: QUALCOMM Incorporated
G06V10/82
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,657,900
App. No.
18/467,017
Granted
Jun 16, 2026
Kind
B2
Abstract

Certain aspects of the present disclosure provide techniques and apparatus for training a machine learning model based on transfer learning and weight divergence constraints. The method generally includes receiving weight information associated with a machine learning model, wherein the machine learning model comprises a model trained to perform a first task; updating the machine learning model to perform a second task based on the received weight information and a weight divergence constraint between weights defined for the first task and weights updated for the second task; and deploying the updated machine learning model.

Claims (57)

1 . A processing system, comprising:

at least one memory having executable instructions stored thereon; and

one or more processors configured to execute the executable instructions in order to cause the processing system to:

receive weight information associated with a machine learning model, wherein the machine learning model comprises a model trained to perform a first task;

update the machine learning model to perform a second task based on the received weight information and a weight divergence constraint between weights defined for the first task and weights updated for the second task; and

deploy the updated machine learning model;

wherein to update the machine learning model, the one or more processors are configured to cause the processing system to minimize a sum of a first loss function and a second loss function, wherein the first loss function comprises a task-specific loss for the second task, and wherein the second loss function comprises a similarity loss between the weights defined for the first task and the weights updated for the second task;

wherein the similarity loss comprises a loss function measuring a normalized loss based on the weights updated for the second task and the received weight information; and

wherein the loss function is based on a sum of a difference between the weights updated for the second task and the received weight information calculated over each layer in a portion of the machine learning model.

2 . The processing system of claim 1 , wherein the machine learning model comprises an encoder-decoder model including an encoder trained to map an input into a latent space and a decoder trained to make predictions based on a latent space representation of the input.

3 . The processing system of claim 1 , wherein the portion of the machine learning model comprises an encoder portion of an encoder-decoder model.

4 . The processing system of claim 1 , wherein the normalized loss is normalized based on a sum of the received weight information.

5 . The processing system of claim 1 , wherein:

the first task comprises a semantic segmentation task for performing on image data, and

the second task comprises an object detection task for performing on the image data.

6 . The processing system of claim 1 , wherein:

the first task comprises an object detection task for performing on image data, and

the second task comprises a semantic segmentation task for performing on the image data.

7 . The processing system of claim 1 , wherein the weight divergence constraint comprises a product of a similarity loss and a task-specific constant.

8 . A processor-implemented method, comprising:

receiving weight information associated with a machine learning model, wherein the machine learning model comprises a model trained to perform a first task;

updating the machine learning model to perform a second task based on the received weight information and a weight divergence constraint between weights defined for the first task and weights updated for the second task; and

deploying the updated machine learning model;

wherein updating the machine learning model comprises minimizing a sum of a first loss function and a second loss function, wherein the first loss function comprises a task-specific loss for the second task, and wherein the second loss function comprises a similarity loss between the weights defined for the first task and the weights updated for the second task;

wherein the similarity loss comprises a loss function measuring a normalized loss based on the weights updated for the second task and the received weight information; and

wherein the loss function is based on a sum of a difference between the weights updated for the second task and the received weight information calculated over each layer in a portion of the machine learning model.

9 . The method of claim 8 , wherein the machine learning model comprises an encoder-decoder model including an encoder trained to map an input into a latent space and a decoder trained to make predictions based on a latent space representation of the input.

10 . The method of claim 8 , wherein the portion of the machine learning model comprises an encoder portion of an encoder-decoder model.

11 . The method of claim 8 , wherein the normalized loss is normalized based on a sum of the received weight information.

12 . The method of claim 8 , wherein:

the first task comprises a semantic segmentation task for performing on image data, and

the second task comprises an object detection task for performing on the image data.

13 . The method of claim 8 , wherein:

the first task comprises an object detection task for performing on image data, and

the second task comprises a semantic segmentation task for performing on the image data.

14 . The method of claim 8 , wherein the weight divergence constraint comprises a product of a similarity loss and a task-specific constant.

15 . A processing system, comprising:

means for receiving weight information associated with a machine learning model, wherein the machine learning model comprises a model trained to perform a first task;

means for updating the machine learning model to perform a second task based on the received weight information and a weight divergence constraint between weights defined for the first task and weights updated for the second task; and

means for deploying the updated machine learning model;

wherein the means for updating the machine learning model comprises minimizing a sum of a first loss function and a second loss function, wherein the first loss function comprises a task-specific loss for the second task, and wherein the second loss function comprises a similarity loss between the weights defined for the first task and the weights updated for the second task;

wherein the similarity loss comprises a loss function measuring a normalized loss based on the weights updated for the second task and the received weight information; and

wherein the loss function is based on a sum of a difference between the weights updated for the second task and the received weight information calculated over each layer in a portion of the machine learning model.

16 . The processing system of claim 15 , wherein the machine learning model comprises an encoder-decoder model including an encoder trained to map an input into a latent space and a decoder trained to make predictions based on a latent space representation of the input.

17 . The processing system of claim 15 , wherein the portion of the machine learning model comprises an encoder portion of an encoder-decoder model.

18 . The processing system of claim 15 , wherein the normalized loss is normalized based on a sum of the received weight information.

19 . The processing system of claim 15 , wherein:

the first task comprises a semantic segmentation task for performing on image data, and

the second task comprises an object detection task for performing on the image data.

20 . The processing system of claim 15 , wherein the weight divergence constraint comprises a product of a similarity loss and a task-specific constant.

21 . A non-transitory computer-readable medium having executable instructions stored thereon which, when executed by one or more processors, perform an operation comprising:

receiving weight information associated with a machine learning model, wherein the machine learning model comprises a model trained to perform a first task;

updating the machine learning model to perform a second task based on the received weight information and a weight divergence constraint between weights defined for the first task and weights updated for the second task; and

deploying the updated machine learning model;

wherein to update the machine learning model, the one or more processors are configured to cause the processing system to minimize a sum of a first loss function and a second loss function, wherein the first loss function comprises a task-specific loss for the second task, and wherein the second loss function comprises a similarity loss between the weights defined for the first task and the weights updated for the second task;

wherein the similarity loss comprises a loss function measuring a normalized loss based on the weights updated for the second task and the received weight information; and

wherein the loss function is based on a sum of a difference between the weights updated for the second task and the received weight information calculated over each layer in a portion of the machine learning model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 9, 2023
From: EKMAN VON HUTH, SIMON; MALEK-MOHAMMADI, MOHAMMADREZA; DABBAGHCHIAN, SAEED
To: QUALCOMM INCORPORATED
Reel/Frame 065162/0786 →
Continuity (2)
Provisional Application 63507542 · Jun 12, 2023
Related Publication 20240412499A1 · Dec 12, 2024
References Cited (20)
US 11798180B2 · Yin · 2023 [cited by examiner]
US 12217361B1 · Chunduru · 2025 [cited by examiner]
US 20210390270A1 · Fei · 2021 [cited by examiner]
US 20230368147A1 · Adeli-Nadjafi · 2023 [cited by examiner]
US 20240144019A1 · Weinzaepfel · 2024 [cited by examiner]
US 20240296208A1 · Abbassi · 2024 [cited by examiner]
US 20250006346A1 · Kanakatte Gurumurthy · 2025 [cited by examiner]
US 20250117929A1 · Barve · 2025 [cited by examiner]
US 20250131281A1 · Lord · 2025 [cited by examiner]
EP 4354353A1 · 2024 [cited by applicant]
Bengio Y., et al., “Curriculum Learning”, International Conference on Machine Learning, Jun. 14, 2009, Association for Computing Machinery, pp. 41-48. [cited by applicant]
Li X., et al., “Explicit Inductive Bias for Transfer Learning with Convolutional Networks”, Proceedings of the 35th International Conference on Machine Learning, Jul. 2018, ISSN: 2640-3498, 10 Pages. [cited by applicant]
Parisi G.I., “Continual Lifelong Learning with Neural Networks: A Review”, arXiv:1802.07569v4 [cs.LG], Feb. 11, 2019, Neural Networks (2019), pp. 1-29. [cited by applicant]
Yosinski J., et al., “How Transferable are Features in Deep Neural Networks?”, Advances in Neural Information Processing Systems, vol. 27, Nov. 6, 2014, pp. 1-9, XP055277610. [cited by applicant]
Zamir A.R., et al., “Taskonomy: Disentangling Task Transfer Learning,” Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR), Apr. 23, 2018, 12 Pages, ArXiv:1804.08328v1 [cs.CV], pp. 13 an… [cited by applicant]
Zilly J., et al., “On Plasticity, Invariance, and Mutually Frozen Weights in Sequential Task Learning”, Advances in Neural Information Processing Systems, vol. 34, Dec. 2021, pp. 1-14. [cited by applicant]
Crawshaw M., “Multi-Task Learning with Deep Neural Networks: A Survey”, arXiv.org, Sep. 10, 2020, XP093115219, 43 Pages, Ithaca, The Whole Document. [cited by applicant]
International Search Report and Written Opinion - PCT/US2024/027084—ISA/EPO—Aug. 27, 2024. [cited by applicant]
Kumar V.R., “Surround-View Cameras Based Holistic Visual Perception for Automated Driving”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Jun. 11, 2022, 245 Pages, XP091245… [cited by applicant]
Li X., et al., “Explicit Inductive Bias for Transfer Learning with Convolutional Networks”, arxiv.org, Cornell University Library, 201 Olin Library Cornell University Ithaca, NY 14853, Feb. 5, 2018, XP081215039, Jun. 6,… [cited by applicant]