IP Library › Granted Patent US 11,989,871
Granted Patent B2
US 11,989,871 · App. 17/336,968 · Granted May 21, 2024

Model training apparatus and method

Inventors: Joseph Henry (Edinburgh, GB); Aneta Lisowska (Edinburgh, GB)
Assignee: CANON MEDICAL SYSTEMS CORPORATION
G06T7/0004G06N3/045G06N3/08G06T7/11G06T2207/20081G06T2207/20084G06T2207/30016
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,989,871
App. No.
17/336,968
Granted
May 21, 2024
Kind
B2
Abstract

An apparatus comprises processing circuitry configured to receive a first model and a second model; determine difference information that is representative of a difference between the first model and the second model and/or between the first task and the second task and/or between the first domain and the second domain; and generate a third model using the first model, the second model and the difference information, wherein the generating of the third model comprises training the third model to perform both of the first task and the second task and/or to operate on both the first domain and the second domain.

Claims (36)

1. An apparatus comprising processing circuitry configured to:

receive a first model and a second model, wherein:

the first model is trained to perform a first task;

the first model is trained on first training data of a first domain;

the second model is trained on second, different training data; and at least one of a) and b):—

a) the second model is trained to perform a second task that is different from the first task;

b) the second training data is data of a second domain that is different from the first domain;

determine difference information that is representative of a difference between the first model and the second model and/or between the first task and the second task and/or between the first domain and the second domain; and

generate a third model using the first model, the second model and the difference information, wherein the generating of the third model comprises training the third model to perform both of the first task and the second task and/or to operate on both the first domain and the second domain.

2. An apparatus according to claim 1 , wherein the third model is a distil model.

3. An apparatus according to claim 1 , wherein the training of the third model comprises using a loss function that comprises or is derived from the determined difference information.

4. An apparatus according to claim 3 , wherein the loss function comprises a regularization term that is configured to balance a difference between the first model and the third model with a difference between the second model and the third model, so as to discourage bias in the training towards either the first task or the second task and/or towards either the first domain or the second domain.

5. An apparatus according to claim 4 , wherein the regularization term comprises a weighting based on model parameters of the first model, second model and third model.

6. An apparatus according to claim 5 , wherein the weighting is a decaying weighting.

7. An apparatus according to claim 3 , wherein the loss function further comprises a distillation loss term that comprises a difference between outputs of the first model and outputs of the third model, and a difference between outputs of the second model and outputs of the third model.

8. An apparatus according to claim 1 , wherein the training of the third model is performed using the second training data.

9. An apparatus according to claim 1 , wherein the training of the third model is performed without access to the first training data.

10. An apparatus according to claim 1 , wherein the second model is trained to perform the second task that is different from the first task, and wherein the second model is not trained to perform the first task.

11. An apparatus according to claim 1 , wherein the determining of the difference information is based at least partly on feature extractors of the first model and the second model.

12. An apparatus according to claim 1 , wherein the determining of the difference information is based at least partly on a difference in distribution of task intensities between the first task and the second task and/or between the first domain and the second domain.

13. An apparatus according to claim 1 , wherein the determining of the difference information is based at least partly on a learned distance metric.

14. An apparatus according to claim 13 , wherein the user input comprises information relating to anatomical similarity or intensity distribution.

15. An apparatus according to claim 1 , wherein the determining of the difference information is based at least partly on user input.

16. An apparatus according to claim 1 , wherein the determining of the difference information is based at least partly on a prediction by a further trained model.

17. An apparatus according to claim 1 , wherein the first training data and the second training data each comprise medical imaging data.

18. An apparatus according to claim 1 , wherein the first task comprises segmentation of a first anatomical feature or pathology and the second task comprises segmentation of a second, different anatomical feature or pathology.

19. An apparatus according to claim 1 , wherein the first domain and the second domain relate to different locations, and/or different scanners, and/or different imaging modalities.

20. A method comprising:

receiving a first model and a second model, wherein:

the first model is trained to perform a first task;

the first model is trained on first training data of a first domain;

the second model is trained on second, different training data; and at least one of a) and b):—

a) the second model is trained to perform a second task that is different from the first task;

b) the second training data is data of a second domain that is different from the first domain;

determining difference information that is representative of a difference between the first model and the second model and/or between the first task and the second task and/or between the first domain and the second domain; and

generating a third model using the first model, the second model and the difference information, wherein the generating of the third model comprises training the third model to perform both of the first task and the second task and/or to operate on both the first domain and the second domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2021
From: HENRY, JOSEPH; LISOWSKA, ANETA; CANON MEDICAL RESEARCH EUROPE, LTD.
To: CANON MEDICAL SYSTEMS CORPORATION
Reel/Frame 057000/0213 →
Continuity (1)
Related Publication 20220392048A1 · Dec 8, 2022