IP Library › Granted Patent US 12,645,990
Granted Patent B2
US 12,645,990 · App. 18/051,419 · Granted Jun 2, 2026

Continual learning techniques for training models

Inventors: Sandeep Jana (Bengaluru, IN); Edwin Thomas (Bengaluru, IN); Kulbhushan Pachauri (Bengaluru, IN)
Assignee: Oracle International Corporation
G06N20/00G06V10/774
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,645,990
App. No.
18/051,419
Granted
Jun 2, 2026
Kind
B2
Abstract

Continual learning techniques are described for extending the capabilities of a base model, which is trained to predict a set of existing or base classes, to generate a target model that is capable of making predictions for both the existing or base classes and additionally for making predictions for new or custom classes. The techniques described herein enable the target model to be trained such that the model can make predictions involving both base classes and custom classes with high levels of accuracy.

Claims (145)

1 . A method comprising:

training, by a model training system, a model using a base classes training dataset comprising datapoints associated with a plurality of base classes, to generate a base model;

training, by the model training system, the base model to generate a custom model, the training the base model including:

(i) training the base model using a custom classes training dataset including training datapoints associated with a plurality of custom classes different from the plurality of base classes, to generate an intermediate custom model version,

(ii) evaluating whether a performance metric of the intermediate custom model version for the plurality of custom classes meets a custom model related performance threshold,

upon determining, based on the evaluating in (ii), that the performance metric does not meet the custom model related performance threshold, repeating (i) and (ii) until the performance metric of the intermediate custom model version meets the custom model related performance threshold, and

in response to the performance metric of the intermediate custom model version meeting the custom model related performance threshold for the plurality of custom classes, designating the intermediate custom model version as the custom model; and

training, by the model training system, the custom model to generate a target model, the training the custom model including:

(iii) training the custom model using a target model training dataset including a first plurality of datapoints from the base classes training dataset and a second plurality of datapoints from the custom classes training dataset, to generate an intermediate target model version,

(iv) evaluating whether corresponding performance metrics of the intermediate target model version for the plurality of base classes and the plurality of custom classes meet respective target model related performance thresholds for the plurality of base classes and the plurality of custom classes,

upon determining, based on the evaluating in (iv), that the performance metrics do not meet the target model related performance thresholds, repeating one or more of (iii) and (iv) until the performance metrics of the intermediate target model version meet the target model related performance thresholds, and

in response to the performance metrics of the intermediate target model version meeting the target model related performance thresholds for the plurality of base classes and the plurality of custom classes, designating the intermediate target model version as the target model,

wherein the target model is configured to, based on provided first input, make at least one prediction involving one or more classes from the plurality of base classes, and, based on provided second input, make at least one prediction involving one or more classes from the plurality of custom classes.

2 . The method of claim 1 , wherein the base model, the custom model, and the target model are neural network models.

3 . The method of claim 1 , wherein:

the training the base model using the custom classes training dataset in (i) further comprises (a) generating the intermediate custom model version by training the base model using one or more first datapoints selected from the training datapoints of the custom classes training dataset,

the evaluating in (ii) further comprises:

(b) determining the performance metric for the intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(c) comparing the performance metric to the custom model related performance threshold,

and

the repeating (i) and (ii) further comprises, upon determining, based on the comparing in (c), that the performance metric does not meet the custom model related performance threshold, performing (d) which comprises:

(e) generating a subsequent intermediate custom model version by training the intermediate custom model version, which was previously generated, using the one or more first datapoints and additional first datapoints that are selected from the training datapoints of the custom classes training dataset,

(f) determining a performance metric for the subsequent intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(g) comparing the performance metric for the subsequent intermediate custom model version to the custom model related performance threshold;

upon determining, based on the comparing in (g), that the performance metric meets the custom model related performance threshold, designating the subsequent intermediate custom model version as the custom model; and

upon determining, based on the comparing in (g), that the performance metric does not meet the custom model related performance threshold, repeating (d) until determining that the performance metric for a particular subsequent intermediate custom version generated as a result of one of repetitions of (d) meets the custom model related performance threshold, and designating a particular subsequent intermediate custom model version as the custom model.

4 . The method of claim 1 , wherein:

the model training system is provided by a cloud services provider (CSP),

the custom model related performance threshold is specified by a customer of the CSP, and

the custom classes training dataset is provided by the customer.

5 . The method of claim 1 , wherein:

the training the custom model using the target model training dataset in (iii) further comprises (a) generating the intermediate target model version by training the custom model using one or more first datapoints selected from the target model training dataset, and

the evaluating in (iv) further comprises:

(b) determining the performance metrics for the intermediate target model version using one or more second datapoints selected from the target model training dataset, and

(c) comparing the performance metrics to the target model related performance thresholds.

6 . The method of claim 1 , wherein:

the target model related performance thresholds comprise a first threshold related to a performance of the target model for the plurality of base classes and a second threshold related to a performance of the target model for the plurality of custom classes,

the model training system is provided by a cloud services provider (CSP),

at least one from among the first threshold and the second threshold is specified by a customer of the CSP, and

the custom classes training dataset is provided by the customer.

7 . The method of claim 1 , wherein the training the custom model to generate the target model further comprises:

prior to the evaluating in (iv), receiving a first threshold related to a performance of the target model with respect to the plurality of base classes among the target model related performance thresholds, and a second threshold related to a performance of the target model with respect to the plurality of custom classes among the target model related performance thresholds.

8 . The method of claim 7 , wherein the receiving further comprises:

receiving at least one from among the first threshold and the second threshold through a user interface subsystem of the model training system.

9 . The method of claim 7 , wherein;

the training the custom model using the target model training dataset in (iii) further comprises (a) generating the intermediate target model version of a current epoch by training the custom model using one or more first datapoints among the first plurality of datapoints and one or more first datapoints among the second plurality of datapoints,

the evaluating in (iv) further comprises:

(b) determining, for the intermediate target model version, a first performance metric, among the performance metrics, for base classes datapoints using one or more second datapoints among the first plurality of datapoints,

(c) determining, for the intermediate target model version, a second performance metric, among the performance metrics, for custom classes datapoints using one or more second datapoints among the second plurality of datapoints,

(d) determining, for the intermediate target model version, an overall performance metric for mixed classes datapoints including at least one datapoint selected from the first plurality of datapoints and at least one datapoint selected from the second plurality of datapoints,

(e) comparing the first performance metric to the first threshold, the second performance metric to the second threshold, and the overall performance metric to a previously determined overall performance metric which was determined for a previously generated intermediate target model version generated in a previous epoch; and

(f) based on the comparing in (e), determining whether at least one condition from among conditions including the first performance metric being not less than the first threshold, the second performance metric being not less than the second threshold, and the overall performance metric exceeding the previously determined overall performance metric is not satisfied;

the repeating the one or more (iii) and (iv) further comprises, upon the determining that the at least one condition is not satisfied, repeating one or more of (a), (b), (c), (d), (e), and (f); and

upon the determining that the conditions including the first performance metric being not less than the first threshold, the second performance metric being not less than the second threshold, and the overall performance metric exceeding the previously determined overall performance metric are satisfied, designating the intermediate target model version of the current epoch as the target model.

10 . The method of claim 9 , further comprising determining that the at least one condition is not satisfied,

wherein:

the second threshold has a first value, and

the repeating the one or more of (a), (b), (c), (d), (e), and (f) comprises:

receiving a second value for the second threshold that is lower than the first value, the second threshold being related to the performance of the target model with respect to the plurality of custom classes; and

repeating at least (e) and (f) using the second value for the second threshold.

11 . The method of claim 10 , wherein the receiving the second value for the second threshold further comprises receiving the second value for the second threshold through a user interface subsystem of the model training system.

12 . The method of claim 9 , further comprising determining that the at least one condition is not satisfied,

wherein the repeating the one or more of (a), (b), (c), (d), (e), and (f) comprises:

repeating (a), by modifying at least one from among a set comprising the one or more first datapoints among the first plurality of datapoints and a set comprising the one or more first datapoints among the second plurality of datapoints; and

subsequently repeating (b), (c), (d), (e), and (f).

13 . The method of claim 12 , further comprising determining that the at least one condition is not satisfied by determining that the first performance metric is less than the first threshold or the second performance metric is less than the second threshold,

wherein the repeating (a) further comprises:

adding more datapoints corresponding to the plurality of base classes to the set comprising the one or more first datapoints among the first plurality of datapoints if the first performance metric is less than the first threshold, or

adding more training datapoints corresponding to the plurality of custom classes to the set comprising the one or more first datapoints among the second plurality of datapoints if the second performance metric is less than the second threshold.

14 . The method of claim 9 , further comprising determining that the at least one condition is not satisfied,

wherein the repeating the one or more of (a), (b), (c), (d), (e), and (f) comprises:

determining a first difference between a value associated with the first performance metric for the base classes datapoints and the first threshold;

determining a second difference between a value associated with the second performance metric for the custom classes datapoints and the second threshold;

based on the first difference and the second difference, determining whether a performance of the intermediate target model version is better for the base classes datapoints or the custom classes datapoints, wherein the performance of the intermediate target model version is better for the custom classes datapoints if the first difference exceeds the second difference, and the performance of the intermediate target model version is better for the base classes datapoints if the second difference exceeds the first difference;

repeating (a) by performing one from among adding more datapoints to the one or more first datapoints of the second plurality of datapoints, upon determining that the performance of the intermediate target model version is better for the base classes datapoints, and adding more datapoints to the one or more first datapoints of the first plurality of datapoints, upon determining that the performance of the intermediate target model version is better for the custom classes datapoints; and

subsequently repeating (b), (c), (d), (e), and (f).

15 . A non-transitory computer-readable medium storing computer-executable instructions that, when executed by one or more computer systems of a model training system, cause the model training system to perform a method including:

training a model using a base classes training dataset comprising datapoints associated with a plurality of base classes, to generate a base model;

training the base model to generate a custom model, the training the base model including:

(i) training the base model using a custom classes training dataset including training datapoints associated with a plurality of custom classes different from the plurality of base classes, to generate an intermediate custom model version,

(ii) evaluating whether a performance metric of the intermediate custom model version for the plurality of custom classes meets a custom model related performance threshold,

upon determining, based on the evaluating in (ii), that the performance metric does not meet the custom model related performance threshold, repeating (i) and (ii) until the performance metric of the intermediate custom model version meets the custom model related performance threshold, and

in response to the performance metrics of the intermediate custom model version meeting the custom model related performance threshold for the plurality of custom classes, designating the intermediate custom model version as the custom model; and

training the custom model to generate a target model, the training the custom model including:

(iii) training the custom model using a target model training dataset including a first plurality of datapoints from the base classes training dataset and a second plurality of datapoints from the custom classes training dataset, to generate an intermediate target model version,

(iv) evaluating whether corresponding performance metrics of the intermediate target model version for the plurality of base classes and the plurality of custom classes meet respective target model related performance thresholds for the plurality of base classes and the plurality of custom classes,

upon determining, based on the evaluating in (iv), that the performance metrics do not meet the target model related performance thresholds, repeating one or more of (iii) and (iv) until the performance metrics of the intermediate target model version meet the target model related performance thresholds, and

in response to the performance metrics of the intermediate target model version meeting the target model related performance thresholds for the plurality of base classes and the plurality of custom classes, designating the intermediate target model version as the target model,

wherein the target model is configured to, based on provided first input, make at least one prediction involving one or more classes from the plurality of base classes, and, based on provided second input, make at least one prediction involving one or more classes from the plurality of custom classes.

16 . The non-transitory computer-readable medium of claim 15 , wherein:

the training the base model using the custom

classes training dataset in (i) further includes (a) generating the intermediate custom model version by training the base model using one or more first datapoints selected from the training datapoints of the custom classes training dataset,

the evaluating in (ii) further includes:

(b) determining the performance metric for the intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(c) comparing the performance metric to the custom model related performance threshold,

and

the repeating (i) and (ii) further includes, upon determining, based on the comparing in (c), that the performance metric does not meet the custom model related performance threshold, performing (d) which includes:

(e) generating a subsequent intermediate custom model version by training the intermediate custom model version, which was previously generated, using the one or more first datapoints and additional first datapoints that are selected from the training datapoints of the custom classes training dataset,

(f) determining a performance metric for the subsequent intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(g) comparing the performance metric for the subsequent intermediate custom model version to the custom model related performance threshold;

upon determining, based on the comparing in (g), that the performance metric meets the custom model related performance threshold, designating the subsequent intermediate custom model version as the custom model; and

upon determining, based on the comparing in (g), that the performance metric does not meet the custom model related performance threshold, repeating (d) until determining that the performance metric for a particular subsequent intermediate custom version generated as a result of one of repetitions of (d) meets the custom model related performance threshold, and designating a particular subsequent intermediate custom model version as the custom model,

wherein:

the training the custom model using the target model training dataset in (iii) further includes (h) generating the intermediate target model version by training the custom model using one or more first datapoints selected from the target model training dataset,

the evaluating in (iv) further includes:

(j) determining the performance metrics for the intermediate target model version using one or more second datapoints selected from the target model training dataset, and

(k) comparing the performance metrics to the target model related performance thresholds, and

the repeating of the one or more of (iii) and (iv) further includes, based on the comparing in (k), that a performance metric does not meet the target model related performance thresholds, repeating one or more of (h), (j), and (k) until the performance metrics of the intermediate target model version meet the target model related performance thresholds.

17 . The non-transitory computer-readable medium of claim 15 , wherein the training the custom model to generate the target model further includes:

prior to the evaluating in (iv), receiving a first threshold related to a performance of the target model with respect to the plurality of base classes among the target model related performance thresholds, and a second threshold related to a performance of the target model with respect to the plurality of custom classes among the target model related performance thresholds, and

wherein at least one from among the first threshold and the second threshold is received through a user interface subsystem of the model training system.

18 . A system comprising:

one or more computer systems configured to perform a method including:

training a model using a base classes training dataset comprising datapoints associated with a plurality of base classes, to generate a base model;

training the base model to generate a custom model, the training the base model including:

(i) training the base model using a custom classes training dataset including training datapoints associated with a plurality of custom classes different from the plurality of base classes, to generate an intermediate custom model version,

(ii) evaluating whether a performance metric of the intermediate custom model version for the plurality of custom classes meets a custom model related performance threshold,

upon determining, based on the evaluating in (ii), that the performance metric does not meet the custom model related performance threshold, repeating (i) and (ii) until the performance metric of the intermediate custom model version meets the custom model related performance threshold, and

in response to the performance metrics of the intermediate custom model version meeting the custom model related performance threshold for the plurality of custom classes, designating the intermediate custom model version as the custom model; and

training the custom model to generate a target model, the training the custom model including:

(iii) training the custom model using a target model training dataset including a first plurality of datapoints from the base classes training dataset and a second plurality of datapoints from the custom classes training dataset, to generate an intermediate target model version,

(iv) evaluating whether corresponding performance metrics of the intermediate target model version for the plurality of base classes and the plurality of custom classes meet respective target model related performance thresholds for the plurality of base classes and the plurality of custom classes,

upon determining, based on the evaluating in (iv), that the performance metrics do not meet the target model related performance thresholds, repeating one or more of (iii) and (iv) until the performance metrics of the intermediate target model version meet the target model related performance thresholds, and

in response to the performance metrics of the intermediate target model version meeting the target model related performance thresholds for the plurality of base classes and the plurality of custom classes, designating the intermediate target model version as the target model,

wherein the target model is configured to, based on provided first input, make at least one prediction involving one or more classes from the plurality of base classes, and, based on provided second input, make at least one prediction involving one or more classes from the plurality of custom classes.

19 . The system of claim 18 , wherein:

the training the base model using the custom classes training dataset in (i) further includes (a) generating the intermediate custom model version by training the base model using one or more first datapoints selected from the training datapoints of the custom classes training dataset,

the evaluating in (ii) further includes:

(b) determining the performance metric for the intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(c) comparing the performance metric to the custom model related performance threshold,

and

the repeating (i) and (ii) further includes, upon determining, based on the comparing in (c), that the performance metric does not meet the custom model related performance threshold, performing (d) which includes:

(e) generating a subsequent intermediate custom model version by training the intermediate custom model version, which was previously generated, using the one or more first datapoints and additional first datapoints that are selected from the training datapoints of the custom classes training dataset,

(f) determining a performance metric for the subsequent intermediate custom model version using one or more second datapoints selected from the training datapoints of the custom classes training dataset, and

(g) comparing the performance metric for the subsequent intermediate custom model version to the custom model related performance threshold;

upon determining, based on the comparing in (g), that the performance metric meets the custom model related performance threshold, designating the subsequent intermediate custom model version as the custom model; and

upon determining, based on the comparing in (g), that the performance metric does not meet the custom model related performance threshold, repeating (d) until determining that the performance metric for a particular subsequent intermediate custom version generated as a result of one of repetitions of (d) meets the custom model related performance threshold, and designating a particular subsequent intermediate custom model version as the custom model,

wherein the training the custom model using the target model training dataset in (iii) further includes (h) generating the intermediate target model version by training the custom model using one or more first datapoints selected from the target model training dataset,

the evaluating in (iv) further includes:

(j) determining the performance metrics for the intermediate target model version using one or more second datapoints selected from the target model training dataset, and

(k) comparing the performance metrics to the target model related performance thresholds, and

the repeating of the one or more of (iii) and (iv) further includes, based on the comparing in (k), that a performance metric does not meet the target model related performance thresholds, repeating one or more of (h), (j), and (k) until the performance metrics of the intermediate target model version meet the target model related performance thresholds.

20 . The system of claim 18 , wherein the training the custom model to generate the target model further includes:

prior to the evaluating in (iv), receiving a first threshold related to a performance of the target model with respect to the plurality of base classes among the target model related performance thresholds, and a second threshold related to a performance of the target model with respect to the plurality of custom classes among the target model related performance thresholds, and

wherein at least one from among the first threshold and the second threshold is received through a user interface subsystem of the system.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 1, 2022
From: JANA, SANDEEP; THOMAS, EDWIN; PACHAURI, KULBHUSHAN
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 061609/0223 →
Continuity (1)
Related Publication 20240144081A1 · May 2, 2024
References Cited (24)
US 11989626B2 · Arnold · 2024 [cited by examiner]
US 20210097407A1 · Pan · 2021 [cited by examiner]
US 20210183367A1 · Sharifi · 2021 [cited by examiner]
US 20210350274A1 · Pfitzmann · 2021 [cited by examiner]
US 20210365793A1 · Surya · 2021 [cited by examiner]
US 20220044149A1 · Rand · 2022 [cited by examiner]
US 20220164667A1 · Saki · 2022 [cited by examiner]
“OCI Vision”, Available Online at: https://www.oracle.com/in/artificial-intelligence/vision/, Accessed from Internet on May 6, 2022, 7 pages. [cited by applicant]
Baylor et al., “Continuous Training for Production ML in the TensorFlow Extended (TFX) Platform”, 2019 USENIX Conference on Operational Machine Learning (OpML 19), May 20, 2019, pp. 51-53. [cited by applicant]
De Lange et al., “A Continual Learning Survey: Defying Forgetting in Classification Tasks”, IEEE Transactions on Pattern Analysis and Machine Intelligence, Available Online at: https://arxiv.org/pdf/1909.08383.pdf, Apr.… [cited by applicant]
Diethe et al., “Continual Learning in Practice”, 32nd Conference on Neural Information Processing Systems (NIPS 2018), Available Online at: arXiv preprint arXiv:1903.05202, Mar. 18, 2019, 9 pages. [cited by applicant]
Hung et al., “Compacting, Picking and Growing for Unforgetting Continual Learning”, Advances in Neural Information Processing Systems 32 (2019), 11 pages. [cited by applicant]
Hung et al., “Compacting, Picking and Growing for Unforgetting Continual Learning”, Available Online at: arXiv preprint arXiv:1910.06562, Oct. 30, 2019, 12 pages. [cited by applicant]
Joseph et al., “Towards Open World Object Detection”, Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition, Available Online at: https://arxiv.org/pdf/2103.02603.pdf, May 9, 2021, 16 pages. [cited by applicant]
Kaushik et al., “Understanding Catastrophic Forgetting and Remembering in Continual Learning with Optimal Relevance Mapping”, Available Online at: arXiv preprint arXiv:2102.11343, Feb. 22, 2021, 17 pages. [cited by applicant]
Kirkpatrick et al., “Overcoming Catastrophic Forgetting in Neural Networks”, Proceedings of the National Academy of Sciences (PNAS), vol. 114, No. 13, Available Online at: https://www.pnas.org/doi/pdf/10.1073/pnas.16118… [cited by applicant]
Li et al., “Learning without Forgetting”, IEEE Transactions on Pattern Analysis and Machine Intelligence, vol. 40, No. 12, Available Online at: https://arxiv.org/pdf/1606.09282.pdf, Feb. 14, 2017, pp. 1-13. [cited by applicant]
Masana et al., “Class-Incremental Learning: Survey and Performance Evaluation on Image Classification”, Available Online at: arXiv preprint arXiv:2010.15277, May 6, 2021, pp. 1-26. [cited by applicant]
Prabhu et al., “GDumb: A Simple Approach that Questions Our Progress in Continual Learning”, European Conference on Computer Vision, Available Online at: https://www.ecva.net/papers/eccv_2020/papers_ECCV/papers/12347051… [cited by applicant]
Rebuffi et al., “iCaRL: Incremental Classifier and Representation Learning”, Proceedings of the IEEE conference on Computer Vision and Pattern Recognition, Available Online at: https://arxiv.org/pdf/1611.07725.pdf, Apr.… [cited by applicant]
Swaminathan et al., “Now Easily Perform Incremental Learning on Amazon SageMaker”, Available Online at: https://aws.amazon.com/blogs/machine-learning/now-easily-perform-incremental-learning-on-amazon-sagemaker/, Nov. 7,… [cited by applicant]
Tu et al., “Extending Conditional Convolution Structures For Enhancing Multitasking Continual Learning”, 2020 Asia-Pacific Signal and Information Processing Association Annual Summit and Conference (APSIPA ASC). IEEE, 2… [cited by applicant]
Wang et al., “Wanderlust: Online Continual Object Detection in the Real World”, Proceedings of the IEEE/CVF International Conference on Computer Vision, Available Online at: https://arxiv.org/pdf/2108.11005.pdf, Sep. 7,… [cited by applicant]
“Incremental Training in Amazon SageMaker”, Available Online at: https://docs.aws.amazon.com/sagemaker/latest/dg/incremental-training.html, Accessed from Internet on May 6, 2022, 3236 pages. [cited by applicant]