IP Library Granted Patent US 11,775,812
Granted Patent B2
US 11,775,812 · App. 16/379,704 · Granted Oct 3, 2023

Multi-task based lifelong learning

Inventors: Jie Zhang (San Jose, CA); Junting Zhang (Los Angeles, CA); Shalini Ghosh (Menlo Park, CA); Dawei Li (San Jose, CA); Jingwen Zhu (Santa Clara, CA)
Assignee: Samsung Electronics Co., Ltd.
G06N3/08G06N3/04
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,775,812
App. No.
16/379,704
Granted
Oct 3, 2023
Kind
B2
Abstract

Methods, devices, and computer-readable media for multi-task based lifelong learning. A method for lifelong learning includes identifying a new task for a machine learning model to perform. The machine learning model trained to perform an existing task. The method includes adaptively training a network architecture of the machine learning model to generate an adapted machine learning model based on incorporating inherent correlations between the new task and the existing task. The method further includes using the adapted machine learning model to perform both the existing task and the new task.

Claims (56)

1. A method for lifelong learning, the method comprising:

identifying a new task for a machine learning model to perform, the machine learning model trained to perform an existing task;

adaptively training a network architecture of the machine learning model to generate an adapted machine learning model based on incorporating inherent correlations between the new task and the existing task, wherein adaptively training the network architecture includes:

generating a plurality of child network architectures, wherein each of the plurality of child network architectures is expanded from a size of the network architecture by at least one of: adding one or more new layers to the network architecture or expanding one or more existing layers of the network architecture; and

determining an optimal child network architecture from the plurality of child network architectures for the adapted machine learning model; and

using the adapted machine learning model to perform both the existing task and the new task.

2. The method of claim 1 , wherein, for each of the plurality of child network architectures, the size of the network architecture is expanded using AutoML.

3. The method of claim 2 , wherein expanding the size of the network architecture using AutoML comprises at least one of:

using a deeper operator to add the one or more new layers to the network architecture; and

using a wider operator to expand the one or more existing layers of the network architecture.

4. The method of claim 1 , further comprising:

identifying the one or more new layers as at least one task-specific layer for the new task.

5. The method of claim 1 , further comprising:

compressing the optimal child network architecture to reduce a size of the optimal child network architecture;

wherein the size of the optimal child network architecture is not compressed smaller than the size of the network architecture.

6. The method of claim 1 , wherein the machine learning model is a compressed model.

7. The method of claim 1 , wherein adaptively training the network architecture further comprises:

training the machine learning model to perform the new task using training data for the new task; and

compressing the optimal child network architecture of the trained machine learning model using the training data for the new task.

8. An electronic device for lifelong learning, the electronic device comprising:

a memory configured to store a machine learning model trained to perform an existing task; and

a processor operably connected to the memory, the processor configured to:

identify a new task for the machine learning model to perform;

adaptively train a network architecture of the machine learning model to generate an adapted machine learning model based on incorporating inherent correlations between the new task and the existing task, wherein, to adaptively train the network architecture, the processor is configured to:

generate a plurality of child network architectures, wherein each of the plurality of child network architectures is expanded from a size of the network architecture by at least one of: adding one or more new layers to the network architecture or expanding one or more existing layers of the network architecture; and

determine an optimal child network architecture from the plurality of child network architectures for the adapted machine learning model; and

use the adapted machine learning model to perform both the existing task and the new task.

9. The electronic device of claim 8 , wherein, for each of the plurality of child network architectures, the processor is configured to expand the size of the network architecture using AutoML.

10. The electronic device of claim 9 , wherein, to expand the size of the network architecture using AutoML, the processor is configured to at least one of:

use a deeper operator to add the one or more new layers to the network architecture; and

use a wider operator to expand the one or more existing layers of the network architecture.

11. The electronic device of claim 8 , wherein the processor is further configured to identify the one or more new layers as at least one task-specific layer for the new task.

12. The electronic device of claim 8 , wherein:

the processor is further configured to compress the optimal child network architecture to reduce a size of the optimal child network architecture; and

the size of the optimal child network architecture is not compressed smaller than the size of the network architecture.

13. The electronic device of claim 8 , wherein the machine learning model is a compressed model.

14. The electronic device of claim 8 , wherein, to adaptively train the network architecture, the processor is further configured to:

train the machine learning model to perform the new task using training data for the new task; and

compress the optimal child network architecture of the trained machine learning model using the training data for the new task.

15. A non-transitory, computer-readable medium comprising program code for lifelong learning that, when executed by a processor of an electronic device, causes the electronic device to:

identify a new task for a machine learning model to perform, the machine learning model trained to perform an existing task;

adaptively train a network architecture of the machine learning model to generate an adapted machine learning model based on incorporating inherent correlations between the new task and the existing task; and

use the adapted machine learning model to perform both the existing task and the new task;

wherein the program code that, when executed by the processor, causes the electronic device to adaptively train the network architecture comprises program code that, when executed by the processor, causes the electronic device to:

generate a plurality of child network architectures, wherein each of the plurality of child network architectures is expanded from a size of the network architecture by at least one of: adding one or more new layers to the network architecture or expanding one or more existing layers of the network architecture; and

determine an optimal child network architecture from the plurality of child network architectures for the adapted machine learning model.

16. The non-transitory, computer-readable medium of claim 15 , wherein the program code that, when executed by the processor, causes the electronic device to generate the plurality of child network architectures comprises program code that, when executed by the processor, causes the electronic device to, for each of the plurality of child network architectures, expand the size of the network architecture using AutoML.

17. The non-transitory, computer-readable medium of claim 16 , wherein the program code that, when executed by the processor, causes the electronic device to expand the size of the network architecture using AutoML comprises program code that, when executed by the processor, causes the electronic device to at least one of:

use a deeper operator to add the one or more new layers to the network architecture; and

use a wider operator to expand the one or more existing layers of the network architecture.

18. The non-transitory, computer-readable medium of claim 15 , further comprising program code that, when executed by the processor, causes the electronic device to identify the one or more new layers as at least one task-specific layer for the new task.

19. The non-transitory, computer-readable medium of claim 15 , further comprising program code that, when executed by the processor, causes the electronic device to compress the optimal child network architecture of the machine learning model to reduce a size of the optimal child network architecture;

wherein the size of the optimal child network architecture is not compressed smaller than the size of the network architecture.

20. The non-transitory, computer-readable medium of claim 15 , wherein the program code that, when executed by the processor, causes the electronic device to adaptively train the network architecture further comprises program code that, when executed by the processor, causes the electronic device to:

train the machine learning model to perform the new task using training data for the new task; and

compress the optimal child network architecture of the trained machine learning model using the training data for the new task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 15, 2019
From: ZHANG, JIE; ZHANG, JUNTING; GHOSH, SHALINI; LI, DAWEI; ZHU, JINGWEN
To: SAMSUNG ELECTRONICS CO., LTD
Reel/Frame 048889/0565 →
Continuity (3)
Provisional Application 62792052 · Jan 14, 2019
Provisional Application 62774043 · Nov 30, 2018
Related Publication 20200175362A1 · Jun 4, 2020