IP Library › Granted Patent US 12,236,352
Granted Patent B2
US 12,236,352 · App. 18/349,902 · Granted Feb 25, 2025

Transfer learning based on cross-domain homophily influences

Inventors: Craig M. Trim (Glendale, CA); Aaron K. Baughman (Research Triangle Park, NC); Garfield W. Vaughn (Southbury, CT); Micah Forster (Austin, TX)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/086G06N3/045
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,236,352
App. No.
18/349,902
Granted
Feb 25, 2025
Kind
B2
Abstract

Methods, computer program products, and systems are presented. The methods can include, for instance: generating a plurality of deep transfer learning networks. Further, the methods can include, for instance: encoding one or more transfer layers.

Claims (62)

1. A computer implemented method comprising:

generating a plurality of deep transfer learning networks based on a plurality of first exemplars and a plurality of second exemplars;

encoding one or more transfer layers to a chromosome for genetic operators, where the one or more transfer layers are to be transferred from a source deep transfer learning network corresponding to a source domain to a target deep transfer learning network corresponding to a target domain;

diversifying concurrently both the source deep transfer learning network and the target deep transfer learning network by use of the genetic operators; and

producing the target deep transfer learning network that integrates a result from the diversifying.

2. The computer implemented method of claim 1 , wherein a first subset of the first exemplars represents the source domain, a second subset of the first exemplars represents the target domain, and the second exemplars represent both the source domain and the target domain.

3. The computer implemented method of claim 1 , wherein the plurality of deep transfer learning networks include both the source deep transfer learning network and the target deep transfer learning network.

4. The computer implemented method of claim 1 , wherein the producing the target deep transfer learning network that integrates a result from the diversifying includes producing the target deep transfer learning network that integrates a result from the diversifying on the source deep transfer learning network and that passes a predefined fitness threshold condition for the target deep transfer learning network.

5. The computer implemented method of claim 4 , wherein the predefined fitness threshold condition is determined based on the accuracy of the target deep transfer learning network as determined from an average between ground truths specific to the target domain and ground truths applicable for both the source domain and the target domain.

6. The computer implemented method of claim 1 , further comprising:

prior to the encoding, training the source deep transfer learning network with a first subset of the first exemplars and the second exemplars; and

prior to the encoding, and concurrently with the training the source deep transfer learning network, training the target deep transfer learning network with a second subset of the first exemplars and the second exemplars.

7. The computer implemented method of claim 1 , further comprising:

prior to the encoding, measuring a homophily influence of the source deep transfer learning network to the target deep transfer learning network.

8. The computer implemented method of claim 1 , the diversifying comprising:

mutating randomly the transfer layers of the source deep transfer learning network from the encoding.

9. The computer implemented method of claim 1 , the diversifying comprising:

ascertaining that respective transfer layers of first and second source deep transfer learning networks have their respective weights migrated over to the target deep transfer learning network for respectively passing a fitness threshold condition for the source deep transfer learning network; and

crossing over the transfer layers of the target deep transfer learning network with the transfer layers of the first and second source deep transfer learning networks.

10. The computer implemented method of claim 1 , the diversifying comprising:

ascertaining that respective transfer layers of more than three source deep transfer learning networks have their respective weights migrated over to the target deep transfer learning network for respectively passing a fitness threshold condition for the source deep transfer learning network;

selecting first and second source deep transfer learning networks from the more than three source deep transfer learning network from the ascertaining; and

crossing over the transfer layers of the target deep transfer learning network with the transfer layers of the first and second source deep transfer learning networks from the selecting.

11. A computer program product comprising:

a computer readable storage medium readable by one or more processor and storing instructions for execution by the one or more processor for performing a method comprising:

generating a plurality of deep transfer learning networks based on a plurality of first exemplars and a plurality of second exemplars;

encoding one or more transfer layers to a chromosome for genetic operators, where the one or more transfer layers are to be transferred from a source deep transfer learning network corresponding to a source domain to a target deep transfer learning network corresponding to a target domain;

diversifying concurrently both the source deep transfer learning network and the target deep transfer learning network by use of the genetic operators; and

producing the target deep transfer learning network that integrates a result from the diversifying.

12. A system comprising:

a memory;

one or more processor in communication with the memory; and

program instructions executable by the one or more processor via the memory to perform a method comprising:

generating a plurality of deep transfer learning networks based on a plurality of first exemplars and a plurality of second exemplars;

encoding one or more transfer layers to a chromosome for genetic operators, where the one or more transfer layers are to be transferred from a source deep transfer learning network corresponding to a source domain to a target deep transfer learning network corresponding to a target domain;

diversifying concurrently both the source deep transfer learning network and the target deep transfer learning network by use of the genetic operators; and

producing the target deep transfer learning network that integrates a result from the diversifying.

13. The system of claim 12 , wherein a first subset of the first exemplars represents the source domain, a second subset of the first exemplars represents the target domain, and the second exemplars represent both the source domain and the target domain.

14. The system of claim 12 , wherein the plurality of the deep transfer learning networks include both the source deep transfer learning network and the target deep transfer learning network.

15. The system of claim 12 , wherein the producing the target deep transfer learning network that integrates a result from the diversifying includes producing the target deep transfer learning network that integrates a result from the diversifying on the source deep transfer learning network and that passes a predefined fitness threshold condition for the target deep transfer learning network.

16. A computer implemented method comprising:

generating a plurality of deep transfer learning networks based on a plurality of first exemplars and a plurality of second exemplars;

training a source deep transfer learning network with a subset of the first exemplars;

encoding one or more transfer layers to a chromosome for genetic operators, where the one or more transfer layers are to be transferred from a source deep transfer learning network corresponding to a source domain to a target deep transfer learning network corresponding to a target domain;

mutating randomly the transfer layers of the source deep transfer learning network from the encoding; and

ascertaining that a fitness threshold condition for the source deep transfer learning network is satisfied.

17. The computer implemented method of claim 16 , further comprising:

prior to the encoding, measuring a homophily influence of the source deep transfer learning network to the target deep transfer learning network.

18. The computer implemented method of claim 16 , wherein a first subset of the first exemplars represents the source domain, a second subset of the first exemplars represents the target domain, and the second exemplars represent both the source domain and the target domain.

19. The computer implemented method of claim 16 , wherein the plurality of deep transfer learning networks include both the source deep transfer learning network and the target deep transfer learning network.

20. The computer implemented method of claim 16 , wherein the method includes migrating weights of respective transfer layers of first and second or more source deep transfer learning networks over to the target deep transfer learning network respectively.

21. A computer implemented method comprising:

generating a plurality of deep transfer learning networks based on a plurality of first exemplars and a plurality of second exemplars;

training a target deep transfer learning network with a subset of the first exemplars and the second exemplars;

encoding one or more transfer layers to a chromosome for genetic operators, where the one or more transfer layers are to be transferred from a source deep transfer learning network corresponding to a source domain to a target deep transfer learning network corresponding to a target domain;

ascertaining that respective transfer layers of first and second source deep transfer learning networks have their respective weights migrated over to the target deep transfer learning network; and

crossing over the transfer layers of the target deep transfer learning network with the transfer layers of the first and second source deep transfer learning networks from the ascertaining.

22. The computer implemented method of claim 21 , wherein the method includes producing, the target deep transfer learning network that integrates a result from the crossing over on the target deep transfer learning network and that passes a predefined fitness threshold condition for the target deep transfer learning network.

23. The computer implemented method of claim 21 , wherein the plurality of deep transfer learning networks include both the source deep transfer learning network and the target deep transfer learning network.

24. The computer implemented method of claim 21 , further comprising:

prior to the crossing over, randomly selecting the first and second source deep transfer learning network from a pool of source deep transfer learning networks that had migrated over respective weights of the transfer layers.

25. The computer implemented method of claim 21 , wherein a first subset of the first exemplars represents the source domain, a second subset of the first exemplars represents the target domain, and the second exemplars represent both the source domain and the target domain.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 3, 2023
From: TRIM, CRAIG M.; BAUGHMAN, AARON K.; VAUGHN, GARFIELD W.; FORSTER, MICAH
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 064485/0599 →
Continuity (2)
Continuation 16553823 · Aug 28, 2019
Related Publication 20230359899A1 · Nov 9, 2023
References Cited (51)
US 5140530A · Guha · 1992 [cited by examiner]
US 6241069B1 · Mazur · 2001 [cited by examiner]
US 8795523B2 · Su · 2014 [cited by examiner]
US 9542626B2 · Martinson · 2017 [cited by examiner]
US 11475297B2 · Trim et al. · 2022 [cited by applicant]
US 20020143598A1 · Scheer · 2002 [cited by examiner]
US 20040131998A1 · Marom · 2004 [cited by examiner]
US 20050119919A1 · Eder · 2005 [cited by examiner]
US 20050246297A1 · Chen · 2005 [cited by examiner]
US 20090248488A1 · Shah · 2009 [cited by examiner]
US 20160142266A1 · Carroll · 2016 [cited by examiner]
US 20160171398A1 · Eder · 2016 [cited by examiner]
US 20170024641A1 · Wierzynski · 2017 [cited by examiner]
US 20170193400A1 · Bhaskar · 2017 [cited by examiner]
US 20170337464A1 · Rabinowitz · 2017 [cited by examiner]
US 20180302306A1 · Carroll · 2018 [cited by examiner]
US 20190049127A1 · Shin · 2019 [cited by examiner]
US 20190370645A1 · Lee · 2019 [cited by examiner]
US 20210064982A1 · Trim · 2021 [cited by examiner]
US 20210065013A1 · Trim · 2021 [cited by examiner]
US 20220267756A1 · Lande · 2022 [cited by examiner]
WO WO2019049127 · 2019 [cited by applicant]
Maitrei Kohli, “Evolving Neural Networks Using Behavioural Genetic Principles”, Mar. 1, 2017, 290 pgs., XP055557613, Retrieved from the Internet URL <http://www.dcs.bbk.ac.uk/site/assets/files/1025/mkohli.pdf>. [cited by applicant]
Tian Haiman et al., “Automated Neural Network Construction with Similarity Sensitive Evolutionary Algorithms”, 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), IEEE, J… [cited by applicant]
Tian Haiman et al., “Genetic Algorithm Based Deep Learning Model Selection for Visual Data Classification”, 2019 IEEE 20th International Conference on Information Reuse and Integration for Data Science (IRI), IEEE, Jul.… [cited by applicant]
International Search Report and Written Opinion for PCT/EP2020/073727, completed Jan. 1, 2021, 17 pgs. [cited by applicant]
P. Mell, et al. “ [cited by applicant]
D. Roy et al. “ [cited by applicant]
M. Wang et al. “ [cited by applicant]
Anonymous, “ [cited by applicant]
Anonymous, “ [cited by applicant]
J. Zhang et al. “ [cited by applicant]
O. Mayer et al., “ [cited by applicant]
A. Wang et al., “ [cited by applicant]
“Smarter Supply Chain of the Future: Insights from the Global Chief Supply Chain Officer Study.” IBM Institute for Business Value. 2010. https://www-935.ibm.com/services/US/gbs/bus/html/gbs-csco-study.html. [cited by applicant]
IBM press release. “Aerialtronics Commercial Drones Give IBM Watson Internet of Things a Bird's Eye View.” 2016. https://www.ibm.com/press/us/en/pressrelease/50688.wss. [cited by applicant]
IBM case study. “Jabil Circuit implements a larger-scale analytics solution using IBM Analytics to reduce monthly close time.” 2015. https://www-03.ibm.com/software/businesscasestudies/us/en/corp?synkey=M200424F25312E29. [cited by applicant]
IBM press release. “Local Motors Debuts ‘Olli,’ the First Self-driving Vehicle to Tap the Power of IBM Watson.” 2016. http://www-03.ibm.com/press/us/en/pressrelease/49957.wss. [cited by applicant]
Lewis, Karen E. “Watson makes building management as a service possible.” IBM Cloud computing news. 2017. https://www.ibm.com/blogs/cloud-computing/2017/02/watson-building-management-service/. [cited by applicant]
Butner, Karen, Dave Lubowe and Louise Skordby. “Who's leading the cognitive pack in digital operations?” IBM Institute for Business Value. Nov. 2016. https://www.ibm.com/services/us/gbs/thoughtleadership/cognitiveops. [cited by applicant]
Butner, Karen and Dave Lubowe. “Thinking out of the toolbox: How digital technologies are powering the operations revolution.” IBM Institute for Business Value. Nov. 2015. http://www.ibm.com/services/us/gbs/thoughtleade… [cited by applicant]
Butner, Karen and Dave Lubowe. “The digital overhaul: Redefining manufacturing in a digital age.” IBM Institute for Business Value. May 2015. http://www.ibm.com/services/us/gbs/thoughtleadership/digitalmanufacturing/. [cited by applicant]
List of IBM Patent and/or Patent Applications treated as related for U.S. Appl. No. 18/349,902, filed Jul. 10, 2023, dated Aug. 22, 2023. [cited by applicant]
A. Sharma et al. “ [cited by applicant]
La Fond et al. “ [cited by applicant]
Q. Han et al. “ [cited by applicant]
Y. Sun, “Automatically Designing CNN Architectures Using Genetic Algorithm for Image Classification.” (Submitted on Aug. 11, 2018). https://arxiv.org/abs/1808.03818. [cited by applicant]
Y. Kanada, “Optimizing neural-network learning rate by using a genetic algorithm with per-epoch mutations,” 2016 International Joint Conference on Neural Networks (IJCNN), Vancouver, BC, 2016, pp. 1472-1479. [cited by applicant]
I. Athanasiadis, A Framework of Transfer Learning in Object Detection for Embedded Systems (Submitted on Nov. 12, 2018 (v1), last revised Nov. 24, 2018 (this version, v2)) https://arxiv.org/abs/1811.04863. [cited by applicant]
F. Assunção, “DENSER: Deep Evolutionary Network Structured Representation.” (Submitted on Jan. 4, 2018 (v1), last revised Jun. 1, 2018 (this version, v3)), https://arxiv.org/abs/1801.01563. [cited by applicant]
C. Fernando, “PathNet: Evolution Channels Gradient Descent in Super Neural Networks.” (Submitted on Jan. 30, 2017) https://arxiv.org/abs/1701.08734. [cited by applicant]