IP Library › Granted Patent US 12,664,424
Granted Patent B2
US 12,664,424 · App. 17/955,055 · Granted Jun 23, 2026

Contrastive learning by dynamically selecting dropout ratios and locations based on reinforcement learning

Inventors: Zhong Fang Yuan (Xi'an, CN); Si Tong Zhao (Beijing, CN); Tong Liu (Xi'An, CN); Yi Chen Zhong (Shanghai, CN); Yuan Yuan Ding (Shanghai, CN); Hai Bo Zou (Beijing, CN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06N3/08G06N3/047
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,664,424
App. No.
17/955,055
Granted
Jun 23, 2026
Kind
B2
Abstract

A method for contrastive learning by selecting dropout ratios and locations based on reinforcement learning includes receiving training data having a positive sample corresponding to a target and negative samples not corresponding to the target. A dropout policy for a neural network is produced based on the training data, where the dropout policy identifies at least one connection between neurons in the neural network to dropout. The training data is encoded, based on the dropout policy, to form embeddings, where the embeddings include multiple positive sample embeddings corresponding to the positive sample and multiple negative sample embedding corresponding to the negative samples.

Claims (46)

1 . A method, comprising:

receiving, by a processor, training data comprising a positive sample corresponding to a target and negative samples not corresponding to the target;

producing, by a dropout controller and based on the training data, a dropout policy for a neural network, wherein the dropout policy identifies at least one connection between neurons in the neural network to dropout, wherein producing the dropout policy further comprises,

producing, by a policy network and based on the training data, a set of dropout policies corresponding to different dropout policies,

selecting, by the dropout controller and based on a short-term reward, a portion of the set of dropout policies, and

selecting, by the dropout controller and based on a long-term reward, one policy from the portion of the set of dropout policies as the dropout policy; and

encoding, by an encoder and based on the dropout policy, the training data to form embeddings,

wherein the embeddings comprise multiple positive sample embeddings corresponding to the positive sample and multiple negative sample embeddings corresponding to the negative samples, and wherein the embeddings are generated based on the dropout policy.

2 . The method of claim 1 , wherein producing the dropout policy comprises:

representing connections between the neurons in the neural network with dropout matrices; and

selecting, based on a reinforcement learning method, one or more of the connections to disconnect to produce the dropout policy.

3 . The method of claim 2 , wherein representing the connections comprises producing N−1 dropout matrices, wherein N comprises any number and corresponds to a number of layers in the neural network, and wherein dimensions of each of the N−1 dropout matrices correspond to a number of neurons in adjacent layers in the neural network.

4 . The method of claim 1 , further comprising calculating both the short-term reward and the long-term reward using a same loss function.

5 . The method of claim 1 , wherein selecting the one policy comprises comparing policies within the portion of the set of dropout policies using a Monte Carlo search tree.

6 . The method of claim 1 , wherein producing the dropout policy comprises producing the dropout policy using transformer networks, wherein the transformer networks comprise a current network, a history network, and a composite network, wherein the current network comprises current parameter information of the current network, wherein the history network comprises previous parameter information of previous networks, and wherein the composite network is configured to merge the current parameter information with the previous parameter information to produce composite parameter information.

7 . A system, comprising:

a non-transitory computer-readable storage memory configured to store instructions; and

a processor coupled to the non-transitory computer-readable storage memory and configured to execute the instructions to cause the system to:

receive, by the processor, training data comprising a positive sample corresponding to a target and negative samples not corresponding to the target;

produce, by a dropout controller and based on the training data, a dropout policy for a neural network, wherein the dropout policy identifies at least one connection between neurons in the neural network to dropout, wherein producing the dropout policy further comprises,

produce, by a policy network and based on the training data, a set of dropout policies corresponding to different dropout policies,

select, by the dropout controller and based on a short-term reward, a portion of the set of dropout policies, and

select, by the dropout controller and based on a long-term reward, one policy from the portion of the set of dropout policies as the dropout policy; and

encode, by an encoder and based on the dropout policy, the training data to form embeddings,

wherein the embeddings comprise multiple positive sample embeddings corresponding to the positive sample and multiple negative sample embeddings corresponding to the negative samples, and wherein the embeddings are generated based on the dropout policy.

8 . The system of claim 7 , wherein the processor is further configured to execute the instructions to cause the system to:

represent connections between the neurons in the neural network with dropout matrices; and

select, based on a reinforcement learning method, one or more of the connections to disconnect to produce the dropout policy.

9 . The system of claim 8 , wherein the processor is further configured to execute the instructions to cause the system to produce N−1 dropout matrices, wherein N comprises any number and corresponds to a number of layers in the neural network, and wherein dimensions of each of the N−1 dropout matrices correspond to a number of neurons in adjacent layers in the neural network.

10 . The system of claim 7 , wherein the processor is further configured to execute the instructions to cause the system to calculate both the short-term reward and the long-term reward using a same loss function.

11 . The system of claim 7 , wherein the processor is further configured to execute the instructions to cause the system to compare policies within the portion of the set of dropout policies using a Monte Carlo search tree.

12 . The system of claim 7 , wherein the processor is further configured to execute the instructions to cause the system to produce the dropout policy using transformer networks, wherein the transformer networks comprise a current network, a history network, and a composite network, wherein the current network comprises current parameter information of the current network, wherein the history network comprises previous parameter information of previous networks, and wherein the composite network is configured to merge the current parameter information with the previous parameter information to produce composite parameter information.

13 . A computer program product comprising instructions stored on a non-transitory computer-readable medium that, when executed by a processor, cause a system to:

receive, by the processor, training data comprising a positive sample corresponding to a target and negative samples not corresponding to the target;

produce, by a dropout controller and based on the training data, a dropout policy for a neural network, wherein the dropout policy identifies at least one connection between neurons in the neural network to dropout, wherein producing the dropout policy further comprises,

produce, by a policy network and based on the training data, a set of dropout policies corresponding to different dropout policies,

select, by the dropout controller and based on a short-term reward, a portion of the set of dropout policies, and

select, by the dropout controller and based on a long-term reward, one policy from the portion of the set of dropout policies as the dropout policy; and

encode, by an encoder and based on the dropout policy, the training data to form embeddings,

wherein the embeddings comprise multiple positive sample embeddings corresponding to the positive sample and multiple negative sample embeddings corresponding to the negative samples, and wherein the embeddings are generated based on the dropout policy.

14 . The computer program product of claim 13 , wherein the instructions further cause the system to:

represent connections between the neurons in the neural network with dropout matrices; and

select, based on a reinforcement learning method, one or more of the connections to disconnect to produce the dropout policy.

15 . The computer program product of claim 14 , wherein the instructions further cause the system to produce N−1 dropout matrices, wherein N comprises any number and corresponds to a number of layers in the neural network, and wherein dimensions of each of the N−1 dropout matrices correspond to a number of neurons in adjacent layers in the neural network.

16 . The computer program product of claim 13 , wherein the instructions further cause the system to calculate both the short-term reward and the long-term reward using a same loss function.

17 . The computer program product of claim 13 , wherein the instructions further cause the system to compare policies within the portion of the set of dropout policies using a Monte Carlo search tree.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 28, 2022
From: YUAN, ZHONG FANG; ZHAO, SI TONG; LIU, TONG; ZHONG, YI CHEN; DING, YUAN YUAN; ZOU, HAI BO
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 061246/0044 →
Continuity (1)
Related Publication 20240119275A1 · Apr 11, 2024
References Cited (15)
US 10839260B1 · Khanna · 2020 [cited by applicant]
US 20210279511A1 · Gordon · 2021 [cited by examiner]
US 20220156585A1 · Leng · 2022 [cited by examiner]
US 20220179419A1 · Romeres · 2022 [cited by examiner]
US 20220309774A1 · Uzkent · 2022 [cited by examiner]
KR 102061615B1 · 2020 [cited by applicant]
WO 20202205441W · 2020 [cited by applicant]
Huang, et al., “A Two-Stage Contrastive Learning Framework For Imbalanced Aerial Scene Recognition,” ICASSP 2002, 2022 IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP) 978-1-6654-0540-9/… [cited by applicant]
Pham, et al., “AutoDropout: Learning Dropout Patterns to Regularize Deep Networks,” Association for the Advancement of Artificial Intelligence (www.aaai.org), arXiv:2101.01761v1, [cs.LG] Jan. 5, 2021, 15 pages. [cited by applicant]
Wu, et al., “BlockDrop: Dynamic Inference Paths in Residual Networks,” Proceedings / CVPR, IEEE Computer Society Conference on Computer Vision and Pattern Recognition. IEEE Computer Society Conference on Computer Vision… [cited by applicant]
Srinivas, et al., “CURL: Contrastive Unsupervised Representations for Reinforcement Learning,” Proceedings of the 37th International Conference on Machine Learning, Vienna, Austria, PMLR 119, 2020, arXiv:2004.04136v4 [c… [cited by applicant]
Gao, et al., “SimCSE: Simple Contrastive Learning of Sentence Embeddings,” arXiv:2104.08821v4 [cs.CL] May 18, 2022, 17 pages. [cited by applicant]
Ghiasi et al., “DropBlock: A regularization method for convolutional networks”, 32nd Conference on Neural Information Processing Systems NeurIPS, Oct. 30, 2018, 11 pages. [cited by applicant]
Srivastava et al., “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”, Journal of Machine Learning Research 15, Jan. 1, 2014, pp. 1929-1958. [cited by applicant]
Tompson et al. “Efficient Object Localization Using Convolutional Networks”, arXiv: 1411.4280v3 [cs.CV], Jun. 9, 2015, 9 pages. [cited by applicant]