IP Library › Granted Patent US 12,093,645
Granted Patent B2
US 12,093,645 · App. 17/474,364 · Granted Sep 17, 2024

Inter-training of pre-trained transformer-based language models using partitioning and classification

Inventors: Eyal Shnarch (Tel Aviv, IL); Ariel Gera (Haifa, IL); Alon Halfon (Rishon Lezion, IL); Lena Dankin (Haifa, IL); Leshem Choshen (Haifa, IL); Ranit Aharonov (Ramat Hasharon, IL); Noam Slonim (Beit HaKerem, IL)
Assignee: International Business Machines Corporation
G06F40/279G10L25/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,093,645
App. No.
17/474,364
Granted
Sep 17, 2024
Kind
B2
Abstract

An example system includes a processor to pre-train a transformer-based language model on a general domain. The processor can inter-train the pre-trained transformer-based language model using partitioning and classification to generate an inter-trained transformer-based pre-trained language model. The processor can then fine-tune the inter-trained transformer-based pre-trained language model on a target task to generate a fine-tuned transformer-based language model.

Claims (26)

1. A system, comprising a processor to:

pre-train a transformer-based language model on a general domain;

inter-train the pre-trained transformer-based language model using mask language modeling to generate a mask language modeling (MLM) inter-trained transformer-based pre-trained language model;

inter-train the MLM inter-trained transformer-based pre-trained language model using partitioning and classification to generate a doubly inter-trained transformer-based pre-trained language model, wherein the partitioning comprises a clustering based on bag of words representations on a stemmed text to partition, according to class labels, unlabeled training data into clusters of text instances and wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises using the clusters of text instances as labeled data for an intermediate training task; and

fine-tune the doubly inter-trained transformer-based pre-trained language model on a target task to generate a fine-tuned transformer-based language model.

2. The system of claim 1 , wherein the classification comprises an unsupervised classification of training samples into a plurality of labels corresponding to partitions.

3. The system of claim 1 , wherein the classification comprises a binary classification in which the MLM inter-trained transformer-based pre-trained language model is trained to predict whether a pair of training samples is from a same partition or from a different partition.

4. The system of claim 1 , wherein the processor is to pre-train the transformer-based language model on the general domain using mask language modeling.

5. The system of claim 1 , wherein the pre-trained transformer-based language model comprises a Bidirectional Encoder Representations from Transformers (BERT) model.

6. A computer-implemented method, comprising:

pre-training, via a processor, a transformer-based language model on a general domain;

inter-training the pre-trained transformer-based language model using mask language modeling to generate a mask language modeling (MLM) inter-trained transformer-based pre-trained language model;

inter-training, via the processor, the MLM inter-trained transformer-based pre-trained language model using partitioning and classification to generate a doubly inter-trained transformer-based language model, wherein the partitioning comprises a clustering based on bag of words representations on a stemmed text to partition, according to class labels, unlabeled training data into clusters of text instances and wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises using the clusters of text instances as labeled data for an intermediate training task; and

fine-tuning, via the processor, the doubly inter-trained transformer-based language model on a target task to generate a fine-tuned transformer-based language model.

7. The computer-implemented method of claim 6 , wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises sampling a pair of training samples from two different partitions and training the MLM inter-trained transformer-based pre-trained language model to classify the pair of training samples differently.

8. The computer-implemented method of claim 6 , wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises sampling a pair of training samples from one partition and training the MLM inter-trained transformer-based pre-trained language model to classify the pair of training samples similarly.

9. The computer-implemented method of claim 6 , wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises inter-training the MLM inter-trained transformer-based pre-trained language model over pseudo-labels generated via an unsupervised sequential Information Bottleneck (sIB) clustering.

10. The computer-implemented method of claim 6 , comprising executing, via the processor, the target task on the fine-tuned transformer-based language model.

11. A computer program product for inter-training transformer-based language models, the computer program product comprising a computer-readable storage medium having program code embodied therewith, wherein the computer-readable storage medium is not a transitory signal per se, the program code executable by a processor to cause the processor to:

pre-train a transformer-based language model on a general domain;

inter-train the pre-trained transformer-based language model using mask language modeling to generate a mask language modeling (MLM) inter-trained transformer-based pre-trained language model;

inter-train the MLM inter-trained transformer-based pre-trained language model using partitioning and classification to generate a doubly inter-trained transformer-based pre-trained language model, wherein the partitioning comprises a clustering based on bag of words representations on a stemmed text to partition, according to class labels, unlabeled training data into clusters of text instances and wherein the inter-training the MLM inter-trained transformer-based pre-trained language model comprises using the clusters of text instances as labeled data for an intermediate training task; and

fine-tune the doubly inter-trained transformer-based pre-trained language model on a target task to generate a fine-tuned transformer-based language model.

12. The computer program product of claim 11 , further comprising program code executable by the processor to sample a pair of training samples and inter-train the MLM inter-trained transformer-based pre-trained language model to classify the pair of training samples as belonging to a same partition or a different partition.

13. The computer program product of claim 11 , further comprising program code executable by the processor to inter-train the MLM inter-trained transformer-based pre-trained language model over pseudo-labels generated via unsupervised sequential Information Bottleneck (sIB) clustering.

14. The computer program product of claim 11 , further comprising program code executable by the processor to execute the target task on the fine-tuned transformer-based language model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 14, 2021
From: SHNARCH, EYAL; GERA, ARIEL; HALFON, ALON; DANKIN, LENA; CHOSHEN, LESHEM; AHARONOV, RANIT; SLONIM, NOAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 057474/0832 →
Continuity (1)
Related Publication 20230078698A1 · Mar 16, 2023