IP Library › Granted Patent US 12,724,966
Granted Patent B2
US 12,724,966 · App. 18/623,064 · Granted Sep 1, 2026

Text augmentation using dataset reconstruction

Inventors: Guy Uziel (Rishon Lezion, IL); Esther Goldbraich (Kiryat Ata, IL); Ateret Anaby-Tavor (Givat Ada, IL); Adir Rahamim (Netanya, IL)
Assignee: International Business Machines Corporation
G06F40/284G06N3/02G06N3/08G06N3/096
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,724,966
App. No.
18/623,064
Granted
Sep 1, 2026
Kind
B2
Abstract

A computer-implemented method comprising: receiving a source dataset comprising a plurality of textual data instances and corresponding labels in two or more classes; training a machine learning classifier on the source dataset; performing inference by the trained machine learning classifier over a subset of the data instances in the source dataset, to extract a hidden representation for each of said data instances in said subset; applying a trained multilayer perceptron (MLP) network to the extracted hidden representations, to generate a set of corresponding soft prompts; and feeding the generated set of soft-prompts as prompts for a trained language model, to tune the trained language model to reconstruct the data instances in the subset.

Claims (37)

1 . A computer-implemented method comprising:

receiving a source dataset comprising: a plurality of textual data instances, and corresponding labels in two or more classes;

training a machine learning classifier on the source dataset;

performing inference by the trained machine learning classifier over a subset of said data instances in the source dataset, to extract a hidden representation for each of said data instances in said subset;

applying a trained multilayer perceptron (MLP) network to the extracted hidden representations, to generate a set of corresponding soft prompts; and

feeding the generated set of soft prompts as prompts for a trained language model, to tune said trained language model to reconstruct said data instances in said subset.

2 . The computer-implemented method of claim 1 , wherein said machine learning classifier is a pre-trained classifier configured for text classifications tasks.

3 . The computer-implemented method of claim 1 , wherein each of said hidden representations represents a contextual embedding of a corresponding one of said data instances in said subset.

4 . The computer-implemented method of claim 1 , wherein said machine learning classifier is based on the Bidirectional Encoder Representations from Transformers (BERT) family of classifiers.

5 . The computer-implemented method of claim 4 , wherein each of said hidden representations is a last hidden representation.

6 . The computer-implemented method of claim 1 , further comprising averaging selected two of said soft prompts associated with one of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

7 . The computer-implemented method of claim 1 , further comprising averaging selected two of said soft prompts associated, respectively, with two different ones of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

8 . A system comprising:

at least one hardware processor; and

a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by said at least one hardware processor to:

receive a source dataset comprising a plurality of textual data instances and corresponding labels in two or more classes,

train a machine learning classifier on the source dataset,

perform inference by the trained machine learning classifier over a subset of said data instances in the source dataset, to extract a hidden representation for each of said data instances in said subset,

apply a trained multilayer perceptron (MLP) network to the extracted hidden representations, to generate a set of corresponding soft prompts, and

feed the generated set of soft prompts as prompts for a trained language model, to tune said trained language model to reconstruct said data instances in said subset.

9 . The system of claim 8 , wherein said machine learning classifier is a pre-trained classifier configured for text classifications tasks.

10 . The system of claim 8 , wherein each of said hidden representations represents a contextual embedding of each a corresponding one of said data instances in said subset.

11 . The system of claim 8 , wherein said machine learning classifier is based on the Bidirectional Encoder Representations from Transformers (BERT) family of classifiers.

12 . The system of claim 11 , wherein s each of said hidden representations is a last hidden representation.

13 . The system of claim 8 , wherein said program code is further executable to average selected two of said soft prompts associated with the same one of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

14 . The system of claim 8 , wherein said program code is further executable to average selected two of said soft prompts associated, respectively, with two different ones of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

15 . A computer program product comprising a non-transitory computer-readable storage medium having program code embodied therewith, the program code executable by at least one hardware processor to:

receive a source dataset comprising a plurality of textual data instances and corresponding labels in two or more classes;

train a machine learning classifier on the source dataset;

perform inference by the trained machine learning classifier over a subset of said data instances in the source dataset, to extract a hidden representation for each of said data instances in said subset;

apply a trained multilayer perceptron (MLP) network to the extracted hidden representations, to generate a set of corresponding soft prompts; and

feed the generated set of soft prompts as prompts for a trained language model, to tune said trained language model to reconstruct said data instances in said subset.

16 . The computer program product of claim 15 , wherein said machine learning classifier is a pre-trained classifier configured for text classifications tasks.

17 . The computer program product of claim 15 , wherein each of said hidden representations represents a contextual embedding of each a corresponding one of said data instances in said subset.

18 . The computer program product of claim 15 , wherein said machine learning classifier is based on the Bidirectional Encoder Representations from Transformers (BERT) family of classifiers, and wherein each of said hidden representations is a last hidden representation.

19 . The computer program product of claim 15 , wherein said program code is further executable to average selected two of said soft prompts associated with one of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

20 . The computer program product of claim 15 , wherein said program code is further executable to average selected two of said soft prompts associated, respectively, with two different ones of said two or more classes, to obtain an aggregated soft prompt, and wherein said aggregated soft prompt is used as one of said prompts for said trained language model.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 1, 2024
From: UZIEL, GUY; GOLDBRAICH, ESTHER; ANABY - TAVOR, ATERET; RAHAMIM, ADIR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 066959/0369 →
Continuity (1)
Related Publication 20250307550A1 · Oct 2, 2025
References Cited (28)
US 11526667B2 · Kantor · 2022 [cited by applicant]
US 20200372404A1 · Mahmud · 2020 [cited by examiner]
US 20210117718A1 · Badjatiya · 2021 [cited by examiner]
US 20210350076A1 · Kantor · 2021 [cited by examiner]
US 20230419049A1 · Chen · 2023 [cited by examiner]
US 20230419164A1 · Mrini · 2023 [cited by examiner]
US 20240020546A1 · Vu · 2024 [cited by examiner]
US 20240144651A1 · Bulat · 2024 [cited by examiner]
US 20250131212A1 · Yu · 2025 [cited by examiner]
US 20250307550A1 · Uziel · 2025 [cited by examiner]
CN 107526725B · 2021 [cited by applicant]
Ding et al., “DAGA: Data augmentation with a generation approach for low-resource tagging tasks.” arXiv preprint arXiv:2011.01549 (Year: 2020). [cited by examiner]
Yang et al., “Generative data augmentation for commonsense reasoning.” Findings of the Association for Computational Linguistics: EMNLP 2020 (Year: 2020). [cited by examiner]
Akari Asai et al., “Attempt: Parameter-Efficient Multi-task Tuning via Attentional Mixtures of Soft Prompts”; Online at: https://arxiv.org/abs/2205.11961, Dec. 1, 2022. [cited by applicant]
Ateret Anaby-Tavor et al., “Do Not Have Enough Data? Deep Learning to the Rescue!”; Online at: https://arxiv.org/pdf/1911.03118.pdf, Nov. 27, 2019. [cited by applicant]
Brian Lester et al, “The Power of Scale for Parameter-Efficient Prompt Tuning”; In Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, pp. 3045-3059, Nov. 2021. [cited by applicant]
Hugo Queiroz Abonizio et al, “Pre-trained Data Augmentation for Text Classification”; Intelligent Systems: 9th Brazilian Conference, BRACIS 2020, Rio Grande, Brazil, Proceedings, Part I, Oct. 2020, pp. 551-565, Oct. 20-… [cited by applicant]
Jason Wei et al, EDA: Easy Data Augmentation Techniques for Boosting Performance on Text Classification Tasks; Online at: https://arxiv.org/abs/1901.11196, Jan. 31, 2019. [cited by applicant]
Lukas Fromme et al, “ContextGen: Targeted Data Generation for Low Resource Domain Specific Text Classification”; Proceedings of The 25th International Conference on Artificial Intelligence and Statistics, PMLR 151:3016-… [cited by applicant]
Maria Tsimpoukelli et al, “Multimodal Few-Shot Learning with Frozen Language Models”; Online at: https://arxiv.org/abs/2106.13884, Jun. 25, 2021. [cited by applicant]
Ruibo Liu et al, “Data Boost: Text Data Augmentation Through Reinforcement Learning Guided Conditional Generation”; In Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), pp. … [cited by applicant]
Tianyu Gao et al, “Making Pre-trained Language Models Better Few-shot Learners”; Online at: https://arxiv.org/abs/2012.15723, Jun. 2, 2021. [cited by applicant]
Tom B. Brown et al, “Language Models are Few-Shot Learners”; Online at: https://arxiv.org/abs/2005.14165, Jul. 22, 2020. [cited by applicant]
Tu Vu et al, “SPOT: Better Frozen Model Adaptation through Soft Prompt Transfer”; Online at: https://arxiv.org/abs/2110.07904, Oct. 15, 2021. [cited by applicant]
Xiang Dai et al, “An Analysis of Simple Data Augmentation for Named Entity Recognition”; Online at: https://arxiv.org/abs/2010.11683, Oct. 22, 2020. [cited by applicant]
Xiang Lisa Li et al, “Prefix-Tuning: Optimizing Continuous Prompts for Generation”; Online at: https://arxiv.org/abs/2101.00190, Jan. 1, 2021. [cited by applicant]
Xing Wu et al, “Conditional BERT Contextual Augmentation”; Online at: https://arxiv.org/abs/1812.06705, Dec. 17, 2018. [cited by applicant]
Yufei Wang et al, “PromDA: Prompt-based Data Augmentation for Low-Resource NLU Tasks”; Online at: https://arxiv.org/abs/2202.12499, Feb. 25, 2022. [cited by applicant]