IP Library › Granted Patent US 12,008,330
Granted Patent B2
US 12,008,330 · App. 17/510,640 · Granted Jun 11, 2024

Apparatus and method for augmenting textual data

Inventors: Na Un Kang (Seoul, KR); Geon Yi (Seoul, KR); Min Young Lee (Seoul, KR); Min Soo Kim (Seoul, KR)
Assignee: SAMSUNG SDS CO., LTD.
G06F40/40G06F18/2193G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,008,330
App. No.
17/510,640
Granted
Jun 11, 2024
Kind
B2
Abstract

An apparatus for augmenting textual data according to an embodiment includes a data augmenter configured to generate augmented data by augmenting input textual data according to a data augmentation scheme decided based on a type of natural language processing task of the input textual data and a data classifier configured to classify the augmented data into a positive sample or a negative sample by determining whether or not the augmented data maintains label information of the input textual data based on one or more data classification criteria.

Claims (44)

1. An apparatus for augmenting textual data, the apparatus comprising:

a data augmenter configured to generate augmented data by augmenting input textual data according to a data augmentation scheme decided based on a type of natural language processing task of the input textual data; and

a data classifier configured to classify the augmented data into a positive sample or a negative sample by determining whether or not the augmented data maintains label information of the input textual data based on one or more data classification criteria,

wherein the data classifier comprises at least one of:

a first analyzer configured to decide whether or not the augmented data is the positive sample or the negative sample using a mapping table preset according to the data augmentation scheme and the type of natural language processing task of the input textual data;

a second analyzer configured to analyze whether or not the augmented data satisfies grammar to decide whether or not the augmented data is the positive sample or the negative sample; or

a third analyzer configured to compare a predicted value of user input label with a label of the augmented data to decide whether or not the augmented data is the positive sample or the negative sample.

2. The apparatus of claim 1 , further comprising:

a consistency determinator configured to decide whether or not to use the augmented data based on a result classified according the one or more data classification criteria.

3. The apparatus of claim 2 , wherein the consistency determinator is further configured to decide whether or not to use the augmented data based on a ratio or number of results decided as positive samples among the results of at least one of the first analyzer, the second analyzer, or the third analyzer.

4. The apparatus of claim 2 , wherein the consistency determinator is further configured to:

decide to use the augmented data when it is determined that the augmented data is the positive sample based on results classified according to the one or more data classification criteria; and

decide not to use the augmented data when it is determined that the augmented data is the negative sample based on the result classified according to the one or more data classification criteria.

5. The apparatus of claim 4 , wherein the consistency determinator is further configured to decide whether or not to use the augmented data further based on the type of natural language processing task of the input textual data when it is determined that the augmented data is the negative sample based on the result classified according to the one or more data classification criteria.

6. The apparatus of claim 1 , wherein the data augmenter is further configured to decide an augmentation scale based on at least one of the type of natural language processing task, the data augmentation scheme, whether or not it is a key sentence, or a type of the input textual data.

7. The apparatus of claim 1 , further comprising:

a preprocessor configured to preprocess the input textual data using at least one of tokenization, stopword removing, stemming, or lemmatization and transmit the preprocessed input textual data to the data augmenter.

8. The apparatus of claim 1 , further comprising:

an input data analyzer configured to decide at least one of whether or not the input textual data satisfies a predetermined requirement for data augmentation; a type of a dominant language of the input textual data; whether or not the input textual data corresponds to any one of a single sentence, a single document, and a corpus; or the type of natural language processing task corresponding to the input textual data.

9. The apparatus of claim 8 , wherein the input data analyzer is further configured to decide that the input textual data satisfies the predetermined requirement for data augmentation when the input textual data includes one or more sentences in which one or more sentence elements are combined.

10. The apparatus of claim 8 , wherein the input data analyzer is further configured to decide the type of the dominant language of the input textual data based on Unicode for each language.

11. The apparatus of claim 8 , wherein the input data analyzer is further configured to decide the type of natural language processing task based on the label of the input textual data.

12. A method for augmenting textual data comprising:

generating augmented data by augmenting input textual data according to a data augmentation scheme decided based on a type of natural language processing task of the input textual data; and

classifying the augmented data into a positive sample or a negative sample by determining whether or not the augmented data maintains label information of the input textual data based on one or more data classification criteria,

wherein the classifying of the augmented data comprises classifying the augmented data using at least one of:

a first analysis method for deciding whether or not the augmented data is the positive sample or the negative sample using a mapping table preset according to the data augmentation scheme and the type of natural language processing task of the input textual data;

a second analysis method for analyzing whether or not the augmented data satisfies grammar to decide whether or not the augmented data is the positive sample or the negative sample; or

a third analysis method for comparing a predicted value of user input label with a label of the augmented data to decide whether or not the augmented data is the positive sample or the negative sample.

13. The method of claim 12 , further comprising:

deciding whether or not to use the augmented data based on a result classified according to the one or more data classification criteria.

14. The method of claim 13 , wherein the deciding whether or not to use the augmented data comprises deciding whether or not to use the augmented data based on a ratio or number of results decided as positive samples among the results of at least one of the first analysis method, the second analysis method, or the third analysis method.

15. The method of claim 13 , wherein the deciding whether or not to use the augmented data comprises:

deciding to use the augmented data when it is determined that the augmented data is the positive sample based on the result classified according to the one or more data classification criteria; and

deciding not to use the augmented data when it is determined that the augmented data is the negative sample based on the result classified according to the one or more data classification criteria.

16. The method of claim 15 , wherein the deciding whether or not to use the augmented data comprises deciding whether or not to use the augmented data further based on the type of natural language processing task of the input textual data when it is determined that the augmented data is the negative sample based on the result classified according to the one or more data classification criteria.

17. The method of claim 12 , wherein the generating of the augmented data comprises deciding an augmentation scale based on at least one of the type of natural language processing task, the data augmentation scheme, whether or not it is a key sentence, or a type of the input textual data.

18. The method of claim 12 , further comprising:

preprocessing the input textual data using at least one of tokenization, stopword removing, stemming, or lemmatization.

19. The method of claim 12 , further comprising:

deciding at least one of whether or not the input textual data satisfies a predetermined requirement for data augmentation; a type of a dominant language of the input textual data; whether or not the input textual data corresponds to any one of a single sentence, a single document, and a corpus; or the type of natural language processing task corresponding to the input textual data.

20. The method of claim 19 , wherein the deciding comprises deciding that the input textual data satisfies the predetermined requirement for data augmentation when the input textual data includes one or more sentences in which one or more sentence elements are combined.

21. The method of claim 19 , wherein the deciding comprises deciding the type of the dominant language of the input textual data based on Unicode for each language.

22. The method of claim 19 , wherein the deciding comprises deciding the type of natural language processing task based on the label of the input textual data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 9, 2024
From: KANG, NA UN; YI, GEON; LEE, MIN YOUNG; KIM, MIN SOO
To: SAMSUNG SDS CO., LTD.
Reel/Frame 067362/0872 →
Priority Claims (1)
KR 10-2020-0139566 · Oct 26, 2020 · national
Continuity (1)
Related Publication 20220129644A1 · Apr 28, 2022
Cited By (1)
US 12,517,960