IP Library › Granted Patent US 11,657,227
Granted Patent B2
US 11,657,227 · App. 17/148,020 · Granted May 23, 2023

Corpus data augmentation and debiasing

Inventors: Shikhar Kwatra (San Jose, CA); Nishtha Madaan (Gurgaon, IN); Sushain Pandit (Austin, TX); Kuntal Dey (Rampurhat, IN)
Assignee: INTERNATIONAL BUSINESS MACHINES CORPORATION
G06F40/289G06F16/285G06F40/58G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,657,227
App. No.
17/148,020
Granted
May 23, 2023
Kind
B2
Abstract

Machine learning model training corpus debiasing includes identifying an attribute of input text selected from the training corpus, the attribute including word(s) of the input text, and the attribute corresponding to an attribute class encompassing different possible class values, recognizing bias in the input text with respect to the attribute class, and generating output text corresponding to the attribute and imparting diversity with respect to the attribute class and relative to the input text, where generating the output text uses an optimization function based on loss objectives to minimize loss in the generated output text as compared to the input text.

Claims (32)

1. A method comprising:

identifying an attribute of input text, the input text selected from a machine learning model training corpus, the attribute comprising one or more words of the input text, and the attribute corresponding to an attribute class encompassing a plurality of different possible class values;

recognizing bias in the input text with respect to the attribute class; and

generating output text corresponding to the attribute and imparting diversity with respect to the attribute class and relative to the input text, wherein generating the output text uses an optimization function based on a plurality of loss objectives to minimize loss in the generated output text as compared to the input text, the plurality of loss objectives comprising a label loss, a proximal loss, an attribute loss, and a diversity loss.

2. The method of claim 1 , wherein recognizing the bias comprises identifying that one possible class value of the plurality of different possible class values of the attribute class is represented by the input text and at least one other possible class value of the plurality of different possible class values of the attribute class is not represented by the input text.

3. The method of claim 1 , wherein the input text comprises a sentence and wherein the generated output text comprises one or more output sentences.

4. The method of claim 3 , wherein each output sentence of the one or more output sentences represents a respective different possible class value of the plurality of different possible class values of the attribute class.

5. The method of claim 1 , further comprising including the generated output text in the machine learning model training corpus to facilitate debiased training of a machine learning model using the machine learning model training corpus.

6. The method of claim 1 , further comprising using the generated output text as a baseline for comparison against output of a training corpus debiasing model to evaluate performance of the training corpus debiasing model in debiasing the machine learning model training corpus.

7. The method of claim 1 , wherein the label loss corresponds to cross entropy in the generated output text as compared to the input text, the proximity loss corresponds to loss of proximity of the generated output text as compared to the input text, the attribute loss corresponds to loss of attentiveness to the identified attribute in the generated output text as compared to the input text, and the diversity loss corresponds to overlap in diversity with respect to the attribute class of the generated output text as compared to the input text.

8. The method of claim 7 , wherein the optimization function to minimize loss in the generated output text as compared to the input text minimizes a composite of the label loss, the proximity loss, the attribute loss, and the diversity loss in selecting candidate output text to output as the generated output text.

9. The method of claim 7 , wherein the optimization function encourages diversity in representation of the plurality of different possible class values of the attribute class by minimizing the diversity loss to minimize overlap in diversity with respect to the attribute class of the generated output text as compared to the input text, thereby maximizing diversity in representation of the plurality of different possible class values of the attribute class across the input text and the generated output text.

10. A computer system comprising:

a memory; and

a processor in communication with the memory, wherein the computer system is configured to perform a method comprising:

identifying an attribute of input text, the input text selected from a machine learning model training corpus, the attribute comprising one or more words of the input text, and the attribute corresponding to an attribute class encompassing a plurality of different possible class values;

recognizing bias in the input text with respect to the attribute class; and

generating output text corresponding to the attribute and imparting diversity with respect to the attribute class and relative to the input text, wherein generating the output text uses an optimization function based on a plurality of loss objectives to minimize loss in the generated output text as compared to the input text, the plurality of loss objectives comprising a label loss, a proximity loss, an attribute loss, and a diversity loss.

11. The computer system of claim 10 , wherein recognizing the bias comprises identifying that one possible class value of the plurality of different possible class values of the attribute class is represented by the input text and at least one other possible class value of the plurality of different possible class values of the attribute class is not represented by the input text.

12. The computer system of claim 10 , wherein the input text comprises a sentence, wherein the generated output text comprises one or more output sentences, and wherein each output sentence of the one or more output sentences represents a respective different possible class value of the plurality of different possible class values of the attribute class.

13. The computer system of claim 10 , wherein the method further comprises including the generated output text in the machine learning model training corpus to facilitate debiased training of a machine learning model using the machine learning model training corpus.

14. The computer system of claim 10 , wherein the label loss corresponds to cross entropy in the generated output text as compared to the input text, the proximity loss corresponds to loss of proximity of the generated output text as compared to the input text, the attribute loss corresponds to loss of attentiveness to the identified attribute in the generated output text as compared to the input text, and the diversity loss corresponds to overlap in diversity with respect to the attribute class of the generated output text as compared to the input text.

15. The computer system of claim 14 , wherein the optimization function encourages diversity in representation of the plurality of different possible class values of the attribute class by minimizing the diversity loss to minimize overlap in diversity with respect to the attribute class of the generated output text as compared to the input text, thereby maximizing diversity in representation of the plurality of different possible class values of the attribute class across the input text and the generated output text.

16. A computer program product comprising:

a computer readable storage medium readable by a processing circuit and storing instructions for execution by the processing circuit for performing a method comprising:

identifying an attribute of input text, the input text selected from a machine learning model training corpus, the attribute comprising one or more words of the input text, and the attribute corresponding to an attribute class encompassing a plurality of different possible class values;

recognizing bias in the input text with respect to the attribute class; and

generating output text corresponding to the attribute and imparting diversity with respect to the attribute class and relative to the input text, wherein generating the output text uses an optimization function based on a plurality of loss objectives to minimize loss in the generated output text as compared to the input text, the plurality of loss objectives comprising a label loss, a proximity loss, an attribute loss, and a diversity loss.

17. The computer program product of claim 16 , wherein recognizing the bias comprises identifying that one possible class value of the plurality of different possible class values of the attribute class is represented by the input text and at least one other possible class value of the plurality of different possible class values of the attribute class is not represented by the input text.

18. The computer program product of claim 16 , wherein the input text comprises a sentence, wherein the generated output text comprises one or more output sentences, and wherein each output sentence of the one or more output sentences represents a respective different possible class value of the plurality of different possible class values of the attribute class.

19. The computer program product of claim 16 , wherein the method further comprises including the generated output text in the machine learning model training corpus to facilitate debiased training of a machine learning model using the machine learning model training corpus.

20. The computer program product of claim 16 , wherein the label loss corresponds to cross entropy in the generated output text as compared to the input text, the proximity loss corresponds to loss of proximity of the generated output text as compared to the input text, the attribute loss corresponds to loss of attentiveness to the identified attribute in the generated output text as compared to the input text, and the diversity loss corresponds to overlap in diversity with respect to the attribute class of the generated output text as compared to the input text.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 13, 2021
From: KWATRA, SHIKHAR; MADAAN, NISHTHA; PANDIT, SUSHAIN; DEY, KUNTAL
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 054908/0168 →
Continuity (1)
Related Publication 20220222438A1 · Jul 14, 2022