IP Library Granted Patent US 11,314,950
Granted Patent B2
US 11,314,950 · App. 16/830,106 · Granted Apr 26, 2022

Text style transfer using reinforcement learning

Inventors: Lingfei Wu (Elmsford, NY); Jinjun Xiong (Goldens Bridge, NY); Hongyu Gong (Urbana, IL); Suma Bhat (Urbana, IL); Wen-Mei Hwu (Urbana, IL)
Assignees: INTERNATIONAL BUSINESS MACHINES CORPORATION; THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
G06F40/56G06F40/253G06F40/35G06N3/049
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,314,950
App. No.
16/830,106
Granted
Apr 26, 2022
Kind
B2
Abstract

A computer-implemented method is provided for transferring a target text style using Reinforcement Learning (RL). The method includes pre-determining, by a Long Short-Term Memory (LSTM) Neural Network (NN), the target text style of a target-style natural language sentence. The method further includes transforming, by a hardware processor using the LSTM NN, a source-style natural language sentence into the target-style natural language sentence that maintains the target text style of the target-style natural language sentence. The method also includes calculating an accuracy rating of a transformation of the source-style natural language sentence into the target-style natural language sentence based upon rewards relating to at least the target text style of the source-style natural language sentence.

Claims (33)

1. A computer-implemented method for transferring a target text style using Reinforcement Learning (RL), comprising:

pre-determining, by a Long Short-Term Memory (LSTM) Neural Network (NN), the target text style of a target-style natural language sentence;

transforming, by a hardware processor using the LSTM NN, a source-style natural language sentence into the target-style natural language sentence that maintains the target text style of the target-style natural language sentence; and

calculating an accuracy rating of a transformation of the source-style natural language sentence into the target-style natural language sentence based upon rewards relating to at least the target text style of the source-style natural language sentence,

wherein the rewards comprise style rewards which are determined using a style classifier built upon a bidirectional recurrent neural network with an attention mechanism.

2. The computer-implemented method of claim 1 , wherein the rewards further comprise, semantic rewards and fluency rewards, and wherein said calculating step comprises evaluating the target-style natural language sentence with respect to content preservation, target text style, and fluency using the style rewards, the semantic rewards, and the fluency rewards, respectively.

3. The computer-implemented method of claim 2 , wherein the semantic rewards are determining using a semantic module configured to determine a Word Mover's Distance (WMD) as an embedding-based similarity metric calculated as a sum of distances between co-occurring words in the source-style natural language sentence relative to the target-style natural language sentence.

4. The computer-implemented method of claim 2 , wherein the fluency rewards are determined using a Recurrent Neural Network (RNN)-based language model.

5. The computer-implemented method of claim 1 , wherein the style classifier is pre-trained on a source and target corpus in style classification.

6. The computer-implemented method of claim 1 , wherein the style classifier is adversarially trained on target-style natural language sentences generated by said transforming step.

7. The computer-implemented method of claim 1 , further comprising guaranteeing, by a LSTM-based style discriminator, a style transfer strength of the transformation above a threshold amount.

8. The computer-implemented method of claim 1 , further comprising guaranteeing, by a sentence discriminator, a content preservation between the source-style natural language sentence and the target-style natural language sentence.

9. The computer-implemented method of claim 1 , further comprising guaranteeing, by a Recurrent Neural Network (RNN)-based language model, a fluency of the target-style natural language sentence.

10. The computer-implemented method of claim 1 , guiding, based on the rewards, the generator in performing a subsequent transformation that achieves a higher accuracy rating, responsive to the accuracy rating being below a threshold amount.

11. The computer-implemented method of claim 1 , repeating said guiding step until the accuracy rating is equal to or greater than the threshold amount.

12. A computer program product for transferring a target text style using Reinforcement Learning (RL), the computer program product comprising a non-transitory computer readable storage medium having program instructions embodied therewith, the program instructions executable by a computer to cause the computer to perform a method comprising:

pre-determining, by a Long Short-Term Memory (LSTM) Neural Network (NN), the target text style of a target-style natural language sentence;

transforming, by the LSTM NN, a source-style natural language sentence into the target-style natural language sentence that maintains the target text style of the target-style natural language sentence; and

calculating an accuracy rating of a transformation of the source-style natural language sentence into the target-style natural language sentence based upon rewards relating to at least the target text style of the source-style natural language sentence,

wherein the rewards comprise style rewards which are determined using a style classifier built upon a bidirectional recurrent neural network with an attention mechanism.

13. The computer program product of claim 12 , wherein the rewards further comprise semantic rewards and fluency rewards, and wherein said calculating step comprises evaluating the target-style natural language sentence with respect to content preservation, target style, and fluency using the style rewards, the semantic rewards, and the fluency rewards, respectively.

14. The computer program product of claim 12 , further comprising guaranteeing, by a LSTM-based style discriminator, a style transfer strength of the transformation above a threshold amount.

15. The computer program product of claim 12 , further comprising guaranteeing, by a sentence discriminator, a content preservation between the source-style natural language sentence and the target-style natural language sentence.

16. The computer program product of claim 12 , further comprising guaranteeing, by a Recurrent Neural Network (RNN)-based language model, a fluency of the target-style natural language sentence.

17. The computer program product of claim 12 , further comprising guiding, based on the rewards, the generator in performing a subsequent transformation that achieves a higher accuracy rating, responsive to the accuracy rating being below a threshold amount.

18. The computer program product of claim 17 , further comprising repeating said guiding step until the accuracy rating is equal to or greater than the threshold amount.

19. A computer processing system for transferring a target text style using Reinforcement Learning (RL), comprising:

a memory device including program code stored thereon;

a hardware processor, operatively coupled to the memory device, and configured to run the program code stored on the memory device to

pre-determine, using a Long Short-Term Memory (LSTM) Neural Network (NN), the target text style of a target-style natural language sentence;

transform, using the LSTM NN, a source-style natural language sentence into the target-style natural language sentence that maintains the target text style of the target-style natural language sentence; and

calculate an accuracy rating of a transformation of the source-style natural language sentence into the target-style natural language sentence based upon rewards relating to at least the target text style of the source-style natural language sentence,

wherein the rewards comprise style rewards which are determined using a style classifier built upon a bidirectional recurrent neural network with an attention mechanism.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2020
From: WU, LINGFEI; XIONG, JINJUN
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 052228/0924 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2020
From: GONG, HONGYU; BHAT, SUMA; HWU, WEN-MEI
To: THE BOARD OF TRUSTEES OF THE UNIVERSITY OF ILLINOIS
Reel/Frame 052228/0943 →
Continuity (1)
Related Publication 20210303803A1 · Sep 30, 2021
Cited By (4)
US 12,210,848 US 12,216,983 US 12,333,247 US 12,549,347