IP Library › Granted Patent US 12,210,848
Granted Patent B2
US 12,210,848 · App. 17/682,282 · Granted Jan 28, 2025

Techniques and models for multilingual text rewriting

Inventors: Xavier Eduardo Garcia (New York, NY); Orhan Firat (New York, NY); Noah Constant (Cupertino, CA); Xiaoyue Guo (San Francisco, CA); Parker Riley (Mountain View, CA)
Assignee: Google LLC
G06F40/58G06F40/166G06F40/197G06F40/253G06F40/56G06N3/045G06N3/047G06N3/08G06N3/084
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,210,848
App. No.
17/682,282
Filed
Feb 28, 2022
Granted
Jan 28, 2025
Kind
B2
Art Unit
2172
USPC
715/229
Abstract

The technology provides a model-based approach for multilingual text rewriting that is applicable across many languages and across different styles including formality levels or other textual attributes. The model is configured to manipulate both language and textual attributes jointly. This approach supports zero-shot formality-sensitive translation, with no labeled data in the target language. An encoder-decoder architectural approach with attribute extraction is used to train rewriter models that can thus be used in “universal” textual rewriting across many different languages. A cross-lingual learning signal can be incorporated into the training approach. Certain training processes do not employ any exemplars. This approach enables not just straight translation, but also the ability to create new sentences with different attributes.

Claims (62)

1. A system configured for a multilingual text rewriting model, the system comprising:

memory configured to store a set of text exemplars in a source language and a set of rewritten texts in a plurality of languages different from the source language; and

one or more processing elements operatively coupled to the memory, the one or more processing elements implementing a multilingual text rewriter as a neural network having:

a corruption module configured to generate a corrupted version of an input text sequence obtained based on the set of text exemplars in the source language stored in the memory, the corruption module employing a corruption function that uses negation of an attribute vector associated with the input text sequence;

an encoder module comprising an encoder neural network configured to receive the corrupted version of the input text sequence from the corruption module and to generate a set of encoded representations of the corrupted version of the input text sequence;

a style extractor module configured to extract a set of style vector representations associated with the input text sequence; and

a decoder module comprising a decoder neural network configured to receive the set of encoded representations of the corrupted version of the input text sequence and the set of style vector representations and to output the set of rewritten texts in the plurality of languages, each style vector representation of the set being added element-wise to one of the set of encoded representations of the corrupted version of the input text sequence;

wherein a set of model weights is shared by the encoder module and the style extractor module, and a unique token is appended to the input text sequence for style extraction instead of mean-pooling all of the style vector representations in the set; and

wherein the system is configured to provide rewritten texts in selected ones of the plurality of languages according to a change in a least a sentiment or a formality of the input text sequence.

2. The system of claim 1 , wherein the encoder module, the style extractor module and the decoder module are configured as transformer stacks initialized from a common pretrained language model.

3. The system of claim 1 , wherein both the encoder module and the decoder module are attention-based neural network modules.

4. The system of claim 1 , wherein the corruption function is a corruption function C for a given pair of non-overlapping spans (s1, s2) of the input text sequence, so that the multilingual text rewriting model is capable of reconstructing span s2 from C(s2) and Vs1, where Vs1 is the attribute vector of span s1.

5. The system of claim 4 , wherein the corruption function C is as follows:

C:=f (⋅,source language,− Vs 1).

6. The system of claim 1 , wherein the style extractor module includes a set of style extractor elements including a first subset configured to operate on different style exemplars and a second subset configured to operate on the input text sequence prior to corruption by the corruption module.

7. The system of claim 1 , wherein during training the corruption module is configured to employ at least one of token-level corruption or style-aware back-translation corruption.

8. The system of claim 1 , wherein the set of model weights is initialized with the weights of a pretrained text-to-text model.

9. The system of claim 1 , wherein the model weights for the style extractor module are not tied to the weights of the encoder module during training.

10. The system of claim 1 , wherein the system is configured to extract pairs of random non-overlapping spans of tokens from each line of text in a given text exemplar, and configured to use a first-occurring span in the text as an exemplar of attributes of a second span in the text.

11. The system of claim 1 , wherein a cross-lingual learning signal is added to a training objective for training the model.

12. The system of claim 1 , wherein the encoder module is configured to use negation of a true exemplar vector associated with the input text sequence as the attribute vector in a forward pass operation.

13. The system of claim 1 , wherein the decoder module is configured to receive a set of stochastic tuning ranges that provide conditioning for the decoder module.

14. The system of claim 1 , wherein:

the set of style vector representations corresponds to a set of attributes associated with the input text sequence; and

the system is configured to input a set of exemplars illustrating defined attributes and to extract corresponding attribute vectors for use in rewriting the input text sequence into the plurality of languages according to the defined attributes.

15. The system of claim 14 , wherein:

the system is configured to form an attribute delta vector including a scale factor; and

the attribute delta vector is added to the set of encoded representations of the corrupted version of the input text sequence before processing by the decoder module.

16. A computer-implemented method for providing multilingual text rewriting according to a machine learning model, the method comprising:

obtaining an input text sequence based on a set of text exemplars in a source language stored in a memory;

generating, by a corruption module, a corrupted version of the input text sequence, the corruption module employing a corruption function that uses negation of an attribute vector associated with the input text sequence;

receiving, by an encoder neural network of an encoder module, the corrupted version of the input text sequence from the corruption module;

generating, by the encoder module, a set of encoded representations of the corrupted version of the input text sequence;

extracting, by a style extractor module, a set of style vector representations associated with the input text sequence;

receiving, by a decoder neural network of a decoder module, the set of encoded representations of the corrupted version of the input text sequence and the set of style vector representations, in which each style vector representation of the set is added element-wise to one of the set of encoded representations of the corrupted version of the input text sequence;

outputting, by the decoder module, a set of rewritten texts in a plurality of languages different from the source language; and

storing the rewritten texts in selected ones of the plurality of languages according to a change in a least a sentiment or a formality of the input text sequence;

wherein a set of model weights is shared by the encoder module and the style extractor module, and a unique token is appended to the input text sequence for style extraction instead of mean-pooling all of the style vector representations in the set.

17. The method of claim 16 , wherein the corruption function is a corruption function C for a given pair of non-overlapping spans (s1, s2) of the input text sequence, so that the machine learning model is capable of reconstructing span s2 from C(s2) and Vs1, where Vs1 is the attribute vector of span s1.

18. The method of claim 16 , further comprising:

extracting pairs of random non-overlapping spans of tokens from each line of text in a given text exemplar; and

using a first-occurring span in the text as an exemplar of attributes of a second span in the text.

19. The method of claim 16 , further comprising adding a cross-lingual learning signal to a training objective for training the model.

20. The method of claim 16 , further comprising applying a set of stochastic tuning ranges to selectively condition the decoder module.

21. The method of claim 16 , wherein:

the set of style vector representations corresponds to a set of attributes associated with the input text sequence; and

outputting the set of rewritten texts includes generating one or more versions of the input text sequence in selected ones of the plurality of languages according to the set of attributes.

22. A non-transitory computer-readable medium storing instructions which, when executed, cause a computing device to perform a method of providing multilingual text rewriting according to a machine learning model, the method comprising:

obtaining an input text sequence based on a set of text exemplars in a source language stored in a memory;

generating, by a corruption module, a corrupted version of the input text sequence, the corruption module employing a corruption function that uses negation of an attribute vector associated with the input text sequence;

receiving, by an encoder neural network of an encoder module, the corrupted version of the input text sequence from the corruption module;

generating, by the encoder module, a set of encoded representations of the corrupted version of the input text sequence;

extracting, by a style extractor module, a set of style vector representations associated with the input text sequence;

receiving, by a decoder neural network of a decoder module, the set of encoded representations of the corrupted version of the input text sequence and the set of style vector representations, in which each style vector representation of the set is added element-wise to one of the set of encoded representations of the corrupted version of the input text sequence;

outputting, by the decoder module, a set of rewritten texts in a plurality of languages different from the source language; and

storing the rewritten texts in selected ones of the plurality of languages according to a change in a least a sentiment or a formality of the input text sequence;

wherein a set of model weights is shared by the encoder module and the style extractor module, and a unique token is appended to the input text sequence for style extraction instead of mean-pooling all of the style vector representations in the set.

23. The non-transitory computer-readable medium of claim 22 , wherein the corruption function is a corruption function C for a given pair of non-overlapping spans (s1, s2) of the input text sequence, so that the machine learning model is capable of reconstructing span s2 from C(s2) and Vs1, where Vs1 is the attribute vector of span s1.

24. The non-transitory computer-readable medium of claim 22 , wherein the method further comprises:

extracting pairs of random non-overlapping spans of tokens from each line of text in a given text exemplar; and

using a first-occurring span in the text as an exemplar of attributes of a second span in the text.

25. The non-transitory computer-readable medium of claim 22 , wherein the method further comprises adding a cross-lingual learning signal to a training objective for training the model.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 23, 2024
From: RILEY, PARKER
To: GOOGLE LLC
Reel/Frame 067195/0601 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 3, 2022
From: GARCIA, XAVIER EDUARDO; FIRAT, ORHAN; CONSTANT, NOAH; GUO, XIAOYUE
To: GOOGLE LLC
Reel/Frame 059159/0838 →
Continuity (1)
Related Publication 20230274100A1 · Aug 31, 2023
References Cited (98)
US 7085707B2 · Milner · 2006 [cited by examiner]
US 8032355B2 · Narayanan · 2011 [cited by examiner]
US 9009027B2 · Lehman · 2015 [cited by examiner]
US 9201866B2 · Lehman · 2015 [cited by examiner]
US 9342509B2 · Meng · 2016 [cited by examiner]
US 10418023B2 · Jiang · 2019 [cited by examiner]
US 10423722B2 · Gardner · 2019 [cited by examiner]
US 10606943B2 · Cunico · 2020 [cited by examiner]
US 10789430B2 · Lev Tov · 2020 [cited by examiner]
US 10891435B1 · Ruiz · 2021 [cited by examiner]
US 10902188B2 · Ragan, Jr. · 2021 [cited by examiner]
US 10977439B2 · Mishra · 2021 [cited by examiner]
US 11030990B2 · Jiang · 2021 [cited by examiner]
US 11037028B2 · Bojar · 2021 [cited by examiner]
US 11068654B2 · Beller · 2021 [cited by examiner]
US 11138392B2 · Chen · 2021 [cited by examiner]
US 11151334B2 · Rezagholizadeh · 2021 [cited by examiner]
US 11194971B1 · Dobranic · 2021 [cited by examiner]
US 11200881B2 · Witherspoon · 2021 [cited by examiner]
US 11202131B2 · Zass · 2021 [cited by examiner]
US 11210477B2 · Srinivasan · 2021 [cited by examiner]
US 11314950B2 · Wu · 2022 [cited by examiner]
US 11429785B2 · Kwon · 2022 [cited by examiner]
US 11475223B2 · Chhaya · 2022 [cited by examiner]
US 11487971B2 · Goyal · 2022 [cited by examiner]
US 11551011B2 · Lev-Tov · 2023 [cited by examiner]
US 11562148B2 · Lev-Tov · 2023 [cited by examiner]
US 11574132B2 · Jain · 2023 [cited by examiner]
US 11586828B2 · Lev-Tov · 2023 [cited by examiner]
US 11586830B2 · Maheswaran · 2023 [cited by examiner]
US 11587561B2 · Weir · 2023 [cited by examiner]
US 11610061B2 · Eisenschlos · 2023 [cited by examiner]
US 11630959B1 · Dobranic · 2023 [cited by examiner]
US 11645478B2 · Tambi · 2023 [cited by examiner]
US 11694042B2 · Fei · 2023 [cited by examiner]
US 11741190B2 · Goyal · 2023 [cited by examiner]
US 11763085B1 · Alikaniotis · 2023 [cited by examiner]
US 20030203343A1 · Milner · 2003 [cited by examiner]
US 20050075880A1 · Pickover · 2005 [cited by examiner]
US 20070294077A1 · Narayanan · 2007 [cited by examiner]
US 20100114556A1 · Meng · 2010 [cited by examiner]
US 20100217582A1 · Waibel · 2010 [cited by examiner]
US 20130325437A1 · Lehman · 2013 [cited by examiner]
US 20150120283A1 · Lehman · 2015 [cited by examiner]
US 20150254238A1 · Waibel · 2015 [cited by examiner]
US 20190108212A1 · Cunico · 2019 [cited by examiner]
US 20190108213A1 · Cunico · 2019 [cited by examiner]
US 20190115008A1 · Jiang · 2019 [cited by examiner]
US 20190392813A1 · Jiang · 2019 [cited by examiner]
US 20200057798A1 · Ragan, Jr. · 2020 [cited by examiner]
US 20200097554A1 · Rezagholizadeh · 2020 [cited by examiner]
US 20200110797A1 · Melnyk · 2020 [cited by examiner]
US 20200159826A1 · Lev Tov · 2020 [cited by examiner]
US 20200210772A1 · Bojar · 2020 [cited by examiner]
US 20200311195A1 · Mishra · 2020 [cited by examiner]
US 20200356634A1 · Srinivasan · 2020 [cited by examiner]
US 20200387674A1 · Lev-Tov · 2020 [cited by examiner]
US 20200410171A1 · Lev-Tov · 2020 [cited by examiner]
US 20200410172A1 · Lev-Tov · 2020 [cited by examiner]
US 20210020161A1 · Gao · 2021 [cited by examiner]
US 20210027761A1 · Witherspoon · 2021 [cited by examiner]
US 20210034705A1 · Chhaya · 2021 [cited by examiner]
US 20210165960A1 · Eisenschlos · 2021 [cited by examiner]
US 20210303803A1 · Wu · 2021 [cited by examiner]
US 20210383074A1 · Maheswaran · 2021 [cited by examiner]
US 20210390270A1 · Fei · 2021 [cited by examiner]
US 20210397793A1 · Li · 2021 [cited by examiner]
US 20220075965A1 · Srinivasan · 2022 [cited by examiner]
US 20220121879A1 · Goyal · 2022 [cited by examiner]
US 20220198157A1 · Li · 2022 [cited by examiner]
US 20220414400A1 · Goyal · 2022 [cited by examiner]
US 20230026945A1 · Friedlander · 2023 [cited by examiner]
US 20230137209A1 · Nangi · 2023 [cited by examiner]
US 20230289529A1 · Alikaniotis · 2023 [cited by examiner]
US 20240104294A1 · Shen · 2024 [cited by examiner]
Sandeep Subramanian, “Multiple-Attribute Text Style Transfer”, Nov. 1, 2018, 20 pages https://doi.org/10.48550/arXiv.1811.00552 (Year: 2018). [cited by examiner]
Lample, Guillaume et al. “Multiple-Attribute Text Rewriting.” International Conference on Learning Representations, Sep. 27, 2018, 20 pages, https://openreview.net/pdf?id=H1g2NhC5KQ. (Year: 2018). [cited by examiner]
Bandel, Elron et al. “SimpleStyle: An Adaptable Style Transfer Approach.” ArXiv abs/2212.10498 (Dec. 2022), 12 pages (Year: 2022). [cited by examiner]
Jin, Di et al. “Deep Learning for Text Style Transfer: A Survey.” Computational Linguistics 48 (2020): p. 155-205. https://arxiv.org/pdf/2011.00416.pdf (Year: 2020). [cited by examiner]
Sudhakar, Akhilesh et al. ““Transforming” Delete, Retrieve, Generate Approach for Controlled Text Style Transfer.” Conference on Empirical Methods in Natural Language Processing, Aug. 25, 2019, 11 pages, DOI:10.18653/v1… [cited by examiner]
MohammadiBaghmolaei, Rezvan and Ali Ahmadi. “TET: Text emotion transfer.” Knowledge-Based Systems, Dec. 1, 2022, DOI:10.1016/j.knosys.2022.110236 (Year: 2022). [cited by examiner]
Parker Riley et al., “TextSETTR: Label-Free Text Style Extraction and Tunable Targeted Restyling”, Version 2, Dec. 29, 2020, 16 pages, https://doi.org/10.48550/arXiv.2010.03802, arXiv:2010.03802v2 (Year: 2020). [cited by examiner]
Syed, Bakhtiyar et al. “Adapting Language Models for Non-Parallel Author-Stylized Rewriting.” AAAI Conference on Artificial Intelligence, Sep. 22, 2019, 8 pages; DOI:10.1609/AAAI.V34l05.6433 (Year: 2019). [cited by examiner]
Zhiqiang Hu, “Text Style Transfer: A Review and Experimental Evaluation” Oct. 24, 2020, 32 pages, https://doi.org/10.48550/arXiv.2010.12742 (Year: 2020). [cited by examiner]
Riley, Parker et al. “FRMT: A Benchmark for Few-Shot Region-Aware Machine Translation.” ArXiv abs/2210.00193 (Oct. 1, 2022): 15 pages (Year: 2022). [cited by examiner]
Shrimai Prabhumoye, “Style Transfer Through Back-Translation”, In Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), pp. 866-876, Melbourne, Australia. Apr. 24… [cited by examiner]
Juncen Li; “Delete, Retrieve, Generate: A Simple Approach to Sentiment and Style Transfer”; Apr. 17, 2018, 12 pages, https://doi.org/10.48550/arXiv.1804.06437 (Year: 2018). [cited by examiner]
Anumanchipalli, et al., “Intent Transfer in Speech-to-Speech Machine Translation”, In IEEE Spoken Language Technology Workshop, Dec. 2, 2012, 6 pages. (Year: 2012). [cited by examiner]
Parker Riley et al., “TextSETTR: Label-Free Text Style Extraction and Tunable Targeted Restyling”, Version 1, Oct. 8, 2020, 15 pages, https://doi.org/10.48550/arXiv.2010.03802, arXiv:2010.03802v1 (Year: 2020). [cited by examiner]
Reif, Emily et al. “A Recipe for Arbitrary Text Style Transfer with Large Language Models.” ArXiv abs/2109.03910 ( Sep. 8, 2021): n. pag. 12 pages (Year: 2021). [cited by examiner]
Alex York, The Best 10 Wordtune Alternatives for Rewriting Content in 2024, Dec. 19, 2023, 38 pages, https://clickup.com/blog/wordtune-alternatives/ (Year: 2023). [cited by examiner]
Parker Riley, “Text Rewriting with Missing Supervision”, PhD Thesis; 2021, 129 pages file:///C:/Users/bsmith3/Downloads/Riley_rochester_0188E_12322.pdf (Year: 2021). [cited by examiner]
Mitamura, Teruko , “Controlled Language for Multilingual Machine Translation”, Proceedings of Machine Translation Summit VII, Singapore, Sep. 13-17, 1999. 8 pages. [cited by applicant]
Rosca, Mihaela , et al., “Sequence-to-sequence neural network models for transliteration”, arXiv:1610.09565v1 [cs.CL] Oct. 29, 2016, 5 pages. [cited by applicant]
Sproat, Richard , “Multilingual Text Analysis for Text-to-Speech Synthesis”, ECAI 96. 12th European Conference on Artificial Intelligence, cmp-lg/9608012 Aug. 19, 1996, 6 pages. [cited by applicant]
Riley, Parker , “Text Rewriting with Missing Supervision”, Department of Computer Science, University of Rochester, 2021, pp. 1-129. [cited by applicant]
Riley, Parker , et al., “TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling”, 59th Annual Meeting of the Association for Computational Linguistics and the 11th International Joint Conference on Nat… [cited by applicant]
Riley, Parker , et al., “TextSETTR: Few-Shot Text Style Extraction and Tunable Targeted Restyling”, arXiv:2010.03802v3, Jun. 23, 2021, pp. 1-15. [cited by applicant]