IP Library › Granted Patent US 12,619,835
Granted Patent B2
US 12,619,835 · App. 17/521,204 · Granted May 5, 2026

Adapters for zero-shot multilingual neural machine translation

Inventors: Matthias Galle (Eybens, FR); Alexandre Berard (Grenoble, FR); Laurent Besacier (Seyssinet Pariset, FR); Jerin Philip (Kannur, IN)
Assignee: NAVER CORPORATION
G06F40/58G06N3/04G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,619,835
App. No.
17/521,204
Granted
May 5, 2026
Kind
B2
Abstract

Multilingual neural machine translation systems having monolingual adapter layers and bilingual adapter layers for zero-shot translation include an encoder configured for encoding an input sentence in a source language into an encoder representation and a decoder configured for processing output of the encoder adapter layer to generate a decoder representation. The encoder includes an encoder adapter selector for selecting, from a plurality of encoder adapter layers, an encoder adapter layer for the source language to process the encoder representation. The decoder includes a decoder adapter selector for selecting, from a plurality of decoder adapter layers, a decoder adapter layer for a target language for generating a translated sentence of the input sentence in the target language from the decoder representation.

Claims (38)

1 . A multilingual neural machine translation system comprising at least one processor for translating an input sequence from a source language to a target language, comprising:

an encoder configured for encoding the input sequence in the source language into an encoder representation, wherein the encoder comprises an encoder adapter selector for selecting, from a plurality of encoder adapter layers, an encoder adapter layer corresponding to the source language for processing the encoder representation; and

a decoder configured for processing output of the encoder adapter layer to generate a decoder representation, wherein the decoder comprises a decoder adapter selector for selecting, from a plurality of decoder adapter layers, a decoder adapter layer corresponding to the target language for generating a translation of the input sequence in the target language from the decoder representation;

wherein (i) at least one of each encoder adapter layer and each decoder adapter layer corresponding to a language in a set of languages is trained with parallel data of at least one other language in the set of languages, and (ii) at least another one of an encoder adapter layer and a decoder adapter layer corresponding to a language in the set of languages is not trained with parallel data of at least one other language in the set of languages; and

wherein the multilingual neural machine translation system is configured to perform zero-shot translation using the encoder representation and the decoder representation that are produced by the encoder adapter layer and the decoder adapter layer, respectively, for the language in the set of languages that is not trained with parallel data of at least one other language in the set of languages.

2 . The multilingual neural machine translation system of claim 1 , wherein the at least one of each encoder adapter layer and each decoder adapter layer is a monolingual adapter layer implemented by the processor and trained using parallel data for the set of languages.

3 . The multilingual neural machine translation system of claim 1 , wherein the encoder comprises a plurality of transformer encoder layers forming an encoder pipeline, wherein each transformer encoder layer comprises a respective encoder adapter layer for the source language.

4 . The multilingual neural machine translation system of claim 3 , wherein the decoder comprises a plurality of transformer decoder layers forming a decoder pipeline, wherein each transformer decoder layer comprises a respective decoder adapter layer for the target language.

5 . The multilingual neural machine translation system of claim 1 , wherein the encoder and the decoder comprise transformers, and wherein the encoder adapter layer and the decoder adapter layer are adapter layers comprising a feed-forward network with a bottleneck layer.

6 . The multilingual neural machine translation system of claim 5 , wherein each adapter layer of the plurality of encoder adapter layers and the plurality of decoder adapter layers comprises a residual connection between input of each adapter layer and output of each adapter layer.

7 . The multilingual neural machine translation system of claim 1 , further comprising:

a source pre-processing unit, implemented by the processor, with an initial source embedding layer trained on the plurality of languages and one or more language-specific source embedding layers that are each trained on languages that are not one of the plurality of languages; one of the initial source embedding layer and the one or more language-specific source embedding layers being configured to pre-process the input sequence to generate representations for input to the encoder; and

a target pre-processing unit, implemented by the processor, with an initial target embedding layer trained on the plurality of languages and one or more language-specific target embedding layers that are each trained on the languages that are not one of the plurality of languages; one of the initial target embedding layer and the one or more language-specific target embedding layers being configured to pre-process the input sequence to generate representations for input to the decoder.

8 . The multilingual neural machine translation system of claim 7 , wherein the encoder and the decoder are configured with language-specific parameters that correspond to the one or more language-specific embedding layers, independent of the parameters that correspond to the plurality of languages.

9 . The multilingual neural machine translation system of claim 8 , wherein the source pre-processing unit is configured with language codes that are associated with the one or more language-specific target embedding layers, independent of the initial embedding layers that are associated with the plurality of languages.

10 . A multilingual neural machine translation method for translating an input sequence from a source language to a target language, comprising:

storing in a memory an encoder having a plurality of encoder adapter layers and a decoder having a plurality of decoder adapter layers;

selecting, from the plurality of encoder adapter layers, an encoder adapter layer for the source language;

processing, using the selected encoder adapter layer corresponding to the source language, the input sequence in the source language to generate an encoder representation;

selecting, from the plurality of decoder adapter layers, a decoder adapter layer for the target language; and

processing, using the selected decoder adapter layer corresponding to the target language, the encoder representation to generate a translation of the input sequence in the target language;

wherein (i) the encoder adapter layers and the decoder adapter layers are trained using parallel data for a set of languages, (ii) at least one of each encoder adapter layer and each decoder adapter layer corresponding to a language in the set of languages is trained with parallel data of at least one other language in the set of languages, and (iii) at least another one of an encoder adapter layer and a decoder adapter layer corresponding to a language in the set of languages is not trained with parallel data of at least one other language in the set of languages; and

wherein the multilingual neural machine translation system is configured to perform zero-shot translation using the encoder representation and the decoder representation that are produced by the encoder adapter layer and the decoder adapter layer, respectively, for the language in the set of languages that is not trained with parallel data of at least one other language in the set of languages.

11 . The multilingual neural machine translation method of claim 10 , wherein the at least one of each encoder adapter layer and each decoder adapter layer is a monolingual adapter layer trained using parallel data for the set of languages.

12 . The multilingual neural machine translation method of claim 10 , wherein the at least one of each encoder adapter layer and each decoder adapter layer is a bilingual adapter layer trained using parallel data for the set of languages.

13 . The multilingual neural machine translation method of claim 10 , wherein said storing the encoder further comprises storing a plurality of transformer encoder layers forming an encoder pipeline, wherein each transformer encoder layer comprises a respective encoder adapter layer for the source language.

14 . The multilingual neural machine translation method of claim 13 , wherein said storing the decoder further comprises storing a plurality of transformer decoder layers forming a decoder pipeline, wherein each transformer decoder layer comprises a respective decoder adapter layer for the target language.

15 . The multilingual neural machine translation method of claim 10 , wherein the encoder and the decoder stored in said memory are transformers.

16 . The multilingual neural machine translation method of claim 15 , wherein the encoder adapter layer and the decoder adapter layer stored in said memory are adapter layers comprising a feed-forward network with a bottleneck layer.

17 . The multilingual neural machine translation method of claim 16 , wherein each adapter layer of the plurality of encoder adapter layers and the plurality of decoder adapter layers stored in said memory has a residual connection between input of each adapter layer and output of each adapter layer.

18 . The multilingual neural machine translation method of claim 10 , further comprising:

storing in the memory a source pre-processing unit with an initial source embedding layer trained on the plurality of languages and one or more language-specific source embedding layers that are each trained on languages that are not one of the plurality of languages;

selecting one from the initial source embedding layer and the one or more language-specific source embedding layers to pre-process the input sequence in the source language to generate representations for input to the encoder;

storing in the memory a target pre-processing unit with an initial target embedding layer trained on the plurality of languages and one or more language-specific target embedding layers that are each trained on the languages that are not one of the plurality of languages; and

selecting one from the initial target embedding layer and the one or more language-specific target embedding layers to pre-process the input sequence to generate representations for input to the decoder.

19 . The multilingual neural machine translation method of claim 18 , further comprising:

storing in the memory language-specific parameters for the encoder or the decoder that correspond to the one or more language-specific embedding layers, independent of parameters stored in the memory that correspond to the plurality of languages, wherein said selecting the source embedding layer for the source language selects the language-specific parameters for the encoder or decoder when the source language is not one of the plurality of languages; and

storing in the memory language-specific parameters for the encoder or the decoder that correspond to the one or more language-specific embedding layers, independent of parameters stored in the memory that correspond to the plurality of languages, wherein said selecting the target embedding layer for the target language selects the language-specific parameters for the encoder or the decoder when the target language is not one of the plurality of languages.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2023
From: GALLE, MATTHIAS; BERARD, ALEXANDRE; BESACIER, LAURENT; PHILIP, JERIN
To: NAVER CORPORATION
Reel/Frame 062268/0594 →
Continuity (3)
Provisional Application 63253698 · Oct 8, 2021
Provisional Application 63111863 · Nov 10, 2020
Related Publication 20220147721A1 · May 12, 2022
References Cited (49)
US 10452978B2 · Shazeer et al. · 2019 [cited by applicant]
US 20200380215A1 · Kannan · 2020 [cited by examiner]
Guo et al. “Incorporating bert into parallel sequence decoding with adapters.” Advances in Neural Information Processing Systems 33, p. 10843-10854 (Year: 2020). [cited by examiner]
Aharoni, R., et al., “Massively Multilingual Neural Machine Translation,” Proceedings of the 2019 Conference of the North, Minneapolis, Minnesota, Association for Computational Linguistics, 2019, pp. 3874-3884. [cited by applicant]
Anastasopoulas, A., et al., “Should All Cross-Lingual Embeddings Speak English?” published on ArXiv org as 1911.03058, Nov. 8, 2019, 22 pages. [cited by applicant]
Arivazhagen, N., et al., “Massively Multilingual Neural Machine Translation in the Wild: Findings and Challenges,” published on ArXiv org as:1907.05019, Jul. 11, 2019, 27 pages. [cited by applicant]
Artetxe, M., et al., “On the Cross-Lingual Transferability of Monolingual Representations,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online: Association for Computational … [cited by applicant]
Bañón, M., et al., “ParaCrawl: Web-Scale Acquisition of Parallel Corpora,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Online: Association for Computational Linguistics, Jul.… [cited by applicant]
Bapna, A., et al., “Simple, Scalable Adaptation for Neural Machine Translation,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natur… [cited by applicant]
Berard, A., “Continual Learning in Multilingual NMT via Language-Specific Embeddings,” published on arXiv as 2110.10478, Oct. 20, 2021, 24 pages. [cited by applicant]
Berard, A., et al., “Naver Labs Europe's Systems for the WMT19 Machine Translation Robustness Task,” Proceedings of the Fourth Conference on Machine Translation (vol. 2: Shared Task Papers, Day 1), Florence, Italy, Asso… [cited by applicant]
Berard, A., et al., “Efficient Inference for Multilingual Neural Machine Translation,” published on arXiv as 2109.06679, Sep. 14, 2021, 20 pages. [cited by applicant]
Dabre, R., et al., “A Comprehensive Survey of Multilingual Neural Machine Translation,” published on ArXiv org as 2001.01115, Jan. 7, 2020, 34 pages. [cited by applicant]
Escolano, C., et al., “Multilingual Machine Translation: Closing the Gap between Shared and Language-Specific Encoder-Decoders,” published on ArXiv org as 2004.06575, Apr. 14, 2020, 5 pages. [cited by applicant]
Escolano, C., et al., “From Bilingual to Multilingual Neural Machine Translation by Incremental Training,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research Worksh… [cited by applicant]
Escolano, C., et al., “Training Multilingual Machine Translation by Alternately Freezing Language-Specific Encoders-Decoders,” published on arXiv:2006.01594v1, May 29, 2020, 10 pages. [cited by applicant]
Fan, A., et al., “Beyond English-Centric Multilingual Machine Translation,” published on arXiv:2010.11125v1, Oct. 21, 2020, 38 pages. [cited by applicant]
Firat, O., et al., “Multi-Way, Multilingual Neural Machine Translation with a Shared Attention Mechanism,” Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistic… [cited by applicant]
Freitag, M., et al., “Complete Multilingual Neural Machine Translation,” Proceedings of the Fifth Conference on Machine Translation, Nov. 19-20, 2020, pp. 550-560. [cited by applicant]
Garcia, X., et al., “Towards Continual Learning for Multilingual Machine Translation via Vocabulary Substitution,” Proceedings of the 2021 Conference of the North American Chapter of the Association for Computational Li… [cited by applicant]
Houlsby, N., et al., “Parameter-Efficient Transfer Learning for NLP,” International Conference on Machine Learning, 2019, pp. 2790-2799. [cited by applicant]
Johnson, M., et al., “Google's Multilingual Neural Machine Translation System: Enabling Zero-Shot Translation,” published on ArXiv org as 1611.04558, Nov. 14, 2016, 16 pages. [cited by applicant]
Kingma, D., et al., “Adam: A Method for Stochastic Optimization,” published on arXiv:1412.6980v7, Jul. 20, 2015, 13 pages. [cited by applicant]
Kudugunta, S., et al., “Investigating Multilingual NMT Representations at Scale,” Proceedings of the 2019 Conference on Empirical Methods in Natural Language Processing and the 9th International Joint Conference on Natu… [cited by applicant]
Lakew, S., et al., “Adapting Multilingual Neural Machine Translation to Unseen Languages,” published on arXiv:1910.13998v1, Oct. 30, 2019, 9 pages. [cited by applicant]
Lui, M., et al., “langid.py: An Off-the-shelf Language Identification Tool,” Proceedings of the ACL 2012 System Demonstrations, Jeju Island, Korea, Jul. 8-14, 2021, pp. 25-30. [cited by applicant]
Lyu, S., et al., “Revisiting Modularized Multilingual NMT to Meet Industrial Demands,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Association for Computationa… [cited by applicant]
Neubig, G., et al., “Rapid Adaptation of Neural Machine Translation to New Languages,” published on ArXiv org as 1808.04189, Aug. 13, 2018, 6 pages. [cited by applicant]
Ott, M., et al., “Fairseq: A Fast, Extensible Toolkit for Sequence Modeling,” published on ArXiv org as 1904.01038, Apr. 1, 2019, 6 pages. [cited by applicant]
Ott, M., et al., “Scaling Neural Machine Translation,” Proceedings of the Third Conference on Machine Translation: Research Papers, Belgium, Brussels: Association for Computational Linguistics, Oct. 31-Nov. 1, 2018, 9 p… [cited by applicant]
Pfeiffer, J., et al., “AdapterHub: A Framework for Adapting Transformers,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing: System Demonstrations, Online: Association for Computati… [cited by applicant]
Pfeiffer, J., et al., “Mad-X: An Adapter-Based Framework for Multi-Task Cross-Lingual Transfer,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Association for Co… [cited by applicant]
Pfeiffer, J., et al., “UNKs Everywhere: Adapting Multilingual Language Models to New Scripts,” published on arXiv:2012.15562v1, Dec. 31, 2020, 19 pages. [cited by applicant]
Philip, J., et al., “Monolingual Adapters for Zero-Shot Neural Machine Translation,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Association for Computational … [cited by applicant]
Qi, Y., et al., “When and Why Are Pre-trained Word Embeddings Useful for Neural Machine Translation?” Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Hu… [cited by applicant]
Rebuffi, S., et al., “Learning Multiple Visual Domains with Residual Adapters,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017, 11 pages. [cited by applicant]
Rebuffi, S., et al., “Efficient Parametrization of Multi-Domain Deep Neural Networks,” 2018 IEEE/CVF Conference on Computer Vision and Pattern Recognition, Salt Lake City, UT: IEEE, 2018, pp. 8119-8127. [cited by applicant]
Reimers, N., et al., “Making Monolingual Sentence Embeddings Multilingual Using Knowledge Distillation,” Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing (EMNLP), Online: Associatio… [cited by applicant]
Sennrich, R., et al., “Neural Machine Translation of Rare Words with Subword Units,” Proceedings of the 54th Annual Meeting of the Association for Computational Linguistics (vol. 1: Long Papers), Berlin, Germany: Associ… [cited by applicant]
Stickland, A., et al., “Recipes for Adapting Pre-trained Monolingual and Multilingual Models to Machine Translation,” published on arXiv:2004.14911v1, Apr. 30, 2020, 13 pages. [cited by applicant]
Tan, X., et al., “Multilingual Neural Machine Translation with Knowledge Distillation,” published on ArXiv org as 1902.10461v3, Apr. 30, 2019, 14 pages. [cited by applicant]
Thompson, B., et al., “Freezing Subnetworks to Analyze Domain Adaptation in Neural Machine Translation,” Proceedings of the Third Conference on Machine Translation: Research Papers, Belgium, Brussels: Association for Co… [cited by applicant]
Üstün, A., et al., “Multilingual Unsupervised Neural Machine Translation with Denoising Adapters,” Proceedings of the 2021 Conference on Empirical Methods in Natural Language Processing, Online and Punta Cana, Dominican… [cited by applicant]
Üstün, A., et al., “UDapter: Language Adaptation for Truly Universal Dependency Parsing,” published on arXiv:2004.14327v1, Apr. 29, 2020, 14 pages. [cited by applicant]
Variš, D., et al., “Unsupervised Pretraining for Neural Machine Translation Using Elastic Weight Consolidation,” Proceedings of the 57th Annual Meeting of the Association for Computational Linguistics: Student Research … [cited by applicant]
Vaswani, A., et al., “Attention Is All You Need,” 31st Conference on Neural Information Processing Systems (NIPS 2017), Long Beach, CA, USA, 2017, pp. 5998-6008. [cited by applicant]
Zhang, B., et al., “Improving Massively Multilingual Neural Machine Translation and Zero-Shot Translation,” published on ArXiv org as 2004.11867v1, Apr. 24, 2020, 13 pages. [cited by applicant]
Ziemski, M., et al., “The United Nations Parallel Corpus v1.0,” Proceedings of the Tenth International Conference on Language Resources and Evaluation (LREC'16), Portorož, Slovenia, 2016, pp. 3530-3534. [cited by applicant]
Ba, J., et al., “Layer Normalization,” published on arXiv: 1607.06450v1, Jul. 21, 2016, 14 pages. [cited by applicant]