IP Library › Granted Patent US 12,393,795
Granted Patent B2
US 12,393,795 · App. 18/089,684 · Granted Aug 19, 2025

Modeling ambiguity in neural machine translation

Inventors: Felix Stahlberg (Berlin, DE); Shankar Kumar (New York, NY)
Assignee: GOOGLE LLC
G06F40/58G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,393,795
App. No.
18/089,684
Granted
Aug 19, 2025
Kind
B2
Abstract

The technology addresses ambiguity in neural machine translation. An encoder module receives a given text exemplar and generates an encoded representation of it. A decoder module receives the encoded representation and a set of translation prefixes. The decoder module outputs an unbounded function corresponding to a set of tokens associated with each pair of the given text exemplar and translation prefix from the set of translation prefixes. Each token is assigned a probability between 0 and 1 in a vocabulary of the exemplar at each time step. A logits module generates, based on the unbounded function, a corresponding bounded conditional probability for each token, wherein the probabilities are not normalized over the vocabulary at each time step. A loss function module having a positive loss component and a scaled negative loss component identifies whether each target text of a set of target texts is a valid translation of the exemplar.

Claims (34)

1. A system configured for a machine translation model, the system comprising:

memory configured to store a set of text exemplars in a source language and a set of rewritten texts in one or more languages different from the source language; and

one or more processing elements operatively coupled to the memory, the one or more processing elements implementing the machine translation model as a neural network having:

an encoder module comprising an encoder neural network configured to receive a given text exemplar (x=<x 1 , . . . x |x| >) and to generate an encoded representation of the given text exemplar:

a decoder module comprising a decoder neural network configured to receive the encoded representation and a set of translation prefixes (y <i =<y 1 , . . . , y 1-1 >) and to output an unbounded function ƒ(x, y <a ) corresponding to a set of tokens associated with each pair of the given text exemplar x and translation prefix from the set of translation prefixes ye, wherein each token is assigned a probability between 0 and 1 in a vocabulary of the given text exemplar at each time step:

a logits module configured to act on the unbounded function to generate a corresponding bounded conditional probability for each token, wherein the conditional probabilities are not normalized over the vocabulary at each time step; and

a loss function module having a positive loss component and a scaled negative loss component, in which the loss function module is configured to identify whether each target text of a set of target texts is a valid translation of the given text exemplar.

2. The system of claim 1 , wherein the set of text exemplars comprises a set of input sentences, and the system is configured to learn binary classifiers for each sentence pair (x,y) that indicate whether or not y is a valid translation of x.

3. The system of claim 2 , wherein intrinsic uncertainty is represented by setting the conditional probabilities of at least two correct translations y 1 and y 2 to a maximum probability simultaneously.

4. The system of claim 2 , wherein the logits module is configured to generate probabilities for each translation using separate binary classifiers.

5. The system of claim 1 , wherein a probability of a complete translation of the given text exemplar is decomposed into a product of token-level probabilities.

6. The system of claim 1 , wherein the logits module is configured to apply sigmoid activations to the unbounded function at each time step.

7. The system of claim 1 , wherein the positive loss component applies a log function to the bounded conditional probability for each reference token, and the scaled negative loss component applies a log function to the bounded conditional probability for each non-reference token.

8. The system of claim 1 , wherein during inference the system is configured to search for a translation that has a highest probability of being a translation of a text segment to be translated.

9. The system of claim 1 , wherein the system is configured to express intrinsic uncertainty associated with the machine translation model.

10. The system of claim 1 , wherein the loss function module is configured to adjust scaling of the negative loss component to maximize translation performance.

11. The system of claim 1 , wherein the encoder module and the decoder module comprise a self-attention neural network encoder-decoder architecture.

12. The system of claim 1 , wherein the encoder module and the decoder module comprise a sequence to sequence model architecture.

13. A machine translation method employing a neural network, the method comprising:

storing, in a memory, a set of text exemplars:

receiving, by an encoder module comprising an encoder neural network, a given text exemplar (x=<x 1 , . . . , x x| >):

generating, by the encoder module, an encoded representation of the given text exemplar:

receiving, by a decoder module, the encoded representation and a set of translation prefixes (y <i =<y 1 , . . . , y i-1 >):

outputting, by the decoder module, an unbounded function ƒ(x, y <i ) corresponding to a set of tokens associated with each pair of the given text exemplar x and translation prefix from the set of translation prefixes y <i , wherein each token is assigned a probability between 0 and 1 in a vocabulary of the given text exemplar at each time step:

generating, by a logits module based on the unbounded function, a corresponding bounded conditional probability for each token, wherein the conditional probabilities are not normalized over the vocabulary at each time step; and

identifying, by a loss function module having a positive loss component and a scaled negative loss component, whether each target text of a set of target texts is a valid translation of the given text exemplar.

14. The method of claim 13 , wherein the set of text exemplars comprises a set of input sentences, and the method includes learning binary classifiers for each sentence pair (x,y) that indicate whether or not y is a valid translation of x.

15. The method of claim 14 , wherein intrinsic uncertainty is represented by setting the conditional probabilities of at least two correct translations y 1 and y 2 to a maximum probability simultaneously.

16. The method of claim 14 , wherein generating the corresponding bounded conditional probability for each token comprises generating probabilities for each translation using separate binary classifiers.

17. The method of claim 13 , wherein a probability of a complete translation of the given text exemplar is decomposed into a product of token-level probabilities.

18. The method of claim 13 , wherein generating the corresponding bounded conditional probability for each token comprises applying sigmoid activations to the unbounded function at each time step.

19. The method of claim 13 , wherein the positive loss component applies a log function to the bounded conditional probability for each token, and the scaled negative loss component applies a log function to the bounded conditional probability for each token.

20. The method of claim 13 , further comprising, during inference searching for a translation that has a highest probability of being a translation of a text segment to be translated.

21. The method of claim 13 , further comprising adjusting scaling of the negative loss component to maximize translation performance.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2022
From: STAHLBERG, FELIX; KUMAR, SHANKAR
To: GOOGLE LLC
Reel/Frame 062221/0458 →
Continuity (2)
Continuation PCTUS2022026683 · Apr 28, 2022
Related Publication 20230351125A1 · Nov 2, 2023
References Cited (34)
US 10346548B1 · Wuebker · 2019 [cited by examiner]
US 10452978B2 · Shazeer · 2019 [cited by examiner]
US 20080162111A1 · Bangalore · 2008 [cited by examiner]
US 20110282643A1 · Chatterjee · 2011 [cited by examiner]
US 20140019113A1 · Wu · 2014 [cited by examiner]
US 20140288914A1 · Shen · 2014 [cited by examiner]
US 20190129947A1 · Shin · 2019 [cited by examiner]
US 20200034436A1 · Chen et al. · 2020 [cited by applicant]
US 20200104371A1 · Ma · 2020 [cited by examiner]
US 20200117715A1 · Lee · 2020 [cited by examiner]
US 20200184020A1 · Hashimoto et al. · 2020 [cited by applicant]
US 20210103704A1 · Li · 2021 [cited by examiner]
US 20210286955A1 · Scharnbacher · 2021 [cited by examiner]
US 20210390269A1 · Rezagholizadeh · 2021 [cited by examiner]
US 20220147721A1 · Galle · 2022 [cited by examiner]
US 20220366152A1 · Zenkel · 2022 [cited by examiner]
US 20230169281A1 · Zheng · 2023 [cited by examiner]
US 20230267285A1 · Wu · 2023 [cited by examiner]
Stahlberg and Kumar, “Jam or Cream First? Modeling Ambiguity in Neural Network Translation with SCONES”, Proceedings of the 2022 Conference of the North American Chapter of the Association for Computational Linguistics:… [cited by examiner]
Ro et al., “Transformer-based Models of Text Normalization for Speech Applications”, arXiv:2202.00153v1, Feb. 1, 2022, 5 Pages. (Year: 2022). [cited by examiner]
Sia et al., “Prefix Embeddings for In-context Machine Translation”, Proceedings of the 15th Biennial Conference of the Association for Machine Translation in the Americas (vol. 1: Research Track), Sep. 2022, pp. 45-57. … [cited by examiner]
Stahlberg et al., “Seq2Edits: Sequence Transduction Using Span-level Edit Operations”, Proceedings of the 2020 Conference on Empirical Methods in Natural Language Processing, Nov. 16-20, 2020, pp. 5147-5159. (Year: 2020… [cited by examiner]
Kano et al., “Simultaneous Neural Machine Translation with Prefix Alignment”, Proceedings of the 19th International Conference on Spoken Language Translation (IWSLT 2022), May 26-27, 2022, pp. 22-31. (Year: 2022). [cited by examiner]
International Search Report and Written Opinion for Application No. PCT/US2022/026683 dated Dec. 1, 2022 (11 pages). [cited by applicant]
Bradbury , et al., “Composable transformations of Python+NumPy programs: differentiate, vectorize, JIT to GPU/TPU, and more”, 2018, https://github.com/jax-ml/jax?tab=readme-ov-file#readme, downloaded from internet on Fe… [cited by applicant]
Brown, Peter F., et al., “The Mathematics of Statistical Machine Translation: Parameter Estimation”, IBM T.J. Watson Research Center, Yorktown Heights, NY 10598, Computational Linguistics, vol. 19, No. 2, 1993, pp. 263-… [cited by applicant]
Gao, Qin , et al., “Parallel Implementations ofWord Alignment Tool”, Software Engineering, Testing, and Quality Assurance for Natural Language Processing, pp. 49-57, Columbus, Ohio, USA, Jun. 2008. [cited by applicant]
Gutmann, Michael , et al., “Noise-contrastive estimation: A new estimation principle for unnormalized statistical models”, Appearing in Proceedings of the 13th International Conference on Artificial Intelligence and Sta… [cited by applicant]
Koehn, Philipp , et al., “Six Challenges for Neural Machine Translation”, arXiv:1706.03872v1 [cs.CL] Jun. 12, 2017, 12 pages. [cited by applicant]
Kudo, Taku , et al., “SentencePiece: A simple and language independent subword tokenizer and detokenizer for Neural Text Processing”, arXiv:1808.06226v1 [cs.CL] Aug. 19, 2018, 6 pages. [cited by applicant]
Post, Matt , “A Call for Clarity in Reporting BLEU Scores”, Amazon Research, Berlin, Germany, arXiv:1804.08771v2 [cs.CL] Sep. 12, 2018, 6 pages. [cited by applicant]
Stahlberg, Felix , et al., “On NMT Search Errors and Model Errors: Cat Got Your Tongue?”, University of Cambridge, Department of Engineering, Trumpington St, Cambridge CB2 1PZ, UK, arXiv:1908.10090v1 [cs.CL] Aug. 27, 20… [cited by applicant]
Sutskever, Ilya, et al., “Sequence to Sequence Learning with Neural Networks”, arXiv:1409.3215v3 [cs.CL] Dec. 14, 2014, pp. 1-9. [cited by applicant]
You, Yang , et al., “Large Batch Optimization for Deep Learning: Training Bert in 76 Minutes”, arXiv:1904.00962v5 [cs.LG] Jan. 3, 2020, Published as a conference paper at ICLR 2020, pp. 1-37. [cited by applicant]