IP Library › Granted Patent US 12,591,745
Granted Patent B2
US 12,591,745 · App. 17/937,629 · Granted Mar 31, 2026

Method and system for fine-tuning neural conditional language models using constraints

Inventors: Tomasz Korbak (Warsaw, PL); Hady Elsahar (Grenoble, FR); German Kruszewski (Saint Mandé, FR); Marc Dymetman (Grenoble, FR)
Assignee: NAVER CORPORATION
G06F40/30G06F40/166G06F40/47G06N3/08
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,591,745
App. No.
17/937,629
Granted
Mar 31, 2026
Kind
B2
Abstract

A processor-implemented method for fine-tuning a pre-trained neural conditional language model to perform a downstream task. A pre-trained conditional language model and at least one target constraint for satisfying a task-related control objective are received. A neural model is trained to approximate a target conditional model that optimally reconciles a distance from the pre-trained conditional language model and the control objective across multiple contexts.

Claims (84)

1 . A method implemented by a processor for fine-tuning a pre-trained neural conditional language model to perform a downstream task, the method comprising:

receiving the pre-trained conditional language model having attributes with existing probability distributions conditioned on contexts;

receiving at least one target constraint for satisfying a task-related control objective associated with the downstream task, the target constraint specifying an expectation of a target attribute;

using the processor, training a neural model to approximate a target conditional model that reconciles a distance from the pre-trained conditional language model and the control objective across multiple contexts; and

outputting the trained neural model representing the pre-trained neural conditional language model fine-tuned to perform the downstream task;

wherein said training the neural model comprises:

for each of one or more iterations:

sampling contexts from a distribution of contexts;

computing an unconditional energy-based model (EBM) for each sampled context using the pre-trained conditional language model and the target constraint; and

updating the neural model over the sampled contexts using the computed unconditional EBMs.

2 . The method of claim 1 , wherein the trained neural model comprises one of a sequence-to-sequence (seq2seq) model, an encoder and a decoder.

3 . The method of claim 1 , wherein the trained neural model is trained to approximate a target conditional model that optimally reconciles a distance from the pre-trained conditional language model and the control objective on average across multiple contexts.

4 . The method of claim 1 , wherein the trained neural model comprises a sequence-to-sequence (seq2seq) model, and wherein said training the seq2seq model comprises:

initializing a reference policy; and

training the reference policy by stochastic gradient descent using a loss gradient that minimizes a distance from the pre-trained conditional language model and the control objective across multiple contexts to provide the trained neural seq2seq model.

5 . The method of claim 4 , wherein the loss gradient minimizes a distance from the pre-trained conditional language model and the control objective on average across multiple contexts to provide the trained neural seq2seq model.

6 . The method of claim 4 , wherein the loss gradient minimizes an expected cross-entropy CE between the reference policy and multiple target distributions p c 's, where each target distribution p c is a normalization of an unconditional energy-based model (EBM) P c mapped by the target conditional model to a context c, and where the expected cross-entropy is over the distribution of contexts, the distribution of contexts being a distribution τ(c) of contexts c over a set of contexts C.

7 . The method of claim 6 , wherein the trained neural model comprises a sequence-to-sequence (seq2seq) model

wherein said sampling contexts comprises sampling N contexts from the distribution τ(c) of contexts c;

wherein said computing an unconditional energy-based model comprises computing the unconditional energy-based model (EBM) P c for each sampled context c over the distribution τ(c) of contexts c using the pre-trained conditional language model and the target constraint; and

wherein said updating the neural model comprises updating the neural model over the N contexts by importance sampling using, for each context, M samples from the reference policy.

8 . The method of claim 7 , wherein the distribution τ(c) of contexts c is provided from a set of source documents.

9 . The method of claim 7 , wherein the N sample contexts are source documents from the set of contexts C.

10 . The method of claim 9 , wherein the source documents are provided from one or more of a dataset, an encoder encoding an input sequence, or a portion of the pretrained model.

11 . The method of claim 6 , wherein the unconditional EBM P c corresponds to the target distribution p c ; and

wherein the target distribution p c is conditioned on the context c based on probabilities provided by the pre-trained conditional language model for the context a(x/c) and the target constraint for the context b(x/c).

12 . The method of claim 11 , wherein the target constraint is specified by a constraint satisfaction score that is combined with the probabilities provided by the pre-trained conditional language model for the context.

13 . The method of claim 12 , wherein the target constraint comprises one or more of a pointwise constraint, a binary constraint and a distributional constraint.

14 . The method of claim 4 ,

wherein said updating the neural model comprises:

for each of N sampled contexts, sampling M samples x from the reference policy based on the context;

estimating a normalization for the unconditional energy-based model (EBM) P c for each context c over the M samples using the computed unconditional EBM for the context; and

updating parameters of the reference policy by applying the estimated normalization for each context-sample pair to the loss gradient;

where M and N are hyperparameters.

15 . The method of claim 14 , wherein:

the estimated normalization for each context-sample pair is stored in a buffer during each iteration; and

said applying the estimated normalization for each context-sample pair to the loss gradient comprises:

shuffling the buffer; and

iterating over the shuffled buffer.

16 . The method of claim 15 , wherein said estimating a normalization comprises: computing a score for each sample x using the computed unconditional EBM for the context c; and

computing the estimated normalization using the computed scores.

17 . The method of claim 16 , wherein the estimated normalization comprises a normalizing constant or partition function.

18 . The method of claim 14 , wherein said estimating a normalization further comprises:

computing the estimated normalization over the M obtained samples by importance sampling; and/or

reweighting the obtained samples by their likelihood according to the reference policy.

19 . The method of claim 1 , wherein the reference policy is initialized using the pre-trained conditional language model.

20 . The method of claim 1 , further comprising:

determining if the reference policy has converged with the target conditional model; and

ending the training if the reference policy has converged with the target conditional model.

21 . The method of claim 1 , wherein the downstream task is one of a summarization task, a code generation task, a dialogue task, and a translation task.

22 . The method of claim 1 , wherein the downstream task is a summarization task and the context is provided by processing a document to be summarized, and the constraint is based on factual correctness of the summarized document.

23 . The method of claim 1 , wherein the downstream task is a translation task and the context is provided by processing a document to be translated, and the constraint is based on consistency of terminology.

24 . A method for generating text comprising:

receiving a context by a trained neural language model trained according to the method of claim 1 ;

the trained neural language model generating text in response to the received context.

25 . A method for generating text comprising:

receiving an input sequence;

processing the input sequence to determine a context; and

processing the context by a trained neural model trained according to the method of claim 1 to generate text in response to the context.

26 . A non-transitory computer-readable medium having executable instructions stored thereon for causing a processor and a memory to implement a method for fine-tuning a pre-trained neural conditional language model to perform a downstream task, the method comprising:

receiving the pre-trained conditional language model having attributes with existing probability distributions conditioned on contexts;

receiving at least one target constraint for satisfying a task-related control objective associated with the downstream task, the target constraint specifying an expectation of a target attribute;

training a neural model using the processor to approximate a target conditional model that reconciles a distance from the pre-trained conditional language model and the control objective across multiple contexts; and

outputting the trained neural model;

wherein said training the neural model comprises:

for each of one or more iterations:

sampling contexts from a distribution of contexts;

computing an unconditional energy-based model (EBM) for each sampled context using the pre-trained conditional language model and the target constraint; and

updating the neural model over the sampled contexts using the computed unconditional EBMs.

27 . A processor-based system for fine-tuning a pre-trained neural conditional language model to perform a downstream task, comprising:

an energy-based model (EBM) computing block configured to compute using one or more processors an energy-based model based on a received pre-trained conditional language model having attributes with existing probability distributions conditioned on contexts and at least one received target constraint for satisfying a task-related control objective, the target constraint specifying an expectation of a target attribute;

a context sampling block configured to sample using one or more processors a plurality of contexts;

a model output sampling block configured to sample using one or more processors a plurality of model outputs associated with each sampled context;

a normalization estimating block configured to estimate using one or more processors a normalization based on said computed energy-based model and said sampled plurality of model outputs; and

a model updating block for training a neural model using one or more processors to approximate using one or more processors a target conditional model that reconciles a distance from the pre-trained conditional language model and the control objective across multiple contexts based on a loss gradient determined using said estimated normalization; the model updating block being further configured to output the trained neural model.

28 . The system of claim 27 , wherein said model updating block is configured to iteratively update parameters of the neural model.

29 . The system of claim 28 , further comprising:

a memory for storing said sampled plurality of contexts, said sample plurality of model outputs, and said estimated normalizations;

wherein the stored estimated normalizations are associated with the sampled plurality of contexts and the sample plurality of model outputs in the memory.

30 . The system of claim 29 , wherein said model updating block is configured to iteratively update parameters of the neural model updating block by:

shuffling at least the stored estimated normalizations; and

computing the loss gradient using the shuffled estimated normalizations.

31 . The system of claim 30 , wherein the trained neural model comprises at least a portion of a sequence-to-sequence (seq2seq) model.

32 . The system of claim 31 , wherein the trained neural model is trained to approximate a target conditional model that optimally reconciles a distance from the pre-trained conditional language model and the control objective on average across multiple contexts.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 3, 2022
From: KORBAK, TOMASZ; ELSAHAR, HADY; KRUSZEWSKI, GERMAN; DYMETMAN, MARC
To: NAVER CORPORATION
Reel/Frame 061291/0985 →
Continuity (2)
Provisional Application 63369576 · Jul 27, 2022
Related Publication 20240054338A1 · Feb 15, 2024
References Cited (53)
US 20210357187A1 · Clement · 2021 [cited by examiner]
US 20220083852A1 · Parshakova · 2022 [cited by examiner]
US 20220108081A1 · Dymetman · 2022 [cited by examiner]
US 20220261555A1 · Lebanoff · 2022 [cited by examiner]
EP 3979121A1 · 2022 [cited by applicant]
Alnajjar, K., “When Word Embeddings Become Endangered,” Multilingual Facilitation, University of Helsinki, Mar. 2021, pp. 275-288 published on arXiv:2103.13275v1, Mar. 24, 2021, 14 pages. [cited by applicant]
Austin, J., et al., “Program Synthesis with Large Language Models,” published on arXiv, as 210807732, Aug. 16, 2021, 34 pages. [cited by applicant]
Bielik, P., et al., “PHOG: Probabilistic Model for Code,” ICML'16, JMLR.org, 2016, pp. 2933-2942. [cited by applicant]
Black, S., et al., “GPT-Neo: Large Scale Autoregressive Language Modeling with Mesh-Tensoflow,” Tensorflow (1.0), Zenodo, Mar. 21, 2021, 4 pages. [cited by applicant]
Bommasani, R., et al. “On the Opportunities and Risks of Foundation Models,” published on arXiv as 210807258, Jul. 12, 2022, 214 pages. [cited by applicant]
Brown, T., et al., “Language Models Are Few-Shot Learners,” published on arXiv as 200214165, Jul. 22, 2020, 75 pages. [cited by applicant]
Chen, M., et al., “Evaluating Large Language Models Trained on Code,” published on arXiv as 210703374, Jul. 14, 2021, 35 pages. [cited by applicant]
Csiszar, I., et al., “Information Theory and Statistics: A Tutorial,” Foundations and Trends™ in Communications and Information Theory, 2004, pp. 417-528. [cited by applicant]
Dathathri, S., et al., “Plug and play language models: A simple approach to controlled text generation,” ICLR 2020, Addis Ababa, Ethiopia, Apr. 26-30, 2020, 34 pages. [cited by applicant]
Devlin, J., et al., “Bert: Pre-Training of Deep Bidirectional Transformers for Language Understanding,” Proceedings of NAACL-HLT 2019, Minneapolis, Minnesota, Jun. 2-7, 2019, pp. 4171-4186. [cited by applicant]
Falke, T. et al., “Ranking Generated Summaries by Correctness: An Interesting but Challenging Application for Natural Language Inference,” Proceedings of the 57th Annual Meeting of the Association for Computational Ling… [cited by applicant]
Gao, L., et al., “The Pile: An 800GB Dataset of Diverse Text for Language Modeling,” published on arXiv as 202100027, Dec. 31, 2020, 39 pages. [cited by applicant]
Ghazvininejad, M., et al., “Hafez: An Interactive Poetry Generation System,” Proceedings of ACL 2017, System Demonstrations, Vancouver, Canada: Association for Computational Linguistics, 2017, pp. 43-48. [cited by applicant]
Hinton, G., “Training Products of Experts by Minimizing Contrastive Divergence,” Neural Computation 14, No. 8, Aug. 1, 2002, pp. 1771-1800. [cited by applicant]
Holtzman, A., et al., “Learning to Write with Cooperative Discriminators,” Proceedings of the 56th Annual Meeting of the Association for Computational Linguistics, Melbourne, Australia, Jul. 15-20, 2018, pp. 1638-1649. [cited by applicant]
Holtzman, A., et al., “The Curious Case of Neural Text Degeneration,” 8th International Conference on Learning Representations, ICLR 2020, Addis Ababa, Ethiopia, Apr. 26-30, 2020, 16 pages. [cited by applicant]
Karpathy, A., “Visualizing and Understanding Recurrent Networks,” published on arXiv as 150602078, Nov. 16, 2015. [cited by applicant]
Khalifa, M., et al., “A Distributional Approach to Controlled Text Generation,” International Conference on Learning Representations, 2021, 69 pages. [cited by applicant]
Kingma, D., et al., “Adam: A Method for Stochastic Optimization,” Published as a conference paper at International Conference on Learning Representations, 2015, 15 pages. [cited by applicant]
Koehn, P., “A parallel corpus for statistical machine translation,” Proceedings of Machine Translation Summit X: Papers, Phuket, Thailand, Sep. 13-15, 2005, pp. 79-86. [cited by applicant]
Korbak, T., et al., “Energy-Based Models for Code Generation under Compilability Constraints,” published on arXiv as 210604985, Jun. 9, 2021, 18 pages. [cited by applicant]
Li, J., et al., “A Diversity-Promoting Objective Function for Neural Conversation Models,” Proceedings of the 2016 Conference of the North American Chapter of the Association for Computational Linguistics: Human Languag… [cited by applicant]
Lin, C., Rouge: A Package for Automatic Evaluation of Summaries, Text Summarization Branches Out, Association for Computational Linguistics, Barcelona, Spain, Jul. 2004, pp. 74-81. [cited by applicant]
Lu, S., et al. “CodeXGlue: A Machine Learning Benchmark Dataset for Code Understanding and Generation,” published on arXiv as 210204664, Mar. 16, 2021, 14 pages. [cited by applicant]
Maddison, C. et al., “Structured Generative Models of Natural Source Code,” Proceedings of the 31st International Conference on International Conference on Machine Learning, vol. 32, ICML'14, 2014, pp. II-649-II-657. [cited by applicant]
Maynez, J., et al., “On Faithfulness and Factuality in Abstractive Summarization,” Proceedings of the 58th Annual Meeting of the Association for Computational Linguistics, Association for Computational Linguistics, Jul.… [cited by applicant]
Nallapati, R., et al., “Abstractive Text Summarization Using Sequence-to-Sequence RNNs and Beyond,” Proceedings of the 20th SIGNLL Conference on Computational Natural Language Learning, Berlin, Germany: Association for … [cited by applicant]
Nan, F., et al., “Entity-Level Factual Consistency of Abstractive Text Summarization,” Proceedings of the 16th Conference of the European Chapter of the Association for Computational Linguistics: Main vol. Association f… [cited by applicant]
Nguyen, T., et al., “A Statistical Semantic Language Model for Source Code,” Proceedings of the 2013 9th Joint Meeting on Foundations of Software Engineering, Saint Petersburg Russia: ACM, Aug. 18-26, 2013, pp. 532-542. [cited by applicant]
Parshakova, T., et al., “Distributional Reinforcement Learning for Energy-Based Sequential Models,” published on arXiv as 191208517, Dec. 18, 2019, 17 pages. [cited by applicant]
Parshakova, T., et al., “Global Autoregressive Models for Data-Efficient Sequence Learning,” Proceedings of the 23rd Conference on Computational Natural Language Learning (CoNLL), Hong Kong, China: Association for Compu… [cited by applicant]
Pasunuru, R., et al., “Multi-Reward Reinforced Summarization with Saliency and Entailment,” Proceedings of the 2018 Conference of the North American Chapter of the Association for Computational Linguistics: Human Langua… [cited by applicant]
Paszke, A., et al., “Pytorch: An imperative style, high-performance deep learning library,” 33rd Conference on Neural Information Processing Systems, Vancouver, Canada, 2019, 12 pages. [cited by applicant]
Radford, A., et al., “Learning Transferable Visual Models From Natural Language Supervision,” Proceedings of the 38th International Conference on Machine Learning, 2021, 16 pages. [cited by applicant]
Radford, A., et al., “Language Models Are Unsupervised Multitask Learners,” OpenAI Blog, 1(8):9, 2019, 24 pages. [cited by applicant]
Raffel, C., et al., “Exploring the Limits of Transfer Learning with a Unified Text-to-Text Transformer,” Journal of Machine Learning Research 21, No. 140, 2020, pp. 1-67. [cited by applicant]
Raychev, V., et al., “Probabilistic Model for Code with Decision Trees,” Proceedings of the 2016 ACM SIGPLAN International Conference on Object-Oriented Programming, Systems, Languages, and Applications, Amsterdam Nethe… [cited by applicant]
Raychev, V., et al., “Code Completion with Statistical Language Models,” Proceedings of the 35th ACM SIGPLAN Conference on Programming Language Design and Implementation, Edinburgh United Kingdom: ACM, Jun. 9-11, 2014, … [cited by applicant]
Roziere, B., et al., “Unsupervised Translation of Programming Languages,” 34th Conference on Neural Information Processing Systems Vancouver, Canada, 2020, 11 pages. [cited by applicant]
See, A., et al., “What Makes a Good Conversation? How Controllable Attributes Affect Human Judgments,” Proceedings of the 2019 Conference of the North, Minneapolis, Minnesota: Association for Computational Linguistics, … [cited by applicant]
Sutton, R., et al., “Policy Gradient Methods for Reinforcement Learning with Function Approximation,” Advances in Neural Information Processing Systems 12, MIT Press, 2000, pp. 1057-1063. [cited by applicant]
Von Rossum, G., et al., “PEP 8—Style Guide for Python Code,” Python Enhanacement Proposals, PEP Index, 2001, Retrieved from the Internet: https://github.com/python/peps/blob/main/peps/pep-0008.rst, 37 pages. [cited by applicant]
Williams, R., “Simple Statistical Gradient-Following Algorithms for Connectionist Reinforcement Learning,” Machine Learning 8, 1992, pp. 229-256. [cited by applicant]
Wolf, T., et al., “HuggingFace's Transformers: State-of-the-Art Natural Language Processing,” published on arXiv as 191003771, Jul. 14, 2020, 8 pages. [cited by applicant]
Zhong, V., et al., “Seq2SQL: Generating Structured Queries from Natural Language Using Reinforcement Learning,” Published on arXiv as 170900130, Nov. 9, 2017, 12 pages. [cited by applicant]
Ziegler, D., et al., “Fine-Tuning Language Models from Human Preferences,” published on arXiv as 190908593, Sep. 18, 2019, 26 pages. [cited by applicant]
Montani, I., et al., “explosion/spaCy: v3.4.1: Fix compatibility with CuPy v9.x,” explosion/spaCy: v2.3.5: Bug fixes and simpler source installs (v2.3.5). Zenodo, Jul. 26, 2022, 4 pages. [cited by applicant]
Montani, I., et al., “explosion/spaCy: v2.3.5: Bug fixes and simpler source installs,” explosion/spaCy: v2.3.5: Bug fixes and simpler source installs (v2.3.5). Zenodo, Dec. 11, 2020, 4 pages. [cited by applicant]