IP Library › Granted Patent US 12,153,896
Granted Patent B2
US 12,153,896 · App. 17/391,178 · Granted Nov 26, 2024

Method and system for controlling distributions of attributes in language models for text generation

Inventors: Marc Dymetman (Grenoble, FR); Hady Elsahar (Grenoble, FR); Muhammad Khalifa (Grenoble, FR)
Assignee: Naver Corporation
G06F40/40G06F40/10G06N7/01G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,153,896
App. No.
17/391,178
Filed
Aug 2, 2021
Granted
Nov 26, 2024
Kind
B2
Art Unit
2695
USPC
704/9
Abstract

A method for generating a language model for text generation by receiving a pre-trained language model having attributes with existing probability distributions over the pre-trained language model; receiving at least one target constraint; the target constraint specifying an expectation of a target attribute over a language model that approximates the pre-trained language model; computing parameters of an energy based model by applying the target constraint to the pre-trained language model; obtaining samples from a reference policy; updating parameters of a target policy using the obtained samples and the energy based model; updating the reference policy with the target policy if the target policy is superior to the reference policy; and outputting the target policy as a target language model. The target language model is adapted to generate text with the target attribute over a probability distribution that approximates the desired probability distribution specified by the target constraint.

Claims (40)

1. A method for generating, from a pre-trained language model, a target language model for controlled text generation, the target language model having minimal divergence with pre-trained language model distribution, comprising:

(a) receiving a pre-trained language model having attributes with existing probability distributions over the pre-trained language model;

(b) receiving at least one target constraint, the received target constraint specifying an expectation of a target attribute over the target language model, the target language model approximating the pre-trained language model;

(c) computing parameters of an energy based model by applying the received target constraint to the pre-trained language model;

(d) obtaining samples from a reference policy;

(e) updating parameters of a target policy using the obtained samples from the reference policy and the energy based model;

(f) updating the reference policy with the target policy if a first distance between the target policy and an implicit probability distribution, the implicit probability distribution being represented by the energy based model, is smaller than a second distance between the reference policy and the implicit probability distribution represented by the energy based model, the first and second distances being calculated as a divergence;

(g) repeating (d), (e) and (f) until the target policy converges with the target constraint; and

(h) outputting the target policy as the target language model having minimal divergence with pre-trained language model distribution and configured to generate controlled text with the target attribute over a probability distribution approximating a probability distribution specified by the target constraint.

2. The method as claimed in claim 1 , wherein the parameters of the target policy are updated using a distributional policy gradient computed using the obtained samples from the reference policy and the energy based model.

3. The method as claimed in claim 1 , wherein the parameters of the target policy are updated using a distributional policy gradient computed using the obtained samples from the reference policy and samples obtained from the energy based model using a Monte Carlo method.

4. The method as claimed in claim 1 , wherein the target policy and the reference policy are initialized with the pre-trained language model.

5. The method as claimed in claim 1 , wherein the pre-trained language model is an autoregressive model.

6. The method as claimed in claim 1 , wherein the pre-trained language model is GPT-2.

7. The method as claimed in claim 1 , wherein the target constraint is one of a pointwise constraint, a distributional constraint and a hybrid constraint.

8. The method as claimed in claim 1 , wherein the target constraint is two of a pointwise constraint, a distributional constraint and a hybrid constraint.

9. The method as claimed in claim 1 , wherein the target constraint is a distributional constraint which is verified using a plurality of outputs from the reference policy.

10. The method as claimed in claim 1 , wherein the target language model generates text with the target attribute.

11. The method as claimed in claim 1 , wherein the divergence is calculated as a KL divergence.

12. A non-transitory computer-readable medium, on which is stored a computer program product comprising:

code instructions, when the computer program product is executed on a computer, to execute a method for generating, from a pre-trained language model, a target language model for controlled text generation, the target language model having minimal divergence with distribution of the pre-trained language model;

said code instructions executing

(a) receiving a pre-trained language model having attributes with existing probability distributions over the pre-trained language model;

(b) receiving at least one target constraint, the received target constraint specifying an expectation of a target attribute over the target language model, the target language model approximating the pre-trained language model;

(c) computing parameters of an energy based model by applying the received target constraint to the pre-trained language model;

(d) obtaining samples from a reference policy;

(e) updating parameters of a target policy using the obtained samples from the reference policy and the energy based model;

(f) updating the reference policy with the target policy if a first distance between the target policy and an implicit probability distribution, the implicit probability distribution being represented by the energy based model, is smaller than a second distance between the reference policy and the implicit probability distribution represented by the energy based model, the first and second distances being calculated as a divergence;

(g) repeating (d), (e) and (f) until the target policy converges with the target constraint; and

(h) outputting the target policy as the target language model having minimal divergence with pre-trained language model distribution and configured to generate controlled text with the target attribute over a probability distribution approximating a probability distribution specified by the target constraint.

13. The non-transitory computer-readable medium as claimed in claim 12 , wherein the parameters of the target policy are updated using a distributional policy gradient computed using the obtained samples from the reference policy and the energy based model.

14. The non-transitory computer-readable medium as claimed in claim 12 , wherein the parameters of the target policy are updated using a distributional policy gradient computed using the obtained samples from the reference policy and samples obtained from the energy based model using a Monte Carlo method.

15. The non-transitory computer-readable medium as claimed in claim 12 , wherein the target policy and the reference policy are initialized with the pre-trained language model.

16. The non-transitory computer-readable medium as claimed in claim 12 , wherein the pre-trained language model is an autoregressive model.

17. The non-transitory computer-readable medium as claimed in claim 12 , wherein the pre-trained language model is GPT-2.

18. The non-transitory computer-readable medium as claimed in claim 12 , wherein the target constraint is one of a pointwise constraint, a distributional constraint and a hybrid constraint.

19. The non-transitory computer-readable medium as claimed in claim 12 , wherein the target constraint is two of a pointwise constraint, a distributional constraint and a hybrid constraint.

20. The non-transitory computer-readable medium as claimed in claim 12 , wherein the target constraint is a distributional constraint which is verified using a plurality of outputs from the reference policy.

21. The non-transitory computer-readable medium as claimed in claim 12 , wherein the target language model generates text with the target attribute.

22. The non-transitory computer-readable medium as claimed in claim 12 , wherein the divergence is calculated as a KL divergence.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: DYMETMAN, MARC; KHALIFA, MUHAMMAD
To: NAVER CORPORATION
Reel/Frame 064996/0575 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 22, 2023
From: NAVER FRANCE
To: NAVER CORPORATION
Reel/Frame 064996/0704 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 26, 2023
From: DYMETMAN, MARC; KHALIFA, MUHAMMAD
To: NAVER CORPORATION
Reel/Frame 064385/0225 →
EMPLOYMENT AGREEMENT WITH ASSIGNMENT PROVISION Recorded Jul 26, 2023
From: ELSAHAR, HADY
To: NAVER FRANCE
Reel/Frame 064603/0799 →
Priority Claims (2)
FR 2010054 · Oct 1, 2020 · national
EP 21305835 · Jun 17, 2021 · regional
Continuity (1)
Related Publication 20220108081A1 · Apr 7, 2022
Cited By (1)
US 12,639,531