IP Library Granted Patent US 12,361,227
Granted Patent B2
US 12,361,227 · App. 17/799,588 · Granted Jul 15, 2025

Sequence conversion apparatus, machine learning apparatus, sequence conversion method, machine learning method, and program

Inventors: Mana Ihori (Tokyo, JP); Ryo Masumura (Tokyo, JP)
Assignee: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
G06F40/40
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,361,227
App. No.
17/799,588
Granted
Jul 15, 2025
Kind
B2
Abstract

A device, to perform highly accurate sequence conversion consistent with the context, uses a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text to obtain a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ, and estimates information corresponding to a t-th word string Y t of the second text by copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 based on the probability of copying.

Claims (104)

1. A sequence conversion device, comprising processing circuitry configured to estimate information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein

the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and the processing circuitry configured to:

execute conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ v on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,

execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,

execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

2. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the sequence conversion device according to claim 1 .

3. A sequence conversion device, comprising processing circuitry configured to estimate information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein

the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and the processing circuitry configured to:

execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,

execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,

execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,

execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P(y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

4. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the sequence conversion device according to claim 3 .

5. A machine learning device, comprising processing circuitry

configured to execute machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a t-th word string X t of the first text, on the basis of the model parameter θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and

the model is configured to:

execute conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, i , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ v on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,

execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,

execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

6. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the machine learning device according to claim 5 .

7. A sequence conversion method, comprising:

estimating information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ by using, as inputs, the t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein

the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ p , θ w , θ w , θ o , and

the sequence conversion method comprises:

executing conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

executing conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

executing conversion based on the model parameter θ y on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,

executing conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

executing conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word firm the t-th word string X t , on the basis of the model parameters θ,

executing conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and

executing conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

8. A machine learning method, comprising:

executing machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a word string X t of the t-th first text, on the basis of the model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and

the model is configured to:

execute conversion based on the model parameter σ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ y on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,

execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the pobability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,

execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

9. A machine learning device, comprising processing circuitry

configured to execute machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a t-th word string X t of the first text, on the basis of the model parameter θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and

the model is configured to:

execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,

execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,

execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector g n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,

execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P (y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

10. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the machine learning device according to claim 9 .

11. A sequence conversion method, comprising:

estimating information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ by using, as inputs, the t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein

the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and

the sequence conversion method comprises:

executing conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

executing conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

executing conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,

executing conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

executing conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,

executing conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,

executing conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and

executing conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P(y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

12. A machine learning method, comprising:

executing machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a word string X t of the t-th first text, on the basis of the model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein

a language of the first text is same as a language of the second text,

the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and

the model is configured to:

execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,

execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,

execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,

execute conversion based on the model parameter θ s on a word string y i t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,

execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,

execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector g n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,

execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and

execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P (y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.

Assignments (2)
CHANGE OF NAME Recorded Jan 1, 2026
From: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
To: NTT, INC.
Reel/Frame 074164/0623 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 12, 2022
From: IHORI, MANA; MASUMURA, RYO
To: NIPPON TELEGRAPH AND TELEPHONE CORPORATION
Reel/Frame 060798/0707 →
Continuity (1)
Related Publication 20230072015A1 · Mar 9, 2023
References Cited (12)
US 20180300400A1 · Paulus · 2018 [cited by applicant]
US 20210019479A1 · Tu · 2021 [cited by examiner]
US 20210110254A1 · Hoang · 2021 [cited by examiner]
US 20210406483A1 · Suzuki et al. · 2021 [cited by applicant]
Yamagishi et al., Improving Context-aware Neural Machine Translation with Target-side Context, Sep. 2, 2019, arxiv: 1909.00531, pp. 1-12. (Year: 2019). [cited by examiner]
Yamagishi et al. (2019) “Contextual Neural Machine Translation with Consideration for Natural Language Intersentential Context” Proceedings of the 25th Annual Meeting of the Natural Language Processing Society, with Eng… [cited by applicant]
Luong et al. (2015) “Effective Approaches to Attention-based Neural Machine Translation,” in Proc. EMNLP, pp. 1412-1421. [cited by applicant]
See et al. (2017) “Get to the point: Summarization with pointer-generator networks,” in Proc. Annual Meeting of the Association for Computational Linguistics (ACL), pp. 73-83. [cited by applicant]
Godfrey et al. (1992) “Switchboard: Telephone speech corpus for research and development,” in Proc. ICASSP, pp. 517-520. [cited by applicant]
Ueffing et al. (20130 “Improved models for automatic punctuation prediction for spoken and written text,” in Proc. Interspeech, pp. 3097-31. [cited by applicant]
Maekawa et al. (2000) “Spontaneous speech corpus of Japanese,” in Proc. LREC, pp. 947-9520. [cited by applicant]
Maekawa et al. (2000) “Spontaneous speech corpus of Japanese,” in Proc. International Conference on Language Resources and Evaluation (LREC) ACL Anthology, pp. 947-9520. [cited by applicant]