Sequence conversion apparatus, machine learning apparatus, sequence conversion method, machine learning method, and program
View Patent ↗A device, to perform highly accurate sequence conversion consistent with the context, uses a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text to obtain a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ, and estimates information corresponding to a t-th word string Y t of the second text by copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 based on the probability of copying.
1. A sequence conversion device, comprising processing circuitry configured to estimate information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein
the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and the processing circuitry configured to:
execute conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ v on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,
execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,
execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
2. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the sequence conversion device according to claim 1 .
3. A sequence conversion device, comprising processing circuitry configured to estimate information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein
the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and the processing circuitry configured to:
execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,
execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,
execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,
execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P(y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
4. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the sequence conversion device according to claim 3 .
5. A machine learning device, comprising processing circuitry
configured to execute machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a t-th word string X t of the first text, on the basis of the model parameter θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and
the model is configured to:
execute conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, i , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ v on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,
execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,
execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
6. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the machine learning device according to claim 5 .
7. A sequence conversion method, comprising:
estimating information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ by using, as inputs, the t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein
the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ p , θ w , θ w , θ o , and
the sequence conversion method comprises:
executing conversion based on the model parameter θ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
executing conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
executing conversion based on the model parameter θ y on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,
executing conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
executing conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the probability p n t , operations of copying the word firm the t-th word string X t , on the basis of the model parameters θ,
executing conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and
executing conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
8. A machine learning method, comprising:
executing machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a word string X t of the t-th first text, on the basis of the model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ p , θ w , θ o , and
the model is configured to:
execute conversion based on the model parameter σ y on the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the first (n−1) word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ y on the sequence u Y, 1 , . . . , u Y, t−1 to obtain a (t−1)-th second text sequence embedded vector v t−1 ,
execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ p on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a copy probability p n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability p n t represents a probability of copying a word from the t-th word string X t , thereby deleting, based on the pobability p n t , operations of copying the word from the t-th word string X t , on the basis of the model parameters θ,
execute conversion based on the model parameter θ w on the second text sequence embedded vector v t−1 and the context vector s n t to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ o on the t-th word string X t of the first text, the posterior probability P(y n t ), and the copy probability p n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
9. A machine learning device, comprising processing circuitry
configured to execute machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a t-th word string X t of the first text, on the basis of the model parameter θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and
the model is configured to:
execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,
execute conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,
execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector g n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,
execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P (y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
10. A non-transitory computer-readable recording medium storing a program for causing a computer to operate as the machine learning device according to claim 9 .
11. A sequence conversion method, comprising:
estimating information corresponding to a t-th word string Y t of a second text, which is a conversion result of a t-th word string X t of a first text, on the basis of one or more model parameters θ by using, as inputs, the t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, wherein
the model parameters θ are obtained by machine learning of a model estimating the information corresponding to the t-th word string Y t , using training data which is a set of multiple couples of a sequence of word strings of the first text and a sequence of word strings of the second text,
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and
the sequence conversion method comprises:
executing conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
executing conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
executing conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,
executing conversion based on the model parameter θ s on a word string y 1 t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
executing conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,
executing conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,
executing conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and
executing conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P(y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.
12. A machine learning method, comprising:
executing machine learning with training data being a sequence of couples of a word string A i of a second text and a word string B i of a first text to obtain one or more model parameters θ of a model that estimates information corresponding to a t-th word string Y t of the second text, which is a conversion result of a word string X t of the t-th first text, on the basis of the model parameters θ, by using, as inputs, a t-th word string X t of the first text and a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of first to (t−1)-th word strings of the second text, which is a conversion result of a sequence X 1 , . . . X t−1 of first to (t−1)-th word strings of the first text, where t is an integer of 2 or greater, X i is a word string of the first text, and Y i is a word string of the second text obtained by rewriting X i , wherein
a language of the first text is same as a language of the second text,
the model parameters θ include model parameters θ y , θ x , θ s , θ v , θ q , θ d , θ m , θ a , and
the model is configured to:
execute conversion based on the model parameter θ y on a sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text to obtain a sequence u Y, 1 , . . . , u Y, t−1 of a text vector u Y, i of a word string Y{circumflex over ( )} i of the second text for i=1, . . . , t−1,
execute conversion based on the model parameter θ x on the t-th word string X t of the first text to obtain a text vector u X, t of the t-th word string X t of the first text,
execute conversion based on the model parameter θ v on a sequence u Y, 1 , . . . , u Y, t−2 to obtain a (t−2)-th second text sequence embedded vector v t−2 ,
execute conversion based on the model parameter θ s on a word string y i t , . . . , y n-1 t previous to an n-th word y n t included in a t-th word string Y{circumflex over ( )} t of the second text and the text vector u X, t to obtain a context vector s n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account the first text, where n is a positive integer equal to or less than a number of words included in the t-th word string Y{circumflex over ( )} t of the second text,
execute conversion based on the model parameter θ q on the word string y 1 t , . . . , y n-1 t , the second text sequence embedded vector v t−2 , and the sequence u Y, t−1 to obtain a context vector g n t for an n-th word of the t-th word string Y{circumflex over ( )} t of the second text taking into account previous word string of the second text,
execute conversion based on the model parameter θ m on the context vector s n t , a (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, and the context vector g n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a copy probability M n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text where the copy probability M n t represents a probability of copying a word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , thereby deleting, based on the probability M n t , operations of copying the word from the t-th word string X t or the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , on the basis of the model parameters θ,
execute conversion based on the model parameter θ d on the context vector s n t and the context vector q n t for the n-th word of the t-th word string Y{circumflex over ( )} t of the second text to obtain a posterior probability P(y n t ) for the n-th word y n t , and
execute conversion based on the model parameter θ a on the t-th word string X t of the first text, the (t−1)-th word string Y{circumflex over ( )} t−1 of the second text, the posterior probability P (y n t ), and the copy probability M n t to obtain a posterior probability P(y n t |y 1 t , . . . , y n-1 t , Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) corresponding to a posterior probability P(Y t |Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 , X t , θ) of the t-th word string Y t of the second text given the t-th word string X t of the first text, the sequence Y{circumflex over ( )} 1 , . . . , Y{circumflex over ( )} t−1 of the word strings of the second text, and the model parameters θ.