IP Library › Granted Patent US 10,460,726
Granted Patent B2
US 10,460,726 · App. 15/426,727 · Granted Oct 29, 2019

Language processing method and apparatus

Inventor: Jihyun Lee (Suwon-si, KR)
Assignee: Samsung Electronics Co., Ltd.
G10L15/197G10L15/02G10L15/063G10L15/16G10L15/183
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,460,726
App. No.
15/426,727
Granted
Oct 29, 2019
Kind
B2
Abstract

A language processing method and apparatus is disclosed. The language processing method includes obtaining, using an encoder, a first feature vector representing an input word based on an input sequence of first characters included in the input word. The method also generates, using a word estimator, a second feature vector representing a predicted word associated with the input word by processing the first feature vector using a language model, and decodes, using a decoder, the second feature vector to an output sequence of second characters included in the predicted word using the second feature vector.

Claims (64)

1. A language processing method, comprising:

obtaining, using an encoder, a first feature vector representing an input word based on an input sequence of first characters of alphabet letters included in the input word;

generating, using a word estimator, a second feature vector representing a predicted word associated with the input word by processing the first feature vector using a language model, the first feature vector being output from the encoder and input to the language model; and

decoding, using a decoder, the second feature vector to an output sequence of second characters included in the predicted word using the second feature vector.

2. The method of claim 1 , wherein either one or both of the first feature vector and the second feature vector is a d-dimension real-valued vector formed by a combination of real values,

wherein d is a natural number smaller than a number of words predictable by the language model.

3. The method of claim 1 , wherein the obtaining of the first feature vector comprises:

encoding the input sequence to the first feature vector at a neural network of the encoder.

4. The method of claim 3 , wherein the obtaining of the first feature vector comprises:

sequentially inputting first vectors representing, respectively, the first characters to an input layer of the neural network of the encoder; and

generating the first feature vector based on activation values to be sequentially generated by an output layer of the neural network of the encoder,

wherein a dimension of the first feature vector corresponds to a number of nodes of the output layer.

5. The method of claim 3 , wherein a dimension of a first vector representing a first character, among the first characters, corresponds to either one or both of a number of nodes of an input layer of the neural network of the encoder and a number of types of characters included in a language of the input word.

6. The method of claim 3 , wherein the neural network of the encoder comprises a recurrent neural network, and

the obtaining of the first feature vector comprises:

generating the first feature vector based on the input sequence and an output value previously generated by the recurrent neural network.

7. The method of claim 1 , wherein the obtaining of the first feature vector comprises:

obtaining the first feature vector representing the input word from a lookup table in which input sequences and feature vectors corresponding to the input sequences are recorded.

8. The method of claim 1 , wherein the obtaining of the first feature vector further comprises:

encoding the input sequence to the first feature vector by inputting the input sequence to a neural network of the encoder.

9. The method of claim 1 , wherein the predicted word comprises a word subsequent to the input word.

10. The method of claim 1 , wherein a neural network of the language model comprises a recurrent neural network, and

the generating of the second feature vector comprises:

generating the second feature vector based on the first feature vector and an output value previously generated by the recurrent neural network.

11. The method of claim 1 , wherein a dimension of the first feature vector corresponds to a number of nodes of an input layer of a neural network of the language model, and

a dimension of the second feature vector corresponds to a number of nodes of an output layer of the neural network of the language model.

12. A language processing method, comprising:

obtaining, using an encoder, a first feature vector representing an input word based on an input sequence of first characters of alphabet letters included in the input word;

generating, using a word estimator, a second feature vector representing a predicted word associated with the input word by processing the first feature vector using a language model; and

decoding, using a decoder, the second feature vector to an output sequence of second characters included in the predicted word using the second feature vector, wherein a dimension of the second feature vector corresponds to a number of nodes of an input layer of a neural network of the decoder, and

a dimension of a second vector representing a second character, among the second characters, corresponds to a number of nodes of an output layer of the neural network of the decoder,

wherein the dimension of the second vector corresponds to a number of types of characters included in a language of the predicted word.

13. The method of claim 12 , wherein the second vector is generated based on activation values generated by the output layer of the neural network of the decoder,

wherein one of the activation values indicates a probability corresponding to the second character.

14. The method of claim 1 , wherein a neural network of the decoder comprises a recurrent neural network, and

the decoding comprises:

generating the output sequence based on the second feature vector and an output value previously generated by the recurrent neural network.

15. The method of claim 1 , wherein the decoding comprises:

generating the output sequence based on activation values to be sequentially generated by an output layer of a neural network of the decoder.

16. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 1 .

17. A training method, comprising:

processing a training sequence of training characters of alphabet letters included in a training word at a neural network of an encoder to produce a training feature vector representing the training word;

processing the training feature vector at a neural network of a language model to produce an output feature vector, the training feature vector being output from the neural network of the encoder and input to the neural network of the language model;

processing the output feature vector at a neural network of a decoder to produce an output sequence corresponding to the training sequence; and

training the neural network of the language model based on a label of the training word and the output sequence.

18. The training method of claim 17 , wherein either one or both of the training feature vector and the output feature vector is a d-dimension real-valued vector formed by a combination of real values, wherein d is a natural number smaller than a number of words predictable by the language model, and

the training of the neural network of the language model comprises:

training the neural network of the language model to represent, through the output feature vector, a predicted word associated with the training word.

19. A training method, comprising:

processing a training sequence of training characters of alphabet letters included in a training word at a neural network of an encoder to produce a training feature vector representing the training word;

processing the training feature vector at a neural network of a language model to produce an output feature vector;

processing the output feature vector at a neural network of a decoder to produce an output sequence corresponding to the training sequence; and

training the neural network of the language model based on a label of the training word and the output sequence, wherein the training method further comprises:

obtaining a second output sequence corresponding to the training sequence by inputting the training feature vector to the neural network of the decoder; and

training the neural network of the encoder and the neural network of the decoder based on the label of the training word and the second output sequence.

20. The training method of claim 19 , wherein a training vector included in the training sequence is a c-dimension one-hot vector, wherein c is a natural number, and a value of a vector element corresponding to the training character is 1 and a value of a remaining vector element is 0, and

the training comprises:

training the neural network of the encoder and the neural network of decoder such that the training sequence and the second output sequence are identical to each other.

21. A non-transitory computer-readable storage medium storing instructions that, when executed by a processor, cause the processor to perform the method of claim 17 .

22. A language processing apparatus, comprising:

a processor configured to:

generate, using an encoder, a first feature vector representing an input word based on an input sequence of first characters of alphabet letters included in the input word;

generate a second feature vector representing a predicted word, associated with the input word, based on the first feature vector using a language model, the first feature vector being output from the encoder and input to the language model; and

generate, using a decoder, an output sequence of second characters included in the predicted word based on the second feature vector.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 7, 2017
From: LEE, JIHYUN
To: SAMSUNG ELECTRONICS CO., LTD.
Reel/Frame 041196/0220 →
Priority Claims (1)
KR 10-2016-0080955 · Jun 28, 2016 · national
Continuity (1)
Related Publication 20170372696A1 · Dec 28, 2017
Cited By (1)
US 12,566,807