IP Library › Granted Patent US 10,817,650
Granted Patent B2
US 10,817,650 · App. 15/982,841 · Granted Oct 27, 2020

Natural language processing using context specific word vectors

Inventors: Bryan McCann (Menlo Park, CA); Caiming Xiong (Mountain View, CA); Richard Socher (Menlo Park, CA)
Assignee: salesforce.com, inc.
G06F40/126G06F40/205G06F40/289G06F40/30G06F40/47G06N3/0445G06N3/0454G06N3/08G06F40/44G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,817,650
App. No.
15/982,841
Granted
Oct 27, 2020
Kind
B2
Abstract

A system is provided for natural language processing. In some embodiments, the system includes an encoder for generating context-specific word vectors for at least one input sequence of words. The encoder is pre-trained using training data for performing a first natural language processing task. A neural network performs a second natural language processing task on the at least one input sequence of words using the context-specific word vectors. The first natural language process task is different from the second natural language processing task and the neural network is separately trained from the encoder. In some embodiments, the first natural processing task can be machine translation, and the second natural processing task can be one of sentiment analysis, question classification, entailment classification, and question answering.

Claims (31)

1. A system for natural language processing, the system comprising:

a non-transitory memory; and

one or more hardware processors coupled to the non-transitory memory and configured to execute instructions to cause the system to perform operations comprising:

converting, using an encoder, at least one input sequence of words in a first language to a sequence of word vectors, wherein the encoder is pre-trained using training data for translating phrases of a first language to phrases of a second language;

generating, using the encoder, a sequence of context-specific word vectors for the at least one input sequence of words in the first language;

concatenating, using the encoder, the sequence of word vectors and the sequence of context-specific word vectors into a sequence of concatenated vectors;

receiving, at a multi-layer neural network from the encoder, the sequence of concatenated vectors; and

performing, using the multi-layer neural network, a first natural language processing task on the at least one input sequence of words in the first language using the sequence of concatenated vectors.

2. The system of claim 1 , wherein the first natural processing task is one of sentiment analysis, question classification, entailment classification, and question answering.

3. The system of claim 1 , wherein the encoder comprises at least one bidirectional long-term short-term memory configured to process at least one word vector in the sequence of word vectors.

4. The system of claim 3 , wherein the encoder comprises an attention mechanism configured to compute an attention weight based on an output of the at least one bidirectional long-term short-term memory.

5. The system of claim 1 , wherein the operations further comprise pre-training the encoder using a decoder, wherein the decoder is initialized with hidden vectors generated by the encoder.

6. The system of claim 5 , wherein the decoder comprises at least one bidirectional long-term short-term memory configured to process at least one word vector in the second language during training of the encoder.

7. The system of claim 1 , wherein the operations further comprise generating, using a biattentive classification network, attention weights based on the sequence of concatenated vectors.

8. A system for natural language processing, the system comprising:

an encoder for generating context-specific word vectors for at least one input sequence of words, wherein the encoder is pre-trained using training data for performing a first natural language processing task; and

a neural network for performing a second natural language processing task on the at least one input sequence of words using a received sequence of concatenated vectors comprising a concatenation of word vectors and the context-specific word vectors, wherein the first natural language process task is different from the second natural language processing task and the neural network is separately trained from the encoder.

9. The system of claim 8 , wherein the first natural language processing task includes machine translation.

10. The system of claim 8 , wherein the second natural language processing task is one of sentiment analysis, question classification, entailment classification, and question answering.

11. The system of claim 8 , wherein the encoder is pre-trained using a machine translation dataset.

12. The system of claim 8 , wherein the neural network is trained using a dataset for one of sentiment analysis, question classification, entailment classification, and question answering.

13. The system of claim 8 , wherein the first natural language processing task is different from the second natural language processing task.

14. The system of claim 8 , wherein the encoder comprises at least one bidirectional long-term short-term memory.

15. A method comprising:

using an encoder, generating context-specific word vectors for at least one input sequence of words, wherein the encoder is pre-trained using training data for performing a first natural language processing task; and

using a neural network, performing a second natural language processing task on the at least one input sequence of words using a received sequence of concatenated vectors comprising a concatenation of word vectors and the context-specific word vectors, wherein the first natural language process task is different from the second natural language processing task and the neural network is separately trained from the encoder.

16. The method of claim 15 , wherein the first natural language processing task is machine translation.

17. The method of claim 15 , wherein the second natural language processing task is one of sentiment analysis, question classification, entailment classification, and question answering.

18. The method of claim 15 , wherein the encoder is pre-trained using a machine translation dataset.

19. The method of claim 15 , wherein the neural network is trained using a dataset for one of sentiment analysis, question classification, entailment classification, and question answering.

20. The method of claim 15 , wherein the first natural language processing task is different from the second natural language processing task.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 1, 2018
From: MCCANN, BRYAN; XIONG, CAIMING; SOCHER, RICHARD
To: SALESFORCE.COM, INC.
Reel/Frame 045963/0376 →
Continuity (3)
Provisional Application 62536959 · Jul 25, 2017
Provisional Application 62508977 · May 19, 2017
Related Publication 20180373682A1 · Dec 27, 2018
Cited By (5)
US 12,265,909 US 12,299,020 US 12,430,515 US 12,481,834 US 12,530,560