IP Library › Granted Patent US 10,699,060
Granted Patent B2
US 10,699,060 · App. 16/000,638 · Granted Jun 30, 2020

Natural language processing using a neural network

Inventors: Bryan McCann (Menlo Park, CA); Caiming Xiong (Mountain View, CA); Richard Socher (Menlo Park, CA)
Assignee: salesforce.com, inc.
G06F40/126G06F40/205G06F40/289G06F40/30G06F40/47G06N3/0445G06N3/0454G06N3/08G06F40/44G06F40/58
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,699,060
App. No.
16/000,638
Granted
Jun 30, 2020
Kind
B2
Abstract

A system includes a neural network for performing a first natural language processing task. The neural network includes a first rectifier linear unit capable of executing an activation function on a first input related to a first word sequence, and a second rectifier linear unit capable of executing an activation function on a second input related to a second word sequence. A first encoder is capable of receiving the result from the first rectifier linear unit and generating a first task specific representation relating to the first word sequence, and a second encoder is capable of receiving the result from the second rectifier linear unit and generating a second task specific representation relating to the second word sequence. A biattention mechanism is capable of computing, based on the first and second task specific representations, an interdependent representation related to the first and second word sequences. In some embodiments, the first natural processing task performed by the neural network is one of sentiment classification and entailment classification.

Claims (48)

1. A system comprising:

a neural network for performing a first natural language processing task, the neural network comprising:

a first rectifier linear unit capable of executing an activation function on a first input related to a first word sequence;

a second rectifier linear unit capable of executing an activation function on a second input related to a second word sequence;

a first encoder capable of receiving the result from the first rectifier linear unit and generating a first task specific representation relating to the first word sequence;

a second encoder capable of receiving a result from the second rectifier linear unit and generating a second task specific representation relating to the second word sequence; and

a biattention mechanism capable of computing, based on the first and second task specific representations, an interdependent representation related to the first and second word sequences.

2. The system of claim 1 , wherein the first input is a first concatenated word vector and its corresponding context-specific word vector generated from the first word sequence, and wherein the second input is a second concatenated word vector and its corresponding context-specific word vector generated from the second word sequence.

3. The system of claim 1 , wherein the first natural language processing task performed by the neural network is one of sentiment classification and entailment classification.

4. The system of claim 1 , wherein the neural network is trained using a dataset for one of sentiment analysis, question classification, entailment classification, and question answering.

5. The system of claim 1 , wherein the first word sequence is different from the second word sequence.

6. The system of claim 1 , further comprising an encoder that is pre-trained on a second natural language processing task, the encoder capable of generating a context-specific word vector for one of the first and second word sequences, the context-specific word vector forming at least a part of one of the first and second inputs.

7. The system of claim 6 , wherein the second natural language processing task is machine-translation.

8. The system of claim 6 , wherein the first natural language processing task is different from the second natural language processing task.

9. The system of claim 1 , wherein computing by the biattention mechanism comprises:

computing an affinity matrix based on the first and second task specific representations;

extracting, based on the affinity matrix, a first attention weight relating to the first task specific representation and a second attention weight relating to the second task specific representation; and

generating, based on the first and second attention weights, a first and a second context summaries to condition the first and second task specific representations.

10. The system of claim 9 , wherein the neural network further comprises:

a first integrator capable of integrating the first context summary to generate a first integrated output; and

a second integrator capable of integrating the second context summary to generate a second integrated output.

11. The system of claim 10 , wherein the neural network further comprises:

a first pool mechanism capable of aggregating the first integrated output to generate a first pooled representation relating to the first task specific representation; and

a second pool mechanism capable of aggregating the second integrated output to generate a second pooled representation relating to the second task specific representation.

12. The system of claim 11 , wherein the neural network further comprises a maxout layer capable of combining the first and second pooled representations to generate a result for the first natural language processing task.

13. A method for performing a first natural language processing task, the method comprising:

executing an activation function on a first input related to a first word sequence;

executing an activation function on a second input related to a second word sequence;

generating, based on the execution of the activation function on the first input, a first task specific representation relating to the first word sequence;

generating, based on the execution of the activation function on the second input, a second task specific representation relating to the second word sequence; and

computing, based on the first and second task specific representations, an interdependent representation related to the first and second word sequences.

14. The method of claim 13 , wherein the first input is a first concatenated word vector and its corresponding context-specific word vector generated from the first word sequence, and wherein the second input is a second concatenated word vector and its corresponding context-specific word vector generated from the second word sequence.

15. The method of claim 13 , wherein the first natural language processing task is one of sentiment classification and entailment classification.

16. The method of claim 13 , wherein the first word sequence is different from the second word sequence.

17. The method of claim 13 , further comprising generating, using an encoder that is pre-trained on a second natural language processing task, a context-specific word vector for one of the first and second word sequences, the context-specific word vector forming at least a part of one of the first and second inputs.

18. The method of claim 17 , wherein the second natural language processing task is machine-translation.

19. The method of claim 17 , wherein the first natural language processing task is different from the second natural language processing task.

20. The method of claim 13 , wherein computing the interdependent representation comprises:

computing an affinity matrix based on the first and second task specific representations;

extracting, based on the affinity matrix, a first attention weight relating to the first task specific representation and a second attention weight relating to the second task specific representation; and

generating, based on the first and second attention weights, a first and a second context summaries to condition the first and second task specific representations.

21. The method of claim 13 , further comprising:

integrating a first context summary to generate a first integrated output; and

integrating a second context summary to generate a second integrated output.

22. The method of claim 21 , further comprising:

aggregating the first integrated output to generate a first pooled representation relating to the first task specific representation; and

aggregating the second integrated output to generate a second pooled representation relating to the second task specific representation.

23. The method of claim 22 , further comprising combining the first and second pooled representations to generate a result for the first natural language processing task.

Assignments (2)
CHANGE OF NAME Recorded Aug 4, 2026
From: SALESFORCE.COM, INC.
To: SALESFORCE, INC.
Reel/Frame 076118/0548 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 14, 2019
From: MCCANN, BRYAN; XIONG, CAIMING; SOCHER, RICHARD
To: SALESFORCE.COM, INC.
Reel/Frame 047988/0527 →
Continuity (4)
Continuation 15982841 · May 17, 2018
Provisional Application 62536959 · Jul 25, 2017
Provisional Application 62508977 · May 19, 2017
Related Publication 20180349359A1 · Dec 6, 2018
Cited By (4)
US 12,265,909 US 12,299,982 US 12,530,560 US 12,632,672