IP Library › Granted Patent US 10,599,885
Granted Patent B2
US 10,599,885 · App. 16/010,156 · Granted Mar 24, 2020

Utilizing discourse structure of noisy user-generated content for chatbot learning

Inventor: Boris Galitsky (San Jose, CA)
Assignee: Oracle International Corporation
G06F40/30G06F40/205G06F40/216G06F40/253G06F40/289G06F40/35G06N3/006G06N5/003G06N5/02G06N5/022G06N20/10
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,599,885
App. No.
16/010,156
Filed
Jun 15, 2018
Granted
Mar 24, 2020
Kind
B2
Art Unit
2657
USPC
704/9
Abstract

Systems, devices, and methods of the present invention uses noisy-robust discourse trees to determine a rhetorical relationship between one or more sentences. In an example, a rhetoric classification application creates a noisy-robust communicative discourse tree. The application accesses a document that includes a first sentence, a second sentence, a third sentence, and a fourth sentence. The application identifies that syntactic parse trees cannot be generated for the first sentence and the second sentence. The application further creates a first communicative discourse tree from the second, third, and fourth sentences and a second communicative discourse tree from the first, third, and fourth sentences. The application aligns the first communicative discourse tree and the second communicative discourse tree and removes any elementary discourse units not corresponding to a relationship that is in common between the first and second communicative discourse trees.

Claims (77)

1. A method of creating a noisy-text robust communicative discourse tree, comprising:

accessing a document comprising a first sentence, a second sentence, a third sentence, and a fourth sentence;

identifying that syntactic parse trees cannot be generated for the first sentence and the second sentence;

creating a first communicative discourse tree from the second, third, and fourth sentences;

creating a second communicative discourse tree from the first, third, and fourth sentences; and

aligning the first communicative discourse tree and the second communicative discourse tree by:

determining a mapping between elementary discourse units in the first communicative discourse tree and the second communicative discourse tree; and

identifying which rhetorical relationships are common between the first communicative discourse tree and the second communicative discourse tree; and

removing, from the first communicative discourse tree and the second communicative discourse tree, any elementary discourse units not corresponding to a relationship that is in common, thereby creating a noisy-text robust communicative discourse tree.

2. The method of claim 1 , wherein identifying that a syntactic parse tree cannot be generated for a particular sentence comprises:

accessing a first confidence score for the sentence, wherein the first confidence score represents a confidence of a first parse;

accessing a second confidence score for the sentence, wherein the second confidence score represents a confidence of a second parse; and

determining that a difference between the first confidence score and the second score confidence is above a threshold.

3. The method of claim 1 , wherein each sentence comprises a plurality of fragments and a verb, and wherein generating a communicative discourse tree comprises:

generating a discourse tree that represents rhetorical relationships between the plurality of fragments, wherein the discourse tree comprises a plurality of nodes, each nonterminal node representing a rhetorical relationship between two of the plurality of fragments, each terminal node of the nodes of the discourse tree is associated with one of the plurality of fragments; and

matching each fragment that has a verb to a verb signature, thereby creating a communicative discourse tree, the matching comprising:

accessing a plurality of verb signatures, wherein each verb signature comprises the verb of the fragment and a sequence of thematic roles, wherein thematic roles describe the relationship between the verb and related words;

determining, for each verb signature of the plurality of verb signatures, a plurality of thematic roles of the respective signature that match a role of a word in the fragment;

selecting a particular verb signature from the plurality of verb signatures based on the particular verb signature comprising a highest number of matches; and

associating the particular verb signature with the fragment.

4. The method of claim 3 , wherein the verb is a communicative verb and each verb signature of the plurality of verb signatures comprises one of (i) an adverb, (ii) a noun phrase, or (iii) a noun.

5. The method of claim 3 , wherein the associating further comprises:

identifying each of the plurality of thematic roles in the particular verb signature; and

matching, for each of the plurality of thematic roles in the particular verb signature, a corresponding word in the fragment to the thematic role.

6. The method of claim 1 , wherein the document comprises text received from an Internet-based source.

7. The method of claim 1 , wherein the document comprises text that is grammatically incorrect.

8. A computer-implemented method for determining a complementarity of a pair of two documents by analyzing noisy-text robust communicative discourse trees, the method comprising:

determining, for a first document, a first noisy-text robust communicative discourse tree comprising a first root node, wherein the noisy-text robust communicative discourse tree is a discourse tree that includes communicative actions and wherein the first document comprises a first sentence for which a parse tree cannot be reliably generated;

determining, for a second document, a second noisy-text robust communicative discourse tree comprising a second root node, wherein the second document comprises a second sentence for which a parse tree cannot be reliably generated;

merging the noisy-text robust communicative discourse trees by identifying that the first root node and the second root node are identical;

computing a level of complementarity between the first noisy-text robust communicative discourse tree and the second noisy-text robust communicative discourse tree by applying a predictive model to the merged communicative discourse tree; and

responsive to determining that the level of complementarity is above a threshold, identifying the first and second documents as complementary.

9. The method of claim 8 , further comprising:

accessing a set of training data comprising a set of training pairs, wherein each training pair comprises a communicative discourse tree that represents text and an expected value; and

training the predictive model by iteratively:

providing one of the training pairs to the predictive model,

receiving, from the predictive model, a predicted value;

calculating a loss function by calculating a difference between the predicted value and the respective expected value; and

adjusting internal parameters of the predictive model to minimize the loss function, wherein the expected value and the predicted value comprise (i) a particular text style, (ii) a presence of an argument, (iii) a validity of the text, (iv) a truthfulness of the text, or (v) an authenticity of the text.

10. The method of claim 8 , further comprising:

accessing a set of training data comprising a set of training pairs, wherein each training pair comprises (i) a communicative discourse tree that represents a question and an answer and (ii) an expected level of complementarity; and

training the predictive model by iteratively:

providing one of the training pairs to the predictive model,

receiving, from the predictive model, a determined level of complementarity;

calculating a loss function by calculating a difference between the determined level of complementarity and the respective expected level of complementarity; and

adjusting internal parameters of the predictive model to minimize the loss function.

11. The method of claim 8 , wherein the predictive model is trained to determine a level of complementarity of sub-trees of two noisy-text robust communicative discourse trees.

12. The method of claim 8 , wherein the predictive model is a support vector model.

13. The method of claim 10 , wherein the communicative discourse tree comprises an answer that is relevant but is rhetorically incorrect when compared to the question.

14. A system comprising:

a non-transitory computer-readable medium storing computer-executable program instructions for creating a noisy-text robust communicative discourse tree; and

a processing device communicatively coupled to the non-transitory computer-readable medium for executing the computer-executable program instructions, wherein executing the non-transitory computer-executable program instructions configures the processing device to perform operations comprising:

accessing a document comprising a first sentence, a second sentence, a third sentence, and a fourth sentence;

identifying that syntactic parse trees cannot be generated for the first sentence and the second sentence;

creating a first communicative discourse tree from the second, third, and fourth sentences;

creating a second communicative discourse tree from the first, third, and fourth sentences; and

aligning the first communicative discourse tree and the second communicative discourse tree by:

determining a mapping between elementary discourse units in the first communicative discourse tree and the second communicative discourse tree; and

identifying which rhetorical relationships are common between the first communicative discourse tree and the second communicative discourse tree; and

removing, from the first communicative discourse tree and the second communicative discourse tree, any elementary discourse units not corresponding to a relationship that is in common, thereby creating a noisy-text robust communicative discourse tree.

15. The system of claim 14 , wherein identifying that a syntactic parse tree cannot be generated for a particular sentence comprises:

accessing a first confidence score for the sentence, wherein the first confidence score represents a confidence of a first parse;

accessing a second confidence score for the sentence, wherein the second confidence score represents a confidence of a second parse; and

determining that a difference between the first confidence score and the second score confidence is above a threshold.

16. The system of claim 14 , wherein each sentence comprises a plurality of fragments and a verb, and wherein generating a communicative discourse tree comprises:

generating a discourse tree that represents rhetorical relationships between the plurality of fragments, wherein the discourse tree comprises a plurality of nodes, each nonterminal node representing a rhetorical relationship between two of the plurality of fragments, each terminal node of the nodes of the discourse tree is associated with one of the plurality of fragments; and

matching each fragment that has a verb to a verb signature, thereby creating a communicative discourse tree, the matching comprising:

accessing a plurality of verb signatures, wherein each verb signature comprises the verb of the fragment and a sequence of thematic roles, wherein thematic roles describe the relationship between the verb and related words;

determining, for each verb signature of the plurality of verb signatures, a plurality of thematic roles of the respective signature that match a role of a word in the fragment;

selecting a particular verb signature from the plurality of verb signatures based on the particular verb signature comprising a highest number of matches; and

associating the particular verb signature with the fragment.

17. The system of claim 16 , wherein the verb is a communicative verb and each verb signature of the plurality of verb signatures comprises one of (i) an adverb, (ii) a noun phrase, or (iii) a noun.

18. The system of claim 16 , wherein the associating further comprises:

identifying each of the plurality of thematic roles in the particular verb signature; and

matching, for each of the plurality of thematic roles in the particular verb signature, a corresponding word in the fragment to the thematic role.

19. The system of claim 16 , wherein the document comprises text received from an Internet-based source.

20. The system of claim 16 , wherein the document comprises text that is grammatically incorrect.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 28, 2018
From: GALITSKY, BORIS
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 046232/0154 →
Continuity (5)
Continuation In Part 15975683 · May 9, 2018
Provisional Application 62680334 · Jun 4, 2018
Provisional Application 62520453 · Jun 15, 2017
Provisional Application 62504377 · May 10, 2017
Related Publication 20180357221A1 · Dec 13, 2018
Cited By (14)
US 12,288,039 US 12,293,158 US 12,356,042 US 12,361,223 US 12,374,136 US 12,423,525 US 12,462,114 US 12,468,694 US 12,505,093 US 12,530,531 US 12,608,416 US 12,614,042 US 12,632,445 US 12,681,997