IP Library › Granted Patent US 11,928,437
Granted Patent B2
US 11,928,437 · App. 17/568,527 · Granted Mar 12, 2024

Machine reading between the lines

Inventor: Boris Galitsky (San Jose, CA)
Assignee: Oracle International Corporation
G06F40/35G06F16/3329G06F16/3344G06F16/36G06F40/284G06F40/289G06F40/30
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,928,437
App. No.
17/568,527
Granted
Mar 12, 2024
Kind
B2
Abstract

Techniques for identifying one or more missing fragments within input text are disclosed. A discourse tree (DT) is generated for the input text (IT) received, the IT having any suitable number of sentence fragments. An indication that the IT is likely missing one or more sentence fragments may be identified based on determining that one or more rhetorical relationships of the DT matches one of a set of predefined rhetorical relationships. A query is generated one or more sentence fragments of the IT and executed against a knowledge base to obtain a set of search results. A most-relevance search result can be utilized to identify a set of candidate sentence fragments. A subset of those candidate sentence fragments can be identified based on comparing them to the sentence fragments provided in the IT, each candidate sentence fragment of the subset being implied but excluded from the IT.

Claims (52)

1. A method of identifying one or more missing natural language expressions that are implied, but missing from input text, the method comprising:

receiving the input text comprising a plurality of sentence fragments;

generating a discourse tree that represents rhetorical relations between the sentence fragments, the discourse tree including a plurality of nodes, each nonterminal node of the plurality of nodes representing a rhetorical relationship between two of the sentence fragments and each terminal node of the plurality of nodes being associated with one of the sentence fragments;

identifying that the input text is likely missing one or more sentence fragments based at least in part on identifying that one or more rhetorical relationships of the discourse tree matches one of a set of predefined rhetorical relationships;

generating a query based at least in part on a subset of the plurality of sentence fragments;

obtaining a set of search results based at least in part on executing the query against a knowledge base;

obtaining, from a search result of the set of search results, a set of candidate sentence fragments for the missing one or more sentence fragments;

identifying a subset of sentence fragments from the set of candidate sentence fragments based at least in part on comparing the sentence fragments of the discourse tree to the set of candidate sentence fragments obtained from the search result; and

performing one or more operations based at least in part on identifying the subset of sentence fragments, the subset of sentence fragments being the natural language expressions that are implied but excluded from the input text.

2. The method of claim 1 , wherein the knowledge base is an online knowledge base.

3. The method of claim 1 , wherein generating the query further comprises:

identifying repetitive sentence fragments from the plurality of sentence fragments of the input text; and

generating a generalized statement from the repetitive sentence fragments, wherein the query is generated from the generalized statement.

4. The method of claim 1 , further comprising selecting the search result of the set of search results based at least in part on identifying that a relevance value between the search result and the query exceeds a predefined threshold.

5. The method of claim 1 , wherein obtaining the set of candidate sentence fragments for the missing one or more sentence fragments comprises generating a respective discourse tree from the search result.

6. The method of claim 5 , wherein identifying the subset of sentence fragments further comprises identifying a second predefined set of rhetorical relations and corresponding fragments of the respective discourse tree generated from the search result.

7. The method of claim 6 , wherein the second predefined set of rhetorical relations comprises at least one of: attribution, condition, background, contrast, cause, or explanation.

8. A computing device, comprising:

one or more processors; and

one or more memories storing computer-executable instructions for identifying one or more missing natural language expressions that are implied, but missing from input text, that, when executed by the one or more processors, cause the computing device to:

receive the input text comprising a plurality of sentence fragments;

generate a discourse tree that represents rhetorical relations between the sentence fragments, the discourse tree including a plurality of nodes, each nonterminal node of the plurality of nodes representing a rhetorical relationship between two of the sentence fragments and each terminal node of the plurality of nodes being associated with one of the sentence fragments;

identify that the input text is likely missing one or more sentence fragments based at least in part on identifying that one or more rhetorical relationships of the discourse tree matches one of a set of predefined rhetorical relationships;

generate a query based at least in part on a subset of the plurality of sentence fragments;

obtain a set of search results based at least in part on executing the query against a knowledge base;

obtain, from a search result of the set of search results, a set of candidate sentence fragments for the missing one or more sentence fragments;

identify a subset of sentence fragments from the set of candidate sentence fragments based at least in part on comparing the sentence fragments of the discourse tree to the set of candidate sentence fragments obtained from the search result; and

perform one or more operations based at least in part on identifying the subset of sentence fragments, the subset of sentence fragments being natural language expressions that are implied but excluded from the input text.

9. The computing device of claim 8 , wherein the knowledge base is an online knowledge base.

10. The computing device of claim 8 , wherein executing the instructions to generate the query further causes the computing device to:

identify repetitive sentence fragments from the plurality of sentence fragments of the input text; and

generate a generalized statement from the repetitive sentence fragments, wherein the query is generated from the generalized statement.

11. The computing device of claim 8 , wherein executing the instructions further causes the computing device to select the search result of the set of search results based at least in part on identifying that a relevance value between the search result and the query exceeds a predefined threshold.

12. The computing device of claim 8 , wherein executing the instructions to obtain the set of candidate sentence fragments for the missing one or more sentence fragments further causes the computing device to generate a respective discourse tree from the search result.

13. The computing device of claim 12 , wherein executing the instructions to identify the subset of sentence fragments further causes the computing device to identify a second predefined set of rhetorical relations and corresponding fragments of the respective discourse tree generated from the search result.

14. The computing device of claim 13 , wherein the second predefined set of rhetorical relations comprises at least one of: attribution, condition, background, contrast, cause, or explanation.

15. A non-transitory computer readable medium storing instructions for identifying one or more missing natural language expressions that are implied, but missing from input text, that, when executed by one or more processors of a computing device, cause the computing device to:

receive the input text comprising a plurality of sentence fragments;

generate a discourse tree that represents rhetorical relations between the sentence fragments, the discourse tree including a plurality of nodes, each nonterminal node of the plurality of nodes representing a rhetorical relationship between two of the sentence fragments and each terminal node of the plurality of nodes being associated with one of the sentence fragments;

identify that the input text is likely missing one or more sentence fragments based at least in part on identifying that one or more rhetorical relationships of the discourse tree matches one of a set of predefined rhetorical relationships;

generate a query based at least in part on a subset of the plurality of sentence fragments;

obtain a set of search results based at least in part on executing the query against a knowledge base;

obtain, from a search result of the set of search results, a set of candidate sentence fragments for the missing one or more sentence fragments;

identify a subset of sentence fragments from the set of candidate sentence fragments based at least in part on comparing the sentence fragments of the discourse tree to the set of candidate sentence fragments obtained from the search result; and

perform one or more operations based at least in part on identifying the subset of sentence fragments, the subset of sentence fragments being natural language expressions that are implied but excluded from the input text.

16. The non-transitory computer readable medium of claim 15 , wherein the knowledge base is an online knowledge base.

17. The non-transitory computer readable medium of claim 15 , wherein executing the instructions to generate the query further causes the computing device to:

identify repetitive sentence fragments from the plurality of sentence fragments of the input text; and

generate a generalized statement from the repetitive sentence fragments, wherein the query is generated from the generalized statement.

18. The non-transitory computer readable medium of claim 15 , wherein executing the instructions further causes the computing device to select the search result of the set of search results based at least in part on identifying that a relevance value between the search result and the query exceeds a predefined threshold.

19. The non-transitory computer readable medium of claim 18 , wherein executing the instructions to obtain the set of candidate sentence fragments for the missing one or more sentence fragments further causes the computing device to generate a respective discourse tree from the search result.

20. The non-transitory computer readable medium of claim 19 , wherein executing the instructions to identify the subset of sentence fragments further causes the computing device to identify a second predefined set of rhetorical relations and corresponding fragments of the respective discourse tree generated from the search result.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 4, 2022
From: GALITSKY, BORIS
To: ORACLE INTERNATIONAL CORPORATION
Reel/Frame 058545/0840 →
Continuity (2)
Provisional Application 63144704 · Feb 2, 2021
Related Publication 20220245360A1 · Aug 4, 2022