IP Library › Granted Patent US 11,741,312
Granted Patent B2
US 11,741,312 · App. 17/008,563 · Granted Aug 29, 2023

Systems and methods for unsupervised paraphrase mining

Inventors: Behzad Golshan (Mountain View, CA); Chen Chen (Sunnyvale, CA); Wang-Chiew Tan (San Jose, CA); Danni Ma (Philadelphia, PA)
Assignee: RECRUIT CO., LTD.
G06F40/35G06F18/2323G06F40/268G06F40/284
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,741,312
App. No.
17/008,563
Granted
Aug 29, 2023
Kind
B2
Abstract

Disclosed embodiments relate to aligning pairs of sentences. Techniques can include receiving a plurality of sentences; generating a graph for each of at least two sentences of the plurality of sentences, wherein generating a graph for each sentence of the at least two sentences comprises: identifying one or more tokens for the sentence; and connecting via edges the one or more tokens; generating a combined graph for the at least two sentences wherein generating a combined graph comprises: aligning the identified tokens of the at least two sentences of the plurality of sentences; identifying matching and non-matching tokens between the at least two sentences based on the alignment; and merging matching tokens into a combined graph node.

Claims (75)

1. A non-transitory computer readable storage medium storing instructions that are executable by a paraphrase mining system that includes one or more processors to cause the paraphrase mining system to perform a method for aligning pairs of sentences, the method comprising:

receiving a plurality of sentences;

determining a group of compatible sentences from the received plurality of sentences based on a first threshold value, wherein a value indicating compatibility between sentences in the group of compatible sentences is above the first threshold value;

generating a graph for each of at least two sentences of the group of compatible sentences, wherein generating a graph for each sentence of the at least two sentences comprises:

identifying one or more tokens for the sentence; and

connecting via edges the one or more tokens;

generating a combined graph for the at least two sentences wherein generating a combined graph comprises:

aligning the identified tokens of the at least two sentences of the group of compatible sentences based on word embeddings of the identified tokens generated using a neural network;

identifying matching and non-matching tokens between the at least two sentences based on the alignment, wherein alignment is above a second threshold value;

merging matching tokens into a combined graph node; and

determining interchangeable paraphrases based on non-matching tokens between the merged tokens of the combined graph.

2. The non-transitory computer readable storage medium of claim 1 , wherein generating a combined graph for the at least two sentences further comprises:

determining compatibility among the at least two sentences; and

removing sentences from the at least two sentences based on the compatibility determination.

3. The non-transitory computer readable storage medium of claim 2 , wherein determining compatibility among the at least two sentences comprises:

determining injectivity among the at least two sentences;

determining monotonicity among the at least two sentences; and

determining transitivity among the at least two sentences.

4. The non-transitory computer readable storage medium of claim 3 , wherein determining monotonicity among the at least two sentences further comprises:

determining a consistency among the at least two sentences; and

removing sentences from the at least two sentences based on the consistency determination.

5. The non-transitory computer readable storage medium of claim 1 , wherein determining a group of compatible sentences from the received plurality of sentences based on a first threshold value further comprises:

determining an intent of each sentence in the received plurality sentences; and

grouping each sentence in the received plurality of sentences into clusters wherein sentences in a cluster share the intent.

6. The non-transitory computer readable storage medium of claim 1 , wherein generating a combined graph for the at least two sentences further comprises:

determining an intent of tokens in the at least two sentences; and

identifying a set of non-matching tokens between the at least two sentences based on the alignment wherein the non-matching tokens in the set share the intent.

7. The non-transitory computer readable storage medium of claim 6 , wherein the non-matching tokens are paraphrases with one or more words that can be used interchangeably.

8. The non-transitory computer readable storage medium of claim 7 , wherein non-aligned tokens are phatic expressions.

9. The non-transitory computer readable storage medium of claim 7 , wherein the phatic expressions in each sentence of the two or more sentences are removed before generating a combined graph.

10. The non-transitory computer readable storage medium of claim 9 , wherein the phatic expressions are added to the combined graph notation of the two more sentences after aligning the identified tokens.

11. The non-transitory computer readable storage medium of claim 1 , further comprises:

generating a new sentence from the matching tokens and a subset of the non-matching tokens.

12. The non-transitory computer readable medium of claim 11 , wherein the new sentence is generated as a response to a question.

13. A method performed by a paraphrase mining system for aligning pairs of sentences, the method comprising:

receiving a plurality of sentences;

determining a group of compatible sentences from the received plurality of sentences based on a first threshold value, wherein a value indicating compatibility between sentences in the group of compatible sentences is above the first threshold value;

generating a graph for each of at least two sentences of the group of compatible of sentences, wherein generating a graph for each sentence of the at least two sentences comprises:

identifying one or more tokens for the sentence; and

connecting via edges the one or more tokens;

generating a combined graph for the at least two sentences wherein generating a combined graph comprises:

aligning the identified tokens of the at least two sentences of the group of compatible sentences based on word embeddings of the identified tokens generated using a neural network;

identifying matching and non-matching tokens between the at least two sentences based on the alignment, wherein alignment is above a second threshold value;

merging matching tokens into a combined graph node; and

determining interchangeable paraphrases based on non-matching tokens between the merged tokens of the combined graph.

14. The method of claim 13 , wherein generating a combined graph for the at least two sentences further comprises:

determining compatibility among the at least two sentences; and

removing sentences from the at least two sentences based on the compatibility determination.

15. The method of claim 14 , wherein determining compatibility among the at least two sentences comprises:

determining injectivity among the at least two sentences;

determining monotonicity among the at least two sentences; and

determining transitivity among the at least two sentences.

16. The method of claim 13 , wherein generating a combined graph for the at least two sentences further comprises:

determining a consistency among the at least two sentences; and

removing sentences from the at least two sentences based on the consistency determination.

17. A paraphrase mining system comprising:

one or more memory devices storing processor executable instructions; and

one or more processors configured to execute the instructions to cause the paraphrase mining system to perform:

receiving a plurality of sentences;

determining a group of compatible sentences from the received plurality of sentences based on a first threshold value, wherein a value indicating compatibility between sentences in the group of compatible sentences is above the first threshold value;

generating a graph for each of at least two sentences of the group of compatible of sentences, wherein generating a graph for each sentence of the at least two sentences comprises:

identifying one or more tokens for the sentence; and

connecting via edges the one or more tokens;

generating a combined graph for the at least two sentences wherein generating a combined graph comprises:

aligning the identified tokens of the at least two sentences of the group of compatible sentences based on word embeddings of the identified tokens generated using a neural network;

identifying matching and non-matching tokens between the at least two sentences based on the alignment, wherein alignment is above a second threshold value;

merging matching tokens into a combined graph node; and

determining interchangeable paraphrases based on non-matching tokens between the merged tokens of the combined graph.

18. The system of claim 17 , wherein the one or more processors are configured to execute the instructions to cause the paraphrase mining system to further perform:

determining an intent of each sentence in the received plurality sentences; and

grouping each sentence in the received plurality of sentences into clusters wherein sentences in a cluster share the intent.

19. The system of claim 17 , wherein the one or more processors are configured to execute the instructions to cause the paraphrase mining system to further perform:

determining an intent of tokens in the at least two sentences.

20. The system of claim 17 , wherein the one or more processors are configured to execute the instructions to cause the paraphrase mining system to further perform:

generating a new sentence from the matching tokens and a subset of the non-matching tokens.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 4, 2020
From: GOLSHAN, BEHZAD; CHEN, CHEN; TAN, WANG-CHIEW; MA, DANNI
To: RECRUIT CO., LTD.
Reel/Frame 053698/0643 →
Continuity (1)
Related Publication 20220067298A1 · Mar 3, 2022