Unidirectional text comparison
View Patent ↗A method, a structure, and a computer system for unidirectional text comparison. The exemplary embodiments may include determining a first similarity score between a first text string and a second text string, and computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string. The exemplary embodiments may further include determining a second similarity score based on the first similarity score and the error term.
1. A computer-implemented method for unidirectional text comparison, the method comprising:
determining a first similarity score between a first text string and a second text string;
computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;
determining a second similarity score based on the first similarity score and the error term.
2. The computer-implemented method of claim 1 , wherein determining the first similarity score further comprises:
training a word to vector model using a domain of interest;
tokenizing the first text string and the second text string; and
calculating a cosine difference between the first text string and the second text string based on the word to vector model.
3. The computer-implemented method of claim 2 , wherein computing the error term further comprises:
tokenizing the first text string and the second text string;
determining an overlap and a symmetric difference of tokens within the first text string and the second text string;
computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and
summing the overlap error term and the non-overlap error term.
4. The computer-implemented method of claim 1 , wherein determining the second similarity score further incorporates a constant.
5. The computer-implemented method of claim 1 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.
6. The computer-implemented method of claim 1 , further comprising:
pre-processing the first text string and the second text string.
7. The computer-implemented method of claim 1 , wherein the first text string is parsed from a first document and the second text string is parsed from a second document.
8. A computer program product for unidirectional text comparison, the computer program product comprising:
one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media capable of performing a method, the method comprising:
determining a first similarity score between a first text string and a second text string;
computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;
determining a second similarity score based on the first similarity score and the error term.
9. The computer program product of claim 8 , wherein determining the first similarity score further comprises:
training a word to vector model using a domain of interest;
tokenizing the first text string and the second text string; and
calculating a cosine difference between the first text string and the second text string based on the word to vector model.
10. The computer program product of claim 9 , wherein computing the error term further comprises:
tokenizing the first text string and the second text string;
determining an overlap and a symmetric difference of tokens within the first text string and the second text string;
computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and
summing the overlap error term and the non-overlap error term.
11. The computer program product of claim 8 , wherein determining the second similarity score further incorporates a constant.
12. The computer program product of claim 8 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.
13. The computer program product of claim 8 , further comprising:
pre-processing the first text string and the second text string.
14. The computer program product of claim 8 , wherein the first text string is parsed from a first document and the second text string is parsed from a second document.
15. A computer system for unidirectional text comparison, the system comprising:
one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising:
determining a first similarity score between a first text string and a second text string;
computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;
determining a second similarity score based on the first similarity score and the error term.
16. The computer system of claim 15 , wherein determining the first similarity score further comprises:
training a word to vector model using a domain of interest;
tokenizing the first text string and the second text string; and
calculating a cosine difference between the first text string and the second text string based on the word to vector model.
17. The computer system of claim 16 , wherein computing the error term further comprises:
tokenizing the first text string and the second text string;
determining an overlap and a symmetric difference of tokens within the first text string and the second text string;
computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and
summing the overlap error term and the non-overlap error term.
18. The computer system of claim 15 , wherein determining the second similarity score further incorporates a constant.
19. The computer system of claim 15 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.
20. The computer system of claim 15 , further comprising:
pre-processing the first text string and the second text string.