IP Library › Granted Patent US 12,182,510
Granted Patent B2
US 12,182,510 · App. 17/653,912 · Granted Dec 31, 2024

Unidirectional text comparison

Inventors: Mehul Thukral (New York, NY); Raj Nagesh (Cary, NC); Saksham Gandhi (New York, NY)
Assignee: International Business Machines Corporation
G06F40/284G06F16/90344G06F40/205G06V30/19093
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,182,510
App. No.
17/653,912
Granted
Dec 31, 2024
Kind
B2
Abstract

A method, a structure, and a computer system for unidirectional text comparison. The exemplary embodiments may include determining a first similarity score between a first text string and a second text string, and computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string. The exemplary embodiments may further include determining a second similarity score based on the first similarity score and the error term.

Claims (55)

1. A computer-implemented method for unidirectional text comparison, the method comprising:

determining a first similarity score between a first text string and a second text string;

computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;

determining a second similarity score based on the first similarity score and the error term.

2. The computer-implemented method of claim 1 , wherein determining the first similarity score further comprises:

training a word to vector model using a domain of interest;

tokenizing the first text string and the second text string; and

calculating a cosine difference between the first text string and the second text string based on the word to vector model.

3. The computer-implemented method of claim 2 , wherein computing the error term further comprises:

tokenizing the first text string and the second text string;

determining an overlap and a symmetric difference of tokens within the first text string and the second text string;

computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and

summing the overlap error term and the non-overlap error term.

4. The computer-implemented method of claim 1 , wherein determining the second similarity score further incorporates a constant.

5. The computer-implemented method of claim 1 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.

6. The computer-implemented method of claim 1 , further comprising:

pre-processing the first text string and the second text string.

7. The computer-implemented method of claim 1 , wherein the first text string is parsed from a first document and the second text string is parsed from a second document.

8. A computer program product for unidirectional text comparison, the computer program product comprising:

one or more non-transitory computer-readable storage media and program instructions stored on the one or more non-transitory computer-readable storage media capable of performing a method, the method comprising:

determining a first similarity score between a first text string and a second text string;

computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;

determining a second similarity score based on the first similarity score and the error term.

9. The computer program product of claim 8 , wherein determining the first similarity score further comprises:

training a word to vector model using a domain of interest;

tokenizing the first text string and the second text string; and

calculating a cosine difference between the first text string and the second text string based on the word to vector model.

10. The computer program product of claim 9 , wherein computing the error term further comprises:

tokenizing the first text string and the second text string;

determining an overlap and a symmetric difference of tokens within the first text string and the second text string;

computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and

summing the overlap error term and the non-overlap error term.

11. The computer program product of claim 8 , wherein determining the second similarity score further incorporates a constant.

12. The computer program product of claim 8 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.

13. The computer program product of claim 8 , further comprising:

pre-processing the first text string and the second text string.

14. The computer program product of claim 8 , wherein the first text string is parsed from a first document and the second text string is parsed from a second document.

15. A computer system for unidirectional text comparison, the system comprising:

one or more computer processors, one or more computer-readable storage media, and program instructions stored on the one or more of the computer-readable storage media for execution by at least one of the one or more processors capable of performing a method, the method comprising:

determining a first similarity score between a first text string and a second text string;

computing an error term between the first text string and the second text string, wherein the error term incorporates a directionality of the first text string and the second text string;

determining a second similarity score based on the first similarity score and the error term.

16. The computer system of claim 15 , wherein determining the first similarity score further comprises:

training a word to vector model using a domain of interest;

tokenizing the first text string and the second text string; and

calculating a cosine difference between the first text string and the second text string based on the word to vector model.

17. The computer system of claim 16 , wherein computing the error term further comprises:

tokenizing the first text string and the second text string;

determining an overlap and a symmetric difference of tokens within the first text string and the second text string;

computing an overlap error term for the overlap and a non-overlap error term for the symmetric difference; and

summing the overlap error term and the non-overlap error term.

18. The computer system of claim 15 , wherein determining the second similarity score further incorporates a constant.

19. The computer system of claim 15 , wherein the second similarity score is not penalized based on the first text string having additional text over the second text string and vice versa.

20. The computer system of claim 15 , further comprising:

pre-processing the first text string and the second text string.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 8, 2022
From: THUKRAL, MEHUL; NAGESH, RAJ; GANDHI, SAKSHAM
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059194/0865 →
Continuity (1)
Related Publication 20230289526A1 · Sep 14, 2023