IP Library Granted Patent US 9,269,028
Granted Patent B2
US 9,269,028 · App. 14/324,755 · Granted Feb 23, 2016

System and method for determining string similarity

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,269,028
App. No.
14/324,755
Granted
Feb 23, 2016
Kind
B2
Abstract

Provided are string similarity assessment techniques. In one embodiment, the techniques include receiving a plurality of input strings comprising characters from a character set and generating hashtables for each respective input string using a hash function that assigns the characters as keys and character positions in the strings as values. The techniques may also include determine a character similarity index for at least two of the input strings relative to each other by comparing a similarity of the values for each key in the their respective hashtables; determining a total disordering index based representative of an alignment of the at least two input strings by determining differences between a plurality of index values for each individual key in their respective hashtables and determining the total disordering index based on the differences; and determining a string similarity metric based on at least one character similarity index and the total disordering index.

Claims (67)

1. A string similarity assessment method, comprising:

using a processor-based device;

generating a first hashtable using a hash function for a first input string comprising characters from a character set;

generating a second hashtable using the hash function for a second input string comprising characters from the character set, wherein, for the first hashtable and the second hashtable, the hash function assigns the characters as keys and character positions in the strings as values;

determining a character similarity index for each of the first input sting and the second input string based on comparing a similarity of the values for each key in the first hashtable and the second hashtable;

determining a string similarity index for transforming the first string into the second string by selecting the character similarity index for the first input sting or the second input string;

determining a total disordering index representative of an alignment score of each key in the first hashtable with its respective key in the second hashtable; and

determining a string similarity metric based on the string similarity index and the total disordering index.

2. The method of claim 1 , comprising storing the first hashtable and the second hashtable.

3. The method of claim 1 , wherein the character set comprises the characters A,C,T, and G.

4. The method of claim 3 , wherein the character set comprises only the characters A,C,T, and G.

5. The method of claim 1 , comprising displaying, using a display, one or more of the string similarity metric, the string similarity index, or the total disordering index.

6. The method of claim 1 , wherein determining the string similarity metric based on the string similarity index and the total disordering index comprises using a length value of the first input string or the second input string.

7. The method of claim 1 , wherein determining the character similarity index for the first input sting comprises determining if each key in the first input string is present in the second input string, and, for each individual key that is not present in the second input string, adding a predetermined score to the character similarity index for the first input string.

8. The method of claim 7 , wherein determining the character similarity index for the second input sting comprises determining if each key in the second input string is present in the first input string, and, for each individual key that is not present in the first input string, adding a predetermined score to the character similarity index for the second input string.

9. The method of claim 8 , comprising selecting a greater of the character similarity index for the first input sting and the character similarity index for the second input sting as the string similarity index.

10. The method of claim 1 , wherein determining the total disordering index comprises comparing a length of a list of index values of each key for the first hashtable and the second hashtable and setting a shortest list of each key as a base list and a longer list of each key as the comparison list.

11. The method of claim 10 , comprising, for each base list of each key, comparing the base list to the comparison list by determining a difference between a first index value in the comparison list that is greater than a first index value in the base list.

12. The method of claim 10 , comprising, determining two difference values for each index value in the base list, wherein each index value in the base list is subtracted from only two index values in the comparison list.

13. The method of claim 10 , comprising, selecting a smaller of the two difference values for each index value in the base list and determining the total disordering index by summing the selected difference values.

14. A string similarity assessment system, comprising:

a memory storing instructions that, when executed, are configured to:

receive a plurality of input strings comprising characters from a character set;

generate hashtables for each respective input string using a hash function that assigns the characters as keys and character positions in the strings as values;

determine a character similarity index for at least two of the input strings relative to each other by comparing a similarity of the values for each key in the their respective hashtables;

determining a total disordering index based representative of an alignment of the at least two input strings by determining differences between a plurality of index values for each individual key in their respective hashtables and determining the total disordering index based on the differences; and

determining a string similarity metric based on at least one character similarity index and the total disordering index; and

one or more processors configured to execute the instructions.

15. The system of claim 14 , wherein the instructions are configured to store the hashtables.

16. The system of claim 14 , wherein the instructions are configured to provide an indication of the string similarity metric.

17. The system of claim 14 , wherein the instructions are configured to provide a security assessment based on the string similarity metric.

18. The system of claim 14 , comprising displaying, using a display, the string similarity metric.

19. The system of claim 14 , wherein the string similarity metric is determined by the following equation:

1

-

(

m

2

2

+

1

-

o

m

2

2

+

1

*

m

-

s

AB

m

)

where m is a length of the longer string of the at least two strings, where o is the total disordering index, and wherein s AB is a greater of the character similarity indices for at least two of the input strings.

20. A string similarity assessment method, comprising:

using a processor-based device:

receiving a plurality of input strings comprising characters from a character set;

generating hashtables for each respective input string using a hash function that assigns the characters as keys and character positions in the strings as values;

determining a character similarity index for at least two of the input strings relative to each other by comparing a similarity of the values for each key in their respective hashtables;

determining a total disordering index based representative of an alignment of the at least two input strings by determining differences between a plurality of index values for each individual key in their respective hashtables and determining the total disordering index based on the differences, and

determining a string similarity metric based on at least one character similarity index and the total disordering index.

Assignments (5)
QUITCLAIM ASSIGNMENT Recorded Sep 18, 2025
From: EDISON INNOVATIONS LLC
To: BLUE RIDGE INNOVATIONS, LLC
Reel/Frame 072938/0793 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 26, 2025
From: GENERAL ELECTRIC COMPANY
To: GE INTELLECTUAL PROPERTY LICENSING, LLC
Reel/Frame 070636/0815 →
CHANGE OF NAME Recorded Mar 26, 2025
From: GE INTELLECTUAL PROPERTY LICENSING, LLC
To: DOLBY INTELLECTUAL PROPERTY LICENSING, LLC
Reel/Frame 070643/0907 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 20, 2025
From: DOLBY INTELLECTUAL PROPERTY LICENSING, LLC
To: EDISON INNOVATIONS, LLC
Reel/Frame 070293/0273 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 7, 2014
From: KURZER, JAKE MATTHEW; CSORBA, BRETT
To: GENERAL ELECTRIC COMPANY
Reel/Frame 033252/0933 →