IP Library Granted Patent US 7,689,638
Granted Patent B2
US 7,689,638 · App. 10/534,007 · Granted Mar 30, 2010

Method and device for determining and outputting the similarity between two data strings

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 7,689,638
App. No.
10/534,007
Granted
Mar 30, 2010
Kind
B2
Abstract

The present invention discloses a method and device for determining and outputting a similarity measure between two data strings each data string comprising data entities, comprising: receiving a first data string, receiving a second data string, which is characterized by determining consecutively following data entities in the first data string, determining the relative positions of the consecutively following data entities in the first data string, determining similar data entities with the same order in the second data string, determining the relative positions of the determined data entities in the second data string, determining a matching measure by determining how far the relative positions of data entities in the second data string match with the relative positions of consecutively following data entities in the first data string, and outputting a similarity measure which corresponds to the matching measure of at least one comparison result.

Claims (53)

1. A method comprising:

receiving a first data string in an electronic component,

receiving a second data string in said electronic component,

determining pairs of consecutively following data entities in said first data string in a processing unit,

determining the relative positions of said pairs of consecutively following data entities in said first data string in said processing unit,

allocating a position label to each of said data entities in the first data string in said processing unit,

numbering same data entities according to their relative position in accordance with the position label in said processing unit,

determining similar data entities with the same order in said second data string in said processing unit,

determining the relative positions of said determined data entities in said second data string in said processing unit,

determining a matching measure by determining how far the relative positions of data entities in said second data string match with the relative positions of consecutively following data entities in said first data string in said processing unit, and

determining a similarity measure which corresponds to the matching measure of at least one comparison result in said processing unit,

repeating said determination of said similarity measure with a number of received second data strings in said processing unit, and

outputting by an interface said determined similarity measures for said data strings according to the amount of similarity to said first data string,

wherein said first data string of entities and said second data string of entities are data strings relating to one of associative text string, genome analysis, speech recognition, and musical melody.

2. The method according to claim 1 , further comprising:

determining at least one error limit for at least one of said entities, and

considering said at least one error limit during said determination of said matching measure.

3. The method according to claim 1 , further comprising:

determining a first distance between said two data entities of consecutively following data entities in said first data string,

determining a second distance of said two data entities determined in said second data string,

determining a difference between said first and second distances, and

considering said difference during said determination of said matching measure.

4. The method according to claim 1 , further comprising:

storing said second string together with said similarity measure.

5. The method according to claim 1 , further comprising:

determining a threshold for said similarity measure, and

outputting said second string, if said determined similarity measure at least equals said threshold.

6. The method according to claim 1 , further comprising:

analyzing the first string for entities not present in the first string, and

suppressing in the second string all said entities not present in said first string.

7. The method according to claim 6 , further comprising:

determining the number of entities that are present in the second string, but are not present in the first string, as a second similarity measure.

8. The method according to claim 7 , further comprising:

determining a section within said second string comprising at least the same number of entities that are simultaneously present in both strings.

9. A computer readable medium stored with code, which when executed by a computer, performs the method of claim 1 .

10. The method according to claim 1 , wherein the first data string and the second data string are pieces of text.

11. The method according to claim 1 , wherein the first data string and the second data string are each a sequence of musical notes.

12. The method according to claim 1 , wherein the first data string and the second data string are sequences of deoxyribonucleic acid.

13. The method according to claim 1 , wherein the first data string and the second data string are each phonetic sounds.

14. An electronic device comprising:

a component configured to receive a first data string of entities and a second data string of entities, said first data string of entities and said second data string of entities being data strings relating to one of associative text string, genome analysis, speech recognition, and musical melody,

a processing unit configured to

determine pairs of consecutively following data entities in said first data string,

determine the relative positions of said pairs of consecutively following data entities in said first data string,

allocate a position label to each of said data entities in the first data string,

number same data entities according to their relative position in accordance with the position label;

determine similar data entities with the same order in said second data string,

determine the relative positions of said determined data entities in said second data string, and

determine a matching measure by determining how far the relative positions of data entities in said second data string match with the relative positions of consecutively following data entities in said first data string, and

repeat said determination of said similarity measure with a number of received second data strings, and

an interface configured to output a similarity measure for said second data string and said number of second data strings according to the amount of similarity to said first data string.

15. An electronic device according to claim 14 , further comprising a storage configured to store received strings and said determined similarity measures.

16. The electronic device according to claim 14 , wherein the electronic device is a mobile terminal device.

Assignments (6)
SECURITY INTEREST Recorded Jun 1, 2021
From: WSOU INVESTMENTS, LLC
To: OT WSOU TERRIER HOLDINGS, LLC
Reel/Frame 056990/0081 →
RELEASE OF SECURITY INTEREST Recorded May 21, 2019
From: OCO OPPORTUNITIES MASTER FUND, L.P. (F/K/A OMEGA CREDIT OPPORTUNITIES MASTER FUND LP
To: WSOU INVESTMENTS, LLC
Reel/Frame 049246/0405 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2017
From: NOKIA TECHNOLOGIES OY
To: WSOU INVESTMENTS, LLC
Reel/Frame 043953/0822 →
SECURITY INTEREST Recorded Sep 21, 2017
From: WSOU INVESTMENTS, LLC
To: OMEGA CREDIT OPPORTUNITIES MASTER FUND, LP
Reel/Frame 043966/0574 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 24, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 035235/0685 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 5, 2005
From: THEIMER, WOLFGANG; ROSS, ANDREE
To: NOKIA CORPORATION
Reel/Frame 017396/0658 →