IP Library › Granted Patent US 8,924,340
Granted Patent B2
US 8,924,340 · App. 12/593,203 · Granted Dec 30, 2014

Method of comparing data sequences

Inventors: Zhan Cui (Colchester, GB); Qiao Tang (Cardiff, GB)
Assignee: BRITISH TELECOMMUNICATIONS public limited company
G06F17/30985
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,924,340
App. No.
12/593,203
Granted
Dec 30, 2014
Kind
B2
Abstract

A method according to the present invention enables the similarity between sequences of symbols to be determined using rules generated from a dictionary-based compression scheme according to the content of the columns from databases. Pairs of symbols can replaced by rules that do not comprise a repeated combination of two symbols and where each rule occurs more than once in the sequence of symbols. The similarity of each set of rules can then be expressed numerically.

Claims (25)

1. A method of determining one or more patterns in a sequence of symbols, wherein the method comprises the steps of:

a) analyzing the sequence of symbols using a processing system comprising a computer processor, such that a rule is generated to replace non-unique diagrams occurring in the sequence, the rule comprising one or more symbols and/or a further rule,

wherein each rule is checked to comply with a requirement that there are multiple instances of the rule occurring in the sequence of symbols such that a rule that occurs only once in the sequence is expanded by reverting the rule to a symbol pattern originally replaced by the rule, and

wherein step a) is repeated until no further non-unique diagrams may be replaced by a rule.

2. A method according to claim 1 comprising the further step of:

b) adding an additional symbol to the sequence of symbols; and then repeating step a).

3. A method of determining the similarity between a first data series and a second data series, wherein the first data series and the second data series have been generated from a respective first sequence of symbols and second sequence of symbols in accordance with claim 1 , and a similarity value is generated which indicates the similarity between the set of rules comprising the first data series and the set of rules comprising the second data series.

4. A method according to claim 3 , wherein the similarity value has a value of 0%, indicating that there are no rules present in the first data series that are present in the second data series.

5. A method according to claim 3 , wherein the similarity value has a value of 100%, indicating that

i) the first data series comprises the same rules as those present in the second data series; and

ii) each rule present in the first data series is present the same number of times in the first data series as in the second data series.

6. A non-transitory computer program product, comprising computer executable code for performing a method according to claim 1 .

7. Apparatus configured to perform a method of determining one or more patterns in a sequence of symbols, the apparatus comprising:

a processing system, comprising a computer processor, the processing system being configured to:

a) analyze the sequence of symbols such that a rule is generated to replace non-unique digrams occurring in the sequence, the rule comprising one or more symbols and/or a further rule,

wherein each rule is checked to comply with a requirement that there are multiple instances of the rule occurring in the sequence of symbols, such that a rule that occurs only once in the sequence is expanded by reverting the rule to a symbol pattern originally replaced by the rule, and

wherein the analysis of a) is repeated until no further non-unique digrams may be replaced by a rule.

8. The apparatus of claim 7 , wherein the processing system is further configured to: b) add an additional symbol to the sequence of symbols; and then repeat the analysis of a).

9. The apparatus of claim 7 , wherein the processing system is further configured to:

determine the similarity between a first data series and a second data series, wherein the first data series and the second data series have been generated from a respective first sequence of symbols and second sequence of symbols analyzed in a); and

generate a similarity value which indicates the similarity between the set of rules comprising the first data series and the set of rules comprising the second data series.

10. The apparatus of claim 9 , wherein the similarity value has a value of 0%, indicating that there are no rules present in the first data series that are present in the second data series.

11. The apparatus of claim 9 , wherein the similarity value has a value of 100%, indicating that

i) the first data series comprises the same rules as those present in the second data series; and

ii) each rule present in the first data series is present the same number of times in the first data series as in the second data series.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 25, 2009
From: CUI, ZHAN; TANG, QIAO
To: BRITISH TELECOMMUNICATIONS PUBLIC LIMITED COMPANY
Reel/Frame 023287/0034 →
Priority Claims (1)
EP 07251307 · Mar 27, 2007 · regional
Continuity (1)
Related Publication 20100121813A1 · May 13, 2010