IP Library Granted Patent US 9,858,280
Granted Patent B2
US 9,858,280 · App. 14/524,373 · Granted Jan 2, 2018

System, apparatus, program and method for data aggregation

Inventors: Vivian Lee (Bracknell Bekshire, GB); Bo Hu (Winchester, GB)
Assignee: FUJITSU LIMITED
G06F17/30091G06F3/0631G06F3/0683G06F17/30412G06F17/30584G06F17/30663G06F3/0607
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 9,858,280
App. No.
14/524,373
Granted
Jan 2, 2018
Kind
B2
Abstract

Embodiments include a method, apparatus, program, and system for distributing data items among a plurality of data storage units, the data items being an aggregation of data from a plurality of data sources. The method comprises generating a semantic description of each of the plurality of data sources; calculating, for each pair of data sources from among the plurality of data sources, a degree of similarity between the semantic descriptions of the pair of data sources; and allocating data items to data storage units in dependence upon the calculated degree of similarity between the data source of a data item being allocated and the or each data source of data items already allocated to the data storage units.

Claims (15)

1. A data storage system comprising: a plurality of data storage units; and

a computer comprising a processor, the processor being configured to perform a method for distributing data items among the plurality of data storage units, the data items being an aggregation of data from a plurality of data sources, the method comprising:

generating a semantic description of each of the plurality of data sources;

calculating, for each pair of data sources from among the plurality of data sources, a degree of similarity between semantic descriptions of the pair of data sources; and

allocating data items to data storage units from of the plurality of data storage units in dependence upon the degree of similarity between the data source of a data item being allocated and the data source of data items already allocated to the data storage units; and

maintaining, for each data source, a record of the identity of each data storage unit storing one or more data items from the data source and an indication of the proportion of the data items from the data source stored on each identified data storage unit;

wherein allocating data items to data storage units in dependence upon the calculated degree of similarity includes allocating a group of data items from the same data source to the data storage unit indicated by the record as storing the largest proportion of the data items from the data source having the highest degree of similarity to the data source of the group of data items to be allocated, and wherein the data storage unit has sufficient storage space for the group of data items.

2. The data storage system according to claim 1 , wherein generating a semantic description of each of the data sources includes extracting most significant terms as a list of weighted terms from a data source which list of weighted terms is a generated semantic description.

3. The storage system according to claim 2 , wherein the most significant terms among terms in the data source are identified by and a weight attributed to each of extracted most significant terms is calculated by a term-frequency method.

4. The storage system according to claim 3 , wherein the degree of similarity is a value obtained by calculating a cosine similarity between generated semantic descriptions of the pair of data sources.

5. The data storage system according to claim 1 , wherein the data items are stored in a unified data format, wherein the unified data format is an Resource Description Framework (RDF) triple format.

6. The data storage system according to claim 5 , further comprising:

reading data from a data source and performing processing to prepare read data for storage as data items having the unified data format.

7. The data storage system according to claim 1 , further comprising, if a data storage unit selected to receive data items being allocated already stores data items from two data sources, and the degree of similarity between the semantic descriptions of the two data sources is less than a higher of the degrees of similarity between the semantic descriptions of each of the data sources and the data source of the data items to be allocated, then the data items from the one of the two data sources having a lowest degree of similarity to the data source of the data items to be allocated are removed from the data storage unit and allocated elsewhere.

8. The data storage system according to claim 1 , wherein the data items each include an identifier identifying a data source of the respective data item from among the plurality of data sources.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2014
From: LEE, VIVIAN; HU, BO
To: FUJITSU LIMITED
Reel/Frame 034050/0293 →
Priority Claims (1)
EP 13193377 · Nov 18, 2013 · regional
Continuity (1)
Related Publication 20150142829A1 · May 21, 2015