IP Library › Granted Patent US 8,799,291
Granted Patent B2
US 8,799,291 · App. 13/601,925 · Granted Aug 5, 2014

Forensic index method and apparatus by distributed processing

Inventors: Joo Young Lee (Daejeon, KR); Youn Hee Gil (Daejeon, KR); Do Won Hong (Daejeon, KR); Keon Woo Kim (Daejeon, KR); Young Soo Kim (Daejeon, KR); Sung Kyong Un (Daejeon, KR); Sang Su Lee (Daejeon, KR); Su Hyung Jo (Daejeon, KR); Woo Yong Choi (Daejeon, KR); Hyun Sook Cho (Daejeon, KR)
Assignee: Electronics and Telecommunications Research Institute
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,799,291
App. No.
13/601,925
Granted
Aug 5, 2014
Kind
B2
Abstract

Provided is a forensic index method by distributed processing, including: generating data to be divided by dividing data to be indexed according to predetermined division setting for distributed processing; allocating the generated data to be divided to a plurality of data processing units according to the predetermined division setting, extracting an index by filtering the allocated data to be divided in the plurality of data processing units, and generating divided index data including the extracted index; and generating an index database by merging the generated divided index data.

Claims (53)

1. A forensic index method by distributed processing, comprising:

generating data to be divided by dividing data to be indexed according to predetermined division setting for distributed processing;

allocating the generated data to be divided to a plurality of data processing units according to the predetermined division setting, extracting an index by filtering the allocated data to be divided in the plurality of data processing units, and generating divided index data including the extracted index using the plurality of data processing units; and

generating an index database by merging the generated divided index data, wherein generating the index database comprises:

when a new index is generated in the index database from the merged generated divided index data, setting an index database identifier as a key in the index database and adding the merged generated divided index data to a list of data corresponding to the key; and

when generated index data corresponds to a file being added to an existing index, generating a temporary index data having a temporary index using the generated index data and adding the temporary index to a list of data corresponding to the key of the existing index.

2. The forensic index method of claim 1 , wherein in the generating of the data to be divided, the data to be divided is generated by dividing the data to be indexed by a number of files included in the data to be indexed and a number of the data processing units based on the predetermined division setting for distributed processing or the data to be divided is generated according to the number and sizes of the files included in the data to be indexed based on the predetermined division setting.

3. The forensic index method of claim 2 , wherein generating of the divided index data includes:

allocating the generated data to be divided to the plurality of data processing units and loading the data to be divided which are allocated to at least one of the plurality of data processing units into a distribution storage unit of the at least one data processing unit;

filtering the loaded data to be divided;

extracting the index with respect to the filtered data to be divided; and

generating the divided index data including the extracted index.

4. The forensic index method of claim 3 , wherein the filtering of the data to be divided includes:

converting the data to be divided into text data by using a file filter including a virtual file system (VFS); and

extracting filtered data from the converted text data.

5. The forensic index method of claim 3 , wherein in the extracting of the index, an index target is extracted by dividing the filtered data to be divided into N syllable units and the index is extracted by comparing the extracted index target with a predetermined pattern.

6. The forensic index method of claim 3 , wherein the generating of the divided index data includes:

matching data of which the index is extracted with the extracted index; and

shuffling, sorting, and merging the matched data, which are generated as the divided index data.

7. The forensic index method of claim 1 , further comprising:

performing retrieval by using the index database by receiving a retrieval word from a user.

8. A forensic index apparatus by distributed processing, comprising:

a computer memory;

a division target data managing unit generating data to be divided by dividing data to be indexed according to predetermined division setting for distributed processing;

a divided index data generating unit, including a plurality of data processing units, allocating the generated data to be divided to each of the plurality of data processing units according to the predetermined division setting, extracting an index by filtering the corresponding allocated data in each of the plurality of data processing units using the respective data processing unit, and generating divided index data including the extracted index using the respective data processing unit; and

an index database managing unit generating an index database by merging the generated divided index data generated by the plurality of data processing units, wherein generating the index database comprises:

when a new index is generated in the index database from the merged generated divided index data, setting an index database identifier as a key in the index database and adding the merged generated divided index data to a list of data corresponding to the key; and

when generated index data corresponds to a file being added to an existing index, generating a temporary index data having a temporary index using generated index data and adding the temporary index to a list of data corresponding to the key of the existing index.

9. The forensic index apparatus of claim 8 , wherein the division target data managing unit generates the data to be divided by dividing the data to be indexed by a number of files included in the data to be indexed and a number of the data processing units based on the predetermined division setting for distributed processing or generates the data to be divided according to the number and sizes of the files included in the data to be indexed based on the predetermined division setting.

10. The forensic index apparatus of claim 9 , wherein at least one of the plurality of data processing units includes:

a distributive storage portion being allocated with the generated data to be divided according to the predetermined division setting and loading the allocated data to be divided;

a filtering portion filtering the data to be divided, which are loaded to the distributive storage portion;

an index extracting portion extracting the index from the filtered data to be divided; and

a divided index data generating portion generating the divided index data including the extracted index.

11. The forensic index apparatus of claim 10 , wherein the filtering portion includes:

a text converting portion converting the data to be divided into text data by using a file filter including a virtual file system (VFS); and

a filtered data extracting portion extracting filtered data from the converted text data.

12. The forensic index apparatus of claim 10 , wherein the index extracting portion includes:

an N-gram analysis portion extracting an index target by dividing the filtered data to be divided into N syllable units; and

a pattern comparison portion extracting the index by comparing the extracted index target with a predetermined pattern.

13. The forensic index apparatus of claim 10 , wherein the divided index data generating portion includes:

an index matching portion matching data of which the index is extracted with the extracted index, and

generates the divided index data by shuffling, sorting, and merging the matched data, which are generated as the divided index data.

14. A forensic index system by distributed processing, comprising:

a terminal unit providing an input interface to a user;

a data storage unit storing data to be indexed; and

a forensic index apparatus dividing the data to be indexed according to predetermined division setting, generating divided index data by using a plurality of data processing units with respect to the divided data to be indexed, and generating an index database by merging the generated divided index data, wherein generating the index database comprises:

when a new index is generated in the index database from the merged generated divided index data, setting an index database identifier as a key in the index database and adding the merged generated divided index data to a list of data corresponding to the key; and

when generated index data corresponds to a file being added to an existing index, generating a temporary index data having a temporary index using the generated index data and adding the temporary index to a list of data corresponding to the key of the existing index.

15. The forensic index system of claim 14 , wherein the forensic index apparatus includes:

a division target data managing unit generating data to be divided by dividing data to be indexed by a number of files included in the data to be indexed and a number of the data processing units or generating the data to be divided according to the number and sizes of the files included in the data to be indexed;

a divided index data generating unit allocating the generated data to be divided to the plurality of data processing units, extracting an index by loading the data to be divided, which are allocated in the plurality of data processing units and filtering the loaded data, and generating the divided index data including the extracted index; and

an index database managing unit generating the index database by merging the generated divided index data.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 6, 2012
From: LEE, JOO YOUNG; GIL, YOUN HEE; HONG, DO WON; KIM, KEON WOO; KIM, YOUNG SOO; UN, SUNG KYONG; LEE, SANG SU; JO, SU HYUNG; CHOI, WOO YONG; CHO, HYUN SOOK
To: ELECTRONICS AND TELECOMMUNICATIONS RESEARCH INSTITUTE
Reel/Frame 028909/0756 →
Priority Claims (1)
KR 10-2011-0114168 · Nov 3, 2011 · national
Continuity (1)
Related Publication 20130117273A1 · May 9, 2013