IP Library › Granted Patent US 11,514,010
Granted Patent B2
US 11,514,010 · App. 17/211,973 · Granted Nov 29, 2022

Deduplication-adapted CaseDB for edge computing

Inventors: Deokhwan Kim (Incheon, KR); Khikmatullo Tulkinbekov (Incheon, KR); Joohwan Kim (Incheon, KR)
Assignee: INHA-INDUSTRY PARTNERSHIP INSTITUTE
G06F16/215G06F16/1744
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,514,010
App. No.
17/211,973
Granted
Nov 29, 2022
Kind
B2
Abstract

Disclosed is a data deduplication method for an edge computer. The method is performed in a key-value store, and may include receiving a compaction request occurred from the key-value store to a metadata layer, checking whether deduplication for removing duplicated data is required when compaction of a metadata file is performed in response to the received compaction request, and removing the duplicated data by checking whether the deduplication is required.

Claims (28)

1. A data deduplication method for an edge computer, which is performed in a key-value store persistent storage, the data deduplication method comprising:

receiving a compaction request occurred from the key-value store persistent storage to a metadata layer having upper levels and lower levels, wherein compaction has previously occurred on the upper levels, wherein the compaction includes a deduplication process for removing duplicating data;

checking whether deduplication for removing duplicated data is required by using a Bloom filter of metadata to check whether a key-value item is present via a processor when compaction of a metadata file is performed in response to the received compaction request; and

removing the duplicated data in response to the deduplication being required, via the processor, wherein removing the duplicated data comprises:

determining a number of duplicated keys of the metadata at only the lower levels when the data duplication check is performed using the Bloom filter,

calculating a duplication ratio, wherein the duplication ratio is determined by a total number of keys divided by the number of duplicated keys, wherein the total number of keys is determined from the metadata layer,

comparing a value of the calculated duplication ratio with a threshold to determine whether a next deduplication process for removing duplicated data is required,

determining that the next deduplication process for removing duplicated data is required in response that the value of the duplication ratio is equal or less than the threshold, and

extending a compaction of the metadata to the lower level, and automatically removing the duplicated data from the persistent storage in the next deduplication process.

2. The data deduplication method of claim 1 , wherein checking whether the deduplication is required comprises performing a first process of performing compaction of the metadata in the key-value store and a second process of relocating string-sorted tables (SSTables) based on the metadata updated through the first process.

3. The data deduplication method of claim 2 , wherein:

checking whether the deduplication is required comprises continuing the compaction of the metadata for levels reaching a size threshold in the first process and then performing a data deduplication check using the Bloom filter of the metadata when the first process and the second process are completed, and

the Bloom filter checks whether the key-value item is present without reading the metadata.

4. The data deduplication method of claim 1 , wherein removing the duplicated data comprises:

returning a return value from the Bloom filter when a data deduplication check is performed using the Bloom filter, the return value being a positive result or a negative result,

skipping a key in response to when the return value is a negative result, and

inserting a key into a duplication log (dup_log) file when the return value is a positive result.

5. A key-value store for data deduplication for an edge computer, comprising:

at least one processor configured to execute a computer-readable instruction included in a memory,

wherein the at least one processor is configured to:

receive a compaction request occurred from the key-value store persistent storage to a metadata layer having upper levels and lower levels, wherein compaction has previously occurred on the upper levels, wherein the compaction includes a deduplication process for removing duplicating data,

check whether deduplication for removing duplicated data is required by using a Bloom filter of metadata to check whether a key-value item is present when compaction of a metadata file is performed in response to the received compaction request, and

remove the duplicated data in response to the deduplication being required, wherein removing the duplicated data comprises:

determining a number of duplicated keys of the metadata at only the lower levels when the data duplication check is performed using the Bloom filter,

calculating a duplication ratio, wherein the duplication ratio is determined by a total number of keys divided by the number of duplicated keys, wherein the total number of keys is determined from the metadata layer,

comparing a value of the calculated duplication ratio with a threshold to determine whether a next deduplication process for removing duplicated data is required,

determining that the next deduplication process for removing duplicated data is required in response that the value of the duplication ratio is equal or less than the threshold, and

extending a compaction of the metadata to the lower level, and automatically removing the duplicated data from the persistent storage in the next deduplication process.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Mar 25, 2021
From: KIM, DEOKHWAN; TULKINBEKOV, KHIKMATULLO; KIM, JOOHWAN
To: INHA-INDUSTRY PARTNERSHIP INSTITUTE
Reel/Frame 055711/0776 →
Priority Claims (1)
KR 10-2020-0053848 · May 6, 2020 · national
Continuity (1)
Related Publication 20210349866A1 · Nov 11, 2021