IP Library › Granted Patent US 12,001,703
Granted Patent B2
US 12,001,703 · App. 17/741,079 · Granted Jun 4, 2024

Data processing method and storage device

Inventors: Ren Ren (Shanghai, CN); Zhongquan Liu (Shenzhen, CN); Hongwei Liu (Shanghai, CN); Fangfang Zhu (Xi'an, CN)
Assignee: HUAWEI TECHNOLOGIES CO., LTD.
G06F3/0641G06F3/0604G06F3/0608G06F3/0674
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,001,703
App. No.
17/741,079
Granted
Jun 4, 2024
Kind
B2
Abstract

This application provides a data processing method and a storage device, and belongs to the field of storage technologies. In this application, the storage device performs deduplication and compression based on different granularities, deduplicates data based on a large granularity, and compresses the data based on a small granularity. Therefore, a limitation that a deduplication granularity and a compression granularity need to be the same is removed. A deduplication ratio decrease caused by an excessively large granularity and a compression ratio decrease caused by an excessively small granularity are avoided to some extent, to improve an overall reduction ratio of deduplication and compression.

Claims (40)

1. A method for processing data performed by a storage device, the method comprising:

obtaining data;

deduplicating the data based on a first granularity;

compressing the deduplicated data based on a second granularity, wherein a size of the second granularity is greater than a size of the first granularity; and

storing data obtained after the deduplication and the compression in a hard disk of the storage device; wherein the storage device stores metadata managed based on a metadata management granularity, a size of the metadata management granularity is less than or equal to a specified largest value and is greater than or equal to a specified smallest value, and the size of the first granularity is equal to an integer multiple of the smallest value.

2. The method according to claim 1 , wherein the size of the second granularity is a product of the smallest value and a compression ratio.

3. The method according to claim 1 , wherein the deduplicating the data based on the first granularity comprises:

dividing the data into a plurality of data blocks;

obtaining a fingerprint of each data block; and

determining a duplicate block and a non-duplicate block from the plurality of data blocks based on the fingerprints.

4. The method according to claim 3 , wherein the compressing the data based on the second granularity comprises:

compressing the non-duplicate block based on the second granularity to obtain a compressed block, wherein the data obtained after the deduplication and the compression comprises the compressed block.

5. The method according to claim 4 , further comprising recording metadata of the compressed block.

6. The method according to claim 5 , wherein the recording metadata of the compressed block comprises:

if there are a plurality of compressed blocks and addresses of the plurality of compressed blocks are consecutive, recording one piece of metadata for the plurality of compressed blocks.

7. The method according to claim 6 , wherein the piece of metadata comprises an address of a first compressed block in the plurality of compressed blocks and a length of each compressed block.

8. The method according to claim 1 , wherein the data is further compressed based on a third granularity before the deduplication and the compression, and a size of the third granularity is less than the size of the second granularity.

9. The method according to claim 1 , wherein the storage device is a part of a storage array.

10. The method according to claim 1 , wherein the storage device is a storage node in a distributed storage system.

11. A storage device, comprising:

a hard disk; and

at least one processor configured to

obtain data,

deduplicate the data based on a first granularity,

compress the deduplicated data based on a second granularity, wherein a size of the second granularity is greater than a size of the first granularity, and

store data obtained after the deduplication and the compression in the hard disk; wherein the storage device stores metadata managed based on a metadata management granularity, a size of the metadata management granularity is less than or equal to a specified largest value and is greater than or equal to a specified smallest value, and the size of the first granularity is equal to an integer multiple of the smallest value.

12. The storage device according to claim 11 , wherein the size of the second granularity is a product of the smallest value and a compression ratio.

13. The storage device according to claim 11 , wherein the at least one processor is configured to:

divide the data into a plurality of data blocks,

obtain a fingerprint of each data block, and

determine a duplicate block and a non-duplicate block from the plurality of data blocks based on the fingerprints.

14. The storage device according to claim 13 , wherein the at least one processor is configured to compress the non-duplicate block based on the second granularity to obtain a compressed block, wherein the data obtained after the deduplication and the compression comprises the compressed block.

15. The storage device according to claim 14 , wherein the at least one processor is further configured to record metadata of the compressed block.

16. The storage device according to claim 15 , wherein the at least one processor is configured to: if there are a plurality of compressed blocks and addresses of the plurality of compressed blocks are consecutive, record one piece of metadata for the plurality of compressed blocks.

17. The storage device according to claim 16 , wherein the piece of metadata comprises an address of the first compressed block in the plurality of compressed blocks and a length of each compressed block.

18. A non-transitory computer-readable medium having instructions stored therein, which when executed by a processor, cause the processor to perform operations of processing data, the operations comprising:

obtaining data;

deduplicating the data based on a first granularity;

compressing the deduplicated data based on a second granularity, wherein a size of the second granularity is greater than a size of the first granularity; and

storing data obtained after the deduplication and the compression in a hard disk of a storage device; wherein the processor stores metadata managed based on a metadata management granularity, a size of the metadata management granularity is less than or equal to a specified largest value and is greater than or equal to a specified smallest value, and the size of the first granularity is equal to an integer multiple of the smallest value.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 18, 2024
From: REN, REN; LIU, ZHONGQUAN; LIU, HONGWEI; ZHU, FANGFANG
To: HUAWEI TECHNOLOGIES CO., LTD.
Reel/Frame 067154/0204 →
Priority Claims (2)
CN 202010526840.7 · Jun 11, 2020 · national
CN 202010784929.3 · Aug 6, 2020 · national
Continuity (2)
Continuation PCTCN2020136106 · Dec 14, 2020
Related Publication 20220269431A1 · Aug 25, 2022