IP Library Granted Patent US 12,585,618
Granted Patent B2
US 12,585,618 · App. 18/390,020 · Granted Mar 24, 2026

Systems and methods for sequence-based data chunking for deduplication

Inventors: Sreeharsha Udayashankar (Waterloo, CA); Abdelrahman Ba'ba' (Waterloo, CA); Samer Al-Kiswany (Waterloo, CA); Serg Bell (Costa del Sol, SG); Stanislav Protasov (Singapore, SG)
Assignee: Acronis International GmbH
G06F16/1752
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,585,618
App. No.
18/390,020
Granted
Mar 24, 2026
Kind
B2
Abstract

Disclosed herein are systems and method for data chunking. In one aspect, a method includes consecutively scanning each respective byte in a byte stream of data. The method includes in response to detecting a first amount of adjacent bytes with values arranged in a decreasing order, marking a cut-off point of a data chunk comprising at least the adjacent bytes. The method includes in response to detecting a second amount of bytes with values arranged in an increasing order, executing a jump mechanism that skips scanning of a fixed amount of bytes after the bytes with values arranged in the increasing order. The method includes subsequent to scanning the byte stream in entirety, outputting a plurality of cut-off points of identified data chunks.

Claims (48)

1 . A method for data chunking, the method comprising:

identifying, by a hardware processor, a desired throughput value for data chunking a byte stream of data by:

determining, by the hardware processor, a size of the byte stream;

receiving, by the hardware processor, a threshold period of time to complete the data chunking; and

calculating, by the hardware processor, the desired throughput value based on the size of the byte stream and the threshold period of time;

setting, by the hardware processor, a first amount of adjacent bytes, a second amount of bytes, and a fixed amount of bytes of a jump mechanism that achieve the desired throughput value;

consecutively scanning, by the hardware processor, each respective byte in the byte stream of data;

in response to detecting the first amount of adjacent bytes with values arranged in a decreasing order, marking, by the hardware processor, a cut-off point of a data chunk comprising at least the adjacent bytes;

in response to detecting the second amount of bytes with values arranged in an increasing order, executing, by the hardware processor, the jump mechanism that skips scanning of the fixed amount of bytes after the bytes with values arranged in the increasing order; and

subsequent to scanning the byte stream in entirety, outputting, by the hardware processor, a plurality of cut-off points of identified data chunks.

2 . The method of claim 1 , wherein setting the first amount of adjacent bytes, the second amount of bytes, and the fixed amount of bytes of the jump mechanism that achieve the desired throughput value comprises executing a machine learning model trained to calculate for an input throughput value: an output amount of adjacent bytes, an output amount of bytes, and an output amount of bytes for the jump mechanism.

3 . The method of claim 2 , wherein the machine learning model is trained on a training dataset comprising output byte amounts for various input throughput values.

4 . A method for data chunking, the method comprising:

identifying, by a hardware processor, a desired throughput value for data chunking a byte stream of data by:

determining, by the hardware processor, a size of the byte stream;

receiving, by the hardware processor, a threshold period of time to complete the data chunking; and

calculating, by the hardware processor, the desired throughput value based on the size of the byte stream and the threshold period of time;

setting, by the hardware processor, a first amount of adjacent bytes, a second amount of bytes, and a fixed amount of bytes of a jump mechanism based on the desired throughput value;

consecutively scanning, by the hardware processor, each respective byte in the byte stream of data;

in response to detecting the first amount of adjacent bytes with values arranged in an increasing order, marking, by the hardware processor, a cut-off point of a data chunk comprising at least the adjacent bytes;

in response to detecting the second amount of bytes with values arranged in a decreasing order, executing, by the hardware processor, the jump mechanism that skips scanning of the fixed amount of bytes after the bytes with values arranged in the decreasing order; and

subsequent to scanning the byte stream in entirety, outputting, by the hardware processor, a plurality of cut-off points of identified data chunks.

5 . The method of claim 4 , wherein the bytes with values arranged in the decreasing order are not all adjacent to one another.

6 . The method of claim 5 , wherein a majority of the second amount of bytes are arranged in the decreasing order.

7 . The method of claim 4 , wherein one or both of the first amount of adjacent bytes and the second amount of bytes is incremented when a respective cut-off point is marked.

8 . The method of claim 4 , wherein one or both of the first amount of adjacent bytes and the second amount of bytes is reset to a default value when the jump mechanism is executed.

9 . The method of claim 4 , further comprising:

subsequent to identifying a data chunk by marking a respective cut-off point, determining an amount of bytes remaining in the byte stream; and

in response to determining that the amount of bytes is less than the first amount of adjacent bytes, identifying the bytes remaining in the byte stream as a last data chunk in the byte stream.

10 . A system for data chunking, comprising:

at least one memory; and

at least one hardware processor coupled with the at least one memory and configured, individually or in combination, to:

identify a desired throughput value for data chunking a byte stream of data by:

determining a size of the byte stream;

receiving a threshold period of time to complete the data chunking; and

calculating the desired throughput value based on the size of the byte stream and the threshold period of time;

set a first amount of adjacent bytes, a second amount of bytes, and a fixed amount of bytes of a jump mechanism that achieve the desired throughput value;

consecutively scan each respective byte in the byte stream of data;

in response to detecting the first amount of adjacent bytes with values arranged in an increasing order, mark a cut-off point of a data chunk comprising at least the adjacent bytes;

in response to detecting the second amount of bytes with values arranged in a decreasing order, execute the jump mechanism that skips scanning of the fixed amount of bytes after the bytes with values arranged in the decreasing order; and

subsequent to scanning the byte stream in entirety, output a plurality of cut-off points of identified data chunks.

11 . The system of claim 10 , wherein the bytes with values arranged in the decreasing order are not all adjacent to one another.

12 . The system of claim 11 , wherein a majority of the second amount of bytes are arranged in the decreasing order.

13 . The system of claim 10 , wherein the second amount of bytes is incremented when a respective cut-off point is marked.

14 . The system of claim 10 , wherein the second amount of bytes is reset to a default value when the jump mechanism is executed.

15 . The system of claim 10 , wherein the at least one hardware processor is further configured to:

subsequent to identifying a data chunk by marking a respective cut-off point, determine an amount of bytes remaining in the byte stream; and

in response to determining that the amount of bytes is less than the first amount of adjacent bytes, identify the bytes remaining in the byte stream as a last data chunk in the byte stream.

Assignments (3)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 23, 2026
From: UDAYASHANKAR, SREEHARSHA; BA`BA`, ABDELRAHMAN; AL-KISWANY, SAMER; BELL, SERG; PROTASOV, STANISLAV
To: ACRONIS INTERNATIONAL GMBH
Reel/Frame 073859/0853 →
CORRECTIVE ASSIGNMENT TO CORRECT THE PATENTS LISTED BY DELETING PATENT APPLICATION NO. 18388907 FROM SECURITY INTEREST PREVIOUSLY RECORDED ON REEL 66797 FRAME 766. ASSIGNOR(S) HEREBY CONFIRMS THE SECURITY INTEREST. Recorded Nov 13, 2024
From: ACRONIS INTERNATIONAL GMBH
To: MIDCAP FINANCIAL TRUST
Reel/Frame 069594/0136 →
SECURITY INTEREST Recorded Mar 14, 2024
From: ACRONIS INTERNATIONAL GMBH
To: MIDCAP FINANCIAL TRUST
Reel/Frame 066797/0766 →
Continuity (1)
Related Publication 20250209043A1 · Jun 26, 2025
References Cited (17)
US 8639669B1 · Douglis · 2014 [cited by examiner]
US 10809945B2 · Patwardhan · 2020 [cited by examiner]
US 10887372B2 · Jang · 2021 [cited by examiner]
US 20130086009A1 · Li · 2013 [cited by examiner]
US 20180196609A1 · Niesen · 2018 [cited by examiner]
US 20210216497A1 · Ramteke · 2021 [cited by examiner]
US 20220004524A1 · Chen · 2022 [cited by examiner]
KR 100796176B1 · 2008 [cited by examiner]
WO WO2014184857A1 · 2014 [cited by examiner]
Saeed, Ahmed Sardar M., et al., “Data Deduplication System Based on Content-Defined Chunking Using Bytes Pair Frequency Occurrence”, Symmetry, vol. 12, Issue 11, Nov. 6, 2020, pp. 1-21. [cited by examiner]
Ajdari, Mohammadamin, et al., “FIDR: A Scalable Storage System for Fine-Grain Inline Data Reduction with Efficient Memory Handling”, MICRO-52, Columbus, OH, Oct. 12-16, 2019, pp. 239-252. [cited by examiner]
Jehlol, Hashem B., et al., “Big Data De-duplication Using Classification Scheme based on Histogram of File Stream”, ITSS-IoE, Hadhramaut, Yemen, Dec. 3-5, 2022, 7 pages. [cited by examiner]
Jehlol, Hashem Bedr, et al., “Big Data Backup: A Survey”, International Journal of Scientific Research in Science, Engineering and Technology, vol. 9, Issue 4, Jul.-Aug. 2022, pp. 174-190. [cited by examiner]
Yu, Chuanshuai, et al., “Leap-based Content Defined Chunking—Theory and Implementation”, MSST 2015, Santa Clara, CA, 30 May-Jun. 5, 2015, 12 pages. [cited by examiner]
Widodo et al., “A new content-defined chunking algorithm for data deduplication in cloud storage,” Future Generation Computer Systems, 2017, vol. 71, pp. 145-156. [cited by applicant]
Xia et al., “FastCDC: A Fast and Efficient Content-Defined Chunking Approach for Data Deduplication,” 2016 USENIX Annual Technical Conference, Jun. 22-24, 2016, pp. 101-114. [cited by applicant]
Zhang et al., “AE: An Asymmetric Extremum content defined chunking algorithm for fast and bandwidth-efficient data deduplication,” 2015 IEEE Conference on Computer Communications (INFOCOM), Apr. 26-May 1, 2015, 9 pages. [cited by applicant]