IP Library › Granted Patent US 12,353,357
Granted Patent B2
US 12,353,357 · App. 17/308,159 · Granted Jul 8, 2025

Systems and methods for formatting and reconstructing raw data

Inventors: Raphael Glon (Guilers, FR); Gaetan Cottereau (Kersaint-Plabennec, FR)
Assignee: OVH
G06F16/1727G06F16/122
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,353,357
App. No.
17/308,159
Granted
Jul 8, 2025
Kind
B2
Abstract

A method for formatting raw data comprises accessing the raw data, the raw data comprising sparse data segments, which are empty of any data, and non-sparse data segments, which comprise data, and generating a formatted data stream comprising one or more atomic blocks, each atomic block corresponding to a metadata file and to a portion of the non-sparse data segments of the raw data. Generating one atomic block comprises browsing the raw data and, upon locating one sparse data segment, populating the corresponding metadata file with offsets indicative of a beginning and an end of the located sparse data segment and populating the atomic block with a concatenation of at least portions of the non-sparse segments of the raw data located before and after the located sparse segment. If the atomic block exceeds a maximum size, another atomic block and another corresponding metadata file are populated.

Claims (36)

1. A method of formatting raw data, comprising:

accessing the raw data from a storage location on a first data storage device, the raw data containing sparse data segments and non-sparse data segments, the sparse data segments being empty of any data and the non-sparse data segments containing data;

reading the accessed raw data to locate the sparse data segments and the non-sparse data segments;

generating a formatted data stream comprising one or more formatted atomic blocks based on the read raw data, the one or more formatted atomic blocks having a maximum block size defined by a size of a destination device memory and associated with a metadata file, each one of the one or more formatted atomic blocks configured with a format that includes a field indicating a data size of the non-sparse data segments of the raw data, a field containing at least a portion of the non-sparse data segments, and a field containing the metadata file identifying specific beginning and end locations, in which the formatting of the one or more atomic blocks comprises:

when one of the atomic blocks does not exceed the maximum block size:

populate the metadata file of the one of the atomic blocks with offsets identifying a beginning of each of the located sparse data segments and an end of each of the located sparse data segments; and

populate the one of the atomic blocks with a concatenation of at least a portion of a first non-sparse segment of the raw data located before the sparse segment with at least a portion of a second non-sparse data segment of the raw data located after the sparse data segment; and

when one of the atomic blocks exceeds the maximum size:

populate the metadata file of another one of the atomic blocks with offsets indicative of the beginning of the located sparse data segment and the end of the located sparse data segment; and

populate the other one of the atomic blocks with a concatenation of the at least a portion of the first non-sparse data segment of the raw data located before the sparse data segment with the at least a portion of the second non-sparse data segment of the raw data located after the sparse data segment; and

reconstructing the raw data on a second data storage device from the formatted data stream, the second data storage device being distinct from the first data storage device, the reconstructing being based on the metadata file, the reconstructing comprising:

browsing the metadata file; and

iteratively reconstructing the raw data by writing a corresponding data content for each non-sparse data segment identified by the metadata file and executing a command to define a segment of zeros for each sparse data segment identified by the metadata file, wherein the segment of zeros comprises multiple zeros,

wherein one atomic block is streamed from the first data storage device to the second data storage device while the generating of the formatted data stream and the iterative reconstructing of the raw data are being executed in parallel.

2. The method of claim 1 , wherein a size of the segment of zeroes corresponds to a size of the corresponding sparse data segment.

3. The method of claim 1 , wherein generating the formatted data stream further comprises generating a footer, the footer comprising a value relating to a number of the one or more formatted atomic blocks and being concatenated to the formatted data stream.

4. A method of formatting raw data, comprising:

accessing the raw data on a first data storage device, the raw data containing sparse data segments and non-sparse data segments, the sparse data segments being empty of any data and the non-sparse data segments containing data;

reading the raw data to locate the sparse data segments and the non-sparse data segments;

generating a formatted data stream based on the read raw data in which a size of the formatted data stream is limited by a size of a destination device memory, the formatted data stream comprising a field indicating a data size of the non-sparse data segments of the raw data, a field containing at least a portion of the non-sparse data segments, and a field containing a metadata file, the metadata file identifying specific locations of located sparse segments, wherein the generating comprises, upon locating one of the sparse data segments in the raw data, assembling the metadata file by:

populating the metadata file with offsets identifying a beginning of each of the located sparse data segments and an end of each of the located sparse data segments; and

populating the metadata file by concatenating each of the located non-sparse segments in successive fashion of the raw data located between each sparse data segments; and

reconstructing the raw data on a second data storage device from the formatted data stream, the second data storage device being distinct from the first data storage device, the reconstructing being based on the metadata file, the reconstructing comprising:

browsing the metadata file; and

iteratively reconstructing the raw data by writing a corresponding data content for each non-sparse data segment identified by the metadata file and executing a command to define a segment of zeros for each sparse data segment identified by the metadata file, wherein the segment of zeros comprises multiple zeros,

wherein the formatted data stream is streamed from the first data storage device to the second data storage device while the generating of the formatted data stream and the reconstructing of the raw data are being executed in parallel.

5. The method of claim 4 , wherein a size of the segment of zeroes corresponds to a size of the corresponding sparse data segment.

6. The method of claim 4 , wherein:

the raw data comprises data blocks and is reconstructed on a block device; and

a command to define a block of zero which size corresponds to the size of the sparse segment is associated with the block device.

7. The method of claim 4 , further comprising, prior to reconstructing the raw data from the formatted data stream, executing a decompression algorithm on the formatted data stream.

8. The method of claim 4 , wherein the formatted data stream is streamed from a first data storage device to a second data storage device while the formatting of the raw data and the reconstructing of the raw data are being executed in parallel.

9. The method of claim 1 , wherein generating the formatted data stream further comprises generating a footer, the footer comprising one or more of the following values: a number of read bytes of the raw data, a number of written bytes including written bytes of the metadata file and the footer, and/or a ratio between the number of read bytes and the number of written bytes; the footer being concatenated to the formatted data stream.

10. The method of claim 1 , wherein the raw data comprises a disk-image containing a content and/or a structure of an entire data storage device.

11. The method of claim 1 , further comprising, subsequent to generating the formatted data stream, executing a compression algorithm on the formatted data stream.

12. A networking device for Operating System (OS) deployment on a bare-metal server, the networking device being configured to perform the method of claim 1 .

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 6, 2022
From: GLON, RAPHAEL; COTTEREAU, GAETAN
To: OVH
Reel/Frame 059841/0192 →
Priority Claims (1)
EP 20315273 · May 28, 2020 · regional
Continuity (1)
Related Publication 20210374102A1 · Dec 2, 2021
References Cited (25)
US 9811545B1 · Bent · 2017 [cited by examiner]
US 10108647B1 · Kumar · 2018 [cited by examiner]
US 10346362B2 · Tao · 2019 [cited by examiner]
US 10831398B2 · Barzik · 2020 [cited by examiner]
US 10860533B1 · Neilsen · 2020 [cited by examiner]
US 20080059398A1 · Tsutsui · 2008 [cited by examiner]
US 20110055273A1 · Resch · 2011 [cited by examiner]
US 20200019535A1 · Parker · 2020 [cited by examiner]
US 20200183886A1 · Wang · 2020 [cited by examiner]
Extended European Search Report with regard to the EP Patent Application No. 20315273.1 mailed Nov. 20, 2020. [cited by applicant]
Wikipedia, “Disk Image”, Apr. 26, 2020, retrieved on Nov. 10, 2020, pdf 7 pages. [cited by applicant]
Wikipedia, “Sparse File”, May 11, 2020, retrieved on Nov. 10, 2020, pdf 4 pages. [cited by applicant]
Wikipedia, “Split (Unix)”, May 11, 2020, retrieved on Nov. 10, 2020, pdf 3 pages. [cited by applicant]
“The .xz File Format”, retrieved on https://tukaani.org/xz/xz-file-format.txt on May 4, 2021, pdf 18 pages. [cited by applicant]
Hocevar, retrieved on https://gist.githubusercontent.com/kempniu/30c7fa2c1825cde80040/raw/7799c6708e89c4646a68f338d2ed3238688d5321/sparsify.c on May 4, 2021, pdf 2 pages. [cited by applicant]
“QEMU Disk Network Block Device Server”, retrieved on https://www.mankier.com/8/qemu-nbd on May 4, 2021, pdf 6 pages. [cited by applicant]
Jones, “Streaming NBD server”, retrieved on https://web.archive.org/web/20191202200236/https://rwmj.wordpress.com/2014/10/14/streaming-nbd-server/ on May 4, 2021, pdf 3 pages. [cited by applicant]
“Enhance importer to use qemu-img pseudo streaming to save /tmp step”, retrieved on https://github.com/kubevirt/containerized-data-importer/issues/254 on May 4, 2021, pdf 3 pages. [cited by applicant]
“Linux manual page”, retrieved on http://man7.org/linux/man-pages/man2/fallocate.2.html on May 4, 2021, pdf 6 pages. [cited by applicant]
“Hypertext Transfer Protocol”, retrieved on https://tools.ietf.org/html/rfc7233 on May 4, 2021, pdf 26 pages. [cited by applicant]
European Communication with regard to the EP Patent Application No. 20315273.1 mailed Aug. 10, 2023. [cited by applicant]
Anonymous: “Converting sparse file to non-sparse in place—Unix & Linux Stack Exchange”, Nov. 24, 2014 (Nov. 24, 2014), XP093070812, Retrieved from the Internet: URL: https://unix.stackexchange.com/questions/169669/conve… [cited by applicant]
Anonymous: “cp(1): copy files/directories—Linux man page”, Jul. 12, 2007 (Jul. 12, 2007), XP093070808, Retrieved from the Internet: URL: https://web.archive.org/web/20070712184940/https://linux.die.net/man/1/cp retrieve… [cited by applicant]
Anonymous: “sparse-fio/README.md at d7efae44cf0aa75be2c2bdc02346549b23579557.anyc/sparse-fio. GitHub”, Jul. 27, 2019 (Jul. 27, 2019), XP093070535, Retrieved from the Internet: URL: https://github.com/anyc/sparse-fio/blo… [cited by applicant]
Anonymous: “History for README.md—anyc/sparse-fio . GitHub”, Jul. 27, 2019 (Jul. 27, 2019), XP093070581—(proof of publication date of D7)—Retrieved from the Internet: URL: https://github.com/anyc/sparse-fio/commits/defa… [cited by applicant]