IP Library › Granted Patent US 12,743,534
Granted Patent B2
US 12,743,534 · App. 17/975,919 · Granted Sep 22, 2026

Version control system using content-based datasets and dataset snapshots

Inventors: Adam Brenner (Mission Viejo, CA); Jehuda Shemer (Kfar Saba, IL); Steven Sadhwani (Round Rock, TX); Valerie Lotosh (Ramat-Gan, IL); Erez Sharvit (Ramat-Gan, IL)
Assignee: Dell Products L.P.
G06F21/6218G06F16/125G06F16/1873G06F2221/2141
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,743,534
App. No.
17/975,919
Filed
Oct 28, 2022
Granted
Sep 22, 2026
Kind
B2
Art Unit
2492
USPC
726/7
Abstract

Managing versioning of data objects for a project revised from a first version to a revised version by producing a dataset representing the data objects as a group by scanning the data objects to identify metadata of the grouped data to be processed similarly within a current version of the lifecycle, and storing the identified metadata in the dataset. Data object changed from the first version to the revised version are identified, and the corresponding metadata for changed data objects in the dataset is updated. A version control operation is then performed on the dataset to update all data objects referenced by the dataset from the first version to the revised version. A commit-map and commit-tree are stored in a repository, and version control operations including commit, checkout, merge, branch and merge-branch are performed on the dataset snapshot.

Claims (61)

1 . A computer-implemented method of managing different versions of data objects for a version control system (VCS) during a lifecycle of the data objects, comprising:

producing a dataset representing the data objects as a group by scanning the data objects to identify metadata of the grouped data to be processed similarly within a current version of the lifecycle, and storing the identified metadata in the dataset, wherein the dataset is generated by executing a user entered query on a catalog comprising a plurality of data catalogs, the dataset comprising both dynamic dataset data and static dataset data, wherein the static dataset data comprises a fixed amount of data set at a time of creation, and the dynamic dataset data comprises an amount of data that changes over time;

storing the dynamic dataset data in a first data catalog of the catalog and the static dataset data in a second data catalog of the catalog, wherein the first data catalog is used when new data is ingested by the VCS;

converting the dynamic dataset data into a static dataset by copying results of a query of the dynamic dataset data into the second data catalog by copying data from a first storage system to a second storage system and adding an entry to the second data catalog to record an identifier of the copied data, and integrating a data mover to the static dataset to move the data as part of an workflow exposed in the catalog to process the user entered query;

identifying data objects that themselves are subject to a change from the current version to a next version during the lifecycle;

updating corresponding metadata for changed data objects in the dataset; and

applying a version control operation on the dataset to update all data objects referenced by the dataset from the current version to the next version so that the dataset itself is re-versioned as a whole.

2 . The method of claim 1 wherein the dataset is distributed across the plurality of storage devices comprise network attached storage (NAS), object storage, local storage, or cloud networks, the method further comprising:

generating, by each provider of a storage device of the plurality of storage devices, a dataset snapshot as a read-only dataset component stored in memory local to the provider, wherein the dataset snapshot comprises a list of snapshot copies provided by each provider; and

copying the dataset to a remote storage location using a dataset backup, wherein the remote storage location is different from the local storage location.

3 . The method of claim 2 wherein the lifecycle of the data objects in the VCS comprises checking out data objects of a project to be modified, modifying the data objects to generate a revised version of the project from a first version, and committing the data objects of the revised version to a repository as a VCS datastore.

4 . The method of claim 3 further comprising storing, in the VCS datastore, a commit-map and commit-tree of the next version of the project, wherein the commit map stores commit records for the data objects from the first version to the revised version, and wherein the commit-tree stores a timeline of the commit operations generating the commit records.

5 . The method of claim 4 further comprising:

assigning a snapshot-ID to each dataset snapshot for tracking a corresponding snapshot through the commit map and commit-tree; and

performing one or more VCS operations on an identified dataset snapshot including at least one of a commit, checkout, merge, branch, or merge-branch operation.

6 . The method of claim 5 further comprising defining a HEAD index that points to a commit operation that the dataset snapshot is based on, and wherein the HEAD index is null at a beginning of a commit-tree for the delete snapshot.

7 . The method of claim 6 wherein, for the commit operation, the method further comprises:

creating the dataset snapshot for storage on either the remote or local storage;

creating a commit record; and

adding a commit identifier in the commit-tree after a position of the HEAD index; and

setting the HEAD index to be the commit identifier.

8 . The method of claim 7 wherein, for the checkout operation, the method further comprises:

retrieving snapshot-ID from the commit record;

copying content of the dataset snapshot to the original dataset; and

setting the HEAD index to be a checkout commit identifier.

9 . The method of claim 8 wherein, for the merge operation, the method further comprises:

retrieving the snapshot-ID from the commit record

merging the content of the dataset snapshot with the original dataset; and

performing the commit operation.

10 . The method of claim 9 wherein, for the branch operation, the method further comprises:

creating a new dataset snapshot from the original dataset; and

creating a new checkout commit-ID to be stored in a new datastore.

11 . The method of claim 10 wherein, for the merge-branch operation, the method further comprises:

merging the original dataset into a target datastore; and

committing the merge in the target datastore.

12 . The method of claim 11 wherein the VCS manages changes to software programs, documents, web sites, and other content data embodying the data objects, and wherein the first version and revised version are each denoted by successive alphanumeric version character, and wherein each identifier of the snapshot-ID and commit-ID reference the version character.

13 . The method of claim 3 wherein the data objects within each version of the project are encompassed by a respective dataset and are subject to same control rules in each stage of a lifecycle of the project as grouped data, wherein the control rules provide access only to authorized users or perform only authorized operations including data storage operations on the dataset referenced data objects based on a current stage of the lifecycle, and wherein the dataset is processed in the system as a single unit based on data content rather than data location.

14 . The method of claim 13 wherein the dataset is produced by:

gathering the identified metadata for storage in a data catalog; and

executing a user entered query comprising metadata selectors as dataset tags for matching against the cataloged metadata to generate the dataset, wherein the metadata selectors comprise tags consisting of alphanumeric strings applied to respective data objects based on user-defined rules, and wherein the tags define at least one of a file type, name, location, creation time, or characteristic.

15 . A computer-implemented method of managing different versions of data objects for a version control system (VCS) during a lifecycle of the data objects, comprising:

identifying data objects that evolve through the different versions during the lifecycle;

producing a dataset for the data objects data as a group by scanning the data objects to identify metadata of the grouped data to be re-versioned together throughout the lifecycle, and storing the identified metadata in the dataset;

executing a user entered query against a catalog comprising a plurality of data catalogs to produce the dataset, wherein the dataset comprises both dynamic dataset data and static dataset data, wherein the static dataset data comprises a fixed amount of data set at a time of creation, and the dynamic dataset data comprises an amount of data that changes over time;

storing the dynamic dataset data in a first data catalog of the catalog and the static dataset data in a second data catalog of the catalog, wherein the first data catalog is used when new data is ingested by the VCS;

converting the dynamic dataset data into a static dataset by copying results of a query of the dynamic dataset data into the second data catalog by copying data from a first storage system to a second storage system and adding an entry to the second data catalog to record an identifier of the copied data, and integrating a data mover to the static dataset to move the data as part of an workflow exposed in the catalog to process the user entered query;

generating dataset snapshots as read-only dataset components for the dataset as it progresses along the lifecycle;

copying the dataset to a remote storage location using a dataset backup;

assigning a snapshot-ID to each dataset snapshot for tracking a corresponding snapshot through the commit map and commit-tree; and

performing one or more VCS operations on an identified dataset snapshot including at least one of a commit, checkout, merge, branch, or merge-branch operation.

16 . The method of claim 15 further comprising storing, in the VCS datastore, a commit-map and commit-tree of the next version of the project, wherein the commit map stores commit records for the data objects from the first version to the revised version, and wherein the commit-tree stores a timeline of the commit operations generating the commit records.

17 . The method of claim 16 wherein the dataset is distributed across the plurality of storage devices comprise network attached storage (NAS), object storage, local storage, or cloud networks, the method further comprising generating by each provider of a storage device of the plurality of storage devices, a dataset snapshot as a read-only dataset component stored in memory local to the provider, wherein the dataset snapshot comprises a list of snapshot copies provided by each provider, the method further comprising converting the dynamic dataset data into static dataset data by copying results of a query of the dynamic dataset data into the second data catalog.

18 . The method of claim 17 wherein the VCS manages changes to software programs, documents, web sites, and other content data embodying the data objects, and wherein the first version and revised version are each denoted by successive alphanumeric version character, and wherein each identifier of the snapshot-ID and commit-ID reference the version character, and further wherein the data objects within each version of the project are encompassed by a respective dataset and are subject to same control rules in each stage of a lifecycle of the project as grouped data, wherein the control rules provide access only to authorized users or perform only authorized operations including data storage operations on the dataset referenced data objects based on a current stage of the lifecycle, and wherein the dataset is processed in the system as a single unit based on data content rather than data location.

19 . A computer-implemented method of managing different versions of data objects for a version control system (VCS) during a lifecycle of the data objects, comprising:

implementing the VCS to manage changes to software programs, documents, web sites, and other content data embodying the data objects, and wherein the first version and revised version are each denoted by successive alphanumeric version character, and wherein each identifier of the snapshot-ID and commit-ID reference the version character;

producing a dataset for the data objects data as a group by scanning the data objects to identify metadata of the grouped data to be re-versioned together, and storing the identified metadata in the dataset;

executing a user entered query against a catalog comprising a plurality of data catalogs to produce the dataset, wherein the dataset comprises both dynamic dataset data and static dataset data, wherein the static dataset data comprises a fixed amount of data set at a time of creation, and the dynamic dataset data comprises an amount of data that changes over time;

storing the dynamic dataset data in a first data catalog of the catalog and the static dataset data in a second data catalog of the catalog, wherein the first data catalog is used when new data is ingested by the VCS;

converting the dynamic dataset data into a static dataset by copying results of a query of the dynamic dataset data into the second data catalog by copying data from a first storage system to a second storage system and adding an entry to the second data catalog to record an identifier of the copied data, and integrating a data mover to the static dataset to move the data as part of an workflow exposed in the catalog to process the user entered query; and

generating dataset snapshots as read-only dataset components for the dataset as it progresses along the lifecycle, wherein the dataset is generated by gathering the identified metadata for storage in a data catalog, and executing a user entered query comprising metadata selectors as dataset tags for matching against the cataloged metadata to generate the dataset.

20 . The method of claim 19 wherein the dataset is distributed across the plurality of storage devices comprise network attached storage (NAS), object storage, local storage, or cloud networks, the method further comprising generating by each provider of a storage device of the plurality of storage devices, a dataset snapshot as a read-only dataset component stored in memory local to the provider, wherein the dataset snapshot comprises a list of snapshot copies provided by each provider.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 28, 2022
From: BRENNER, ADAM; SHEMER, JEHUDA; SADHWANI, STEVEN; LOTOSH, VALERIE; SHARVIT, EREZ
To: DELL PRODUCTS L.P.
Reel/Frame 061578/0929 →
Continuity (1)
Related Publication 20240143815A1 · May 2, 2024
References Cited (27)
US 7158994B1 · Smith · 2007 [cited by applicant]
US 9633023B2 · Kelley · 2017 [cited by examiner]
US 10489518B1 · Ramachandran · 2019 [cited by examiner]
US 11055115B1 · Kudrin · 2021 [cited by examiner]
US 11429353B1 · Liguori · 2022 [cited by examiner]
US 20090198731A1 · Noonan, III · 2009 [cited by examiner]
US 20110093471A1 · Brockway · 2011 [cited by applicant]
US 20120109958A1 · Thakur · 2012 [cited by applicant]
US 20130227089A1 · McLeod · 2013 [cited by examiner]
US 20140289202A1 · Chan · 2014 [cited by applicant]
US 20180027006A1 · Zimmermann · 2018 [cited by applicant]
US 20180145983A1 · Bestler · 2018 [cited by examiner]
US 20180225177A1 · Bhagi · 2018 [cited by applicant]
US 20180260211A1 · Miller · 2018 [cited by examiner]
US 20190182294A1 · Rieke · 2019 [cited by applicant]
US 20190272335A1 · Liu · 2019 [cited by applicant]
US 20200133781A1 · Reddy Av · 2020 [cited by applicant]
US 20200311039A1 · Gupta · 2020 [cited by examiner]
US 20210165782A1 · Deshpande · 2021 [cited by applicant]
US 20210240659A1 · Wee · 2021 [cited by applicant]
US 20220012251A1 · Colcord · 2022 [cited by applicant]
US 20220309046A1 · Qiu · 2022 [cited by examiner]
US 20220318421A1 · Berube · 2022 [cited by applicant]
US 20220335340A1 · Moustafa · 2022 [cited by applicant]
US 20230079486A1 · Yarlagadda et al. · 2023 [cited by applicant]
US 20230409545A1 · Gupta · 2023 [cited by examiner]
Petr Baudis, Current Concepts in Version Control Systems , Sep. 11, 2009, Cornell University. [cited by examiner]