IP Library Granted Patent US 12694107
Granted Patent B1
US 12694107 · App. 19/562,220 · Granted Jul 28, 2026

Metadata-based data object detection and classification in cloud computing environments for data security posture management

Inventors: Alma Raziel (Tel Aviv-Jaffa, IL); Elad Gabay (Tel Aviv, IL); Liron Levin (Kfar Saba, IL); Daniel Lazarev (Tel Aviv, IL); Erez Harush (Tel Aviv, IL); George Pisha (Giv'atayim, IL)
Assignee: Wiz, Inc.
G06F21/554G06F21/577G06F21/552
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12694107
App. No.
19/562,220
Granted
Jul 28, 2026
Kind
B1
Abstract

A system and method for classifying data objects for data security posture management (DSPM) in a cloud computing environment based on metadata are presented. The method includes detecting data objects in data sources of the cloud computing environment, wherein each data object in the data objects comprises metadata; generating metadata-derived features from the metadata, wherein the metadata-derived features include hierarchy and path features; clustering the data objects into data clusters based on a similarity determination applied to the metadata-derived features; generating, for a first data cluster of the data clusters, an aggregated cluster representation and a corresponding prompt based on metadata associated with data objects in the first data cluster; generating a cluster classification for the first data cluster based on a result of processing the corresponding prompt with a language model; and generating data findings based on the cluster classification and classification results.

Claims (60)

1 . A method for classifying data objects for data security posture management (DSPM) in a cloud computing environment based on metadata, comprising:

detecting a plurality of data objects in one or more data sources of the cloud computing environment, wherein each data object in the plurality of data objects comprises metadata;

generating metadata-derived features from the metadata, wherein the metadata-derived features include hierarchy and path features derived from at least one of object keys and file paths;

clustering the plurality of data objects into a plurality of data clusters based on a similarity determination applied to the metadata-derived features;

generating, for a first data cluster of the plurality of data clusters, an aggregated cluster representation and a corresponding prompt based on metadata associated with data objects in the first data cluster;

generating a cluster classification for the first data cluster based on a result of processing the corresponding prompt with a language model; and

generating one or more data findings based on at least one of the cluster classification and a plurality of classification results.

2 . The method of claim 1 , further comprising:

generating a recommended remediation action based on a data finding; and

initiating the recommended remediation action in the cloud computing environment by invoking one or more provider interfaces.

3 . The method of claim 2 , wherein the invoked one or more provider interfaces modify at least one of access controls, encryption settings, tags, and storage configuration associated with a data object implicated by the data finding.

4 . The method of claim 1 , further comprising:

monitoring drift based on changes in at least one of metadata distributions, cluster composition, and rule-based match distributions; and

responsive to detecting drift, triggering at least one of re-characterizing one or more clusters using the language model and re-synthesizing at least one customer-specific classification rule.

5 . The method of claim 1 , further comprising:

selecting, based on at least one of the cluster classification, a confidence score, and a policy constraint, a candidate subset of data objects for optional content scanning; and

performing content scanning only on the selected candidate subset to generate confirmation outputs that refine at least one data finding.

6 . The method of claim 1 , wherein the metadata, for each data object in the plurality of data objects, is obtained via one or more provider interfaces that provide object properties without accessing payload content of the data object.

7 . The method of claim 1 , further comprising:

generating, based on the cluster classification, one or more customer-specific classification rules configured to classify additional data objects using the metadata without accessing payload content.

8 . The method of claim 7 , further comprising:

validating the one or more customer-specific classification rules against a metadata corpus or inventory; and

storing versioned rules based on results of the validating.

9 . The method of claim 8 , further comprising:

applying at least one stored versioned rule to metadata-derived features of newly detected data objects to generate the plurality of classification results.

10 . A system for classifying data objects for data security posture management (DSPM) in a cloud computing environment based on metadata comprising:

a processing circuitry:

a memory, the memory containing instructions that, when executed by the processing circuitry, configure the system to:

detect a plurality of data objects in one or more data sources of the cloud computing environment, wherein each data object in the plurality of data objects comprises metadata;

generate metadata-derived features from the metadata, wherein the metadata-derived features include hierarchy and path features derived from at least one of object keys and file paths;

cluster the plurality of data objects into a plurality of data clusters based on a similarity determination applied to the metadata-derived features;

generate, for a first data cluster of the plurality of data clusters, an aggregated cluster representation and a corresponding prompt based on metadata associated with data objects in the first data cluster;

generate a cluster classification for the first data cluster based on a result of processing the corresponding prompt with a language model; and

generate one or more data findings based on at least one of the cluster classification and a plurality of classification results.

11 . The system of claim 10 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

generate a recommended remediation action based on a data finding; and

initiate the recommended remediation action in the cloud computing environment by invoking one or more provider interfaces.

12 . The system of claim 11 , wherein the invoked one or more provider interfaces modify at least one of access controls, encryption settings, tags, and storage configuration associated with a data object implicated by the data finding.

13 . The system of claim 10 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

monitor drift based on changes in at least one of metadata distributions, cluster composition, and rule-based match distributions; and

responsive to detecting drift, trigger at least one of re-characterizing one or more clusters using the language model and re-synthesizing at least one customer-specific classification rule.

14 . The system of claim 10 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

select, based on at least one of the cluster classification, a confidence score, and a policy constraint, a candidate subset of data objects for optional content scanning; and

perform content scanning only on the selected candidate subset to generate confirmation outputs that refine at least one data finding.

15 . The system of claim 10 , wherein the metadata, for each data object in the plurality of data objects, is obtained via one or more provider interfaces that provide object properties without accessing payload content of the data object.

16 . The system of claim 10 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

generate, based on the cluster classification, one or more customer-specific classification rules configured to classify additional data objects using the metadata without accessing payload content.

17 . The system of claim 16 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

validate the one or more customer-specific classification rules against a metadata corpus or inventory; and

store versioned rules based on results of the validating.

18 . The system of claim 17 , wherein the memory contains further instructions which when executed by the processing circuitry further configure the system to:

apply at least one stored versioned rule to metadata-derived features of newly detected data objects to generate the plurality of classification results.

19 . A non-transitory computer-readable medium storing a set of instructions for classifying data objects for data security posture management (DSPM) in a cloud computing environment based on metadata, the set of instructions comprising:

one or more instructions that, when executed by one or more processing circuitries of a device, cause the device to:

detect a plurality of data objects in one or more data sources of the cloud computing environment, wherein each data object in the plurality of data objects comprises metadata;

generate metadata-derived features from the metadata, wherein the metadata-derived features include hierarchy and path features derived from at least one of object keys and file paths;

cluster the plurality of data objects into a plurality of data clusters based on a similarity determination applied to the metadata-derived features;

generate, for a first data cluster of the plurality of data clusters, an aggregated cluster representation and a corresponding prompt based on metadata associated with data objects in the first data cluster;

generate a cluster classification for the first data cluster based on a result of processing the corresponding prompt with a language model; and

generate one or more data findings based on at least one of the cluster classification and a plurality of classification results.