IP Library Granted Patent US 10,372,731
Granted Patent B1
US 10,372,731 · App. 15/360,612 · Granted Aug 6, 2019

Method of generating a data object identifier and system thereof

Inventors: Yaniv Avidan (Moshav Mahseya, IL); Avner Atias (Kfar Yona, IL)
Assignee: MINEREYE LTD.
G06F16/285G06F16/2219G06F16/2237G06F16/2264
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 10,372,731
App. No.
15/360,612
Granted
Aug 6, 2019
Kind
B1
Abstract

Generating a data object identifier by dividing the data in the data object into a plurality of chunks; processing each chunk using a clustering algorithm to generate, for each chunk, a pair of values characterizing the data in the chunk, thereby giving rise to a plurality of pairs of values (PoV); generating a plurality of nodes in a two dimensional space each corresponding to a respective PoV, wherein, for any given PoV, the values in the given PoV are indicative of location coordinates of the corresponding node in the two dimensional space; generating a plurality of features related to the plurality of nodes, each feature characterizing a spatial relationship between three or more nodes; and generating the data object identifier by arranging the features in a feature vector in accordance with predetermined rules.

Claims (34)

1. A method of generating a data object identifier, the method executed by a computer and comprising:

upon receiving a data object, dividing the data in the data object into a plurality of chunks;

processing each chunk using a clustering algorithm to generate, for each chunk, a pair of values characterizing the data in the chunk, thereby giving rise to a plurality of pairs of values (PoV);

generating a plurality of nodes in a two dimensional space each corresponding to a respective PoV, wherein, for any given PoV, the values in the given PoV are indicative of location coordinates of the corresponding node in the two dimensional space;

generating a plurality of features related to the plurality of nodes, each feature characterizing a spatial relationship between three or more nodes; and

generating the data object identifier by arranging the features in a feature vector in accordance with predetermined rules, wherein the generated data object identifier is usable for detecting similarity with other data objects whereby a level of accuracy of the detecting is increased.

2. The method of claim 1 wherein the data object is a file.

3. The method of claim 1 wherein the data is divided into chunks using a predetermined value n indicative of a maximum chunk size.

4. The method of claim 1 wherein the PoV for a chunk is generated by processing the data in the chunk using a clustering algorithm.

5. The method of claim 4 wherein the clustering algorithm is a self-organizing map algorithm.

6. The method of claim 1 wherein for each node, the first value in the PoV corresponding to the node defines the x-axis coordinate in the two dimensional space and the second value in the PoV corresponding to the node defines the y-axis coordinate in the two dimensional space.

7. The method of claim 1 wherein the spatial relationship comprises at least one of i) an angle formed between a node and two other nodes, and ii) a distance ratio between a given node and two other nodes.

8. The method of claim 1 , wherein detecting similarity with other data objects is usable for at least one of: detecting sensitive data in the data object; identifying data objects with similar data; discovering data objects with duplicate data;

identifying instances of the data object sharing.

9. A system capable of generating a data object identifier comprising a processor and memory block operatively coupled to one or more data repositories, the processor and memory block configured to:

upon receiving a data object stored on the one or more data repositories, divide the data in the data object into a plurality of chunks;

process each chunk using a clustering algorithm to generate, for each chunk, a pair of values characterizing the data in the chunk, thereby giving rise to a plurality of pairs of values (PoV);

generate a plurality of nodes in a two dimensional space each corresponding to a respective PoV, wherein, for any given PoV, the values in the given PoV are indicative of location coordinates of the corresponding node in the two dimensional space;

generate a plurality of features related to the plurality of nodes, each feature characterizing a spatial relationship between three or more nodes; and

generate the data object identifier by arranging the features in a feature vector in accordance with predetermined rules, wherein the generated data object identifier is usable for detecting similarity with other data objects whereby a level of accuracy of the detecting is increased.

10. The system of claim 9 wherein the data object is a file.

11. The system of claim 9 wherein the data is divided into chunks using a predetermined value n indicative of a maximum chunk size.

12. The system of claim 9 wherein the PoV for a chunk is generated by processing the data in the chunk using a clustering algorithm.

13. The system of claim 12 wherein the clustering algorithm is a self-organizing map algorithm.

14. The system of claim 9 wherein for each node, the first value in the PoV corresponding to the node defines the x-axis coordinate in the two dimensional space and the second value in the PoV corresponding to the node defines the y-axis coordinate in the two dimensional space.

15. The system of claim 9 wherein the spatial relationship comprises at least one of i) an angle formed between a node and two other nodes, and ii) a distance ratio between a given node and two other nodes.

16. The system of claim 9 , wherein detecting similarity with other data objects is usable for at least one of: detecting sensitive data in the data object; identifying data objects with similar data; discovering data objects with duplicate data;

identifying instances of the data object sharing.

17. A non-transitory computer-readable memory tangibly embodying a program of instructions executable by a computer for executing a method of generating a data object identifier, the method comprising:

upon receiving a data object, dividing the data in the data object into a plurality of chunks;

processing each chunk using a clustering algorithm to generate, for each chunk, a pair of values characterizing the data in the chunk, thereby giving rise to a plurality of pairs of values (PoV);

generating a plurality of nodes in a two dimensional space each corresponding to a respective PoV, wherein, for any given PoV, the values in the given PoV are indicative of location coordinates of the corresponding node in the two dimensional space;

generating a plurality of features related to the plurality of nodes, each feature characterizing a spatial relationship between three or more nodes; and

generating the data object identifier by arranging the features in a feature vector in accordance with predetermined rules, wherein the generated data object identifier is usable for automated detecting similarity with other data objects whereby a level of accuracy of the detecting is increased.

Assignments (2)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 23, 2024
From: MINEREYE LTD.
To: MINEREYE TECHNOLOGIES LTD.
Reel/Frame 068987/0868 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 7, 2019
From: AVIDAN, YANIV; ATLAS, AVNER
To: MINEREYE LTD.
Reel/Frame 048019/0115 →
Continuity (1)
Provisional Application 62259749 · Nov 25, 2015