IP Library Granted Patent US 12,568,096
Granted Patent B2
US 12,568,096 · App. 18/582,898 · Granted Mar 3, 2026

Systems and methods for structural similarity based hashing

Inventors: Sandeep Paul (Bangalore, IN); Deepen Desai (San Ramon, CA)
Assignee: Zscaler, Inc.
H04L63/1416H04L63/1425H04L63/145
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,568,096
App. No.
18/582,898
Granted
Mar 3, 2026
Kind
B2
Abstract

Systems and methods for structural similarity based hash for sample identification and detection include, monitoring traffic associated with a cloud-based system; identifying a unique file within the traffic and computing a Structural Similarity Hash (SSHash) for the file, wherein the SSHash is based on auxiliary information and a complexity of the file; identifying one or more similar files based on the SSHash; and defining the file as belonging to one or more groups based on the one or more similar files.

Claims (42)

1 . A non-transitory computer-readable medium having instructions stored thereon for programming one or more processors to perform steps of:

monitoring traffic associated with a cloud-based system;

identifying a unique file within the traffic and computing a Structural Similarity Hash (SSHash) for the file, wherein computing the SSHash comprises splitting the filed into a plurality of chunks, computing a complexity of each of the plurality of chunks, rounding the computed complexities to reduce noise, collecting auxiliary information including at least a file type, chunk size, and number of chunks, and generating the SSHash based on both (i) the distribution density of the computed complexities across the plurality of chunks and (ii) the auxiliary information;

identifying one or more similar files based on the SSHash by referencing a database of SSHashes, the database being populated through continuous learning including enrichment from sandbox verdicts of unknown files and curated sets of known clean and known malicious files; and

defining the file as belonging to one of multiple groups comprising at least malware, clean, or unknown based on the one or more similar files.

2 . The non-transitory computer-readable medium of claim 1 , wherein computing the SSHash further comprises:

splitting the file into one or more equal length chunks;

computing a complexity of each of the one or more equal length chunks, including determining a Kolmogorov-style complexity measure of each chunk;

rounding the computed complexity values to mitigate noise and calculating a distribution density of the complexity values across the chunks;

collecting auxiliary information from the file; and

generating the SSHash based on the auxiliary information and the distribution density of the complexity of the file.

3 . The non-transitory computer-readable medium of claim 1 , wherein responsive to determining that no similar files exist, sending the unique file to a sandbox for analysis.

4 . The non-transitory computer-readable medium of claim 1 , wherein identifying one or more similar files includes referencing a SSHash database of files having known groupings.

5 . The non-transitory computer-readable medium of claim 4 , wherein the database is managed and persisted in a Central Database Management System (CDMS) of the cloud-based system.

6 . The non-transitory computer-readable medium of claim 5 , wherein the monitoring and computing are performed via one or more nodes of the cloud-based system communicatively coupled to the CDMS, and wherein the one or more nodes of the cloud-based system are adapted to send query requests to the CDMS for identifying one or more similar files.

7 . The non-transitory computer-readable medium of claim 1 , wherein the steps further comprise:

performing an action based on the defining, wherein the action includes one of blocking the traffic and allowing the traffic.

8 . The non-transitory computer-readable medium of claim 1 , wherein the steps comprise:

building an SSHash database prior to the monitoring.

9 . The non-transitory computer-readable medium of claim 8 , wherein building the SSHash database includes sending a plurality of sample files to a sandbox for grouping the plurality of sample files.

10 . The non-transitory computer-readable medium of claim 1 , wherein the unique file is any of a P file, a Portable Executable (PE) file, a shortcut (LNK) file, a Portable Document Format (PDF) file, an Executable Linkable Format (ELF) file, Android Package Kit (APK) file, Hypertext Markup Language (HTML) file.

11 . A method comprising steps of:

monitoring traffic associated with a cloud-based system;

identifying a unique file within the traffic and computing a Structural Similarity Hash (SSHash) for the file, wherein computing the SSHash comprises splitty the file into a plurality of chunks, computing a complexity of each of the plurality of chunks, rounding the computed complexities to reduce noise, collecting auxiliary information including at least a file type, chunk size, and number of chunks, and generating the SSHash based on both (i) the distribution density of the computed complexities across the plurality of chunks and (ii) the auxiliary information;

identifying one or more similar files based on the SSHash by referencing a database of SSHashes, the database being populated through continuous learning including enrichment from sandbox verdicts of unknown files and curated sets of known clean and known malicious files; and

defining the file as belonging to one of multiple groups comprising at least malware, clean, or unknown based on the one or more similar files.

12 . The method of claim 11 , wherein computing the SSHash further comprises:

splitting the file into one or more equal length chunks;

computing a complexity of each of the one or more equal length chunks, including determining a Kolmogorov-style complexity measure of each chunk;

rounding the computed complexity values to mitigate noise and calculating a distribution density of the complexity values across the chunks;

collecting auxiliary information from the file; and

generating the SSHash based on the auxiliary information and the distribution density of the complexity of the file.

13 . The method of claim 11 , wherein responsive to determining that no similar files exist, sending the unique file to a sandbox for analysis.

14 . The method of claim 11 , wherein identifying one or more similar files includes referencing a SSHash database of files having known groupings.

15 . The method of claim 14 , wherein the database is managed and persisted in a Central Database Management System (CDMS) of the cloud-based system.

16 . The method of claim 15 , wherein the monitoring and computing are performed via one or more nodes of the cloud-based system communicatively coupled to the CDMS, and wherein the one or more nodes of the cloud-based system are adapted to send query requests to the CDMS for identifying one or more similar files.

17 . The method of claim 11 , wherein the steps further comprise:

performing an action based on the defining, wherein the action includes one of blocking the traffic and allowing the traffic.

18 . The method of claim 11 , wherein the steps comprise:

building an SSHash database prior to the monitoring.

19 . The method of claim 18 , wherein building the SSHash database includes sending a plurality of sample files to a sandbox for grouping the plurality of sample files.

20 . The method of claim 11 , wherein the unique file is any of a P file, a Portable Executable (PE) file, a shortcut (LNK) file, a Portable Document Format (PDF) file, an Executable Linkable Format (ELF) file, Android Package Kit (APK) file, Hypertext Markup Language (HTML) file.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Feb 21, 2024
From: PAUL, SANDEEP; DESAI, DEEPEN
To: ZSCALER, INC.
Reel/Frame 066511/0720 →
Priority Claims (1)
IN 202441001556 · Jan 9, 2024 · national
Continuity (1)
Related Publication 20250227116A1 · Jul 10, 2025
References Cited (32)
US 8180916B1 · Nucci · 2012 [cited by examiner]
US 10142362B2 · Weith et al. · 2018 [cited by applicant]
US 10419477B2 · Desai et al. · 2019 [cited by applicant]
US 10904274B2 · Weith et al. · 2021 [cited by applicant]
US 11627148B2 · Desai · 2023 [cited by applicant]
US 11799876B2 · Desai et al. · 2023 [cited by applicant]
US 20070195779A1 · Judge · 2007 [cited by examiner]
US 20080256230A1 · Handley · 2008 [cited by examiner]
US 20100042846A1 · Trotter · 2010 [cited by examiner]
US 20170359220A1 · Weith et al. · 2017 [cited by applicant]
US 20190171665A1 · Navlakha · 2019 [cited by examiner]
US 20210192043A1 · Bhary et al. · 2021 [cited by applicant]
US 20210344693A1 · Azad et al. · 2021 [cited by applicant]
US 20210374121A1 · Paul · 2021 [cited by examiner]
US 20210377301A1 · Desai et al. · 2021 [cited by applicant]
US 20210377303A1 · Bui et al. · 2021 [cited by applicant]
US 20210377304A1 · Ma et al. · 2021 [cited by applicant]
US 20220083661A1 · Ma et al. · 2022 [cited by applicant]
US 20220253430A1 · Paul · 2022 [cited by examiner]
US 20230164182A1 · Kothari et al. · 2023 [cited by applicant]
US 20230164183A1 · Kothari et al. · 2023 [cited by applicant]
US 20230164184A1 · Kothari et al. · 2023 [cited by applicant]
US 20230353587A1 · Bui et al. · 2023 [cited by applicant]
US 20230370495A1 · Zscaler · 2023 [cited by applicant]
US 20230376592A1 · Ma et al. · 2023 [cited by applicant]
US 20240028707A1 · Paul et al. · 2024 [cited by applicant]
US 20240028721A1 · Ma et al. · 2024 [cited by applicant]
US 20240039954A1 · Shete et al. · 2024 [cited by applicant]
Sadowski, Caitlin, and Greg Levin. “Simhash: Hash-based similarity detection.” Dec. 13, 2007 (Year: 2007). [cited by examiner]
NPL Search Terms (Year: 2025). [cited by examiner]
Namanya, Anitta Patience, et al. “Similarity hash based scoring of portable executable files for efficient malware detection in IoT.” Future Generation Computer Systems 110 (2020): 824-832. (Year: 2020). [cited by examiner]
Georg Wicherski, peHash: a novel approach to fast malware clustering, in: LEET'09 Proceedings of the 2nd USENIX Conference on Large-Scale Exploits and Emergent Threats: Botnets, Spyware, Worms, and More, 2009 (Year: 200… [cited by examiner]