IP Library › Granted Patent US 12,393,727
Granted Patent B2
US 12,393,727 · App. 17/601,145 · Granted Aug 19, 2025

Distance preserving hash method

Inventors: Abdelrahman Ali Almahmoud (Abu Dhabi, AE); Ernesto Damiani (Abu Dhabi, AE); Hadi Otrok (Abu Dhabi, AE); Yousof Ali Alhammadi (Abu Dhabi, AE)
Assignees: KHALIFA UNIVERSITY OF SCIENCE AND TECHNOLOGY; BRITISH TELECOMMUNICATIONS PLC; EMIRATES TELECOMMUNICATIONS CORPORATION
G06F21/6254G06F18/22G06F21/602
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,393,727
App. No.
17/601,145
Granted
Aug 19, 2025
Kind
B2
Abstract

A computer-implemented method of preparing an anonymised dataset for use in data analytics. The method includes the steps of: (a) labelling elements of a dataset to be analysed according to a labelling scheme; (b) selecting one or more labelled elements of the dataset to be replaced with a distance preserving hash; and for each selected element: (c) partitioning a data plane including the selected element into a plurality of channels, each channel covering a different distance space of the data plane; (d) hashing, using a cryptographic hash, data associated with the channel of the data plane in which the selected element resides, to form the distance preserving hash; and (e) replacing the selected element with the distance preserving hash.

Claims (43)

1. A computer-implemented method of preparing an anonymised dataset for use in data analytics, the method including the steps of:

(a) labelling elements of a dataset to be analysed according to a labelling scheme;

(b) selecting a subsample from the dataset and deriving therefrom an accuracy threshold indicative of the distance between elements of data within the subsample;

(c) deriving, from the anonymised dataset, an estimated accuracy of distance measurement between elements of the anonymised dataset and comparing this estimated accuracy to the accuracy threshold;

(d) selecting one or more labelled elements of the dataset to be replaced with a distance preserving hash; and for each selected element:

(e) partitioning a data plane including the selected element into a plurality of channels, each channel covering a different distance space of the data plane;

(f) hashing, using a cryptographic hash, data associated with the channel of the data plane in which the selected element resides, to form the distance preserving hash; and

(g) replacing the selected element with the distance preserving hash.

2. The computer-implemented method of claim 1 , further comprising a step of sending, to an analytics server, the anonymised dataset including the distance preserving hash(es) once all of the selected elements have been replaced.

3. The computer-implemented method of claim 1 , wherein if the estimated accuracy is less than the accuracy threshold, the value of ω is increased and steps (e)-(f) are repeated.

4. The computer-implemented method of claim 1 , wherein the data associated with the channel of the data plane within which the selected element resides is a generalised value based on the content of the respective channel.

5. The computer-implemented method of claim 1 , wherein the data associated with the channel of the data plane in which the selected element resides is a channel identifier or code corresponding to the channel.

6. The computer-implemented method of claim 1 , wherein the labelling scheme is derived from a machine-learning classifier which operates over the dataset.

7. A computer-implemented method of analysing a dataset anonymised using the method of claim 1 , the method including the steps of:

receiving an anonymised dataset from each of a plurality of data source computers;

for each distance preserving hash, comparing the distance preserving hash with a respective distance preserving hash of another anonymised dataset; and

estimating, based on the comparison, a distance between one or more element(s) having the same label in each of the anonymised datasets.

8. The computer-implemented method of claim 7 , wherein each dimension of each multi-dimensional hash is compared with a respective dimension of the multi-dimensional hash of another anonymised dataset.

9. The computer-implemented method of claim 8 , wherein each of the dimensions of the multi-dimensional distance hashes is given a different weighting when the distance is estimated.

10. The computer-implemented method of claim 9 , wherein a higher weighting is given to a first dimension which is indicative of a smaller distance space than a lower weighting given to a second dimension which is indicative of a larger distance space.

11. The computer-implemented method of claim 1 , wherein the distance preserving hash is multi-dimensional distance preserving has having ω dimensions, each dimension being formed of C channels, each channel having a size γ, and wherein the steps (e)-(f) are repeated for each dimension of the multi-dimensional distance preserving hash.

12. The computer-implemented method of claim 11 , wherein ω is selected based on any one or more of: an accuracy metric, a privacy metric, and a distance metric.

13. The computer-implemented method of claim 11 , wherein ω is selected based on a quantization resolution metric, which is indicative of a distance from a given distance preserving hash to the associated selected element.

14. A system, including:

one or more data source computers; and

at least one analytics server;

wherein each of the one or more data source computers are configured to prepare an anonymised dataset for the analytics server by:

(a) labelling elements of a dataset to be analysed according to a labelling scheme;

(b) selecting a subsample from the dataset and deriving therefrom an accuracy threshold indicative of the distance between elements of data within the subsample;

(c) deriving, from the anonymised dataset, an estimated accuracy of distance measurement between elements of the anonymised dataset and comparing this estimated accuracy to the accuracy threshold;

(d) selecting one or more labelled elements of the dataset to be replaced with a distance preserving hash; and for each selected element:

(e) partitioning a data plane including the selected element into a plurality of channels, each channel covering a different distance space of the data plane;

(f) hashing, using a cryptographic hash, data associated with the channel of the data plane in which the selected element resides, to form the distance preserving hash;

(g) replacing the selected element with the distance preserving hash; and

(h) sending, to the analytics server, the anonymised dataset including the distance preserving hash once the contents of each selected element has been replaced.

15. A data source computer for preparing an anonymised dataset for an analytics server, the data source computer including memory and a processor, wherein the memory includes instructions which cause the processor to:

(a) label elements of a dataset to be analysed according to a labelling scheme;

(b) selecting a subsample from the dataset and deriving therefrom an accuracy threshold indicative of the distance between elements of data within the subsample;

(c) deriving, from the anonymised dataset, an estimated accuracy of distance measurement between elements of the anonymised dataset and comparing this estimated accuracy to the accuracy threshold;

(d) select one or more labelled elements of the dataset to be replaced with a distance preserving hash; and for each selected element:

(e) partition a data plane including the selected element into a plurality of channels, each channel covering a different distance space of the data plane;

(f) hash, using a cryptographic hash, data associated with the channel of the data plane in which the selected element resides, to form the distance preserving hash; and

(g) replace the selected element with the distance preserving hash.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 28, 2025
From: ALMAHMOUD, ABDELRAHMAN; DAMIANI, ERNESTO; OTROK, HADI; ALHAMMADI, YOUSOF ALI
To: KHALIFA UNIVERSITY OF SCIENCE AND TECHNOLOGY; BRITISH TELECOMMUNICATIONS PIC; EMIRATES TELECOMMUNICATIONS CORPORATION
Reel/Frame 072245/0651 →
Continuity (1)
Related Publication 20220215126A1 · Jul 7, 2022
References Cited (35)
US 9703967B1 · Kothari · 2017 [cited by examiner]
US 9722973B1 · Kothari · 2017 [cited by examiner]
US 9852311B1 · Kothari · 2017 [cited by examiner]
US 10181051B2 · Barday · 2019 [cited by examiner]
US 10318757B1 · Kenthapadi · 2019 [cited by examiner]
US 11893520B2 · Silberman · 2024 [cited by examiner]
US 11921888B2 · Ren · 2024 [cited by examiner]
US 20070156677A1 · Szabo · 2007 [cited by examiner]
US 20070233711A1 · Aggarwal · 2007 [cited by examiner]
US 20080215842A1 · Kerschbaum · 2008 [cited by examiner]
US 20110078143A1 · Aggarwal · 2011 [cited by examiner]
US 20110078779A1 · Liu · 2011 [cited by examiner]
US 20140059695A1 · Parecki · 2014 [cited by examiner]
US 20150339488A1 · Takahashi · 2015 [cited by examiner]
US 20160196577A1 · Reese · 2016 [cited by examiner]
US 20160379011A1 · Koike · 2016 [cited by examiner]
US 20180004976A1 · Davis · 2018 [cited by examiner]
US 20180276415A1 · Barday · 2018 [cited by examiner]
US 20190050465A1 · Khalil · 2019 [cited by examiner]
US 20190087604A1 · Antonatos · 2019 [cited by examiner]
US 20190258824A1 · Gkoulalas-Divanis · 2019 [cited by examiner]
US 20190260730A1 · Mainali · 2019 [cited by examiner]
US 20190279247A1 · Finken · 2019 [cited by examiner]
US 20190287027A1 · Franceschini · 2019 [cited by examiner]
US 20200012797A1 · Adir · 2020 [cited by examiner]
US 20200267936A1 · Tran · 2020 [cited by examiner]
Han et al.; An Anonymization Method to Improve Data Utility for Classification; 2017; retrieved from the Internet: URL https://link.springer.com/chapter/10.1007/978-3-319-69471-9_5; pp. 1-15 as printed. (Year: 2017). [cited by examiner]
Karapiperis et al., (2017), IEEE 33rd International Conference on Data Engineering, pp. 135-138. [cited by applicant]
Vatsalan et al., (2016), Journal of Biomedical Informatics 59, 285-298. [cited by applicant]
Lu et al. “Toward efficient and privacy-preserving computing in big data era,” IEEE Network, vol. 28, No. 4, pp. 46-50, Jul. 2014. [cited by applicant]
Patel et al., “Privacy Preserving Distributed K-Means Clustering in Malicious Model Using Zero Knowledge Proof,” in Distributed Computing and Internet Technology, ser. Lecture Notes in Computer Science, Springer, Berlin… [cited by applicant]
Akhter et al., “Privacy-Preserving Two-Party k-Means Clustering in Malicious Model,” in 2013 IEEE 37th Annual Computer Software and Applications Conference Workshops, Jul. 2013, pp. 121-126. [cited by applicant]
Xiong et al., “Predict: Privacy and Security Enhancing Dynamic Information Collection and Monitoring,” Procedia Computer Science, vol. 18 pp. 1979-1988, Jan. 2013. [cited by applicant]
Yakoubov, et al. “A survey of cryptographic approaches to securing big-data analytics in the cloud”, in 2014 IEEE High Performance Extreme Computing Conference (HPEC), Sep. 2014, pp. 1-6. [cited by applicant]
Gheid et al., “Efficient and Privacy-Preserving k-Means Clustering for Big Data Mining”, in 2016 IEEE Trustcom/BigDataSE/ISPA, Aug. 2016, pp. 791-798. [cited by applicant]