IP Library › Granted Patent US 12,481,788
Granted Patent B2
US 12,481,788 · App. 17/741,521 · Granted Nov 25, 2025

Method to randomize online activity

Inventors: Christian Garcia-Arellano (Richmond Hill, CA); Matthias Seul (Pleasant Hill, CA); Mehran Khan (Ottawa, CA); Daniel Silveira (Toronto, CA); Zvonimir Fras (Markham, CA)
Assignee: International Business Machines Corporation
G06F21/6263G06F21/6254G06N3/08G06N20/00
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 12,481,788
App. No.
17/741,521
Granted
Nov 25, 2025
Kind
B2
Abstract

A computer-implemented method for digital fingerprint obfuscation is disclosed. The computer-implemented method includes training a machine learning model to classify web traffic data into one or more personas. The computer-implemented method further includes identifying, using the trained machine learning model, a particular persona of a user based, at least in part, on a user's real traffic data generated during a current user session. The computer-implemented method further includes generating synthetic traffic data based, at least in part, on the identified particular persona of the user.

Claims (61)

1 . A computer-implemented method for digital fingerprint obfuscation, the computer-implemented method comprising:

sanitizing training data by stripping the training data of information that identifies source individuals, the source of training data being online user activity or synthetically generated online user activity;

training a machine learning model using the training data to classify a user's real traffic data into one or more personas;

collecting and marking a user's real traffic data to run classification algorithms to generate a base-persona, the base persona being main attributes connected to the user, the base persona having similar traits, likes, or dislikes from the user's real traffic data;

identifying, using the trained machine learning model, a particular persona of a user based, at least in part, on a user's real traffic data generated during a current user session;

generating synthetic traffic data based, at least in part, on a different persona than the particular persona of the user identified using the trained machine learning model; and

combining the user's real traffic data of the particular persona with the synthetic traffic data of the different persona to generate combined traffic data and plotting the combined traffic data as a traffic leveling graph to generate a visualization of the flattening of the user's traffic patterns.

2 . The computer-implemented method of claim 1 , further comprising:

transmitting the generated synthetic traffic data after the current user session has ended.

3 . The computer-implemented method of claim 2 , wherein transmitting the generated synthetic traffic data further comprises:

simultaneously transmitting the generated synthetic traffic data simultaneously with the real user traffic data generated during the current user session.

4 . The computer-implemented method of claim 1 , further comprising:

performing traffic leveling during the current user session, wherein performing traffic leveling includes transmitting the generated synthetic traffic data that is complementary to the user traffic data.

5 . The computer-implemented method of claim 4 , wherein traffic leveling is further based, at least in part, on an amount and persona of generated synthetic traffic data which is based, at least in part, on real traffic data.

6 . The computer-implemented method of claim 1 , wherein generating synthetic traffic is further based, at least in part, on:

generating a first score for the identified particular persona of the user;

determining a second persona and a second score for the second persona with a difference in the first score and the second score being above a predetermined threshold; and

generating synthetic traffic data based, at least in part, on the second persona.

7 . The computer-implemented method of claim 1 , further comprising:

generating a first score for the identified particular persona of the user;

determining a second persona and a second score for the second persona with a difference in the first score and the second score being below a predetermined threshold; and

generating synthetic traffic data based, at least in part, on the second persona.

8 . The computer-implemented method of claim 1 , wherein training the machine learning model to classify web traffic data into one or more personas is further based, at least in part on one or more of: active time of day of a user's web browsing history, web pages visited, web browsers accessed, applications accessed, and local contents present in the users' web browser or computer.

9 . The computer-implemented method of claim 1 , wherein generating synthetic traffic data is based, at least in part, on anonymizing a respective traffic pattern by traffic leveling using the synthetic traffic data to create a complimentary frequency of automated traffic to level an original traffic volume curve of a respective user plotted against time vs volume.

10 . The computer-implemented method of claim 1 , wherein generating synthetic traffic data is further based, at least in part on:

generating synthetic traffic data using the identified particular persona of the user to form a curated web surfing pattern based on a mixture of universal resource locators (URLs) accessed by users having different associated personas from the particular persona of the user.

11 . A computer program product for digital fingerprint obfuscation, the computer program product comprising one or more computer readable storage media and program instructions stored on the one or more computer readable storage media, the program instructions including instructions to:

sanitize training data by stripping the training data of information that identifies source individuals, the source of training data being online user activity or synthetically generated online user activity;

train a machine learning model using the training data to classify web a user's real traffic data into one or more personas;

collect and mark a user's real traffic data to run classification algorithms to generate a base-persona, the base persona being main attributes connected to the user, the base persona having similar traits, likes, or dislikes from the user's real traffic data;

identify, using the trained machine learning model, a particular persona of a user based, at least in part, on a user's real traffic data generated during a current user session;

generate synthetic traffic data based, at least in part, on a different persona than the particular persona of the user identified using the trained machine learning model; and

combine the user's real traffic data of the particular persona with the synthetic traffic data of the different persona to generate combined traffic data and plot the combined traffic data as a traffic leveling graph to generate a visualization of the flattening of the user's traffic patterns.

12 . The computer program product of claim 11 , further comprising instructions to:

transmit the generated synthetic traffic data after the current user session has ended.

13 . The computer program product of claim 12 , wherein the instructions to transmit the generated synthetic traffic data further comprises instructions to:

simultaneously transmit the generated synthetic traffic data simultaneously with the real user traffic data generated during the current user session.

14 . The computer program product of claim 11 , further comprising instructions to:

perform traffic leveling during the current user session, wherein performing traffic leveling includes transmitting the generated synthetic traffic data that is complementary to the user traffic data.

15 . The computer program product of claim 14 , wherein traffic leveling is further based, at least in part, on an amount and persona of generated synthetic traffic data which is based, at least in part, on real traffic data.

16 . The computer program product of claim 11 , wherein generating synthetic traffic is further based, at least in part, on instructions to:

generate a first score for the identified particular persona of the user;

determine a second persona and a second score for the second persona with a difference in the first score and the second score being above a predetermined threshold; and

generate synthetic traffic data based, at least in part, on the second persona.

17 . The computer program product of claim 11 , further comprising instructions to:

generate a first score for the identified particular persona of the user;

determine a second persona and a second score for the second persona with a difference in the first score and the second score being below a predetermined threshold; and

generate synthetic traffic data based, at least in part, on the second persona.

18 . The computer program product of claim 11 , wherein training the machine learning model to classify web traffic data into one or more personas is further based, at least in part on one or more of: active time of day of a user's web browsing history, web pages visited, web browsers accessed, applications accessed, and local contents present in the users' web browser or computer.

19 . The computer program product of claim 11 , wherein generating synthetic traffic data is based, at least in part, on anonymizing a respective traffic pattern by traffic leveling using the synthetic traffic data to create a complimentary frequency of automated traffic to level an original traffic volume curve of a respective user plotted against time vs volume.

20 . A computer system for digital fingerprint obfuscation, comprising:

one or more computer processors;

one or more computer readable storage media;

computer program instructions;

the computer program instructions being stored on the one or more computer readable storage media for execution by the one or more computer processors; and the computer program instructions including instructions to:

sanitize training data by stripping the training data of information that identifies source individuals, the source of training data being online user activity or synthetically generated online user activity;

train a machine learning model using the training data to classify a user's real traffic data into one or more personas

collect and mark a user's real traffic data to run classification algorithms to generate a base-persona, the base persona being main attributes connected to the user, the base persona having similar traits, likes, or dislikes from the user's real traffic data;

identify, using the trained machine learning model, a particular persona of a user based, at least in part, on a user's real traffic data generated during a current user session;

generate synthetic traffic data based, at least in part, on a different persona than the particular persona of the user identified using the trained machine learning model; and

combine the user's real traffic data of the particular persona with the synthetic traffic data of the different persona to generate combined traffic data and plot the combined traffic data as a traffic leveling graph to generate a visualization of the flattening of the user's traffic patterns.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded May 11, 2022
From: GARCIA-ARELLANO, CHRISTIAN; SEUL, MATTHIAS; KHAN, MEHRAN; SILVEIRA, DANIEL; FRAS, ZVONIMIR
To: INTERNATIONAL BUSINESS MACHINES CORPORATION
Reel/Frame 059888/0060 →
Continuity (1)
Related Publication 20230367855A1 · Nov 16, 2023
References Cited (41)
US 8468271B1 · Panwar · 2013 [cited by applicant]
US 10332108B2 · Ciurea · 2019 [cited by applicant]
US 10395017B2 · Bender · 2019 [cited by applicant]
US 12039083B1 · Fedele · 2024 [cited by examiner]
US 20110208717A1 · Pierce · 2011 [cited by examiner]
US 20120297017A1 · Livshits · 2012 [cited by applicant]
US 20130282759A1 · Proux · 2013 [cited by examiner]
US 20140287723A1 · Lafever · 2014 [cited by applicant]
US 20150128285A1 · Lafever · 2015 [cited by applicant]
US 20150128287A1 · Lafever · 2015 [cited by applicant]
US 20180121677A1 · Avancha · 2018 [cited by examiner]
US 20210150269A1 · Choudhury · 2021 [cited by examiner]
US 20210232706A1 · Peruski · 2021 [cited by examiner]
EP 3063691B1 · 2020 [cited by applicant]
Ahmad, Wasi Uddin, Kai-Wei Chang, and Hongning Wang. “Intent-aware query obfuscation for privacy protection in personalized web search.” The 41st international ACM SIGIR conference on research & development in informati… [cited by examiner]
Ahmad, Wasi Uddin, Md Masudur Rahman, and Hongning Wang. “Topic model based privacy protection in personalized web search.” Proceedings of the 39th International ACM SIGIR conference on Research and Development in Infor… [cited by examiner]
Sánchez, David, Jordi Castellà-Roca, and Alexandre Viejo. “Knowledge-based scheme to create privacy-preserving but semantically-related queries for web search engines.” Information Sciences 218 (2013): 17-30. (Year: 201… [cited by examiner]
Fröbe, Maik, Eric Oliver Schmidt, and Matthias Hagen. “Efficient query obfuscation with keyqueries.” IEEE/WIC/ACM International Conference on Web Intelligence and Intelligent Agent Technology. 2021. (Year: 2021). [cited by examiner]
Wu, Zongda, et al. “Constructing plausible innocuous pseudo queries to protect user query intention.” Information Sciences 325 (2015): 215-226. (Year: 2015). [cited by examiner]
Yu, Puxuan, Wasi Uddin Ahmad, and Hongning Wang. “Hide-n-seek: An intent-aware privacy protection plugin for personalized web search.” The 41st International ACM SIGIR Conference on Research & Development in Information… [cited by examiner]
“Privacy Badger”, downloaded from the Internet on Mar. 3, 2022, 15 pages, <https://privacybadger.org/>. [cited by applicant]
“Privacy Possum”, downloaded from the Internet on Mar. 3, 2022, 6 pages, <https://github.com/cowlicks/privacypossum>. [cited by applicant]
“Unified ID 2.0”, Industry Initiatives, theTradeDesk, downloaded from the Internet on Mar. 7, 2022, 9 pages, <https://www.thetradedesk.com/us/about-us/industry-initiatives/unified-id-solution-2-0>. [cited by applicant]
“What is a digital footprint? And how to help protect it from prying eyes”, Oct. 15, 2018, <https://us.norton.com/internetsecurity-privacy-clean-up-online-digital-footprint.html>, 5 pages. [cited by applicant]
“Your ISP Is Tracking Every Website You Visit: Here's What We Know”, Last updated on Nov. 15, 2021 by PrivacyPolicies.com Legal Writing Team, 11 pages, <https://www.privacypolicies.com/blog/isp-tracking-you/>. [cited by applicant]
Mell et al., “The NIST Definition of Cloud Computing”, Recommendations of the National Institute of Standards and Technology, NIST Special Publication 800-145, Sep. 2011, 7 pages. [cited by applicant]
Nair, Lenin VJ, “How to Track Link Clicks on Your Website Using Google Analytics & Google Tag Manager,” zoomowl, 2018, (Online), Available: https://www.zoomowl.com/link-click-tracking-using-googleanalytics/, (Accessed N… [cited by applicant]
Rainie, et al., “Anonymity, Privacy, and Security Online”, Pew Research Center; Sep. 5, 2013, 8 pages, <https://www.pewresearch.org/internet/2013/09/05/anonymity-privacy-and-security-online/>. [cited by applicant]
Xia et al., “Mosaic: Quantifying Privacy Leakage in Mobile Networks”, SIGCOMM'13, Aug. 12-16, 2013, Hong Kong, China, 12 pages. [cited by applicant]
Balaban, “DNS Queries and Their Anonymity”, Hackernoon, Aug. 28, 2018, 9 pages, https://hackernoon.com/dns-queries-and-their-anonymity-70cc82fbc60a. [cited by applicant]
Bhalerao et al., “A Survey on User Navigation Pattern Prediction from Web Log Data”, IJCSE International Journal of Computer Sciences and Engineering, vol. 3, Issue-5, May 30, 2015, pp. 133-137. [cited by applicant]
Brave Help Center, “What is “Shields”?”, accessed on Jun. 3, 2024, 3 pages, https://support.brave.com/hc/en-us/articles/360022973471-What-is-Shields. [cited by applicant]
Davies, “WTF are shared identity solutions?”, Digiday, Sep. 23, 2019, 6 pages, https://digiday.com/media/what-are-shared-identity-solutions-and-can-they-really-replace-cookies/. [cited by applicant]
Hasan et al., “Learning and Predicting Key Web Navigation Patterns Using Bayesian Models”, ICCSA 2009, Part II, LNCS 5593, pp. 877-887, 2009. [cited by applicant]
Nield, “It's Time to Switch to a Privacy Browser”, Wired, Apr. 6, 2024, 10 pages. [cited by applicant]
PhantomJS, “Scriptable Headless Browser”, accessed Jun. 3, 2024, 1 page, https://phantomjs.org/. [cited by applicant]
PhantomJS, “Who's using PhantomJS?”, accessed on Jun. 3, 2024, 2 pages, https://phantomjs.org/users.html. [cited by applicant]
Privacy Policies, “Your ISP Is Tracking Every Website You Visit: Here's What We Know”, Blog, Last updated on Jul. 1, 2022, 11 pages, https://www.privacypolicies.com/blog/isp-tracking-you/. [cited by applicant]
ScrambleSuit, “A Polymorphic Network Protocol to Circumvent Censorship”, Last change: Apr. 11, 2016, 1 page. [cited by applicant]
Wikibooks, “Intellectual Property and the Internet/Anonymizers”, accessed on Jun. 3, 2024, 4 pages, https://en.wikibooks.org/wiki/Intellectual_Property_and_the_Internet/Anonymizers. [cited by applicant]
Zoomowl, “How to Track Link Clicks on Your Website Using Google Analytics & Google Tag Manager”, Feb. 17, 2018, 31 pages, https://www.zoomowl.com/link-click-tracking-using-googleanalytics/. [cited by applicant]