IP Library Granted Patent US 11,537,751
Granted Patent B2
US 11,537,751 · App. 17/077,881 · Granted Dec 27, 2022

Using machine learning algorithm to ascertain network devices used with anonymous identifiers

Inventors: Rami Al-Kabra (Bothell, WA); Douglas Galagate (Bellevue, WA); Eric Yatskowitz (Seattle, WA); Chuong Phan (Seattle, WA); Tatiana Dashevskiy (Edmonds, WA); Prem Kumar Bodiga (Bothell, WA); Noah Dahlstrom (Seattle, WA); Ruchir Sinha (Newcastle, WA); Jonathan Morrow (Issaquah, WA); Aaron Drake (Sammamish, WA)
Assignee: T-Mobile USA, Inc.
G06F21/6254H04L63/0421H04L67/02H04L67/535H04L69/22H04W12/02G06Q30/0267G06Q30/0277H04L63/168
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,537,751
App. No.
17/077,881
Granted
Dec 27, 2022
Kind
B2
Abstract

Techniques for identifying certain types of network activity are disclosed, including parsing of a Uniform Resource Locator (URL) to identify a plurality of key-value pairs in a query string of the URL. The plurality of key-value pairs may include one or more potential anonymous identifiers. In an example embodiment, a machine learning algorithm is trained on the URL to determine whether the one or more potential anonymous identifiers are actual anonymous identifiers (i.e., advertising identifiers) that provide advertisers a method to identify a user device without using, for example, a permanent device identifier. In this embodiment, a ranking threshold is used to verify the URL. A verified URL associate the one or more potential anonymous identifiers with the user device as actual anonymous identifiers. Such techniques may be used to identify and eliminate malicious and/or undesirable network traffic.

Claims (37)

1. One or more computer-readable storage media storing computer-executable instructions that upon execution cause one or more processors to perform acts comprising:

monitoring a plurality of uniform resource locators (URLs) by a cellular device over a cellular network;

extracting a query string of a monitored URL;

extracting a plurality of key-value pairs of the query string;

training an algorithm upon an extracted plurality of key-value pairs to generate a number of votes;

ranking the URL based upon the number of votes;

utilizing a ranking threshold to verify the URL, wherein the plurality of key-value pairs of a verified URL is associated with a device identifier of the cellular device; and

storing the plurality of key-value pairs of the verified URL and the associated device identifier.

2. The one or more computer-readable storage media of claim 1 , wherein the plurality of key-value pairs includes at least one key variable that is paired with a value-character string that matches a format of an anonymous identifier.

3. The one or more computer-readable storage media of claim 2 , wherein the value-character string of the verified URL is an actual anonymous identifier.

4. The one or more computer-readable storage media of claim 1 , wherein the plurality of key-value pairs includes at least one key variable that is paired with a value-character string having a format that is different from a format of an anonymous identifier.

5. The one or more computer-readable storage media of claim 4 , wherein the algorithm utilizes the value-character string as a feature for a verification of the URL.

6. The one or more computer-readable storage media of claim 1 , wherein the algorithm includes a Random Forest algorithm that utilizes the plurality of key-value pairs as a feature for a verification of the URL.

7. The one or more computer-readable storage media of claim 1 further comprising: creating the algorithm from dataset URLs that are verified using manual evaluation of each key-value pair.

8. The one or more computer-readable storage media of claim 7 , wherein the dataset URLs are used as training data for the algorithm.

9. The one or more computer-readable storage media of claim 7 , wherein the manual evaluation utilizes stored key variables and device identifiers in a database.

10. A device, comprising:

a processor;

an endpoint detection and response (EDR) mechanism coupled to the processor, the EDR mechanism further comprises:

a uniform resource locator (URL) parser configured to monitor a plurality of URLs by a cellular device over a cellular network, and extract a plurality of key-value pairs in a query string of a monitored URL;

a URL verifier configured to: train an algorithm upon an extracted plurality of key-value pairs to generate a number of votes; rank the number of votes; and use a ranking threshold to verify the URL, wherein the plurality of key-value pairs of a verified URL is associated with a device identifier of the cellular device; and

a key-value pair database that stores the plurality of key-value pairs of the verified URL and the associated device identifier.

11. The device of claim 10 , wherein the plurality of key-value pairs includes at least one key variable that is paired with a value-character string that matches a format of an anonymous identifier.

12. The device of claim 11 , wherein the value-character string of a verified URL is an actual anonymous identifier.

13. The device of claim 10 , wherein the plurality of key-value pairs includes at least one key variable that is paired with a value-character string having a format that is different from a format of an anonymous identifier.

14. The device of claim 13 , wherein the algorithm utilizes the value-character string as a feature for a verification of the URL.

15. The device of claim 10 , wherein the algorithm includes a Random Forest algorithm that utilizes the plurality of key-value pairs as a feature for a verification of the URL.

16. The device of claim 10 , wherein the URL parser further extracts a domain name of the monitored URL, wherein the domain name is used as a feature by the algorithm.

17. The device of claim 10 , wherein the algorithm is created from dataset URLs that are verified using manual evaluation of key-value pairs.

18. A computer-implemented method, comprising:

monitoring a plurality of uniform resource locators (URLs) by a cellular device over a cellular network;

extracting a plurality of key-value pairs in a query string of a monitored URL, wherein the key-value pairs include one or more potential anonymous identifiers;

training a Random Forest algorithm upon an extracted plurality of key-value pairs to verify the URL, wherein a verified URL associates the one or more potential anonymous identifiers as one or more actual anonymous identifiers to a device identifier; and

storing the one or more actual anonymous identifiers and the device identifier.

19. The computer-implemented method of claim 18 further comprising:

creating the Random Forest algorithm from dataset URLs that are verified using manual evaluation of key-value pairs.

20. The computer-implemented method of claim 19 , wherein the manual evaluation utilizes stored key variables and device identifiers in a database.

Assignments (1)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 22, 2020
From: AL-KABRA, RAMI; GALAGATE, DOUGLAS; YATSKOWITZ, ERIC; PHAN, CHUONG; DASHEVSKIY, TATIANA; BODIGA, PREM KUMAR; DAHLSTROM, NOAH; SINHA, RUCHIR; MORROW, JONATHAN; DRAKE, AARON
To: T-MOBILE USA, INC.
Reel/Frame 054144/0037 →
Continuity (3)
Continuation In Part 16932491 · Jul 17, 2020
Continuation 15801971 · Nov 2, 2017
Related Publication 20210042442A1 · Feb 11, 2021