IP Library Granted Patent US 11,537,642
Granted Patent B2
US 11,537,642 · App. 16/007,029 · Granted Dec 27, 2022

Method and system for providing a user agent string database

Inventors: Ling Zhu (Beijing, CN); Min He (Beijing, CN); Fei Yu (Beijing, CN); Minzhang Wei (Beijing, CN)
Assignee: YAHOO ASSETS LLC
G06F16/313G06F16/335G06F16/35G06F16/358G06F16/90344G06N20/00H04L67/02
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,537,642
App. No.
16/007,029
Granted
Dec 27, 2022
Kind
B2
Abstract

Method, system, and programs for determining a keyword from user agent strings are disclosed. In one example, a plurality of user agent strings is received. The plurality of user agent strings is grouped into one or more clusters. The one or more clusters comprise a first cluster that includes two or more user agent strings. The two or more user agent strings in the first cluster are compared. Based on the comparing, a keyword is determined from the first cluster. The keyword represents a type of user agent information.

Claims (58)

1. A method, implemented on at least one computing device, each of which includes at least one processor, storage, and a communication platform connected to a network for recognizing user agent strings, the method comprising:

receiving a plurality of user agent strings in data traffic;

generating a rank of the plurality of user agent strings based on a number of times each of the plurality of user agent strings appeared in the data traffic, wherein the rank indicates popularity of each of the plurality of user agent strings;

selecting, based on the rank, multiple of the plurality of user agent strings;

comparing the multiple of the plurality of user agent strings; and

determining one or more keywords representing a type of user agent information based on a result from the comparing.

2. The method of claim 1 , further comprising:

grouping the plurality of user agent strings into one or more clusters, wherein the multiple of the plurality of user agent strings that are compared are from a same cluster.

3. The method of claim 2 , wherein the one or more clusters comprise a plurality of clusters, the method further comprises:

determining a distance between one or more pairs of clusters from the plurality of clusters;

determining that the distance between at least one of the one or more pairs of clusters is less than or equal to a threshold; and

generating a merged cluster for each of the at least one of the one or more pairs of clusters.

4. The method of claim 3 , wherein the threshold is at least one of predetermined or modified based on a machine learning model.

5. The method of claim 2 , further comprising:

generating a ranked list of the one or more clusters; and

determining a highest ranked cluster based on the ranked list, wherein the multiple of the plurality of user agent strings are selected from the highest ranked cluster.

6. The method of claim 1 , wherein the plurality of user agent strings comprise at least one of:

user agent strings obtained from a user agent string database that have had a detection failure; or

user agent strings unrecognized by a user agent string analyzing engine, wherein the one or more keywords are stored for analyzing subsequently received user agent strings.

7. The method of claim 1 , wherein determining the one or more keywords comprises:

determining the multiple user agent strings from a first cluster of user agent strings; and

determining that the multiple user agent strings comprise a keyword, the keyword being determined based on a longest common subsequence amongst the two or more user agent strings, and a remaining subsequence obtained in response to removing the longest common subsequence.

8. A system including at least one processor, storage, and a communication platform connected to a network for recognizing user agent strings, the system comprising:

a user agent receiver configured to receive a plurality of user agent strings in data traffic;

a count based ranking unit configured to generate a rank of the plurality of user agent strings based on a number of times each of the plurality of user agent strings appeared in the data traffic, wherein the rank indicates popularity of each of the plurality of user agent strings; and

a keyword extractor configured to:

select, based on the rank, multiple of the plurality of user agent strings;

compare the multiple of the plurality of user agent strings; and

determine one or more keywords representing a type of user agent information based on a result from the comparison.

9. The system of claim 8 , further comprising:

a user agent clustering unit configured to group the plurality of user agent strings into one or more clusters, wherein the multiple of the plurality of user agent strings that are compared are from a same cluster.

10. The system of claim 9 , wherein the one or more clusters comprise a plurality of clusters, the user agent clustering unit comprises:

a distance calculation unit configured to determine a distance between one or more pairs of clusters from the plurality of clusters;

a cluster merging determiner configured to determine that the distance between at least one of the one or more pairs of clusters is less than or equal to a threshold; and

a cluster merging unit configured to generate a merged cluster for each of the at least one of the one or more pairs of clusters.

11. The system of claim 10 , wherein the threshold is at least one of predetermined or modified based on a machine learning model.

12. The system of claim 9 , further comprising:

a cluster ranking unit configured to:

generate a ranked list of the one or more clusters, and

determine a highest ranked cluster based on the ranked list, wherein the multiple of the plurality of user agent strings are selected from the highest ranked cluster.

13. The system of claim 8 , wherein the plurality of user agent strings comprise at least one of:

user agent strings obtained from a user agent string database that have had a detection failure; or

user agent strings unrecognized by a user agent string analyzing engine, wherein the one or more keywords are stored for analyzing subsequently received user agent strings.

14. A non-transitory machine-readable medium comprising information for determining a keyword from user agent strings, wherein the information, when read by a machine, causes the machine to perform operations comprising:

receiving a plurality of user agent strings in data traffic;

generating a rank of the plurality of user agent strings based on a number of times each of the plurality of user agent strings appeared in the data traffic, wherein the rank indicates popularity of each of the plurality of user agent strings;

selecting, based on the rank, multiple of the plurality of user agent strings;

comparing the multiple of the plurality of user agent strings; and

determining one or more keywords representing a type of user agent information based on a result from the comparing.

15. The non-transitory machine-readable medium of claim 14 , wherein the operations further comprise:

grouping the plurality of user agent strings into one or more clusters, wherein the multiple of the plurality of user agent strings that are compared are from a same cluster.

16. The non-transitory machine-readable medium of claim 15 , wherein the one or more clusters comprise a plurality of clusters, the operations further comprise:

determining a distance between one or more pairs of clusters from the plurality of clusters;

determining that the distance between at least one of the one or more pairs of clusters is less than or equal to a threshold; and

generating a merged cluster for each of the at least one of the one or more pairs of clusters.

17. The non-transitory machine-readable medium of claim 15 , wherein the operations further comprise:

generating a ranked list of the one or more clusters; and

determining a highest ranked cluster based on the ranked list, wherein the multiple of the plurality of user agent strings are selected from the highest ranked cluster.

Assignments (6)
PATENT SECURITY AGREEMENT (FIRST LIEN) Recorded Sep 29, 2022
From: YAHOO ASSETS LLC
To: ROYAL BANK OF CANADA, AS COLLATERAL AGENT
Reel/Frame 061571/0773 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 16, 2021
From: YAHOO AD TECH LLC (FORMERLY VERIZON MEDIA INC.)
To: YAHOO ASSETS LLC
Reel/Frame 058982/0282 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 26, 2020
From: OATH INC.
To: VERIZON MEDIA INC.
Reel/Frame 054258/0635 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2018
From: ZHU, LING; HE, MIN; YU, FEI; WEI, MINZHANG
To: YAHOO! INC.
Reel/Frame 046070/0240 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2018
From: YAHOO! INC.
To: YAHOO HOLDINGS, INC.
Reel/Frame 046339/0380 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 13, 2018
From: YAHOO HOLDINGS, INC.
To: OATH INC.
Reel/Frame 046352/0431 →
Continuity (2)
Continuation 14410702
Related Publication 20180293297A1 · Oct 11, 2018