IP Library Granted Patent US 11,838,372
Granted Patent B2
US 11,838,372 · App. 18/093,980 · Granted Dec 5, 2023

URL normalization for rendering a service graph

Inventors: Gergely Danyi (San Jose, CA); Joseph Ari Ross (Redwood City, CA)
Assignee: SPLUNK Inc.
H04L67/146G06F16/906G06F16/9566
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 11,838,372
App. No.
18/093,980
Granted
Dec 5, 2023
Kind
B2
Abstract

A method of normalizing URLs associated with a real user session comprises extracting uniform resource locators (URLs) from ingested spans where at least a portion of the URLs comprise unique URL strings. The method also comprises decomposing each of the URLs into a sequence of tokens and grouping together subsets of related URLs. Also, the method comprises representing each subset of related URLs with a normalized URL string.

Claims (55)

1. A method of normalizing uniform resource locators (URLs) associated with a real user session, the method comprising:

extracting a plurality of URLs from a plurality of spans associated with the real user session;

decomposing each of the plurality of URLs into a sequence of tokens by generating a plurality of sequences of tokens;

analyzing each sequence of tokens of the plurality of sequences of tokens to determine retained tokens, the analyzing comprising identifying a number of matching locations between two or more sequence of tokens, and computing a similarity score based on the number of matching locations;

based on the retained tokens, grouping together subsets of related URLs from the plurality of URLs;

representing each subset of related URLs with a normalized URL string; and

rendering a display result for a graphical user interface, wherein each normalized URL string is displayed in an on-screen position adjacent to a node representing a respective subset of related URLs.

2. The method of claim 1 , further comprising:

extracting a new URL from a span; and

responsive to the extracting, determining a subset from the subsets of related URLs in which to group the new URL.

3. The method of claim 1 , wherein each normalized URL string comprises wildcard characters, wherein the wildcard characters are disposed in positions where URLs of a respective subset of related URLs differ from each other.

4. The method of claim 1 , further comprising:

of the subsets of related URLs, computing metrics for each subset by aggregating metrics for each URL of the subset.

5. The method of claim 1 , further comprising:

of the subsets of related URLs, aggregating metrics for each subset; and displaying the metrics on the graphical user interface.

6. The method of claim 1 , further comprising:

of the subsets of related URLs, computing a real-time series of metrics for each subset by aggregating a time series of metrics for each URL of the subset.

7. The method of claim 1 , further comprising:

of the subsets of related URLs, aggregating metrics for each subset; and

filtering the metrics by a user-selected dimension to generate a filtered result for display on the graphical user interface.

8. The method of claim 1 , further comprising:

grouping together URLs responsive to a determination that the similarity score for the URLs is higher than a predetermined threshold.

9. The method of claim 1 , wherein the similarity score is computed based on the number of matching locations, wherein each location in the sequence of tokens is assigned a weight based on the location with respect to the sequence of tokens.

10. The method of claim 1 , wherein the analyzing comprises determining frequently occurring tokens associated with the plurality of URLs, wherein the frequently occurring tokens comprise a subset of frequent tokens present in the plurality of URLs, wherein the subset of frequent tokens being a fixed percentage of frequent tokens present in the plurality of URLs.

11. The method of claim 10 , further comprising:

retaining one or more frequently occurring tokens from the frequently occurring tokens in each sequence of tokens; and

in each subset of related URLs, grouping together URLs, wherein URLs are grouped based on URLs having the one or more frequently occurring tokens occurring in a same position for the subset of related URLs.

12. The method of claim 1 , further comprising:

responsive to a node selection in the graphical user interface, displaying aggregated metrics associated with a respective subset of URLs in the graphical user interface.

13. The method of claim 1 , further comprising:

of the subset of related URLs, computing metrics for each subset of related URLs from the subsets of related URL by aggregating metrics for each URL in a respective subset of related URLs; and

triggering an alert associated with a respective subset of related URLs responsive to a determination that an aggregated metric associated with the respective subset of related URLs satisfies a user-specified threshold.

14. A non-transitory computer-readable medium having computer-readable program code embodied therein for causing a computer system to perform a method of normalizing uniform resource locators (URLs) associated with a real user session, the method comprising:

extracting a plurality of URLs from a plurality of spans associated with the real user session;

decomposing each of the plurality of URLs into a sequence of tokens by generating a plurality of sequences of tokens;

analyzing each sequence of tokens of the plurality of sequences of tokens to determine retained tokens, the analyzing comprising identifying a number of matching locations between two or more sequence of tokens, and computing a similarity score based on the number of matching locations;

based on the retained tokens, grouping together subsets of related URLs from the plurality of URLs;

representing each subset of related URLs with a normalized URL string; and

rendering a display result for a graphical user interface, wherein each normalized URL string is displayed in an on-screen position adjacent to a node representing a respective subset of related URLs.

15. The non-transitory computer-readable medium of claim 14 , wherein each normalized URL string comprises wildcard characters, wherein the wildcard characters are disposed in positions where URLs of a respective subset of related URLs differ from each other.

16. The non-transitory computer-readable medium of claim 14 , further comprising:

of the subsets of related URLs, aggregating metrics for each subset; and

filtering the metrics by a user-selected dimension to generate a filtered result for display on the graphical user interface.

17. The non-transitory computer-readable medium of claim 14 , further comprising:

grouping together URLs responsive to a determination that the similarity score for the URLs is higher than a predetermined threshold.

18. The non-transitory computer-readable medium of claim 14 , wherein the similarity score is computed based on the number of matching locations, wherein each location in the sequence of tokens is assigned a weight based on the location with respect to the sequence of tokens.

19. The non-transitory computer-readable medium of claim 14 , wherein the analyzing comprises determining frequently occurring tokens associated with the plurality of URLs, wherein the frequently occurring tokens comprise a subset of frequent tokens present in the plurality of URLs, wherein the subset of frequent tokens being a fixed percentage of frequent tokens present in the plurality of URLs.

20. A system for normalizing uniform resource locators (URLs) associated with a real user session, the system comprising:

a processing device communicatively coupled with a memory and configured to:

extract a plurality of URLs from a plurality of spans associated with the real user session;

decompose each of the plurality of URLs into a sequence of tokens by generating a plurality of sequences of tokens;

analyze each sequence of tokens of the plurality of sequences of tokens to determine retained tokens by identifying a number of matching locations between two or more sequence of tokens, and computing a similarity score based on the number of matching locations;

based on the retained tokens, group together subsets of related URLs from the plurality of URLs;

represent each subset of related URLs with a normalized URL string; and

render a display result for a graphical user interface, wherein each normalized URL string is displayed in an on-screen position adjacent to a node representing a respective subset of related URLs.

Assignments (3)
CHANGE OF NAME Recorded Jul 22, 2025
From: SPLUNK INC.
To: SPLUNK LLC
Reel/Frame 072170/0599 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jul 22, 2025
From: SPLUNK LLC
To: CISCO TECHNOLOGY, INC.
Reel/Frame 072173/0058 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Oct 24, 2023
From: DANYI, GERGELY; ROSS, JOSEPH ARI
To: SPLUNK INC.
Reel/Frame 065328/0211 →
Continuity (3)
Continuation 17245786 · Apr 30, 2021
Provisional Application 63175455 · Apr 15, 2021
Related Publication 20230156093A1 · May 18, 2023