IP Library Granted Patent US 8,489,605
Granted Patent B2
US 8,489,605 · App. 13/167,170 · Granted Jul 16, 2013

Document object model (DOM) based page uniqueness detection

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,489,605
App. No.
13/167,170
Granted
Jul 16, 2013
Kind
B2
Abstract

DOM based unique ID generation, including receiving a hypertext markup language (HTML) page at a computer, and identifying HTML page elements in response to the receiving, the HTML page elements comprising parent nodes, the parent nodes comprising child nodes. The method further comprising processing each of the HTML page elements, the processing comprising: grouping the child nodes by parent node into a group of child nodes, detecting patterns in the group of child nodes in response to the grouping, reducing the group of child nodes to text strings in response to the detecting, storing the text strings as text values in the parent nodes, and generating a unique identifier (ID) of the HTML page in response to the processing.

Claims (61)

1. A non-transitory computer system comprising:

a host system in communication with at least one client system over a network; a page based unique ID generation application for execution on the host system, the page based for unique ID generation application including logic for implementing a method comprising:

receiving a hypertext markup language (HTML) page at a computer;

identifying HTML page elements in response to the receiving, the HTML page elements comprising parent nodes, the parent nodes comprising child nodes;

processing each of the HTML page elements, the processing comprising:

grouping the child nodes by parent node into a group of child nodes; detecting patterns in the group of child nodes in response to the grouping; reducing the group of child nodes to text strings in response to the detecting; and storing the text strings as text values in the parent nodes; and generating a unique identifier (ID) of the HTML page in response to the processing;

wherein the HTML page is a Web 2.0 page, the Web 2.0 page comprising content, the content being generated dynamically and filtering HTML page elements in response to the identifying, the filtering removing the child nodes and the parent nodes that meet filter criteria, the filter criteria comprising:

extensible markup language path language instructions;

regular expression (regex) instructions; and

a list of html nodes.

2. The system of claim 1 , wherein the processing further comprising:

sorting the group of child nodes in response to the reducing.

3. The system of claim 2 , wherein the HTML page comprises a visible page, the visible page comprising moveable HTML page elements, the moveable HTML page elements configured to occupy any location on the visible page, and the unique ID of the HTML page is the same for all locations of the moveable HTML page elements on the visible page.

4. The system of claim 1 , wherein the HTML page comprises a uniform resource locator (URL), and the unique ID of the HTML page is the same for a plurality of HTML pages, the plurality of HTML pages each comprising a different URL.

5. A computer program product comprising a non-transitory storage medium storing instructions, the computer program product implementing a method, the method comprising:

receiving a hypertext markup language (HTML) page at a computer;

identifying HTML page elements in response to the receiving, the HTML page elements comprising parent nodes, the parent nodes comprising child nodes;

processing each of the HTML page elements, the processing comprising:

grouping the child nodes by parent node into a group of child nodes;

detecting patterns in the group of child nodes in response to the grouping;

reducing the group of child nodes to text strings in response to the detecting;

and storing the text strings as text values in the parent nodes; and

generating a unique identifier (ID) of the HTML page in response to the processing;

wherein the HTML page is a Web 2.0 page, the Web 2.0 page comprising content, the content being generated dynamically and

filtering HTML page elements in response to the identifying, the filtering removing the child nodes and the parent nodes that meet filter criteria, the filter criteria comprising: extensible markup language path language instructions;

regular expression (regex) instructions; and

a list of html nodes.

6. The computer program product of claim 5 , wherein the processing further comprising:

sorting the group of child nodes in response to the reducing.

7. The computer program product of claim 6 , wherein the HTML page comprises a visible page, the visible page comprising moveable HTML page elements, the moveable HTML page elements configured to occupy any location on the visible page, and the unique ID of the HTML page is the same for all locations of the moveable HTML page elements on the visible page.

8. The computer program product of claim 5 , wherein the HTML page comprises a uniform resource locator (URL), and the unique ID of the HTML page is the same for a plurality of HTML pages, the plurality of HTML pages each comprising a different URL.

9. An apparatus comprising:

web indexing application logic communicating with a computer processor to receive a hypertext markup language (HTML) page at a computer and identify HTML page elements, wherein the HTML page elements comprising parent nodes, the parent nodes comprising child nodes, and process each of the HTML page elements, wherein

wherein the computer processor is configured for:

grouping the child nodes by parent node into a group of child nodes;

detecting patterns in the group of child nodes in response to the grouping;

reducing the group of child nodes to text strings in response to the detecting; and

storing the text strings as text values in the parent nodes; and the web indexing application logic further configured to generate a unique identifier (ID) of the HTML page in response to the processing;

wherein the HTML page is a Web 2.0 page, the Web 2.0 page comprising content, the content being generated dynamically and

wherein the web indexing application logic is further configured to filter HTML page elements in response to the identifying, the filtering removing the child nodes and the parent nodes that meet filter criteria, the filter criteria comprising:

extensible markup language path language instructions;

regular expression (reqex) instructions; and a list of html nodes.

10. The apparatus of claim 9 , wherein the processing further comprising:

sorting the group of child nodes in response to the reducing.

11. The apparatus of claim 10 , wherein the HTML page comprises a visible page, the visible page comprising moveable HTML page elements, the moveable HTML page elements configured to occupy any location on the visible page, and the web indexing application logic configured to generate the same unique ID of the HTML page for all locations of the moveable HTML page elements on the visible page.

12. The apparatus of claim 9 , wherein the HTML page comprises a uniform resource locator (URL), and the unique ID of the HTML page is the same for a plurality of HTML pages, the plurality of HTML pages each comprising a different URL.

13. A computing system comprising:

at least one processor;

at least one memory;

a network transceiver;

a bus communicatively linking said at least one processor, memory, and network transceiver to each other, wherein said at least one memory stores an executable instructions, which the at least one processor executes, wherein said network transceiver of the computing system comprising hardware is operable to receive a hypertext markup language (HTML) page from a server over a network, wherein the computing system is configured to function as a client in a client server arrangement with the server, wherein the processor executing the instructions is operable to identify HTML page elements in response to the receiving of the HTML page, the HTML page elements comprising parent nodes, the parent nodes comprising child nodes;

wherein the processor executing the instructions is operable to process each of the HTML page elements, the processing comprising:

grouping the child nodes by parent node into a group of child nodes;

detecting patterns in the group of child nodes in response to the grouping;

reducing the group of child nodes to text strings in response to the detecting; and

storing the text strings as text values in the parent nodes;

and wherein the processor executing the instructions is operable to generate a unique identifier (ID) of the HTML page in response to the processing;

wherein the HTML page is a Web 2.0 page, the Web 2.0 page comprising content, the content being generated dynamically and

filtering HTML page elements in response to the identifying, the filtering removing the child nodes and the parent nodes that meet filter criteria, the filter criteria comprising: extensible markup language path language instructions;

regular expression (regex) instructions; and

a list of html nodes.

Assignments (4)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2022
From: DAEDALUS BLUE LLC
To: TERRACE LICENSING LLC
Reel/Frame 058902/0482 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 28, 2022
From: DAEDALUS BLUE LLC
To: TERRACE LICENSING LLC
Reel/Frame 058895/0322 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jan 29, 2020
From: DAEDALUS GROUP, LLC
To: DAEDALUS BLUE LLC
Reel/Frame 051737/0191 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Nov 14, 2019
From: INTERNATIONAL BUSINESS MACHINES CORPORATION
To: DAEDALUS GROUP LLC
Reel/Frame 051018/0649 →