IP Library Patent Application 13491547
Patent Application
App. No. 13/491,547

DETECTING ERROR PAGES BY ANALYZING SERVER REDIRECTS

Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US None
App. No.
13/491,547
Abstract

A system and method is disclosed for detecting invalid webpages by analyzing server redirects. A storage comprising a set of previously stored target addresses is queried to determine whether one or more of the set of previously stored target addresses result from a redirect initiated from more than a predetermined number of originating addresses. On determining that a target address resulted from a redirect initiated from more than the predetermined number of originating addresses, the originating addresses are analyzed to determine, for each address, a difference between information previously stored for the originating address and information associated with the respective target address. If the difference satisfies a predetermined threshold, the originating address is marked as not valid or is removed.

Claims (55)

1 . A computer-implemented method, comprising:

analyzing previously stored target addresses;

determining one or more of the previously stored target addresses that result from more than a predetermined number of redirected originating addresses; and

on determining a respective target address, determining that one or more corresponding originating addresses are invalid based on a difference between information previously stored for the one or more corresponding originating addresses and information associated with the respective target address.

2 . The computer-implemented method of claim 1 , wherein the one or more corresponding originating address are determined to be invalid when the difference satisfies a predetermined threshold.

3 . The computer-implemented method of claim 1 , further comprising:

analyzing resources corresponding to a plurality of resource addresses, the plurality of resource addresses including the redirected originating addresses,

wherein the previously stored information is derived from resources located at the redirected originating addresses.

4 . The computer-implemented method of claim 3 , wherein a resource address is an internet address, and the analyzed resources include webpages located at respective internet addresses, and wherein analyzing the resources includes performing a web crawling operation on a plurality of webpages.

5 . The computer-implemented method of claim 1 , wherein the information previously stored for an originating address includes content associated with a webpage located at the originating address, and

wherein the information associated with the respective target address includes content associated with a webpage located at the respective target address.

6 . The computer-implemented method of claim 1 , wherein information previously stored for an originating address includes a first set of meta-data associated with the originating address, and the information associated with the respective target address includes a second set of meta-data associated with the respective target address.

7 . The computer-implemented method of claim 1 , further comprising:

determining a first plurality of n-grams based on terms in information previously stored for an originating address;

determining a second plurality of n-grams based on terms in the information associated with the respective target address;

comparing the first plurality and the second plurality; and

determining a number of matching n-grams between the first plurality and the second plurality, wherein the difference is based on the determined number of matching n-grams.

8 . The computer-implemented method of claim 7 , further comprising:

before determining the first plurality of n-grams, excluding terms that are in a group of stop words; and

before determining the second plurality of n-grams, excluding terms that are in the group of stop words.

9 . The computer-implemented method of claim 1 , further comprising:

determining a first semantic content based on terms in the information previously stored for an originating address;

determining a second semantic content based on terms in the information associated with the respective target address; and

comparing the first semantic content with the second semantic content,

wherein the difference is representative of a number of meanings found between the first semantic content and the second semantic content.

10 . The computer-implemented method of claim 1 , further comprising:

storing the one or more corresponding originating addresses, indexed by the respective target address.

11 . The computer-implemented method of claim 1 , wherein the redirected originating addresses include one or more intermediate redirecting addresses between a first redirecting address and a final target address.

12 . The computer-implemented method of claim 1 , further comprising:

providing an indication that the one or more corresponding originating addresses are not valid.

13 . The computer-implemented method of claim 12 , wherein providing the indication includes removing the one or more corresponding originating addresses from a searchable set of originating addresses.

14 . A machine-readable media including instructions thereon that, when executed, perform a method, the method comprising:

determining one or more target addresses that result from a redirection from one or more originating addresses; and

for a target address, storing a plurality of originating addresses, determining that a number of the plurality of originating addresses satisfies a predetermined threshold, and, on determining that the plurality of originating addresses satisfies the predetermined threshold, providing an indication that the plurality of originating addresses is not valid.

15 . The machine-readable media of claim 14 , the method further comprising:

analyzing a plurality of webpage addresses to determine the one or more target addresses.

16 . The machine-readable media of claim 14 , wherein determining the one or more target addresses comprises:

determining one or more intermediary addresses that result from the redirection, the one or more target addresses being a result of a redirection from the one or more intermediary addresses; and

storing the one or more intermediary addresses in the storage location together with the plurality of originating addresses.

17 . The machine-readable media of claim 16 , the method further comprising:

for an intermediary address, if the plurality of originating addresses related to the intermediary address satisfies the predetermined threshold, providing an indication that the intermediary addresses is not valid.

18 . The machine-readable media of claim 14 , the method further comprising:

storing the one or more target addresses in a storage location; and

analyzing the storage location to determine how many originating addresses redirect to each stored target address.

19 . The machine-readable media of claim 14 , wherein providing an indication that an originating address is not valid includes removing the originating address from the plurality of originating addresses, and from a subsequent web crawling operation.

20 . A system, comprising:

a processor; and

a memory, including server instructions that, when executed, cause the processor to:

analyze a plurality of internet addresses;

store information corresponding to the plurality of internet addresses;

from the plurality of internet addresses, determine one or more target addresses redirected from the plurality of internet addresses;

store the one or more target addresses in a storage location; and

for a target address,

store a plurality of originating addresses,

determine a number of the plurality of originating addresses, and, on determining that the number satisfies a first predetermined threshold, identify originating addresses associated with resources that include different information than a resource associated with the target address, and providing an indication that the identified originating addresses are not valid.

Assignments (2)
CERTIFICATE OF CONVERSION Recorded Sep 13, 2021
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 057652/0052 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 8, 2012
From: HYMAN, JOSHUA MARK; WHITE, JOSEPH LAWRENCE; DONNELLY, JUSTIN GABRIEL; BILLOCK, JOSEPH GREGORY
To: GOOGLE INC.
Reel/Frame 028347/0513 →