IP Library Granted Patent US 8,756,688
Granted Patent B1
US 8,756,688 · App. 13/538,013 · Granted Jun 17, 2014

Method and system for identifying business listing characteristics

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,756,688
App. No.
13/538,013
Granted
Jun 17, 2014
Kind
B1
Abstract

Aspects of the disclosure provide for detection of spam attacks. In order to filter spam listings among business listings associated with a particular geographical region, a method and system operate to analyze the frequency of particular characteristics of the business listings, such as words within a listing title, phone numbers, or websites, to identify a normal frequency of each characteristic. Business listings may be periodically analyzed to identify anomalous increases in the frequency of particular characteristics. Characteristics that exhibit these anomalies may be identified as suspicious characteristics, and listings that contain suspicious characteristics may be identified as possible spam listings.

Claims (40)

1. A computer implemented method for identifying business listing characteristics, the method comprising:

determining, using one or more processors, a first frequency value of a business listing characteristic within a first plurality of business listings over a first time period, wherein the first plurality of business listings are associated with a particular geographical region;

comparing, by the one or more processors, the first frequency value with a normal frequency value of the business listing characteristic previously determined from a second plurality of business listings over a second time period;

identifying, by the one or more processors, an anomaly in the first frequency value when a difference between the first frequency value and the normal frequency value is greater than a predetermined threshold; and

identifying, by the one or more processors, the business listing characteristic as a suspicious characteristic in response to identifying the anomaly, wherein the suspicious characteristic includes at least one term selected from a plurality of terms used in the identified business listing;

determining, by the one or more processors, that a ratio of the at least one term to other terms from the plurality of terms is greater than a predetermined ratio; and

identifying, by the one or more processors, the business listing as a spam listing is in response to the ratio being greater than the predetermined ratio.

2. The method of claim 1 , further comprising:

determining that a business listing added to the first plurality of business listings during the first time period contains the suspicious characteristic; and

performing at least one of:

marking the business listing for moderator review, identifying the business listing as a spam listing, or marking the business listing to be analyzed by a spam detection method.

3. The method of claim 1 , further comprising determining normal plurality of frequency values of the business listing characteristic over a respective plurality of time periods equal to the first time period.

4. The method of claim 3 , further comprising determining the normal frequency value of the business listing characteristic by finding at least one of a mean, a median, or a mode of the plurality of frequency values for the respective plurality of time periods.

5. The method of claim 4 , wherein the normal frequency value of the business listing characteristic is determined by finding the mean of the plurality of frequency values for the respective plurality of time periods, and wherein identifying an anomaly in the first frequency value when a difference between the first frequency value and the normal frequency value is greater than a predetermined threshold further comprises:

determining a standard deviation for the normal frequency value; and

identifying the anomaly when the difference between the first frequency value and the normal frequency value is greater than a predetermined number of standard deviations.

6. A processing system for identifying business listing characteristics, the processing system comprising:

one or more processors; and

one or more memories coupled to the one or more processors for storing instructions and a first plurality of business listings, wherein the first plurality of business listings are associated with a particular geographical region;

wherein the one or more processors are configured to execute the instructions stored in the one or more memories in order to:

determine a frequency value for a characteristic associated with the first plurality of business listings over a particular time period;

identify an anomalous frequency value where a difference between the frequency value and a normal frequency value previously determined for the characteristic from a second plurality of business listings is greater than a predetermined threshold; and

in response to identifying the anomalous frequency value, identify the characteristic as a suspicious characteristic, wherein the suspicious characteristic includes at least one term selected from a plurality of terms used in the identified business listing;

determine that a ratio of the at least one term to other terms from the plurality of terms is greater than a predetermined ratio; and

identify the business listing as a spam listing is in response to the ratio being greater than the predetermined ratio.

7. The processing system of claim 6 , wherein the one or more processors are further configured to determine normal plurality of frequency values of the business listing characteristic over a respective plurality of time periods.

8. The processing system of claim 7 , wherein the one or more processor are further configured to determine the normal frequency value of the business listing characteristic by finding at least one of a mean, a median, or a mode of the plurality of frequency values for the respective plurality of time periods.

9. The processing system of claim 8 , wherein the normal frequency value of the business listing characteristic is determined by finding the mean of the plurality of frequency values for the respective plurality of time periods, and wherein identifying an anomalous value in the first frequency value when a difference between the first frequency value and the normal frequency value is greater than a predetermined threshold further comprises:

determining a standard deviation for the normal frequency value; and

identifying the anomalous value when the difference between the first frequency value and the normal frequency value is greater than a predetermined number of standard deviations.

10. A non-transitory computer readable storage medium containing instructions that, when executed by one or more processors, cause the one or more processors to perform a method, the method comprising:

determining a first frequency value of a business listing characteristic within a first plurality of business listings over a first time period, wherein the first plurality of business listings are associated with a particular geographical region;

comparing the first frequency value with a normal frequency value of the business listing characteristic previously determined from a second plurality of business listings over a second time period;

identifying an anomaly in the first frequency value when a difference between the first frequency value and the normal frequency value is greater than a predetermined threshold; and

identifying the business listing characteristic as a suspicious characteristic in response to identifying the anomaly, wherein the suspicious characteristic includes at least one term selected from a plurality of terms used in the identified business listing;

determining that a ratio of the at least one term to other terms from the plurality of terms is greater than a predetermined ratio; and

identifying the business listing as a spam listing is in response to the ratio being greater than the predetermined ratio.

11. The non-transitory computer readable storage medium of claim 10 , wherein the method further comprises:

determining that a business listing added to the first plurality of business listings during the first time period contains a suspicious characteristic; and

performing at least one of: marking the business listing for moderator review, identifying the business listing as a spam listing, or marking the business listing to be analyzed by a spam detection method.

Assignments (2)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044277/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Aug 6, 2012
From: YUKSEL, BARIS
To: GOOGLE INC.
Reel/Frame 028735/0752 →