IP Library Granted Patent US 8,930,361
Granted Patent B2
US 8,930,361 · App. 13/099,685 · Granted Jan 6, 2015

Method and apparatus for cleaning data sets for a search process

View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,930,361
App. No.
13/099,685
Granted
Jan 6, 2015
Kind
B2
Abstract

An approach is provided for cleaning data sets for a search process. The cleanup platform determines one or more reference documents associated with at least one region. Next, the cleanup platform processes and/or facilitates a processing of the one or more reference documents to determine a frequency distribution of one or more candidate stop words with respect to the at least one region. Then, the cleanup platform causes, at least in part, selection of one or more stop words applicable to the at least one region from the one or more candidate stop words based, at least in part, on one or more frequency distribution criteria. Additionally, the cleanup platform processes and/or facilitates a processing of at least one data set associated with a search process to generate at least one enhanced data set by filtering the one or more stop words from the at least one data set.

Claims (67)

1. A method comprising facilitating a processing of and/or processing (1) data and/or (2) information and/or (3) at least one signal, the (1) data and/or (2) information and/or (3) at least one signal based, at least in part, on the following:

one or more reference documents associated with at least one region;

a processing, by a processor, of the one or more reference documents to determine a frequency distribution of one or more candidate stop words with respect to the at least one region;

a selection of one or more stop words applicable to the at least one region from the one or more candidate stop words based, at least in part, on one or more frequency distribution criteria; and

a processing of at least one data set associated with a search process to generate at least one enhanced data set by filtering the one or more stop words from the at least one data set.

2. A method of claim 1 , wherein the at least one data set include, at least in part, (a) at least one content data set over which the search process operates, (b) at least one search log data set for recording activity associated with one or more users of the search process, (c) at least one access log data set for recording access by the one or more users to the at least one content data set, or (d) a combination thereof.

3. A method of claim 1 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

metadata associated with one or more users of the search process, wherein the metadata includes at least in part one or more search histories, context information, or a combination thereof; and

a processing of the metadata to determine one or more alternative tokens to associate with one or more tokens of the at least one data set,

wherein the at least one enhanced data set is further based, at least in part, on the association of the one or more alternative tokens with the one or more tokens.

4. A method of claim 1 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

a processing of the at least one data set to determine one or more prefix tokens, one or more suffix tokens, or a combination thereof,

wherein the at least one enhanced data set is further based, at least in part, on filtering the one or more prefix tokens, the one or more suffix tokens, or a combination thereof from the at least one data set.

5. A method of claim 1 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

a processing of at least one data set to determine one or more synonymous tokens; and

a normalization of the one or more synonymous tokens,

wherein the at least one enhanced data set is further based, at least in part, on the normalization.

6. A method of claim 5 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

respective context information associated with the one or more synonymous tokens;

a processing of the respective context information to determine one or more duplicates among the one or more synonymous tokens,

wherein the normalization comprises, at least in part, removal of the one or more duplicates.

7. A method of claim 5 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

a processing of the at least one data set to determine one or more translations, one or more permutations, or a combination thereof of the one or more synonymous tokens,

wherein the normalization is based, at least in part, on the one or more translation, the one or more permutations, or a combination thereof.

8. A method of claim 1 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

a processing of the at least one data set to determine context information, one or more users, or a combination thereof associated with one or more tokens of the at least one data set; and

at least one determination of a frequency distribution based, at least in part, on the context information, the one or more users, or a combination thereof,

wherein the at least one enhanced data set is further based, at least in part, on filtering of the one or more tokens according to the frequency distribution.

9. A method of claim 1 , wherein the (1) data and/or (2) information and/or (3) at least one signal are further based, at least in part, on the following:

a processing of the at least one data set to determine one or more string tokens associated with at least one test model,

wherein the at least one enhanced data set is further based, at least in part, on filtering of the one or more string tokens.

10. A method of claim 1 , wherein the search process includes, at least in part, a location-based search process, a content search process, a service search process, or a combination thereof.

11. An apparatus comprising:

at least one processor; and

at least one memory including computer program code for one or more programs,

the at least one memory and the computer program code configured to, with the at least one processor, cause the apparatus to perform at least the following,

determine one or more reference documents associated with at least one region;

process and/or facilitate a processing of the one or more reference documents to determine a frequency distribution of one or more candidate stop words with respect to the at least one region;

cause, at least in part, selection of one or more stop words applicable to the at least one region from the one or more candidate stop words based, at least in part, on one or more frequency distribution criteria; and

process and/or facilitate a processing of at least one data set associated with a search process to generate at least one enhanced data set by filtering the one or more stop words from the at least one data set.

12. An apparatus of claim 11 , wherein the at least one data set include, at least in part, (a) at least one content data set over which the search process operates, (b) at least one search log data set for recording activity associated with one or more users of the search process, (c) at least one access log data set for recording access by the one or more users to the at least one content data set, or (d) a combination thereof.

13. An apparatus of claim 11 , wherein the apparatus is further caused to:

determine metadata associated with one or more users of the search process, wherein the metadata includes at least in part one or more search histories, context information, or a combination thereof; and

process and/or facilitate a processing of the metadata to determine one or more alternative tokens to associate with one or more tokens of the at least one data set,

wherein the at least one enhanced data set is further based, at least in part, on the association of the one or more alternative tokens with the one or more tokens.

14. An apparatus of claim 11 , wherein the apparatus is further caused to:

process and/or facilitate a processing of the at least one data set to determine one or more prefix tokens, one or more suffix tokens, or a combination thereof,

wherein the at least one enhanced data set is further based, at least in part, on filtering the one or more prefix tokens, the one or more suffix tokens, or a combination thereof from the at least one data set.

15. An apparatus of claim 11 , wherein the apparatus is further caused to:

process and/or facilitate a processing of at least one data set to determine one or more synonymous tokens; and

cause, at least in part, a normalization of the one or more synonymous tokens,

wherein the at least one enhanced data set is further based, at least in part, on the normalization.

16. An apparatus of claim 15 , wherein the apparatus is further caused to:

determine respective context information associated with the one or more synonymous tokens;

process and/or facilitate a processing of the respective context information to determine one or more duplicates among the one or more synonymous tokens,

wherein the normalization comprises, at least in part, removal of the one or more duplicates.

17. An apparatus of claim 15 , wherein the apparatus is further caused to:

process and/or facilitate a processing of the at least one data set to determine one or more translations, one or more permutations, or a combination thereof of the one or more synonymous tokens,

wherein the normalization is based, at least in part, on the one or more translation, the one or more permutations, or a combination thereof.

18. An apparatus of claim 11 , wherein the apparatus is further caused to:

process and/or facilitate a processing of the at least one data set to determine context information, one or more users, or a combination thereof associated with one or more tokens of the at least one data set; and

determine a frequency distribution based, at least in part, on the context information, the one or more users, or a combination thereof,

wherein the at least one enhanced data set is further based, at least in part, on filtering of the one or more tokens according to the frequency distribution.

19. An apparatus of claim 11 , wherein the apparatus is further caused to:

process and/or facilitate a processing of the at least one data set to determine one or more string tokens associated with at least one test model,

wherein the at least one enhanced data set is further based, at least in part, on filtering of the one or more string tokens.

20. An apparatus of claim 11 , wherein the search process includes, at least in part, a location-based search process, a content search process, a service search process, or a combination thereof.

Assignments (9)
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Dec 28, 2021
From: PROVENANCE ASSET GROUP LLC
To: RPX CORPORATION
Reel/Frame 059352/0001 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: NOKIA US HOLDINGS INC.
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058363/0723 →
RELEASE OF SECURITY INTEREST Recorded Nov 30, 2021
From: CORTLAND CAPITAL MARKETS SERVICES LLC
To: PROVENANCE ASSET GROUP HOLDINGS LLC; PROVENANCE ASSET GROUP LLC
Reel/Frame 058983/0104 →
ASSIGNMENT AND ASSUMPTION AGREEMENT Recorded Feb 14, 2019
From: NOKIA USA INC.
To: NOKIA US HOLDINGS INC.
Reel/Frame 048370/0682 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 13, 2017
From: NOKIA TECHNOLOGIES OY; NOKIA SOLUTIONS AND NETWORKS BV; ALCATEL LUCENT SAS
To: PROVENANCE ASSET GROUP LLC
Reel/Frame 043877/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP LLC
To: NOKIA USA INC.
Reel/Frame 043879/0001 →
SECURITY INTEREST Recorded Sep 13, 2017
From: PROVENANCE ASSET GROUP HOLDINGS, LLC; PROVENANCE ASSET GROUP, LLC
To: CORTLAND CAPITAL MARKET SERVICES, LLC
Reel/Frame 043967/0001 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Apr 20, 2015
From: NOKIA CORPORATION
To: NOKIA TECHNOLOGIES OY
Reel/Frame 035449/0096 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Sep 7, 2012
From: HEINONEN, JARKKO; AGRAWAL, ASHISH KUMAR; TURNER, ROSS
To: NOKIA CORPORATION
Reel/Frame 028919/0144 →