IP Library Granted Patent US 8,185,544
Granted Patent B2
US 8,185,544 · App. 12/420,775 · Granted May 22, 2012

Generating improved document classification data using historical search results

Assignee: Google Inc.
View Patent ↗
Loading inventors, assignments & file history…
Monitor This Case
Get email alerts when status or documents change.
Order Certified Copies
Most orders are placed with the USPTO same day — all within 24 business hours.
Order via The Patent Place →
Pre-filled with this patent's details
Quick Facts
Patent No.
US 8,185,544
App. No.
12/420,775
Granted
May 22, 2012
Kind
B2
Abstract

A server system accesses, respectively, historical query information for queries that have search results corresponding to first information items and second information items and classification data of the first information items. Initially, the first information items are classified and the second information items are unclassified. Based on the classification data of the first information items and the historical query information, the server system generates classification data for the second information items and stores the generated classification data therein. In response to requests for service from client devices, the server system provides customized services to the client devices using the second information items and the corresponding classification data generated for the second information items.

Claims (181)

1. A computer-implemented method, comprising:

at a server system having one or more processors and memory,

respectively accessing historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

accessing classification data of the first information items;

generating classification data for the initially unclassified information items based on the classification data of the first information items and the historical query information;

storing the generated classification data in the server system; and

providing customized services associated with the second information items to a plurality of client devices using the corresponding classification data stored in the server system;

wherein generating classification data for an initially unclassified information item includes:

identifying a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

generating classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

generating classification data for the initially unclassified information item by combining the generated classification data of the subset of the queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein generating classification data for the set of queries includes:

for each of at least a subset of the queries,

identifying a set of search results corresponding to the query and a set of the first information items corresponding to the set of search results;

weighting the classification data of the identified first information items in accordance with at least one of: their respective predefined information retrieval scores, their corresponding search results' positions in the set of search results, and user interaction with the corresponding search results; and

aggregating the weighted classification data of the identified first information items as the query's classification data.

2. The computer-implemented method of claim 1 , further comprising:

updating the historical query information; and

repeating the identification of queries in the historical query information, the generation of classification data for the queries, and the generation of classification data for the initially unclassified information item using the updated historical query information.

3. The computer-implemented method of claim 1 , wherein the historical query information comprises historical query information for queries submitted by a community of users.

4. The computer-implemented method of claim 1 , wherein providing customized services includes:

receiving a query from a user at a respective client device, wherein the user has an associated user profile; and

responding to the query by:

identifying a set of search results corresponding to the query, wherein one of the search results is associated with one of the second information items;

determining a score for the search result by comparing the stored classification data of the second information item with the user profile;

ordering the search result with respect to the other search results in accordance with the determined score; and

providing data representing at least the ordered search result to the client device.

5. The computer-implemented method of claim 1 , wherein providing customized services includes:

identifying a set of queries submitted by a user in the historical query information and the corresponding search results, wherein the search results correspond to one or more of the first and second information items;

generating a user profile for the user by aggregating the classification data of the one or more information items;

storing the generated user profile in the server system; and

in response to a request for service from the user at a client device, customizing the requested service using the stored user profile.

6. The computer-implemented method of claim 5 , wherein customizing the requested service includes:

preparing a user-independent service in response to the service request, wherein the user-independent service includes one or more of the first and second information items;

determining a score for each of the one or more information items by comparing the information item's classification data with the stored user profile; and

re-arranging the one or more information items in the service in accordance with their respective scores.

7. The computer-implemented method of claim 1 , wherein at least one of the information items is a web page.

8. The computer-implemented method of claim 1 , wherein at least one of the information items is a website including multiple web pages.

9. The computer-implemented method of claim 1 , further comprising:

updating the classification data of the first information items based on the historical query information and the generated classification data of the initially unclassified information items.

10. The computer-implemented method of claim 1 , wherein providing customized services associated with the second information items includes:

identifying generated classification data commonly associated with a plurality of users; and

normalizing the identified classification data for each respective user of the plurality of users in accordance with a level of interest demonstrated by the respective user in the identified classification data relative to other users of the plurality of users.

11. A computer-implemented method, comprising:

at a server system having one or more processors and memory,

respectively accessing historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

accessing classification data of the first information items;

generating classification data for the initially unclassified information items based on the classification data of the first information items and the historical query information;

storing the generated classification data in the server system; and

providing customized services associated with the second information items to a plurality of client devices using the corresponding classification data stored in the server system;

wherein generating classification data for an initially unclassified information item includes:

identifying a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

generating classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

generating classification data for the initially unclassified information item by combining the generated classification data of the subset of queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein generating classification data for the initially unclassified information item by combining the generated classification data of the subset of queries includes:

for each query of the subset of queries,

weighting the classification data of the query in accordance with at least one of: the initially unclassified information item's predefined information retrieval score, a search result position of a search result corresponding to the initially unclassified information item in a set of search results for the query, and user interaction with the corresponding search result; and

aggregating the weighted classification data of the subset of queries as the initially unclassified information item's classification data.

12. The computer-implemented method of claim 11 , wherein generating classification data for the initially unclassified information items further includes:

for each query of the subset of queries,

weighting the classification data of the query in accordance with a number of query-terms in the query.

13. A computer system, comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including:

instructions for respectively accessing historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

instructions for accessing classification data of the first information items;

instructions for generating classification data for the second information items based on the classification data of the first information items and the historical query information;

instructions for storing the generated classification data in the server system; and

instructions for providing customized services associated with the second information items to a plurality of client devices using the corresponding classification data stored in the server system;

wherein the instructions for generating classification data for an initially unclassified information item include:

instructions for identifying a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

instructions for generating classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

instructions for generating classification data for the initially unclassified information item by combining the generated classification data of the subset of queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein the instructions for generating classification data for the set of queries include:

for each of at least a subset of the queries,

instructions for identifying a set of search results corresponding to each of at least a subset of the queries and a set of the first information items corresponding to the set of search results;

instructions for weighting the classification data of the identified first information items in accordance with at least one of: their respective predefined information retrieval scores, their corresponding search results' positions in the set of search results, and user interaction with the corresponding search results; and

instructions for aggregating the weighted classification data of the identified first information items as the query's classification data.

14. The computer system of claim 13 , further comprising:

instructions for updating the historical query information; and

instructions for repeating the identification of queries in the historical query information, the generation of classification data for the queries, and the generation of classification data for the initially unclassified information item using the updated historical query information.

15. The computer system of claim 13 , wherein the instructions for providing customized services include:

instructions for receiving a query from a user at a respective client device, wherein the user has an associated user profile;

instructions for identifying a set of search results corresponding to the query, wherein one of the search results is associated with one of the second information items;

instructions for determining a score for the search result by comparing the stored classification data of the second information item with the user profile;

instructions for ordering the search result with respect to the other search results in accordance with the determined score; and

instructions for providing data representing at least the ordered search result to the client device.

16. The computer system of claim 13 , wherein the instructions for providing customized services include:

instructions for identifying a set of queries submitted by a user in the historical query information and the corresponding search results, wherein the search results correspond to one or more of the first and second information items;

instructions for generating a user profile for the user by aggregating the classification data of the one or more information items;

instructions for storing the generated user profile in the server system; and

instructions for customizing the requested service using the stored user profile in response to a request for service from the user at a client device.

17. The computer system of claim 16 , wherein the instructions for customizing the requested service include:

instructions for preparing a user-independent service in response to the service request, wherein the user-independent service includes one or more of the first and second information items;

instructions for determining a score for each of the one or more information items by comparing the information item's classification data with the stored user profile; and

instructions for re-arranging the one or more information items in the service in accordance with their respective scores.

18. The computer system of claim 13 , wherein the historical query information comprises historical query information for queries submitted by a community of users.

19. The computer system of claim 13 , wherein at least one of the information items is a web page.

20. The computer system of claim 13 , wherein at least one of the information items is a website including multiple web pages.

21. The computer system of claim 13 , further comprising:

instructions for updating the classification data of the first information items based on the historical query information and the generated classification data of the initially unclassified information items.

22. The computer system of claim 13 , wherein instructions for providing customized services associated with the second information items include:

instructions for identifying generated classification data commonly associated with a plurality of users; and

instructions for normalizing the identified classification data for each respective user of the plurality of users in accordance with a level of interest demonstrated by the respective user in the identified classification data relative to other users of the plurality of users.

23. A computer system, comprising:

one or more processors;

memory; and

one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, the one or more programs including instructions for:

respectively accessing historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

accessing classification data of the first information items;

generating classification data for the second information items based on the classification data of the first information items and the historical query information;

storing the generated classification data in the server system; and

providing customized services associated with the second information items to a plurality of client devices using the corresponding classification data stored in the server system;

wherein the instructions for generating classification data for an initially unclassified information item include instructions for:

identifying a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

generating classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

generating classification data for the initially unclassified information item by combining the generated classification data of the subset of queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein the instructions for generating classification data for the initially unclassified information item by combining the generated classification data of the subset of queries include instructions for:

for each query of the subset of queries,

weighting the classification data of the query in accordance with at least one of: the initially unclassified information item's predefined information retrieval score, a search result position of a search result corresponding to the initially unclassified information item in a set of search results for the query, and user interaction with the corresponding search result; and

aggregating the weighted classification data of the subset of queries as the initially unclassified information item's classification data.

24. The computer system of claim 23 , wherein the instructions for generating classification data for the initially unclassified information items further include instructions for:

for each query of the subset of queries, weighting the classification data of the query in accordance with a number of query-terms in the query.

25. A non-transitory computer readable storage medium and one or more computer programs embedded therein, the one or more computer programs comprising instructions which, when executed by a computer system, cause the computer system to:

respectively access historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

access classification data of the first information items;

generate classification data for the second information items based on the classification data of the first information items and the historical query information;

store the generated classification data in the server system; and

provide customized services associated with the second information items to a pluralit of client devices using the corresponding classification data stored in the server system;

wherein the instructions to generate classification data for an initially unclassified information item include instructions to:

identify a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

generate classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

generate classification data for the initially unclassified information item by combining the generated classification data of the subset of queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein the instructions to generate classification data for the set of queries include instructions to:

for each of at least a subset of the queries,

identify a set of search results corresponding to each of at least a subset of the queries and a set of the first information items corresponding to the set of search results;

weight the classification data of the identified first information items in accordance with at least one of: their respective predefined information retrieval scores, their corresponding search results' positions in the set of search results, and user interaction with the corresponding search results; and

aggregate the weighted classification data of the identified first information items as the query's classification data.

26. The non-transitory computer readable storage medium of claim 25 , wherein the one or more computer programs further comprise:

instructions to update the historical query information; and

instructions to repeat the identification of queries in the historical query information, the generation of classification data for the queries, and the generation of classification data for the initially unclassified information item using the updated historical query information.

27. The non-transitory computer readable storage medium of claim 25 , wherein the instructions to provide customized services include instructions to:

receive a query from a user at a respective client device, wherein the user has an associated user profile;

identify a set of search results corresponding to the query, wherein one of the search results is associated with one of the second information items;

determine a score for the search result by comparing the stored classification data of the second information item with the user profile;

order the search result with respect to the other search results in accordance with the determined score; and

provide data representing at least the ordered search result to the client device.

28. The non-transitory computer readable storage medium of claim 25 , wherein the instructions to provide customized services include instructions to:

identify a set of queries submitted by a user in the historical query information and the corresponding search results, wherein the search results correspond to one or more of the first and second information items;

generate a user profile for the user by aggregating the classification data of the one or more information items;

store the generated user profile in the server system; and

customize the requested service using the stored user profile in response to a request for service from the user at a client device.

29. The non-transitory computer readable storage medium of claim 28 , wherein the instructions to customize the requested service include instructions to:

prepare a user-independent service in response to the service request, wherein the user-independent service includes one or more of the first and second information items;

determine a score for each of the one or more information items by comparing the information item's classification data with the stored user profile; and

re-arrange the one or more information items in the service in accordance with their respective scores.

30. The non-transitory computer readable storage medium of claim 25 , wherein the historical query information comprises historical query information for queries submitted by a community of users.

31. The non-transitory computer readable storage medium of claim 25 , wherein at least one of the information items is a web page.

32. The non-transitory computer readable storage medium of claim 25 , wherein at least one of the information items is a website including multiple web pages.

33. The non-transitory computer readable storage medium of claim 25 , further comprising:

instructions to update the classification data of the first information items based on the historical query information and the generated classification data of the initially unclassified information items.

34. The non-transitory computer readable storage medium of claim 25 , wherein the instructions to provide customized services associated with the second information items include instructions to:

identify generated classification data commonly associated with a plurality of users; and

normalize the identified classification data for each respective user of the plurality of users in accordance with a level of interest demonstrated by the respective user in the identified classification data relative to other users of the plurality of users.

35. A non-transitory computer readable storage medium and one or more computer programs embedded therein, the one or more computer programs comprising instructions which, when executed by a computer system, cause the computer system to:

respectively access historical query information for queries having search results that correspond to first information items and second information items, wherein the first information items are initially classified and the second information items are initially unclassified;

access classification data of the first information items;

generate classification data for the second information items based on the classification data of the first information items and the historical query information;

store the generated classification data in the server system; and

provide customized services associated with the second information items to a plurality of client devices using the corresponding classification data stored in the server system;

wherein the instructions to generate classification data for an initially unclassified information item include instructions to:

identify a set of queries in the historical query information, wherein at least a subset of the queries each has an associated search result corresponding to the initially unclassified information item;

generate classification data for the set of queries based on the classification data of the first information items and the historical query information for the set of queries; and

generate classification data for the initially unclassified information item by combining the generated classification data of the subset of queries, each of which has an associated search result corresponding to the initially unclassified information item; and

wherein the instructions to generate classification data for the initially unclassified information item by combining the generated classification data of the subset of queries include instructions to:

for each query of the subset of queries,

weight the classification data of the query in accordance with at least one of: the initially unclassified information item's predefined information retrieval score, a search result position of a search result corresponding to the initially unclassified information item in a set of search results for the query, and user interaction with the corresponding search result; and

aggregate the weighted classification data of the subset of queries as the initially unclassified information item's classification data.

36. The non-transitory computer readable storage medium of claim 35 , wherein the instructions to generate classification data for the initially unclassified information items further include instructions to:

for each query of the subset of queries, weight the classification data of the query in accordance with a number of query-terms in the query.

Assignments (4)
CHANGE OF NAME Recorded Oct 2, 2017
From: GOOGLE INC.
To: GOOGLE LLC
Reel/Frame 044101/0405 →
PLEASE RECORD TO CORRECT NAME OF INVENTOR TO READ "PEI-WEN ANDY CHIU" AND TO CORRECT NAME OF ASSIGNEE TO "GOOGLE INC.," PREVIOUSLY RECORDED ON APRIL 22, 2010 REEL 024281; FRAME 0163 Recorded Jun 16, 2010
From: OZTEKIN, BILGEHAN UYGAR; CHIU, PEI-WEN ANDY
To: GOOGLE INC.
Reel/Frame 024549/0296 →
CORRECTIVE ASSIGNMENTTO CORRECT NAME OF INVENTOR TO READ "BILGEHAN UYGAR OZTEKIN" AND "PEI-WEN ANDY CHIU" PREVIOUSLY RECORDED ON JUNE 15, 2009, REEL 022826; FRAME 0056 Recorded Apr 22, 2010
From: OZTEKIN, BILGEHAN UYGAR; CHIU, PEI-WEIN ANDY
To: GOGGLE INC.
Reel/Frame 024281/0163 →
ASSIGNMENT OF ASSIGNOR'S INTEREST Recorded Jun 15, 2009
From: OZTEKIN, BILEGHAN UYGAR; CHIU, PAI-WEN ANDY
To: GOOGLE INC.
Reel/Frame 022826/0056 →
Continuity (1)
Related Publication 20100262615A1 · Oct 14, 2010