METHODS AND APPARATUS FOR DETERMINING IF A SEARCH QUERY SHOULD BE ISSUED
Methods and apparatus assessing, ranking, organizing, and presenting search results associated with a user's current work context are disclosed. The system disclosed assesses, ranks, organizes and presents search results against a user's current work context by comparing statistical and heuristic models of the search results to a statistical and heuristic model of the user's current work context. In this manner, search results are assessed, ranked, organized, and/or presented with the benefit of attributes of the user's current work context that are predictive of relevance, such as words in a user's document (e.g., web page or word processing document) that may not have been included in the search query. In addition, search results from multiple search engines are combined into an organization scheme that best reflects the user's current task. As a result, lists of search results from different search engines can be more usefully presented to the user.
1 . A method of displaying search results, the method comprising:
determining if a length associated with a document being accessed by a user exceeds a first threshold;
determining if a ratio of non-hyperlinked words to hyperlinked words in the document exceeds a second threshold;
determining if a similarity score associated with at least two different segments of the document exceeds a third threshold;
sending a query to a search engine if (a) the length associated with the document exceeds the first threshold, (b) the ratio of non-hyperlinked words to hyperlinked words in the document exceeds the second threshold, and (c) the similarity score associated with at least two different segments of the document exceeds the third threshold;
receiving a plurality of search results from the search engine; and
generating a display indicative of the plurality of search results.
2 . The method of claim 1 , wherein the first threshold is based on a genre associated with the document being accessed by the user.
3 . The method of claim 1 , wherein the second threshold is based on a genre associated with the document being accessed by the user.
4 . The method of claim 1 , wherein the third threshold is based on a genre associated with the document being accessed by the user.
5 . The method of claim 1 , including determining if at least one of an area of interest and an area of disinterest are associated with the document being accessed by the user.
6 . The method of claim 1 , including determining non-document specific context information.
7 . The method of claim 1 , wherein the search query is based on a first aspect of a user context, the first aspect of the user context including data indicative of text being accessed by a user, the query being different than the user context.
8 . The method of claim 7 , wherein the first aspect of the user context includes data indicative of the at least one task.
9 . The method of claim 7 , including comparing data indicative of the plurality of search results to data indicative of a second aspect of the user context to determine a plurality of relevance scores associated with the plurality of search results, the second aspect of the user context including data indicative of at least one task in which the user is engaged out of a plurality of possible user tasks.
10 . The method of claim 9 , wherein the second aspect of the user context is based on at least five of (a) a location of the at least one predetermined word in the text being accessed by the user, (b) a style of the at least one predetermined word in the text being accessed by the user, (c) a presence of at least one specified word in the text being accessed by the user, (d) an absence of the at least one specified word in the text being accessed by the user, (e) metadata attributes of at least a portion of the text being accessed by the user, (f) a field presented by a computer application, (g) an attribute of information being presented in the computer application, (h) an element of the computer application visible to the user, (i) a document genre, (j) a document type, (k) a type associated with the computer application, (l) a method by which the user is accessing the computer application, (m) a role in an organization, (n) a type of the organization, (o) a property of the organization, (p) a stage in a task, (q) a stage in a workflow, (r) a type of task being supported by the computer application, (s) a stage in a task being executed by the computer application, (t) a pervious user behavior, (u) a topical area of interest, (v) a proportion of hyperlinked text to non-hyperlinked text, and (w) an average sentence length in the text being accessed by the user.
11 . The method of claim 9 , wherein the second aspect of the user context is based on (a) a style of the at least one predetermined word in the text being accessed by the user and (b) a type associated with a computer application.
12 . The method of claim 9 , including comparing data indicative of the plurality of search results to data indicative of at least one of the first aspect of the user context and the second aspect of the user context to determine a plurality of organization schemes, the plurality of organization schemes grouping at least a portion of the plurality of search results into at least two genres.
13 . The method of claim 12 , wherein generating the display indicative of the plurality of search results includes generating the display to be indicative of the plurality of organization schemes.
14 . The method of claim 9 , including receiving a plurality of result models from the search engine, the plurality of result models including a plurality of terms associated with the plurality of search results and a plurality weights associated with the plurality of terms.
15 . The method of claim 14 , including comparing data indicative of the plurality of result models to data indicative of a user context to determine a plurality of scores associated with the plurality of search results.
16 . The method of claim 1 , including determining a second search query from a user context and an interim search result.
17 . An apparatus for characterizing a search result as potential spam, the apparatus comprising:
a processor;
a memory device operatively coupled to the processor; and
a network device operatively coupled to the processor; wherein the memory device stores a software program to cause the processor to:
determine if a length associated with a document being accessed by a user exceeds a first threshold;
determine if a ratio of non-hyperlinked words to hyperlinked words in the document exceeds a second threshold;
determine if a similarity score associated with at least two different segments of the document exceeds a third threshold;
send a query to a search engine if (a) the length associated with the document exceeds the first threshold, (b) the ratio of non-hyperlinked words to hyperlinked words in the document exceeds the second threshold, and (c) the similarity score associated with at least two different segments of the document exceeds the third threshold;
receive a plurality of search results from the search engine; and
generate a display indicative of the plurality of search results.
18 . The apparatus of claim 17 , wherein the search query is based on a first aspect of a user context, the first aspect of the user context including data indicative of text being accessed by a user, the query being different than the user context.
19 . The apparatus of claim 18 , wherein the software program is structured to cause the processor to compare data indicative of a plurality of search results to data indicative of a second aspect of the user context to determine a plurality of relevance scores associated with the plurality of search results, the second aspect of the user context including data indicative of at least one task in which the user is engaged out of a plurality of possible user tasks.
20 . The method of claim 19 , wherein the second aspect of the user context is based on at least five of (a) a location of the at least one predetermined word in the text being accessed by the user, (b) a style of the at least one predetermined word in the text being accessed by the user, (c) a presence of at least one specified word in the text being accessed by the user, (d) an absence of the at least one specified word in the text being accessed by the user, (e) metadata attributes of at least a portion of the text being accessed by the user, (f) a field presented by a computer application, (g) an attribute of information being presented in the computer application, (h) an element of the computer application visible to the user, (i) a document genre, (j) a document type, (k) a type associated with the computer application, (l) a method by which the user is accessing the computer application, (m) a role in an organization, (n) a type of the organization, (o) a property of the organization, (p) a stage in a task, (q) a stage in a workflow, (r) a type of task being supported by the computer application, (s) a stage in a task being executed by the computer application, (t) a pervious user behavior, (u) a topical area of interest, (v) a proportion of hyperlinked text to non-hyperlinked text, and (w) an average sentence length in the text being accessed by the user.
21 . The method of claim 19 , wherein the second aspect of the user context is based on (a) a style of the at least one predetermined word in the text being accessed by the user and (b) a type associated with a computer application.
22 . A computer readable medium storing a software program to cause a computing device to:
determine if a length associated with a document being accessed by a user exceeds a first threshold;
determine if a ratio of non-hyperlinked words to hyperlinked words in the document exceeds a second threshold;
determine if a similarity score associated with at least two different segments of the document exceeds a third threshold;
send a query to a search engine if (a) the length associated with the document exceeds the first threshold, (b) the ratio of non-hyperlinked words to hyperlinked words in the document exceeds the second threshold, and (c) the similarity score associated with at least two different segments of the document exceeds the third threshold;
receive a plurality of search results from the search engine; and
generate a display indicative of the plurality of search results.
23 . The computer readable medium of claim 22 , wherein the search query is based on a first aspect of a user context, the first aspect of the user context including data indicative of text being accessed by a user, the query being different than the user context.
24 . The computer readable medium of claim 23 , wherein the software program is structured to cause the computing device to compare data indicative of the plurality of search results to data indicative of a second aspect of the user context to determine a plurality of relevance scores associated with the plurality of search results, the second aspect of the user context including data indicative of at least one task in which the user is engaged out of a plurality of possible user tasks.
25 . The method of claim 24 , wherein the second aspect of the user context is based on at least five of (a) a location of the at least one predetermined word in the text being accessed by the user, (b) a style of the at least one predetermined word in the text being accessed by the user, (c) a presence of at least one specified word in the text being accessed by the user, (d) an absence of the at least one specified word in the text being accessed by the user, (e) metadata attributes of at least a portion of the text being accessed by the user, (f) a field presented by a computer application, (g) an attribute of information being presented in the computer application, (h) an element of the computer application visible to the user, (i) a document genre, (j) a document type, (k) a type associated with the computer application, (l) a method by which the user is accessing the computer application, (m) a role in an organization, (n) a type of the organization, (o) a property of the organization, (p) a stage in a task, (q) a stage in a workflow, (r) a type of task being supported by the computer application, (s) a stage in a task being executed by the computer application, (t) a pervious user behavior, (u) a topical area of interest, (v) a proportion of hyperlinked text to non-hyperlinked text, and (w) an average sentence length in the text being accessed by the user.
26 . The method of claim 24 , wherein the second aspect of the user context is based on (a) a style of the at least one predetermined word in the text being accessed by the user and (b) a type associated with a computer application.