Eyes-on analysis results for improving search quality
Systems are configured for generating and utilizing training data to train learn-to-rank type models in a manner that preserves privacy of client data used for generating the training data. The systems extract features and patterns of the user queries, search results and user interactions with the search results without tracking, storing or transmitting underlying values of the user data to preserve privacy of the user data. Systems are also configured to infer search result quality based on at least the user behavior data, and optionally query intentions, and to generate and label corresponding training data accordingly. This training data is applied to learn-to-rank type models to train the learn-to-rank type model to improve search quality of search results provided by the learn-to-rank type models when new user queries are processed that having features and patterns corresponding to the filtered and labelled training data.
1 . A method implemented by a computing system comprising a processor and a hardware storage device storing computer-executable instructions that are executable by the processor for causing the computing system to train a learn-to-rank type model to improve quality and ranking of search results generated by the learn-to-rank type model, the method comprising:
the computing system accessing client data, which includes a plurality of user queries;
the computing system extracting features and patterns of the client data while keeping values of the user queries private, wherein the extracted features and patterns of the client data include: a first pattern of types of grammar used in the plurality of user queries and a second pattern of formatting of terms used in the plurality of user queries;
the computing system inferring a search result quality of search result content based on user behavior data included in the client data, wherein the inferred search result quality is further based on analyzing click signals made by a user at a user system, and wherein the click signals are further used to determine at least one of:
a frequency of interacting with the search result content,
whether the user navigated to another resource that is referenced in the search result content,
whether the user copied or saved any of the content,
which one or more search results the user interacted with the most out of the different search results in the search result content, or
which search results the user did not interact with at all from the search result content;
the computing system generating labelled training data for a learn-to-rank type model by labelling features and patterns of the search result content with corresponding relevance scores based on the inferred search result quality;
the computing system identifying conflicting relevance scores for common features and patterns of the search result content included in the labelled training data;
the computing system generating filtered and labelled training data by filtering the labelled training data based on omitting or resolving the identified conflicting relevance scores to reduce noise in the labelled training data that is caused by conflicting relevance scores for common features and patterns of the search result content included in the labelled training data;
the computing system applying the filtered and labelled training data to the learn-to-rank type model to train the learn-to-rank type model to improve search quality of search results provided by the learn-to-rank type model when new user queries are provided to the learn-to-rank type model; and
the computing system accessing analysis data that identifies analyst assessments of quality for the different search result content based on the plurality of user queries, wherein the analysis data includes labels and annotations associated with the values of the user queries.
2 . The method of claim 1 , wherein said extracting features and patterns of the client data is performed without storing or transmitting the client data outside of a client enclave and so as to protect confidentiality of the client data.
3 . The method of claim 1 , wherein extracting the features and the patterns of the client data is performed by accessing the analysis data, which identifies features and patterns of queries and the features and the patterns of the search result content.
4 . The method of claim 3 , wherein the inferred search result quality is further based at least in part on determining inferred user intentions of the queries and wherein the inferred user intentions are determined based on the analysis data that further identifies analyst assessments of the inferred user intentions based on the features and patterns of the queries.
5 . The method of claim 4 , wherein the method further includes segmenting the queries into a plurality of different segmented groupings based on query type, the query type being based on the features and patterns of the queries, each segmented grouping of the different segmented groupings corresponding to a different set of inferred intentions and actions to be performed for providing search results for queries of the query type of the corresponding segmented grouping.
6 . The method of claim 5 , wherein the method further includes: applying the filtered and labelled training data to the learn-to-rank type model to train the learn-to-rank type model to infer intention from a query and to distinguish between different intention types for new and different queries based on features and patterns of the new and different queries and to correspondingly perform different actions for the different types of inferred intentions of the new and different queries to improve quality and relevance of the search results generated by the learn-to-rank type model for the new and different queries.
7 . The method of claim 6 , wherein the labelling of the features and patterns of the search result content includes generating labels for relevance of actions corresponding to different query intention types.
8 . The method of claim 7 , wherein generating labels for actions corresponding to different query intention types includes generating labels for actions that includes at least two or more of: (a) correcting typos in a search query, (b) selecting a best fit search result, (c) selecting an exact match search result, (d) identifying all possible results meeting a predetermined relevance threshold, (e) identifying all possible results, or (f) ranking and sorting all possible results by relevance.
9 . The method of claim 1 , wherein omitting conflicting relevance scores further comprises:
deleting a newly assigned relevance score that is conflicting with an existing relevance score; or
deleting an older assigned relevance score that is conflicting with a newer relevance score.
10 . The method of claim 1 , wherein resolving conflicting relevance scores further comprises:
averaging conflicting relevance scores for common features and patterns of search result content.
11 . A computing system that trains a learn-to-rank type model to improve quality and ranking of search results generated by the learn-to-rank type model and to perform certain actions when generating the search results to improve relevance of the search results for corresponding queries, the computing system comprising:
one or more processors; and
one or more storage devices that store instructions that are executable by the one or more processors to cause the computing system to:
access client data, which includes a plurality of user queries;
extract features and patterns of the client data while keeping values of the user queries private, wherein the extracted features and patterns of the client data include: a first pattern of types of grammar used in the plurality of user queries and a second pattern of formatting of terms used in the plurality of user queries;
infer a search result quality of search result content based on user behavior data included in the client data, wherein the inferred search result quality is further based on analyzing click signals made by a user at a user system, and wherein the click signals are further used to determine at least one of:
a frequency of interacting with the search result content,
whether the user navigated to another resource that is referenced in the search result content,
whether the user copied or saved any of the content,
which one or more search results the user interacted with the most out of the different search results in the search result content, or
which search results the user did not interact with at all from the search result content;
generate labelled training data for the learn-to-rank type model by labelling features and patterns of the search result content for correspondingly paired search query features and pattern with corresponding relevance scores based on the inferred search result quality and, additionally, by at least pairing or labelling actions associated with generating the search result content;
identify conflicting relevance scores for common features and patterns of the search result content included in the labelled training data;
generate filtered and labelled training data by filtering the labelled training data based on omitting or resolving the identified conflicting relevance scores to reduce noise in the labelled training data that is caused by conflicting relevance scores for common features and patterns of the search result content included in the labelled training data;
apply the filtered and labelled training data to the learn-to-rank type model to train the learn-to-rank type model to improve search quality of search results provided by the learn-to-rank type model when new user queries are provided to the learn-to-rank type model that have features and patterns corresponding to the filtered and labelled training data and to further train the learn-to-rank type model to perform different actions for different types of queries to improve quality and relevance of the search results generated by the learn-to-rank type model for the new and different queries; and
access analysis data that identifies analyst assessments of quality for the different search result content based on the plurality of user queries, wherein the analysis data includes labels and annotations associated with the values of the user queries.
12 . The computing system of claim 11 , wherein said extracting features and patterns of the client data is performed without storing or transmitting the client data outside of a client enclave and so as to protect confidentiality of the client data.
13 . The computing system of claim 11 , the computer-executable instructions being further executable for further configuring the computing system to implement the following:
extracting features and patterns of the client data by at least accessing the analysis data, which identifies features and patterns of queries and the features and patterns of the search result content.
14 . The computing system of claim 13 , wherein the inferred search result quality is further based at least in part on determining inferred user intentions of the queries and wherein the inferred user intentions are determined based on the analysis data that further identifies analyst assessments of the inferred user intentions based on the features and patterns of the queries.
15 . The computing system of claim 14 , wherein the computer-executable instructions being further executable for further configuring the computing system to implement the following:
segmenting the queries into a plurality of different segmented groupings based on query type, the query type being based on the features and patterns of the queries, each segmented grouping of the different segmented groupings corresponding to a different set of inferred intentions and actions to be performed for providing search results for queries of the query type of the corresponding segmented grouping.
16 . The computing system of claim 15 , wherein the computer-executable instructions being further executable for further configuring the computing system to implement the following:
applying the filtered and labelled training data to the learn-to-rank type model to train the learn-to-rank type model to infer intention from a query and to distinguish between different intention types for new and different queries based on features and patterns of the new and different queries and to correspondingly perform different actions for the different types of inferred intentions of the new and different queries to improve quality and relevance of results generated by the learn-to-rank type model for the new and different queries.
17 . The computing system of claim 16 , wherein the labelling of the features and patterns of the search result content includes generating labels for relevance of actions corresponding to different query intention types.
18 . The computing system of claim 17 , wherein generating labels for actions corresponding to different query intention types includes generating labels for actions that includes at least two or more of: (a) correcting typos in a search query, (b) selecting a best fit search result, (c) selecting an exact match search result, (d) identifying all possible results meeting a predetermined relevance threshold, (e) identifying all possible results, or (f) ranking and sorting all possible results by relevance.