Multi-facet actions for improved conversational item search refinement
Examples provide conversational item search refinement using multi-facet filtering of items in a catalog. A multi-facet filter manager extracts facets and actions corresponding to the facets from a user utterance. The facet-actions include an entity-role and one or more filter actions associated with the facets. A facet-action includes filter actions such as exact, exclude, greater than, less than, etc. Multi-facet filters corresponding to the facet-actions are applied to a plurality of items in the catalog. The candidate items remaining after filtering are scored. The scores indicate relevance of each candidate item. One or more of the candidate items with the highest scores are selected. The selected items are added to search results which are returned to the user in response to the conversational search query. The multi-facet filter manager enables faster and more accurate search results using fewer conversational turns for reduced system resource usage and improved user efficiency.
1 . A system for conversational item search using multi-facet actions, the system comprising:
a processor; and
a computer-readable medium storing instructions that are operative upon execution by the processor to:
obtain a single utterance of a user comprising a conversational search query, the conversational search query comprising a plurality of words identifying an item and a plurality of search refinement terms associated with the item;
convert the single utterance into a bidirectional sequence of tokens using a natural-language model comprising a set of stacked transformer encoder layers and a conditional random field (CRF) decoder, the bidirectional sequence of tokens representing the plurality of words;
extract, from the bidirectional sequence of tokens, a plurality of facet-actions, each facet-action comprising an entity corresponding to an item attribute, and entity-role identifying a semantic refinement value associated with the entity, and a mapped filter action selected from a set of filter operations including exact, exclude, greater-than, less-than, and range;
identify, based on the plurality of facet-actions, a corresponding plurality of multi-facet filters;
apply the plurality of multi-facet filters to a plurality of items in a catalog, the plurality of multi-facet filters removing items that are irrelevant to the conversational search query;
identify a plurality of candidate items from the plurality of items remaining after application of the plurality of multi-facet filters;
score the plurality of candidate items using scoring criteria comprising a semantic-similarity measure between the conversational search query and item-attribute data associated with each candidate item;
select a candidate item from the plurality of candidate items having a highest score; and
generate a search result comprising the selected candidate item, wherein the search result is presented to the user via a user interface.
2 . The system of claim 1 , wherein the instructions are further operative to:
convert the single utterance into a bidirectional tokenized query string comprising a plurality of bidirectional tokens representing the plurality of words, wherein an extractor analyzes the plurality of bidirectional tokens to extract the plurality of facet-actions from the single utterance.
3 . The system of claim 1 , wherein the instructions are further operative to:
identify the entity associated with an attribute of the item;
recognize an entity role corresponding to the entity; and
map the entity and the entity role to a set of actions, wherein a facet-action is associated with the entity, the entity role, and the set of actions mapped to the entity role.
4 . The system of claim 1 , wherein scoring the plurality of candidate items further comprises:
generate two sub-scores associated with each candidate item, wherein the two sub-scores are weighted and summed to obtain a score for each candidate item.
5 . The system of claim 1 , wherein the instructions are further operative to:
identify a first role associated with the entity and a second role associated with the entity;
identify a first action associated with the entity and the first role; and
identify a second action associated with the entity and the second role, wherein a first facet-action corresponds to the first action and a second facet-action corresponds to the second action.
6 . The system of claim 1 , wherein the instructions are further operative to:
identify a price entity, wherein the price entity is associated with a first numeric value role and a second numeric value role, the first numeric value role being lower than the second numeric value role; and
identify a first facet-action of greater than for the first numeric value role and a second facet-action of less than for the second numeric value role, wherein a first filter associated with the first facet-action and a second filter associated with the second facet-action are applied to remove items having a price that falls outside a specified price range.
7 . The system of claim 1 , wherein identifying the entity, recognizing the entity-role, and mapping the entity and the entity-role to the plurality of facet-actions is performed by a named-entity recognition model comprising the set of stacked transformer encoder layers and the conditional random field (CRF) decoder that jointly predicts entity labels and entity-role labels using a BILOU tagging scheme.
8 . A method for conversational item search using multi-facet actions, the method comprising:
obtaining a single utterance of a user comprising a conversational search query, the conversational search query comprising a plurality of words identifying an item and a plurality of search refinement terms associated with the item;
converting the single utterance into a bidirectional sequence of tokens using a natural-language model comprising a set of stacked transformer encoder layers and a conditional random field (CRF) decoder, the bidirectional sequence of tokens representing the plurality of words;
extracting, from the bidirectional sequence of tokens, a plurality of facet-actions corresponding to the plurality of search refinement terms from the plurality of words, a facet comprising an item attribute, wherein a facet-action comprises a set of filter actions corresponding to the item attribute, each facet-action comprising an entity corresponding to the item attribute, an entity-role identifying a semantic refinement value associated with the entity, and a mapped filter action selected from a set of filter operations including exact, exclude, greater-than, less-than, and range;
identifying, based on the plurality of facet-actions, a corresponding plurality of multi-facet filters;
applying the plurality of multi-facet filters to a plurality of items in a catalog, the plurality of multi-facet filters corresponding to the plurality of facet-actions removing items that are irrelevant to the conversational search query;
identifying a plurality of candidate items from the plurality of items remaining after application of the plurality of multi-facet filters;
scoring the plurality of candidate items using a set of scoring criteria comprising a semantic-similarity measure between the conversational search query and item-attribute data associated with each candidate item;
selecting a candidate item from the plurality of candidate items having a highest score; and
generating a search result comprising the selected candidate item, wherein the search result is presented to the user via a user interface.
9 . The method of claim 8 , wherein identifying the entity and the entity-role from the bidirectional sequence of tokens is performed by a named-entity recognition model comprising the set of stacked transformer encoder layers and the conditional random field (CRF) decoder that jointly predicts entity labels and entity-role labels using a BILOU tagging scheme.
10 . The method of claim 8 , wherein training of the natural-language model comprises optimizing a multi-task loss function including an entity-classification loss, an entity-role loss, and a masked-token prediction loss, a total loss being a sum of the entity-classification loss, the entity-role loss, and a mask-modeling loss.
11 . The method of claim 10 , wherein optimizing the masked-token prediction loss comprises masking 15% of tokens in each input sequence such that, for each selected token, 70% are replaced with a mask token, 10% are replaced with a random token, and 20% are left unchanged.
12 . The method of claim 8 , further comprising:
generating two sub-scores associated with each candidate item;
weighting the two sub-scores; and
summing to obtain a score for each candidate item.
13 . The method of claim 8 , further comprising:
identifying a first role associated with the entity forming a first entity-role and a second role associated with the entity forming a second entity-role; and
identifying a first action associated with the first entity-role and a second action associated with the second entity-role, wherein a first facet-action corresponds to the first action, and wherein a second facet-action corresponds to the second action.
14 . The method of claim 8 , wherein a facet-action in the plurality of facet-actions comprises at least one of a brand, price, rating, discount, color, material, and size associated with an item, and wherein a facet-action in the plurality of facet-actions comprises at least one of a greater than action, a less than action, a range action, an exclude action, and an exact match action.
15 . One or more computer storage devices having computer-executable instructions stored thereon, which, upon execution by a computer, cause the computer to perform operations comprising:
obtaining a single utterance of a user comprising a conversational search query, the conversational search query comprising a plurality of words identifying an item and a plurality of search refinement terms associated with the item;
converting the single utterance into a bidirectional sequence of tokens using a natural-language model comprising a set of stacked transformer encoder layers and a conditional random field (CRF) decoder, the bidirectional sequence of tokens representing the plurality of words;
extracting, from the bidirectional sequence of tokens, a plurality of facet-actions corresponding to the plurality of search refinement terms from the plurality of words, a facet comprising an item attribute, wherein a facet-action comprises a set of filter actions corresponding to the item attribute, each facet-action comprising an entity corresponding to the item attribute, an entity-role identifying a semantic refinement value associated with the entity, and a mapped filter action selected from a set of filter operations including exact, exclude, greater-than, less-than, and range;
identifying, based on the plurality of facet-actions, a corresponding plurality of multi-facet filters;
applying the plurality of multi-facet filters to a plurality of items in a catalog, the plurality of multi-facet filters corresponding to the plurality of facet-actions removing items that are irrelevant to the conversational search query;
identifying a plurality of candidate items from the plurality of items remaining after application of the plurality of multi-facet filters;
scoring the plurality of candidate items using a set of scoring criteria comprising a semantic-similarity measure between the conversational search query and item-attribute data associated with each candidate item;
selecting a candidate item from the plurality of candidate items having a highest score; and
generating a search result comprising the selected candidate item, wherein the search result is presented to the user via a user interface.
16 . The one or more computer storage devices of claim 15 , wherein the operations further comprise:
training a machine learning model using a set of customized training data to filter items based on multiple item attributes and multiple filter actions.
17 . The one or more computer storage devices of claim 15 , wherein the operations further comprise:
converting the single utterance into a bidirectional tokenized query string comprising a plurality of bidirectional tokens representing the plurality of words, wherein an extractor analyzes the plurality of bidirectional tokens to extract the plurality of facet-actions from the single utterance.
18 . The one or more computer storage devices of claim 15 , wherein the operations further comprise:
identifying the entity associated with an attribute of the item;
recognizing an entity role corresponding to the entity; and
identifying a set of actions corresponding to the entity and the entity role, wherein the facet-action is associated with the entity, the entity role, and the set of actions mapped to the entity role.
19 . The one or more computer storage devices of claim 15 , wherein the operations further comprise:
generating two sub-scores associated with each candidate item, wherein the two sub-scores are weighted and summed to obtain a score for each candidate item.
20 . The one or more computer storage devices of claim 15 , wherein the operations further comprise:
generating a first facet-action corresponding to a first entity and first role associated with the first entity;
generating a second facet-action corresponding to the first entity and a second role associated with the first entity; and
applying a first filter to a set of items, wherein the first filter corresponds to the first facet-action; and
applying a second filter to the set of items, wherein the second filter corresponds to the second facet-action.