Global explanations of machine learning model predictions for input containing text attributes
A determination is made that an explanatory data set for a common set of predictions generated by a machine learning model for records containing text tokens is to be provided. Respective groups of related tokens are identified from the text attributes of the records, and record-level prediction influence scores are generated for the token groups. An aggregate prediction influence score is generated for at least some of the token groups from the record-level scores, and an explanatory data set based on the aggregate scores is presented.
1 . A computer-implemented method, comprising:
receiving, at a machine learning service of a cloud computing environment, via one or more programmatic interfaces, a request for an explanation for respective predictions generated by a machine learning model with respect to a group of input records, wherein the machine learning model generates predictions for a plurality of different prediction sub-ranges based on attributes of the input records;
subdividing, at the machine learning service, content of at least a portion of the attributes of the input records of the group into respective attribute portions;
determining respective record-level prediction influence scores for individual ones of the attribute portions for input records of the group of input records for which the machine learning model predicts a particular sub-range of the plurality of different prediction sub-ranges; and
providing, by the machine learning service, an explanatory data set for the particular sub-range predicted by the machine learning model, wherein the explanatory data set includes a prediction influence score for a collection of related attribute portions determined from an aggregation of the respective record-level prediction influence scores determined for individual ones of the related attribute portions from different input records of the input records.
2 . The computer-implemented method as recited in claim 1 , wherein a first attribute of a first input record of the group of input records comprises a first text token and a second text token, wherein a particular attribute portion of the first attribute includes the first text token, and wherein the particular attribute portion does not include the second text token.
3 . The computer-implemented method as recited in claim 1 , further comprising:
receiving, at the machine learning service, via the one or more programmatic interfaces, an indication of a granularity at which content of one or more of the attributes of the input records of the group are to be subdivided, wherein said subdividing is performed at the granularity.
4 . The computer-implemented method as recited in claim 1 , wherein the explanatory data set is presented via a visualization in which a visual cue indicates respective magnitudes of prediction influence scores for one or more attribute portions of a first attribute of a first input record of the group of input records, including a particular attribute portion of the first attribute.
5 . The computer-implemented method as recited in claim 4 , wherein the visual cue comprises one or more of: a font size or a color.
6 . The computer-implemented method as recited in claim 1 , further comprising:
generating, at the machine learning service, the explanatory data set using at least a SHAP (Shapley additive explanation) based algorithm.
7 . The computer-implemented method as recited in claim 1 , wherein said subdividing the content of at least the portion of the attributes of the input records of the group into respective attribute portions comprises:
generating tree representations of the first-attributes of the first input records; and
pruning the tree representations.
8 . A system, comprising:
one or more computing devices;
wherein the one or more computing devices include instructions that upon execution on or across the one or more computing devices cause the one or more computing devices to:
receive, at a machine learning service of a cloud computing environment, via one or more programmatic interfaces, a request for an explanation for respective predictions generated by a machine learning model with respect to a group of input records, wherein the machine learning model generates predictions for a plurality of different prediction sub-ranges, based on attributes of the input records;
subdivide, at the machine learning service, content of at least a portion of the attributes of the input records of the group into respective attribute portions;
determine respective record-level prediction influence scores for individual ones of the attribute portions for input records of the group of input records for which the machine learning model predicts a particular sub-range of the plurality of different prediction sub-ranges; and
provide, by the machine learning service, an explanatory data set for the particular sub-range predicted by the machine learning model, wherein the explanatory data set includes a prediction influence score for a collection of related attribute portions determined from an aggregation of the respective record-level prediction influence scores determined for individual ones of the related attribute portions.
9 . The system as recited in claim 8 , wherein a first attribute of a first input record of the group of input records comprises a first text token and a second text token, wherein a particular attribute portion of the first attribute includes the first text token, and wherein the particular attribute portion does not include the second text token.
10 . The system as recited in claim 8 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:
receive, at the machine learning service, via the one or more programmatic interfaces, an indication of a granularity at which content of one or more of the attributes of the input records of the group are to be subdivided, wherein said subdividing is performed at the granularity.
11 . The system as recited in claim 8 , wherein the explanatory data set is presented via a visualization in which a visual cue indicates respective magnitudes of prediction influence scores for one or more attribute portions of a first attribute of a first input records of the group of input records, including a particular attribute portion of the first attribute.
12 . The system as recited in claim 11 , wherein the visual cue comprises one or more of: a font size or a color.
13 . The system as recited in claim 8 , wherein the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:
generate, at the machine learning service, the explanatory data set using at least a SHAP (Shapley additive explanation) based algorithm.
14 . The system as recited in claim 8 , wherein to subdivide the content of at least the portion of the attributes of the input records of the group into respective attribute portions, the one or more computing devices include further instructions that upon execution on or across the one or more computing devices further cause the one or more computing devices to:
generate tree representations of the attributes of the input records; and
prune the tree representations.
15 . One or more non-transitory computer-accessible storage media storing program instructions that when executed on or across one or more processors cause the one or more processors to:
receive, at a machine learning service of a cloud computing environment, via one or more programmatic interfaces, a request for an explanation for respective predictions generated by a machine learning model with respect to a group of input records, wherein the machine learning model generates predictions for a plurality of different prediction sub-ranges, based on attributes of the input records;
subdivide, at the machine learning service, content of at least a portion of the attributes of the input records of the group into respective attribute portions;
determine respective record-level prediction influence scores for individual ones of the attribute portions for input records of the group of input records for which the machine learning model predicts a particular sub-range of the plurality of different prediction sub-ranges; and
provide, by the machine learning service, an explanatory data set for the particular sub-range predicted by the machine learning model, wherein the explanatory data set includes a prediction influence score for collection of related attribute portions determined from an aggregation of the respective record-level prediction influence scores determined for individual ones of the related attribute portions from different input records of the input records.
16 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein a first attribute of a first input record of the group of input records comprises a first text token and a second text token, wherein a particular attribute portion of the first attribute includes the first text token, and wherein the particular attribute portion does not include the second text token.
17 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , storing further program instructions that when executed on or across one or more processors further cause the one or more processors to:
receive, at the machine learning service, via the one or more programmatic interfaces, an indication of a granularity at which content of one or more of the attributes of the input records of the group are to be subdivided, wherein said subdividing is performed at the granularity.
18 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , wherein the explanatory data set is presented via a visualization in which a visual cue indicates respective magnitudes of prediction influence scores for one or more attribute portions of a first attribute of a first input record of the group of input records, including a particular attribute portion of the first attribute.
19 . The one or more non-transitory computer-accessible storage media as recited in claim 18 , wherein the visual cue comprises one or more of: a font size or a color.
20 . The one or more non-transitory computer-accessible storage media as recited in claim 15 , storing further program instructions that when executed on or across one or more processors further cause the one or more processors to:
generate, at the machine learning service, the explanatory data set using at least a SHAP (Shapley additive explanation) based algorithm.