AXE evaluates explanations to avoid misleading Rashomon set model selection.
problem Evaluating explanations for Rashomon set models to avoid false selection.
method Proposed AXE method to evaluate explanation quality.
result AXE detects adversarial fairwashing with 100% success rate.
The paper proposes criteria and methods for evaluating and aggregating feature-based model explanations.
problem Lack of quantitative evaluation criteria for feature-based model explanations.
method Developed quantitative evaluation criteria (low sensitivity, high faithfulness, low complexity), devised a framework for aggregation, and derived a new aggregate Shapley value explanation function.
result A new aggregate Shapley value explanation function that minimizes sensitivity.
Paper proposes metrics to evaluate AI explanations without ground truth.
problem Challenges in evaluating neural network explanations without ground truth.
method Designs four metrics to evaluate explanation results.
result New insights into neural network interpretation methods.
Paper introduces new evaluation criteria for feature-based model explanations.
problem Establishing reliable feature importance explanations for models.
method Robustness analysis using smaller adversarial perturbations.
result New explanations that are necessary and sufficient for predictions.
New framework evaluates model explanations based on decision task improvement.
problem Evaluation of model explanations often misses practical value.
method Decision-theoretic framework quantifying three key values.
result Provides benchmarks and interprets human-AI decision support.
This article evaluates explanations without ground truth in IML.
problem Lack of ground truth for evaluating explanations in IML.
method Rigorously defined problem, reviewed existing efforts, summarized three aspects of explanation, designed a unified evaluation framework.
result Unified evaluation framework for different scenarios in practice.
GRETEL unifies GCE evaluation across various settings.
problem Lack of standardized evaluation for Graph Counterfactual Explanations.
method Unified framework for testing GCE methods in diverse settings.
result GRETEL promotes reproducible evaluations of GCE techniques.
Defines explanations for classifier outcomes using causal concepts.
problem Understanding classifier outcomes in a causal context.
method Proposes a new definition of explanation based on causality, compares it with existing notions, and evaluates it experimentally.
result Experimental evaluation shows the new definition's effectiveness on financial datasets.
Evaluates explanations of LTR models using decision paths and compares their accuracy.
problem Challenges in evaluating local explanations of LTR models due to lack of ground truth feature importance scores.
method Focuses on tree-based LTR models, extracts ground truth feature importance scores using decision paths, and compares them with explanation techniques.
result Explanation accuracy varies depending on the model and data point.
Evaluates local explanations using white-box models and log odds ratios.
problem Costly and subjective evaluation of local explanations.
method Benchmarking explanation techniques using log odds ratios.
result Explanation techniques' performance depends on model, dataset, data point, and normalization.
Unified metric c-Eval evaluates feature-based explanations by perturbation.
problem Lack of consensus on evaluating feature-based local explanations.
method Introduces c-Eval metric and framework to quantify explanation quality.
result c-Eval captures the importance of input features and is applicable in adversarial-robust models.
Study evaluates relevance metrics for similarity-based model explanations.
problem Providing understandable explanations for complex model predictions.
method Evaluated three relevance metrics using three tests.
result Cosine similarity of gradients performs best for explanations.
The study evaluates how well local explanations align with model predictions.
problem Capturing the faithfulness of local explanations to model predictions.
method Introducing consistency and sufficiency as properties, and developing quantitative measures and estimators.
result Quantitative measures of consistency and sufficiency depend on test-time data distribution.
New method to evaluate visual explanations from neural networks.
problem Lack of consensus on measuring effectiveness of visual explanations.
method Proposed a new procedure for evaluating explanations using a range of sources.
result Demonstrated the benefit of combining different sources and the impact of bias parameters.
ALMANACS benchmarks explainability methods on simulatability.
problem Evaluating the effectiveness of explainability methods for language models.
method ALMANACS is a simulatability benchmark that evaluates explainability methods on twelve safety-relevant topics.
result No explainability method outperforms the explanation-free control across all topics.
New method evaluates visual explanations of deep models using adversarial perturbations.
problem Lack of objective evaluation of visual explanations of deep models.
method Proposes an adversarial perturbation approach to evaluate visual explanations of deep models.
result Demonstrates the effectiveness of the proposed approach through comparisons with existing methods.
New definition reveals encoding explanations that retain predictive power.
problem Challenges in evaluating and identifying encoding explanations.
method Developed a definition of encoding based on conditional dependence.
result Existing evaluation scores do not rank non-encoding explanations correctly, but STRIPE-X does.
Two modified tests improve the reliability of evaluating explanation methods.
problem Methodological concerns in evaluating explanation methods for saliency maps.
method Proposed modifications to the Model Parameter Randomisation Test (MPRT): Smooth MPRT and Efficient MPRT.
result Enhanced metric reliability, facilitating more trustworthy deployment of explanation methods.
Study evaluates local explanation methods for time series forecasting.
problem Lack of local interpretability methods for multivariate time series forecasting.
method Proposed two novel evaluation metrics: Area Over the Perturbation Curve for Regression and Ablation Percentage Threshold.
result Comprehensive comparison of local explanation models on two datasets.
Transparency, user trust, and human comprehension are popular ethical motivations for interpretable machine learning. In support of these goals, researchers evaluate model explanation performance using humans and real world applications. This alone presents a challenge in many areas of artificial intelligence. In this …
Study evaluates counterfactual explanations using Pearl's method.
problem Bias in counterfactual explanations generated from machine learning models.
method Evaluates counterfactual explanations using Judea Pearl's counterfactual method.
result Thirty percent of counterfactual explanations conflicted with Pearl's method.
The paper evaluates methods for explaining deep learning in security.
problem Understanding the predictions of deep learning models in security applications.
method Developed criteria to compare and evaluate six explanation methods.
result Significant differences exist between the methods, leading to recommendations.
Study investigates how AI can create and detect deceptive explanations, finding they can fool humans but ML can detect them.
problem The risk of deceptive AI explanations increasing trust issues and economic risks.
method Investigates creation and detection of deceptive explanations using AI models and machine learning methods.
result Deceptive explanations can fool humans, but ML can detect them with high accuracy.
ACE improves counterfactual explanations with fewer model queries.
problem Inefficient sampling for counterfactual explanations in machine learning models.
method Adaptive sampling combining Bayesian estimation and stochastic optimization.
result ACE achieves superior evaluation efficiency compared to state-of-the-art methods.
ManifoldShap improves model explanations by restricting evaluations to the data manifold.
problem Inaccurate and misleading model explanations due to reliance on out-of-distribution data.
method Restricts model evaluations to the data manifold to avoid off-manifold perturbations.
result ManifoldShap provides more accurate and intuitive explanations than existing methods.
LLMs' explanations are often insufficient and vary with input distribution.
problem Evaluating the sufficiency of LLM explanations without predefined biases.
method Generalizing sufficiency to arbitrary explanations, using LLM's input beliefs, and introducing SCSuff metric.
result Explanation sufficiency can vary with input distribution and is weakly correlated with model size, accuracy, or output entropy.
A fast method finds interpretable counterfactual explanations using class prototypes.
problem Finding understandable counterfactual explanations for classifier predictions.
method Using class prototypes, the method speeds up and improves interpretability of counterfactual instances.
result The method significantly speeds up and improves the interpretability of counterfactual explanations.
Despite a growing literature on explaining neural networks, no consensus has been reached on how to explain a neural network decision or how to evaluate an explanation. Our contributions in this paper are twofold. First, we investigate schemes to combine explanation methods and reduce model uncertainty to obtain a sing…
Many methods to explain black-box models, whether local or global, are additive. In this paper, we study global additive explanations for non-additive models, focusing on four explanation methods: partial dependence, Shapley explanations adapted to a global setting, distilled additive explanations, and gradient-based e…
RAW-Explainer generates interpretable subgraph explanations for link predictions in knowledge graphs.
problem Interpreting GNN predictions for link prediction in heterogeneous settings is challenging.
method RAW-Explainer uses random walk objective and neural network to generate connected, concise subgraph explanations.
result RAW-Explainer strikes a balance between explanation quality and computational efficiency.
Post-hoc explanations of machine learning models are crucial for people to understand and act on algorithmic predictions. An intriguing class of explanations is through counterfactuals, hypothetical examples that show people how to obtain a different prediction. We posit that effective counterfactual explanations shoul…
Study shows SHAP explanations impact alert processing decisions but not performance.
problem Utility of SHAP explanations in alert processing by human experts.
method Human-grounded evaluation with three groups of participants, qualitative analysis of reflections.
result SHAP explanations impact decision-making but not alert processing performance.
Differentially private algorithms protect model explanations from leaking training data.
problem Model explanations can leak training data, compromising privacy.
method Adaptive differentially private gradient descent algorithm to produce accurate, private explanations.
result Privacy amplification and reduction of overall privacy loss on explanation data.
We consider objective evaluation measures of saliency explanations for complex black-box machine learning models. We propose simple robust variants of two notions that have been considered in recent literature: (in)fidelity, and sensitivity. We analyze optimal explanations with respect to both these measures, and while…
In many applications, an anomaly detection system presents the most anomalous data instance to a human analyst, who then must determine whether the instance is truly of interest (e.g. a threat in a security setting). Unfortunately, most anomaly detectors provide no explanation about why an instance was considered anoma…
Model explanations can leak sensitive training data information, posing privacy risks.
problem Privacy risks of model explanations that expose training data information.
method Membership inference attacks on feature-based model explanations.
result Backpropagation-based explanations reveal statistical information about decision boundaries, leaking training data membership.
New feature mapping approach improves recommendation accuracy and explainability.
problem Balancing recommendation accuracy and explainability using metadata.
method Maps uninterpretable features to interpretable aspect features, minimizing both prediction and interpretation losses.
result Strong performance in recommendation and explainability, eliminating metadata need.
New saliency evaluations focus on completeness and soundness, improving explanations.
problem Current saliency evaluations focus on completeness but ignore soundness.
method Introduces new intrinsic evaluation metrics based on completeness and soundness.
result Simple saliency method matches or outperforms prior methods in new evaluations.
TED framework teaches AI to explain decisions, improving accuracy.
problem Providing understandable explanations for AI predictions in high-stakes applications.
method Augmenting training data with explanations from domain users, using embeddings and multi-task learning.
result AI models can be taught to provide meaningful explanations, sometimes improving accuracy.
Study finds visual explanations do not significantly improve human accuracy or trust in model predictions.
problem Measuring the impact of visual explanations on human accuracy and trust in model predictions.
method Randomized controlled trial with image-based age prediction task, varying levels of explanation quality.
result Visual explanations do not significantly alter human accuracy or trust in the model.
A new method explains RNNs by decision lists over skipgrams, improving explanation fidelity and interpretability.
problem Lack of understanding how input segments combine to form patterns in neural network outputs.
method Proposes a pipeline to explain RNNs using decision lists over skipgrams, creating synthetic and real-world datasets for evaluation.
result Persistently achieves high explanation fidelity and interpretable rules.
Survey of counterfactual explanations for time series classification.
problem Generating plausible and actionable counterfactuals for time series data.
method Review of various counterfactual generation methods for time series classification.
result Highlight unique challenges and strengths of existing methods.
Paper proposes Coalitional BAE to improve explainability of unsupervised deep learning models.
problem Improving explainability of Autoencoder's predictions.
method Introduces Coalitional BAE, inspired by agent-based system theory, to reduce correlation in explanations.
result Improved quality of explanations using Coalitional BAE on publicly available datasets.
We study the problem of explaining a rich class of behavioral properties of deep neural networks. Distinctively, our influence-directed explanations approach this problem by peering inside the network to identify neurons with high influence on a quantity and distribution of interest, using an axiomatically-justified in…
We explain increases in clinical risk predictions over time.
problem Tackling the challenge of explaining dynamic risk increases in clinical settings.
method Developed methods to extend static attribution techniques to dynamic settings, addressing challenges specific to time-series data.
result Identified and addressed challenges specific to dynamic risk estimation, improving clinical alert explanations.
Paper examines risks of unjustified counterfactual explanations from post-hoc interpretability models.
problem Risks of generating unjustified counterfactual explanations from post-hoc interpretability models.
method Investigates local neighborhoods of instances for justification and compares state-of-the-art approaches.
result High risk of generating unjustified counterfactual examples, leading to less useful explanations.
SEMs fail to provide robust explanations to adversarial inputs.
problem Lack of robustness in interpretability of self-explaining models.
method Evaluation of current SEMs and creation of adversarial inputs.
result Adversarial inputs can cause significant changes in explanations without affecting model outputs.
Attack reveals model details from counterfactual explanations.
problem Extracting model details from counterfactual explanations.
method Adversary uses counterfactual explanations to build high-fidelity model.
result High-fidelity and high-accuracy model extraction possible.