LLMs' explanations are often insufficient and vary with input distribution.
problem Evaluating the sufficiency of LLM explanations without predefined biases.
method Generalizing sufficiency to arbitrary explanations, using LLM's input beliefs, and introducing SCSuff metric.
result Explanation sufficiency can vary with input distribution and is weakly correlated with model size, accuracy, or output entropy.
Unified feature importance for machine learning models tackles sufficiency and necessity limitations.
problem Insufficient and incomplete explanations of machine learning models.
method Formalized sufficiency and necessity notions, proposing a unified importance measure.
result Unified importance measure detects features missed by sufficiency and necessity alone.
P-SE explains model decisions with minimal feature subsets and fast estimators.
problem Explain model decisions in regression and classification.
method Probabilistic Sufficient Explanations (P-SE) with random Forests for conditional probability estimation.
result Consistent and efficient explanations for regression and classification models.
The study evaluates how well local explanations align with model predictions.
problem Capturing the faithfulness of local explanations to model predictions.
method Introducing consistency and sufficiency as properties, and developing quantitative measures and estimators.
result Quantitative measures of consistency and sufficiency depend on test-time data distribution.
This paper introduces a method to find complete and interpretable concept-based explanations for deep neural networks.
problem Lack of complete and interpretable concept-based explanations in deep neural networks.
method Definition of completeness, concept discovery method, and importance score calculation using game-theoretic notions.
result The proposed method finds complete and interpretable concept-based explanations for deep neural networks.
Paper introduces new evaluation criteria for feature-based model explanations.
problem Establishing reliable feature importance explanations for models.
method Robustness analysis using smaller adversarial perturbations.
result New explanations that are necessary and sufficient for predictions.
Study investigates how AI can create and detect deceptive explanations, finding they can fool humans but ML can detect them.
problem The risk of deceptive AI explanations increasing trust issues and economic risks.
method Investigates creation and detection of deceptive explanations using AI models and machine learning methods.
result Deceptive explanations can fool humans, but ML can detect them with high accuracy.
Study shows more data improves model explanations, aiding reliable knowledge extraction.
problem Challenges in deriving reliable knowledge from machine learning models due to the Rashōmon effect.
method Examined the influence of sample size on explanations from models in a Rashōmon set using SHAP.
result Explanations from <128 samples are highly variable, but agreement improves with more data.
New algorithms explain Naive Bayes classifiers in polynomial time and delay.
problem Computing explanations for Naive Bayes classifiers efficiently.
method Developed log-linear time and polynomial delay algorithms for PI-explanations.
result Efficiently computed PI-explanations for linear classifiers.
Regularizes black-box models to improve interpretability.
problem Improving interpretability of black-box models without sacrificing accuracy.
method Regularizes a black-box model at training time to connect model explainability, explanation system, and quality metrics.
result Substantial improvement in explanation fidelity and stability across various datasets and explanation systems, with slight accuracy trade-off.
PiNets provide faithful explanations for neural networks.
problem Lack of true explanations for neural network predictions.
method Pointwise-interpretable Networks (PiNets) that form linear models instance-wise.
result PiNets offer explanations that are meaningful, aligned, robust, and sufficient.
Improves local model explanations using GANs and Linear Model Trees.
problem Need for accurate and intuitive explanations of complex machine learning models.
method Generative Adversarial Network (GAN) for synthetic data generation and Linear Model Trees for surrogate model training.
result Significantly improved local model explanations with contextual information.
Two oppositely charged droplets of (say) water in e.g. oil or air will tend to drift together under the influence of their charges. As they make contact, one might expect them to coalesce and form one large droplet, and this indeed happens when the charge difference is sufficiently small. However, Ristenpart et al disc…
New method explains deep model decisions by adding latent features.
problem Limitations of existing local explanation methods for deep models.
method Leveraging latent features for contrastive local explanations.
result Quantitatively superior explanations on diverse datasets.
Paper offers local neural network explanations by considering model architecture.
problem Creating true neighbourhoods for black box models.
method Penultimate layer decoding for local neighbourhood generation.
result Local explanations are more accurate and relevant to instances.
Bayesian framework for solar magnetogram super-resolution with uncertainty quantification.
problem Uncertainty in super-resolving solar magnetic field images.
method Bayesian decomposition of uncertainties into epistemic and aleatoric.
result Generation of maps measuring the range of possible high-resolution explanations.
Proposes a method to classify with missing features using expected predictions.
problem Challenges of missing feature values in classifier performance.
method Computes expected predictions using geometric programming to learn a naive Bayes distribution.
result Achieves performance similar to full feature classifiers and outperforms imputation techniques.
Triangulation filters spurious circuits in multilingual models.
problem Unreliable explanations of multilingual models across languages.
method Formalizes reference families and introduces triangulation as a causal acceptance rule.
result Triangulation provides a falsifiable standard for mechanistic claims.
Improved credit scoring model with explainability.
problem Making financial decisions based on loan applications.
method Extreme Gradient Boosting (XGBoost) model with 360-degree explanation framework.
result Model achieves state-of-the-art performance and provides understandable explanations.
Generalizes moment-matching for exponential families with conditioning or hidden data.
problem Generalizing moment-matching conditions for exponential families with conditioning or hidden data.
method First-principles explanation and self-contained derivation of generalized moment-matching conditions.
result Derives generalized moment-matching conditions for conditional exponential families and hidden data.
Many methods to explain black-box models, whether local or global, are additive. In this paper, we study global additive explanations for non-additive models, focusing on four explanation methods: partial dependence, Shapley explanations adapted to a global setting, distilled additive explanations, and gradient-based e…
Prediction and explanation are key objects in supervised machine learning, where predictive models are known as black boxes and explanatory models are known as glass boxes. Explanation provides the necessary and sufficient information to interpret the model output in terms of the model input. It includes assessments of…
AXE evaluates explanations to avoid misleading Rashomon set model selection.
problem Evaluating explanations for Rashomon set models to avoid false selection.
method Proposed AXE method to evaluate explanation quality.
result AXE detects adversarial fairwashing with 100% success rate.
The paper proposes criteria and methods for evaluating and aggregating feature-based model explanations.
problem Lack of quantitative evaluation criteria for feature-based model explanations.
method Developed quantitative evaluation criteria (low sensitivity, high faithfulness, low complexity), devised a framework for aggregation, and derived a new aggregate Shapley value explanation function.
result A new aggregate Shapley value explanation function that minimizes sensitivity.
New definition reveals encoding explanations that retain predictive power.
problem Challenges in evaluating and identifying encoding explanations.
method Developed a definition of encoding based on conditional dependence.
result Existing evaluation scores do not rank non-encoding explanations correctly, but STRIPE-X does.
New findings show margins are not sufficient for explaining gradient boosting performance.
problem The inadequacy of margin explanations in explaining the performance of gradient boosting.
method Demonstrated and proved a stronger margin-based generalization bound for boosted classifiers.
result Proved a stronger margin-based generalization bound that explains the performance of modern gradient boosters.
Proposes new measures and methods for evaluating and improving explanations of machine learning models.
problem Evaluating and improving explanations of complex machine learning models.
method Introduces two new measures: infidelity and sensitivity, and proposes methods to optimize these measures.
result Optimal explanations for infidelity involve a novel combination of two methods, and can outperform existing explanations.
Combines neural network explanation methods to improve robustness and accuracy.
problem Lack of consensus on explaining neural network decisions and evaluating explanations.
method Investigates schemes to aggregate explanation methods and reduce model uncertainty.
result Aggregated explanations are better at identifying important features and more robust to adversarial attacks.
Survey on efficient counterfactual explanations for various ML models.
problem Providing understandable explanations for machine learning predictions.
method Review and propose methods for computing counterfactual explanations.
result Efficient methods for various ML models and new methods for unconsidered models.
GRANITE unifies feature-based explanation methods to reduce disagreement.
problem Disagreement among feature-based explanation methods.
method GRANITE partitions feature space into regions minimizing interaction and distribution influences.
result Unified and consistent feature explanations.
Efficiently computes counterfactual explanations for LVQ models.
problem Need to efficiently explain predictions of machine learning models.
method Derives convex and non-convex programs for LVQ models.
result Efficient computation of counterfactual explanations for LVQ models.
Model explanations can leak sensitive training data information, posing privacy risks.
problem Privacy risks of model explanations that expose training data information.
method Membership inference attacks on feature-based model explanations.
result Backpropagation-based explanations reveal statistical information about decision boundaries, leaking training data membership.
Differentially private algorithms protect model explanations from leaking training data.
problem Model explanations can leak training data, compromising privacy.
method Adaptive differentially private gradient descent algorithm to produce accurate, private explanations.
result Privacy amplification and reduction of overall privacy loss on explanation data.
Paper proposes metrics to evaluate AI explanations without ground truth.
problem Challenges in evaluating neural network explanations without ground truth.
method Designs four metrics to evaluate explanation results.
result New insights into neural network interpretation methods.
Paper explores a consumer-friendly approach to explain machine learning decisions.
problem Challenges in providing understandable explanations for machine learning predictions.
method Consumer-driven approach called TED that asks for explanations in training data.
result TED is robust to increasing numbers of explanations, noisy explanations, and missing explanations.
Personalized explanations improve understanding of machine learning models.
problem Improving human understanding of machine learning models and decisions.
method Deriving a conceptualization of personalized explanation, categorizing explainee data, identifying key properties, and introducing new measures.
result Identification of three key properties amendable to personalization: complexity, decision information, and presentation.
Formalizes explanations as blending input and model output.
problem Creating clear and consistent explanations for model predictions.
method Defines properties of explanation functions and links them to model layers.
result Consistency of activations across layers implies consistency of explanations.
The paper investigates the limitations of additive explanations for complex models.
problem The trustworthiness of additive explanations for non-additive models.
method Examine and introduce a new method to detect interactions for instance-level explanations.
result Additive explanations can be misleading for non-additive models.
Improves global counterfactual explanations for model recourse.
problem Inability to provide explanations beyond local instances.
method Investigates and improves Actionable Recourse Summaries (AReS) for global counterfactual explanations.
result Develops more efficient and interactive explainability tools.
New framework evaluates model explanations based on decision task improvement.
problem Evaluation of model explanations often misses practical value.
method Decision-theoretic framework quantifying three key values.
result Provides benchmarks and interprets human-AI decision support.
We define and compute plausible counterfactual explanations using density constraints.
problem Efficiently compute plausible counterfactual explanations for machine learning models.
method Propose and study a formal definition of plausible counterfactual explanations, use density estimators, and introduce convex density constraints.
result Convex density constraints ensure plausible and feasible counterfactual explanations.
G-SHAP generates multiple types of explanations for machine learning models.
problem Understanding model predictions and their differences across groups.
method Generalization of SHAP method to produce additional types of explanations.
result G-SHAP produces explanations for classification, intergroup differences, and model failure.
Study finds visual explanations do not significantly improve human accuracy or trust in model predictions.
problem Measuring the impact of visual explanations on human accuracy and trust in model predictions.
method Randomized controlled trial with image-based age prediction task, varying levels of explanation quality.
result Visual explanations do not significantly alter human accuracy or trust in the model.
Manipulated explanations can be made imperceptible to the naked eye.
problem The trustworthiness and interpretability of neural networks can be compromised by manipulable explanations.
method The paper demonstrates that explanations can be altered by applying small, imperceptible input changes that do not affect the network's output.
result An upper bound on the susceptibility of explanations to manipulation has been derived, and effective mechanisms to enhance explanation robustness have been proposed.
Local explanation frameworks aim to rationalize particular decisions made by a black-box prediction model. Existing techniques are often restricted to a specific type of predictor or based on input saliency, which may be undesirably sensitive to factors unrelated to the model's decision making process. We instead propo…
The paper introduces a new framework for making machine learning explanations more understandable to humans.
problem Making machine learning explanations comprehensible and aligned with human preferences.
method Inspired by philosophy, cognitive science, and social sciences, the paper formalizes a framework using the concept of 'weight of evidence' from information theory.
result The framework produces intuitive and comprehensible explanations that align with human preferences.
Proposes a method to generate counterfactual and contrastive explanations using SHAP.
problem Need for explainable AI and legal requirement for model interpretability.
method Model agnostic method using SHAP to generate contrastive and counterfactual explanations.
result Demonstrates effectiveness of the method on various datasets.
Managing large-scale transportation infrastructure projects is difficult due to frequent misinformation about the costs which results in large cost overruns that often threaten the overall project viability. This paper investigates the explanations for cost overruns that are given in the literature. Overall, four categ…