New method interprets object representations from human behavior.
problem Understanding how mental object representations relate to human behavior.
method Sparse, non-negative representations of objects estimated from behavioral judgments.
result Representations predict latent object similarity and are interpretable.
Study improves understanding of what makes machine learning explanations human-interpretable.
problem Understanding what makes explanations human-interpretable in machine learning systems.
method Controlled human-subject experiments to identify regularizers for interpretability across three tasks.
result Cognitive chunks affect performance more than variable repetitions, suggesting common design principles.
We often desire our models to be interpretable as well as accurate. Prior work on optimizing models for interpretability has relied on easy-to-quantify proxies for interpretability, such as sparsity or the number of operations required. In this work, we optimize for interpretability by directly including humans in the …
New definition of interpretability for deep neural networks.
problem Vague definition of interpretability for deep neural networks.
method Proposed a new definition of human predictability for interpretability of DNNs.
result Our definition will help to the research of interpretability of DNNs.
IRL models human risk decisions based on past outcomes.
problem Understanding human risk decisions under risk.
method Inverse Reinforcement Learning (IRL) with features reflecting state history.
result Human reward function explains risk-prone and risk-averse decisions.
Meta-learning approach to learn interpretable models from human feedback.
problem Tackling the challenge of making machine learning models interpretable.
method A meta-learning approach where a model of non-trivial proxies of human interpretability is learned from human feedback, then incorporated into the ML training process to optimize for interpretability.
result The approach leads to formulas that are either significantly more or equally accurate while being more interpretable.
Quantifies interpretability and trust in ML decisions.
problem Measuring the quality and trustworthiness of ML interpretability methods.
method Proposes a quantitative measure based on information transfer rate and empirical validation.
result Empirical evidence shows the proposed metric differentiates interpretability methods and improves productivity.
This is the Proceedings of the 2017 ICML Workshop on Human Interpretability in Machine Learning (WHI 2017), which was held in Sydney, Australia, August 10, 2017. Invited speakers were Tony Jebara, Pang Wei Koh, and David Sontag.
Tree regularization makes deep models interpretable by approximating them with simple decision trees.
problem Lack of interpretability in deep neural networks.
method Tree regularization to train deep models to resemble compact, axis-aligned decision trees.
result Tree regularized models are easier for humans to interpret without sacrificing accuracy.
The paper introduces a new framework for making machine learning explanations more understandable to humans.
problem Making machine learning explanations comprehensible and aligned with human preferences.
method Inspired by philosophy, cognitive science, and social sciences, the paper formalizes a framework using the concept of 'weight of evidence' from information theory.
result The framework produces intuitive and comprehensible explanations that align with human preferences.
Transparency, user trust, and human comprehension are popular ethical motivations for interpretable machine learning. In support of these goals, researchers evaluate model explanation performance using humans and real world applications. This alone presents a challenge in many areas of artificial intelligence. In this …
This is the Proceedings of the 2018 ICML Workshop on Human Interpretability in Machine Learning (WHI 2018), which was held in Stockholm, Sweden, July 14, 2018. Invited speakers were Barbara Engelhardt, Cynthia Rudin, Fernanda Viégas, and Martin Wattenberg.
New algorithm improves interpretability in sequence classification.
problem Lack of human-independent interpretability metrics in sequence classification.
method Combines linear classifiers with background knowledge embeddings to create a new feature space.
result Preserves predictive power while delivering more interpretable models.
This is the Proceedings of the 2016 ICML Workshop on Human Interpretability in Machine Learning (WHI 2016), which was held in New York, NY, June 23, 2016. Invited speakers were Susan Athey, Rich Caruana, Jacob Feldman, Percy Liang, and Hanna Wallach.
Algorithm learns from human demonstrations to schedule tasks efficiently.
problem Efficient resource scheduling in dynamic environments.
method Personalized apprenticeship learning framework infers decision-making criteria from heterogeneous human demonstrations.
result Achieves high accuracy in synthetic and real-world domains, outperforming baselines.
New framework evaluates model explanations based on decision task improvement.
problem Evaluation of model explanations often misses practical value.
method Decision-theoretic framework quantifying three key values.
result Provides benchmarks and interprets human-AI decision support.
New models predict mobility flows as well as complex machine learning but are simpler and interpretable.
problem Incomplete understanding and modeling of human mobility flows.
method Developed simple machine-learned, closed-form models of mobility.
result These models predict mobility flows more accurately than gravity or complex machine/deep learning models.
LCBM model improves image classification without human supervision.
problem Improving interpretability and generalization of unsupervised concept-based models.
method LCBM models concepts as random variables in a Bernoulli latent space, reducing the number of concepts without sacrificing performance.
result LCBM outperforms existing models in generalization and interpretability.
DAFT models attention as a dynamical system to make neural networks more interpretable.
problem Uninterpretable features learned by neural networks without human priors.
method DAFT models attention as a continuous dynamical system using neural ODEs.
result DAFT reduces the number of reasoning steps while maintaining similar performance.
Geometric framework detects concept frustration between human concepts and machine representations.
problem Aligning human concepts with machine learning representations.
method Geometric framework and similarity measures for detecting concept frustration.
result Concept frustration affects machine learning model performance and reorganizes learned concept representations.
Study assesses human interpretability of machine learning models.
problem Ensuring machine learning models are understandable by humans.
method User study with 1,000 participants testing simulatability and 'what if' local explainability.
result Increased runtime operation count correlates with decreased human accuracy on local interpretability tasks.
Proposes using Dynamic Mode Decomposition with delays for short-term human motion anticipation.
problem Lack of interpretability and explainability in neural network-based motion anticipation methods.
method Dynamic Mode Decomposition with delays for motion representation and prediction.
result Anticipation errors comparable or better than recurrent neural networks for very short times.
Interpretable companion model for black-box classifiers.
problem Dilemma between interpretable and black-box models.
method Trains a companion model from data and black-box model predictions, optimizing a combination of accuracy and complexity.
result Companion model provides interpretable predictions with a slight accuracy loss for user choice.
This research improves interpretability in sequential explanations using mental models.
problem Improving interpretability in sequential explanations between two parties.
method A reinforcement learning framework that selects explanations based on the explainee's mental model.
result Mental model-based policies increase interpretability over random selection in multiple sequential explanations.
VICE embeds concepts in a vector space using human data.
problem Developing numerical models for mental representations of object concepts.
method Variational Interpretable Concept Embeddings (VICE) using variational inference and triplet odd-one-out task data.
result VICE outperforms SPoSE at predicting human behavior and provides more reproducible object representations.
A novel memory mechanism for reinforcement learning agents that stores past events in human-readable language.
problem Lack of interpretability in reinforcement learning agent's memory mechanisms.
method Uses CLIP to associate visual inputs with language tokens, then feeds these tokens to a pretrained language model.
result Significantly faster convergence on challenging continuous recognition tasks.
As artificial intelligence is increasingly affecting all parts of society and life, there is growing recognition that human interpretability of machine learning models is important. It is often argued that accuracy or other similar generalization performance metrics must be sacrificed in order to gain interpretability.…
LFD method improves text classification by making features clearer and less label-leaking.
problem Creating interpretable text representations that are both predictive and understandable.
method LFD method: proposes lexical and semantic features from contrastive text pairs, screens candidates using κ, and selects features by residual gain. result LFD features achieve higher human-human and human-LLM agreement than baseline concepts and are less label-leaking.
Defines interpretability in machine learning and introduces PDR framework.
problem Confusion about interpretability and lack of common evaluation criteria.
method PDR framework: predictive, descriptive, relevancy.
result Provides a common vocabulary for interpretation methods.
We create interpretable word embeddings through sparse coding.
problem Difficult to interpret word embeddings in natural language processing.
method Transform pretrained dense word embeddings into sparse embeddings through sparse coding.
result Sparse embeddings are more interpretable and achieve good performance.
NormLIME improves feature importance explanations for deep neural networks.
problem Improving local feature explanations for deep learning models.
method NormLIME aggregates local models into global and class-specific interpretations.
result NormLIME outperforms other feature importance metrics in human user studies and numerical experiments.
At the core of interpretable machine learning is the question of whether humans are able to make accurate predictions about a model's behavior. Assumed in this question are three properties of the interpretable output: coverage, precision, and effort. Coverage refers to how often humans think they can predict the model…
A human-in-the-loop ML framework for precision dosing reduces expert workload and removes bias.
problem High cost of data annotation and lack of appropriate data for ML models.
method Incorporates human experts into the model learning loop to improve interpretability and reduce bias.
result The approach learns interpretable rules from data and potentially lowers expert workload.
This paper tackles the interpretability gap between deep learning models and human cognition.
problem The gap between deep learning models and human cognitive competence.
method A universal learning framework is proposed to evaluate and improve model interpretability.
result The uniqueness of solution and conditions for unique solution are proved.
This paper refines human labeling as a measurement process, revealing four sources of variation.
problem Systematic variation in human labeling obscures model learning.
method Introduces a statistical framework to decompose labeling outcomes.
result Empirical evidence for four components of labeling variation.
The paper offers a checklist for comparing human and machine visual perception.
problem Comparing human and machine visual perception.
method Designing, conducting, and interpreting experiments to investigate mechanisms.
result Feedback mechanisms may not be necessary for visual reasoning tasks.
While the interpretability of machine learning models is often equated with their mere syntactic comprehensibility, we think that interpretability goes beyond that, and that human interpretability should also be investigated from the point of view of cognitive science. The goal of this paper is to discuss to what exten…
Detects potential local adversarial examples to prevent fraud in credit insurance decisions.
problem Manipulating variables to gain unfair advantages in credit decisions.
method Provides critical features to a human expert to detect and control potential fraud.
result Demonstrates a method to identify and mitigate local adversarial examples.
New framework for explainable AI on high-dimensional data.
problem Challenges in explainability with high-dimensional data.
method Two modules: latent representation and Shapley paradigm adaptation.
result Interpretable model explanations for high-dimensional data.
The paper proposes a method to find interpretable subspaces in node embeddings using a knowledge base.
problem Finding interpretable subspaces in unsupervised node embeddings.
method Using a taxonomy of human-understandable concepts from a knowledge base to identify subspaces in node embeddings.
result Low error in finding fine-grained concepts.
Machine-generated interpretations do not improve users' guessing accuracy in image classifiers.
problem Determining the usefulness of machine-generated explanations for deep neural networks.
method Human evaluation of crowd workers guessing incorrectly predicted labels with and without visual interpretations.
result Showing machine-generated visual interpretations decreased average guessing accuracy by about 10%.
Model learns disease self-representations for drug repositioning.
problem Drug repositioning for disease treatment.
method Enforces proximity in disease self-representations to preserve human phenome network structure.
result Method outperforms state-of-the-art approaches and produces biologically interpretable disease self-representations.
FinHEAR combines LLMs with human expertise for better financial decision-making.
problem Challenges in financial decision-making for language models.
method Multi-agent framework with specialized LLMs for historical analysis, event interpretation, and expert retrieval.
result FinHEAR outperforms baselines in financial tasks with higher accuracy and risk-adjusted returns.
Framework enhances AI explainability by aligning with human cognitive models.
problem Lack of explainability in AI models hinders trust and accountability.
method Integrates explainability techniques with Malle's five category model of behavior explanation.
result Demonstrates practical relevance in credit risk assessment and regulatory analysis.
Develops a learning algorithm for PSL formulas from examples.
problem Learning human-interpretable descriptions of complex systems from examples.
method Reduces learning to propositional logic constraint satisfaction and uses SAT solver.
result Proposed method provides succinct human-interpretable descriptions from examples.
Extracts decision trees from CNNs to explain concept importance.
problem Understanding how CNNs make decisions about human-understandable concepts.
method Inferring labeled concept data from CNN hidden layer activations and creating a shallow decision tree.
result Extracted decision trees accurately represent CNN classifications.
End-to-end learnable network for safer self-driving with interpretable intermediate representations.
problem Safe motion planning for self-driving vehicles.
method Differentiable semantic occupancy representation for cost calculation in motion planning.
result Significantly outperforms state-of-the-art planners in imitating human behaviors and producing safer trajectories.
Study examines human factors in radiographic testing to improve inspection performance.
problem Insufficient consideration of human and organizational factors in NDT.
method CREAM method applied to analyze and model HOF on radiogram interpretation tasks.
result Model CREAM well-adapted for estimating HOF impact on NDT performances.