Debias concept-based explanations by removing confounding information.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Latent representations are the essence of deep generative models and determine their usefulness and power. For latent representations to be useful as generative concept representations, their latent space must support latent space interpolation, attribute vectors and concept vectors, among other things. We investigate …
Concept modulation models unify identifiability and extrapolation in conditional latent variable models.
This paper analyzes the variability of Concept Activation Vectors (CAVs).
Interpretable surrogates of black-box predictors trained on high-dimensional tabular datasets can struggle to generate comprehensible explanations in the presence of correlated variables. We propose a model-agnostic interpretable surrogate that provides global and local explanations of black-box classifiers to address …
LCBM model improves image classification without human supervision.
Formalizes concepts as latent variables in hierarchical models for high-dimensional data.
This work explains how linear representations in large language models arise from training objectives and gradient descent.
Method selects features robust to concept shift using Shapley values.
Common statistical prediction models often require and assume stationarity in the data. However, in many practical applications, changes in the relationship of the response and predictor variables are regularly observed over time, resulting in the deterioration of the predictive performance of these models. This paper …
Framework learns interpretable concepts from data without interventions.
Adapts to shifts in latent subgroup distributions without labeled target data.
PDD detects concept drift using explainable AI, improving model performance in dynamic environments.
We propose a probabilistic model to infer supervised latent variables in the Hamming space from observed data. Our model allows simultaneous inference of the number of binary latent variables, and their values. The latent variables preserve neighbourhood structure of the data in a sense that objects in the same semanti…
A method uses Shapley values and Mahalanobis distances to explain multivariate outliers.
Concept drift is formally defined as the change in joint distribution of a set of input variables X and a target variable y. The two types of drift that are extensively studied are real drift and virtual drift where the former is the change in posterior probabilities p(y|X) while the latter is the change in distributio…
Regression analysis is a standard supervised machine learning method used to model an outcome variable in terms of a set of predictor variables. In most real-world applications we do not know the true value of the outcome variable being predicted outside the training data, i.e., the ground truth is unknown. It is hence…
A new framework detects concept drift in streaming data.
The aim of this paper is to introduce a risk measure that extends the Gini-type measures of risk and variability, the Extended Gini Shortfall, by taking risk aversion into consideration. Our risk measure is coherent and catches variability, an important concept for risk management. The analysis is made under the Choque…
The variability of the clusters generated by clustering techniques in the domain of latitude and longitude variables of fatal crash data are significantly unpredictable. This unpredictability, caused by the randomness of fatal crash incidents, reduces the accuracy of crash frequency (i.e., counts of fatal crashes per c…
A new concept of causality for abstract phenomena.
Enhances topology optimization with multiclass microstructures using latent variable Gaussian process.
A new metric assesses latent variable models using data and model moments.
Study finds flipped classrooms improve student self-concept, enjoyment, but not exam scores.
A fundamental issue for statistical classification models in a streaming environment is that the joint distribution between predictor and response variables changes over time (a phenomenon also known as concept drifts), such that their classification performance deteriorates dramatically. In this paper, we first presen…
CBMs improve interpretability in RUL prediction for aircraft engines.
Proxy methods adapt to distribution shifts without explicitly modeling latent confounders.
Evaluating, explaining, and visualizing high-level concepts in generative models, such as variational autoencoders (VAEs), is challenging in part due to a lack of known prediction classes that are required to generate saliency maps in supervised learning. While saliency maps may help identify relevant features (e.g., p…
Deep neural networks are complex and opaque. As they enter application in a variety of important and safety critical domains, users seek methods to explain their output predictions. We develop an approach to explaining deep neural networks by constructing causal models on salient concepts contained in a CNN. We develop…
Extends partitioned local depth concept with probabilistic considerations.
Tensor decompositions are used in various data mining applications from social network to medical applications and are extremely useful in discovering latent structures or concepts in the data. Many real-world applications are dynamic in nature and so are their data. To deal with this dynamic nature of data, there exis…
Shapley value improves model interpretation but not causal inference.
Clarifies EM algorithm and variational Bayesian inference concepts.
Population growth (or decay) in a country can be due to various f socio-economic constraints, as demonstrated in this paper. For example, sexual intercourse is banned in various religions, during Nativity and Lent fasting periods. Data consisting of registered daily birth records for very long (35,429 points) time seri…
A Ph.D. thesis proposes an efficient approach for optimizing aircraft eco-design with high-dimensional mixed integer variables.
This research generates synthetic data streams for handling concept drifts and novel classes.
New approach for handling uncertain probabilities.
Developing an explainable outlier detection method for interval-valued data using Shapley value-based approach.
The intuition of risk is based on two main concepts: loss and variability. In this paper, we present a composition of risk and deviation measures, which contemplate these two concepts. Based on the proposed Limitedness axiom, we prove that this resulting composition, based on properties of the two components, is a cohe…
Variable importance is central to scientific studies, including the social sciences and causal inference, healthcare, and other domains. However, current notions of variable importance are often tied to a specific predictive model. This is problematic: what if there were multiple well-performing predictive models, and …
VSML unifies meta learning concepts and enables simple backpropagation.
Machine learning models fail due to concept and data drift during pandemic.
New online feature selection method handles streaming data with concept drift.
A framework monitors and diagnoses concept drift in supervised learning models.
We consider the problem of a neural network being requested to classify images (or other inputs) without making implicit use of a "protected concept", that is a concept that should not play any role in the decision of the network. Typically these concepts include information such as gender or race, or other contextual …
We describe a formal approach to identify 'root causes' of outliers observed in variables in a scenario where the causal relation between the variables is a known directed acyclic graph (DAG). To this end, we first introduce a systematic way to define outlier scores. Further, we introduce the concep…
We propose a statistical inference framework for the component-wise functional gradient descent algorithm (CFGD) under normality assumption for model errors, also known as -Boosting. The CFGD is one of the most versatile tools to analyze data, because it scales well to high-dimensional data sets, allows for a very…
Paper introduces new approximations for lognormal sums, matching comonotonicity and moments.