The paper formalizes criteria for non-spurious and disentangled representations using causal methods.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
This paper examines challenges and solutions for solving variational inequalities.
We identify action representations from video data, proving their statistical benefits.
LNK improves uncertainty estimation for molecular dynamics, reducing errors by up to 2.5 times.
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups…
NOMU improves neural network uncertainty estimation.
Machine learning has evolved into an enabling technology for a wide range of highly successful applications. The potential for this success to continue and accelerate has placed machine learning (ML) at the top of research, economic and political agendas. Such unprecedented interest is fuelled by a vision of ML applica…
Unsupervised learning models can be indistinguishable without identifiability, leading to unreliable representations.
We show social events can be accurately predicted, but often undesirably.
New bounds predict deep learning generalization better than existing methods.
Causal Bayesian networks interpret actions as interventions to connect models to real-world outcomes.
Study shows how competition affects learning in matching markets, proving it's possible to balance stability, fairness, and regret.
PACE explains ViTs by modeling patch-level concept distributions, surpassing existing methods.
New bounds improve generalization for deep neural networks in domain adaptation.
Contemporary global optimization algorithms are based on local measures of utility, rather than a probability measure over location and value of the optimum. They thus attempt to collect low function values, not to learn about the optimum. The reason for the absence of probabilistic global optimizers is that the corres…
Anomaly detection has numerous applications and has been studied vastly. We consider a complementary problem that has a much sparser literature: anomaly description. Interpretation of anomalies is crucial for practitioners for sense-making, troubleshooting, and planning actions. To this end, we present a new approach c…
We propose that the Continual Learning desiderata can be achieved through a neuro-inspired architecture, grounded on Mountcastle's cortical column hypothesis. The proposed architecture involves a single module, called Self-Taught Associative Memory (STAM), which models the function of a cortical column. STAMs are repea…
New approach to off-policy evaluation connects causal graph to policy effects.
Most recent work on interpretability of complex machine learning models has focused on estimating explanations for previously trained models around specific predictions. models where interpretability plays a key role already during learning have received much less atte…
Interpretability is an elusive but highly sought-after characteristic of modern machine learning methods. Recent work has focused on interpretability via , which justify individual model predictions. In this work, we take a step towards reconciling machine explanations with those that humans prod…
We propose practical algorithms for entrywise -norm low-rank approximation, for or . The proposed framework, which is non-convex and gradient-based, is easy to implement and typically attains better approximations, faster, than state of the art. From a theoretical standpoint, we show that th…
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
Interpretability has become an important topic of research as more machine learning (ML) models are deployed and widely used to make important decisions. Most of the current explanation methods provide explanations through feature importance scores, which identify features that are important for each individual input. …
In the continual learning setting, tasks are encountered sequentially. The goal is to learn whilst i) avoiding catastrophic forgetting, ii) efficiently using model capacity, and iii) employing forward and backward transfer learning. In this paper, we explore how the Variational Continual Learning (VCL) framework achiev…
We introduce a unified probabilistic framework for solving sequential decision making problems ranging from Bayesian optimisation to contextual bandits and reinforcement learning. This is accomplished by a probabilistic model-based approach that explains observed data while capturing predictive uncertainty during the d…
"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the…
DeepSphere improves spherical CNNs by balancing efficiency and rotation equivariance.
We study model-agnostic copies of machine learning classifiers. We develop the theory behind the problem of copying, highlighting its differences with that of learning, and propose a framework to copy the functionality of any classifier using no prior knowledge of its parameters or training data distribution. We identi…
Explanation in machine learning and related fields such as artificial intelligence aims at making machine learning models and their decisions understandable to humans. Existing work suggests that personalizing explanations might help to improve understandability. In this work, we derive a conceptualization of personali…
Deep Curvature Suite offers a PyTorch package for neural network curvature analysis.
We introduce a novel uncertainty estimation for classification tasks for Bayesian convolutional neural networks with variational inference. By normalizing the output of a Softplus function in the final layer, we estimate aleatoric and epistemic uncertainty in a coherent manner. The intractable posterior probability dis…
Calls to arms to build interpretable models express a well-founded discomfort with machine learning. Should a software agent that does not even know what a loan is decide who qualifies for one? Indeed, we ought to be cautious about injecting machine learning (or anything else, for that matter) into applications where t…
A new method for reinforcement learning scales errors without tuning.
This work frames reward modelling from preferences as a causal problem.
Faster convergence of kernel mean embeddings using variance information.
VSD efficiently learns conditional distributions for combinatorial designs.
Machine learning workflow development is anecdotally regarded to be an iterative process of trial-and-error with humans-in-the-loop. However, we are not aware of quantitative evidence corroborating this popular belief. A quantitative characterization of iteration can serve as a benchmark for machine learning workflow d…
DiSeNE generates interpretable node embeddings without supervision.
Novel flows generate molecules without post-processing.
In this work, we move beyond the traditional complex-valued representations, introducing more expressive hypercomplex representations to model entities and relations for knowledge graph embeddings. More specifically, quaternion embeddings, hypercomplex-valued embeddings with three imaginary components, are utilized to …
IRDs provide local, model-agnostic explanations using hyperboxes.
Novel framework for Bayesian neural networks incorporating task-specific constraints.
ARFs generate plausible counterfactuals for models, improving model understanding.
Paper bridges generative models and explainability.
Policy gradient method proves convergence in imperfect-information games.
Optimizes machine learning models while controlling risks.
A new approach to rationalization identifies true rationales by considering causal relationships.
Ideal attribution mechanisms track model interactions for faithful watermarks.