The paper formalizes criteria for non-spurious and disentangled representations using causal methods.
problem Formalizing criteria for non-spurious and disentangled representations in representation learning.
method Causal perspective, counterfactual quantities, observable consequences of causal assertions.
result Computable metrics for assessing representation learning based on observed data.
This paper examines challenges and solutions for solving variational inequalities.
problem Stability issues in solving variational inequalities, especially in multi-objective scenarios.
method Continuous-time analysis to understand and improve stability of algorithms.
result Understanding continuous-time dynamics can help in designing more stable algorithms for variational inequalities.
We identify action representations from video data, proving their statistical benefits.
problem Identifying latent action policies from video data.
method Entropy-regularized LAPO objective, formalizing desiderata for action representations.
result Entropy-regularized LAPO identifies action representations satisfying desiderata under suitable conditions.
LNK improves uncertainty estimation for molecular dynamics, reducing errors by up to 2.5 times.
problem Uncertainty estimation for molecular force fields to improve model reliability.
method LNK: Gaussian Process-based extension to GNNs addressing six desiderata.
result LNK reduces out-of-equilibrium detection errors by up to 2.5 times compared to existing methods.
Experiments used in current continual learning research do not faithfully assess fundamental challenges of learning continually. Instead of assessing performance on challenging and representative experiment designs, recent research has focused on increased dataset difficulty, while still using flawed experiment set-ups…
NOMU improves neural network uncertainty estimation.
problem Estimating model uncertainty for neural networks with limited data.
method Introduces NOMU, a two-sub-NN architecture with a designed loss function.
result NOMU outperforms state-of-the-art methods in regression and Bayesian optimization.
Machine learning has evolved into an enabling technology for a wide range of highly successful applications. The potential for this success to continue and accelerate has placed machine learning (ML) at the top of research, economic and political agendas. Such unprecedented interest is fuelled by a vision of ML applica…
Unsupervised learning models can be indistinguishable without identifiability, leading to unreliable representations.
problem Unsupervised learning models may be indistinguishable without identifiability, making it impossible to recover a ground truth generative model.
method Construction based on nonlinear independent component analysis theory to illustrate potential failure cases.
result Counterexamples show that identifiability is crucial for reliable unsupervised representation learning.
We show social events can be accurately predicted, but often undesirably.
problem The predictability of social events and its implications.
method Recent ideas from performative prediction and outcome indistinguishability.
result Social events can be efficiently predicted accurately, but often undesirably.
New bounds predict deep learning generalization better than existing methods.
problem Predicting generalization errors for deep learning models.
method Function-based PAC-Bayesian bounds that meet multiple desiderata.
result The new bound performs significantly better than existing parameter-based PAC-Bayes bounds.
Causal Bayesian networks interpret actions as interventions to connect models to real-world outcomes.
problem Connecting causal model predictions to real-world outcomes.
method Formal framework to interpret actions as interventions and prove impossibility results.
result No non-circular interpretation exists that satisfies natural desiderata without violating some.
Study shows how competition affects learning in matching markets, proving it's possible to balance stability, fairness, and regret.
problem How competition affects learning in matching markets and the impossibility of simultaneously guaranteeing stability and low optimal regret.
method Modeling a two-sided matching market with bandit learners and adding components of costs and transfers.
result It is possible to simultaneously guarantee stability, low optimal regret, fairness in the distribution of regret, and high social welfare.
PACE explains ViTs by modeling patch-level concept distributions, surpassing existing methods.
problem Lack of trustworthy post-hoc explanations for Vision Transformers (ViTs)
method Variational Bayesian explanation framework (PACE)
result PACE surpasses state-of-the-art methods in meeting desiderata for ViT explanations.
New bounds improve generalization for deep neural networks in domain adaptation.
problem Deriving tight generalization guarantees for deep neural networks in domain adaptation.
method Combining data-dependent PAC-Bayes analysis with importance weighting.
result A simple importance weighting extension provides the tightest estimable bound.
Contemporary global optimization algorithms are based on local measures of utility, rather than a probability measure over location and value of the optimum. They thus attempt to collect low function values, not to learn about the optimum. The reason for the absence of probabilistic global optimizers is that the corres…
Anomaly detection has numerous applications and has been studied vastly. We consider a complementary problem that has a much sparser literature: anomaly description. Interpretation of anomalies is crucial for practitioners for sense-making, troubleshooting, and planning actions. To this end, we present a new approach c…
We propose that the Continual Learning desiderata can be achieved through a neuro-inspired architecture, grounded on Mountcastle's cortical column hypothesis. The proposed architecture involves a single module, called Self-Taught Associative Memory (STAM), which models the function of a cortical column. STAMs are repea…
New approach to off-policy evaluation connects causal graph to policy effects.
problem Evaluating policies using observational data from different policies.
method Formalizes off-policy evaluation within a causal graph framework.
result Identifies specific causal estimands and highlights necessary experimental data.
Most recent work on interpretability of complex machine learning models has focused on estimating a posteriori explanations for previously trained models around specific predictions. Self-explaining models where interpretability plays a key role already during learning have received much less atte…
We propose practical algorithms for entrywise ℓp-norm low-rank approximation, for p=1 or p=∞. The proposed framework, which is non-convex and gradient-based, is easy to implement and typically attains better approximations, faster, than state of the art. From a theoretical standpoint, we show that th…
New scaling framework for MoE architectures ensures stability and optimal performance at scale.
problem Lack of principled understanding of how hyperparameters should scale in MoE architectures.
method Developed a novel Dynamical Mean Field Theory (DMFT) for three scaling regimes of MoE architectures.
result Derived Maximally Scale-Stable Parameterization (MSSP) for SGD and Adam, providing robust learning rate transfer and monotonic improvement with scale.
Interpretability has become an important topic of research as more machine learning (ML) models are deployed and widely used to make important decisions. Most of the current explanation methods provide explanations through feature importance scores, which identify features that are important for each individual input. …
In the continual learning setting, tasks are encountered sequentially. The goal is to learn whilst i) avoiding catastrophic forgetting, ii) efficiently using model capacity, and iii) employing forward and backward transfer learning. In this paper, we explore how the Variational Continual Learning (VCL) framework achiev…
We introduce a unified probabilistic framework for solving sequential decision making problems ranging from Bayesian optimisation to contextual bandits and reinforcement learning. This is accomplished by a probabilistic model-based approach that explains observed data while capturing predictive uncertainty during the d…
"How much is my data worth?" is an increasingly common question posed by organizations and individuals alike. An answer to this question could allow, for instance, fairly distributing profits among multiple data contributors and determining prospective compensation when data breaches happen. In this paper, we study the…
DeepSphere improves spherical CNNs by balancing efficiency and rotation equivariance.
problem Designing efficient and rotation-equivariant convolutional layers for spherical data.
method Graph-based approach to represent spherical data, focusing on the number of vertices and neighbors.
result DeepSphere achieves state-of-the-art performance and demonstrates efficiency and flexibility.
We study model-agnostic copies of machine learning classifiers. We develop the theory behind the problem of copying, highlighting its differences with that of learning, and propose a framework to copy the functionality of any classifier using no prior knowledge of its parameters or training data distribution. We identi…
Explanation in machine learning and related fields such as artificial intelligence aims at making machine learning models and their decisions understandable to humans. Existing work suggests that personalizing explanations might help to improve understandability. In this work, we derive a conceptualization of personali…
Deep Curvature Suite offers a PyTorch package for neural network curvature analysis.
problem Insufficient use of curvature information in neural networks.
method Implementation of Lanczos algorithm for neural network curvature analysis.
result Our package outperforms existing methods for similar purposes.
We introduce a novel uncertainty estimation for classification tasks for Bayesian convolutional neural networks with variational inference. By normalizing the output of a Softplus function in the final layer, we estimate aleatoric and epistemic uncertainty in a coherent manner. The intractable posterior probability dis…
Calls to arms to build interpretable models express a well-founded discomfort with machine learning. Should a software agent that does not even know what a loan is decide who qualifies for one? Indeed, we ought to be cautious about injecting machine learning (or anything else, for that matter) into applications where t…
A new method for reinforcement learning scales errors without tuning.
problem Error scaling varies across reinforcement learning tasks and stages.
method A simple scaling mechanism for temporal-difference learning.
result The method effectively mitigates interference between learning tasks.
The paper introduces a new framework for making machine learning explanations more understandable to humans.
problem Making machine learning explanations comprehensible and aligned with human preferences.
method Inspired by philosophy, cognitive science, and social sciences, the paper formalizes a framework using the concept of 'weight of evidence' from information theory.
result The framework produces intuitive and comprehensible explanations that align with human preferences.
This work frames reward modelling from preferences as a causal problem.
problem Reward modelling from preference data for AI alignment.
method Causal inference approach to identify challenges and assumptions.
result Causally-inspired approaches improve model robustness.
Faster convergence of kernel mean embeddings using variance information.
problem Speeding up the convergence rate of kernel mean embeddings.
method Leveraging variance information in reproducing kernel Hilbert space and estimating variance from data.
result Efficiently estimate variance information from data to achieve distribution-agnostic convergence bounds.
VSD efficiently learns conditional distributions for combinatorial designs.
problem Learning conditional distributions for rare combinatorial designs.
method Variational Search Distributions (VSD) using variational inference.
result VSD outperforms existing methods on real sequence-design problems.
Machine learning workflow development is anecdotally regarded to be an iterative process of trial-and-error with humans-in-the-loop. However, we are not aware of quantitative evidence corroborating this popular belief. A quantitative characterization of iteration can serve as a benchmark for machine learning workflow d…
DiSeNE generates interpretable node embeddings without supervision.
problem Lack of interpretability in unsupervised node embeddings.
method Disentangled representation learning with novel objective functions and metrics.
result DiSeNE produces interpretable node embeddings aligned with graph structure.
Novel flows generate molecules without post-processing.
problem Generating new molecules efficiently and without post-processing issues.
method Continuous normalizing E(3)-equivariant flows based on node ODEs coupled as a graph PDE.
result Generated samples achieve state-of-the-art performance on QM9 and ZINC250K benchmarks.
In this work, we move beyond the traditional complex-valued representations, introducing more expressive hypercomplex representations to model entities and relations for knowledge graph embeddings. More specifically, quaternion embeddings, hypercomplex-valued embeddings with three imaginary components, are utilized to …
IRDs provide local, model-agnostic explanations using hyperboxes.
problem Local model-agnostic explanations for machine learning predictions.
method Formalizes IRDs as hyperboxes, defines optimization problem, introduces unified framework.
result IRDs offer semi-factual explanations and highlight feature importance.
Novel framework for Bayesian neural networks incorporating task-specific constraints.
problem Task-specific constraints in supervised model deployment.
method Introduces Output-Constrained BNN (OC-BNN) framework.
result OC-BNNs effectively incorporate prior expert knowledge and desiderata like safety and fairness.
ARFs generate plausible counterfactuals for models, improving model understanding.
problem Creating realistic counterfactuals for model analysis.
method Adversarial Random Forests (ARFs) for generating plausible counterfactuals.
result ARFs efficiently generate plausible counterfactuals in a model-agnostic way.
Paper bridges generative models and explainability.
problem Lack of connection between generative models and explainability.
method Proposes a probabilistic framework for example-based explanations.
result Formally defines example-based explanations for deep generative models.
Policy gradient method proves convergence in imperfect-information games.
problem Policy gradient methods in imperfect-information games (EFGs).
method Policy gradient approach with best-iterate convergence.
result Policy gradient leads to provable best-iterate convergence in self-play EFGs.
Optimizes machine learning models while controlling risks.
problem Finding a model configuration that balances multiple conflicting metrics.
method Combines Bayesian Optimization with rigorous risk-controlling procedures.
result Identifies and selects Pareto optimal configurations with guaranteed risk levels.
A new approach to rationalization identifies true rationales by considering causal relationships.
problem Existing rationalization methods struggle with spuriousness, where snippets with similar contributions are hard to distinguish.
method The method leverages causal inference to identify non-spurious rationales, defining probabilities of causation based on a structural causal model.
result The proposed causal rationalization outperforms existing methods on real-world datasets.
Ideal attribution mechanisms track model interactions for faithful watermarks.
problem Ensuring models provide transparent and fair attribution decisions.
method Introducing ideal attribution mechanisms and a ledger for tracking model interactions.
result A unified framework for evaluating watermarking schemes, clarifying attainable guarantees.