Unified framework for fair decision-making across diverse groups.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The paper highlights AI brittleness and the need for robust testing out-of-distribution performance.
Machine learning predicts failure in brittle materials with high accuracy.
New method robust to semi-random sparse recovery, nearly-linear time.
We propose a machine learning approach to address a key challenge in materials science: predicting how fractures propagate in brittle materials under stress, and how these materials ultimately fail. Our methods use deep learning and train on simulation data from high-fidelity models, emulating the results of these mode…
Study identifies key parameters and input dimensions making LLMs and VLMs brittle.
Standard neural network architectures are non-linear only by virtue of a simple element-wise activation function, making them both brittle and excessively large. In this paper, we consider methods for making the feed-forward layer more flexible while preserving its basic structure. We develop simple drop-in replacement…
AI needs causal inference to avoid being just a correlation machine.
Improves efficiency of simulators that fail to return.
In this paper, five different approaches for reduced-order modeling of brittle fracture in geomaterials, specifically concrete, are presented and compared. Four of the five methods rely on machine learning (ML) algorithms to approximate important aspects of the brittle fracture problem. In addition to the ML algorithms…
ARO overfits by making constraints dependent on uncertainty, leading to brittleness.
New framework detects out-of-distribution samples efficiently.
Recent reinforcement learning algorithms, though achieving impressive results in various fields, suffer from brittle training effects such as regression in results and high sensitivity to initialization and parameters. We claim that some of the brittleness stems from variance differences, i.e. when different environmen…
Polynomial-time algorithm finds planted hypercube vectors in Gaussian mixtures.
A model learns causal graphs from summary statistics of synthetic data.
OPAL optimizes labeling strategy for precise inference from uncertain models.
Transformers simulate finite-state automata with fewer layers.
Inference-Time Scaling can be extended to domains prone to systematic failure using intrinsic statistics.
DAFNO learns surrogates for complex systems on irregular geometries.
Adversarial examples have attracted significant attention in machine learning, but the reasons for their existence and pervasiveness remain unclear. We demonstrate that adversarial examples can be directly attributed to the presence of non-robust features: features derived from patterns in the data distribution that ar…
The risks and perils of overfitting in machine learning are well known. However most of the treatment of this, including diagnostic tools and remedies, was developed for the supervised learning case. In this work, we aim to offer new perspectives on the characterization and prevention of overfitting in deep Reinforceme…
Deep neural networks have been increasingly used in software engineering and program analysis tasks. They usually take a program and make some predictions about it, e.g., bug prediction. We call these models neural program analyzers. The reliability of neural programs can impact the reliability of the encompassing anal…
Novel Bayesian neural network method for robustness.
IGNIS uses neural networks to estimate copula parameters robustly.
Optimizes structure topology for ductile and brittle fracture resistance.
While current benchmark reinforcement learning (RL) tasks have been useful to drive progress in the field, they are in many ways poor substitutes for learning with real-world data. By testing increasingly complex RL algorithms on low-complexity simulation environments, we often end up with brittle RL policies that gene…
New method calibrates models under covariate shifts.
Standard stochastic optimization methods are brittle, sensitive to stepsize choices and other algorithmic parameters, and they exhibit instability outside of well-behaved families of objectives. To address these challenges, we investigate models for stochastic minimization and learning problems that exhibit better robu…
Bayesian neural networks (BNNs) have developed into useful tools for probabilistic modelling due to recent advances in variational inference enabling large scale BNNs. However, BNNs remain brittle and hard to train, especially: (1) when using deep architectures consisting of many hidden layers and (2) in situations wit…
Introduces PCG for better counterfactual explanations in vision models.
Recent work has shown that state-of-the-art classifiers are quite brittle, in the sense that a small adversarial change of an originally with high confidence correctly classified input leads to a wrong classification again with high confidence. This raises concerns that such classifiers are vulnerable to attacks and ca…
Deep Reinforcement Learning (DRL) algorithms for continuous action spaces are known to be brittle toward hyperparameters as well as \cut{being}sample inefficient. Soft Actor Critic (SAC) proposes an off-policy deep actor critic algorithm within the maximum entropy RL framework which offers greater stability and empiric…
Variational Proximal Policy Optimization improves reinforcement learning from human feedback.
Maximizes robustness in Bayesian experimental design under model uncertainty.
A robust approach compensates for small-data tasks in mixed linear regression.
A simple method treats heteroscedastic variance variatively, improving model calibration and sample quality.
Modern deep neural networks are well known to be brittle in the face of unknown data instances and recognition of the latter remains a challenge. Although it is inevitable for continual-learning systems to encounter such unseen concepts, the corresponding literature appears to nonetheless focus primarily on alleviating…
K-means clustering improved for robustness to outliers and distribution shifts.
The DoD needs a robust process to evaluate AI/ML model performance and robustness.
The study assesses external validity by evaluating worst-case treatment effects across subpopulations.
Machine learning promises methods that generalize well from finite labeled data. However, the brittleness of existing neural net approaches is revealed by notable failures, such as the existence of adversarial examples that are misclassified despite being nearly identical to a training example, or the inability of recu…
Machine learning is currently dominated by largely experimental work focused on improvements in a few key tasks. However, the impressive accuracy numbers of the best performing models are questionable because the same test sets have been used to select these models for multiple years now. To understand the danger of ov…
We study the problem of using i.i.d. samples from an unknown multivariate probability distribution to estimate the mutual information of . This problem has recently received attention in two settings: (1) where is assumed to be Gaussian and (2) where is assumed only to lie in a large nonparametric smooth…
Recent work in the domain of misinformation detection has leveraged rich signals in the text and user identities associated with content on social media. But text can be strategically manipulated and accounts reopened under different aliases, suggesting that these approaches are inherently brittle. In this work, we inv…
Reinforcement learning offers the promise of automating the acquisition of complex behavioral skills. However, compared to commonly used and well-understood supervised learning methods, reinforcement learning algorithms can be brittle, difficult to use and tune, and sensitive to seemingly innocuous implementation decis…
Extends DRFGP to make GPs more robust and adaptive for dynamic, noisy data.
New method reduces over-pessimism in Bayesian control under parameter uncertainty.
Deconfounding scores improve causal effect estimation with weak overlap.