The MBO scheme for data clustering is analyzed in the large data limit, proving convergence to optimal partition problems.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper detects Trojan neural networks with limited or no data.
Active learning improves SR by proposing experiments in data-limited settings.
Given a data set and a subset of labels the problem of semi-supervised learning on point clouds is to extend the labels to the entire data set. In this paper we extend the labels by minimising the constrained discrete -Dirichlet energy. Under suitable conditions the discrete problem can be connected, in the large da…
Graph Laplacians computed from weighted adjacency matrices are widely used to identify geometric structure in data, and clusters in particular; their spectral properties play a central role in a number of unsupervised and semi-supervised learning algorithms. When suitably scaled, graph Laplacians approach limiting cont…
Derives ideal train/test split for ridge regression in large data limit.
The paper analyzes Laplace learning for Gaussian measure data in infinite dimensions, proving convergence.
Bayesian model explains and improves black-box estimators for class distribution.
Theory predicts neural scaling exponents from language statistics.
Q-CurL optimizes quantum learning with a curriculum design.
We introduce a model of proportional growth to explain the distribution of business firm growth rates. The model predicts that the distribution is exponential in the central part and depicts an asymptotic power-law behavior in the tails with an exponent 3. Because of data limitations, previous studies in this field hav…
New insights into tSNE for large datasets.
New theory explains when pre-trained models can improve downstream tasks.
We introduce a model of proportional growth to explain the distribution of business firm growth rates. The model predicts that is Laplace in the central part and depicts an asymptotic power-law behavior in the tails with an exponent . Because of data limitations, previous studies in this field have b…
This work provides theoretical and empirical evidence that invariance-inducing regularizers can increase predictive accuracy for worst-case spatial transformations (spatial robustness). Evaluated on these adversarially transformed examples, we demonstrate that adding regularization on top of standard or adversarial tra…
Ensembles depend on diversity for improved performance. Many ensemble training methods, therefore, attempt to optimize for diversity, which they almost always define in terms of differences in training set predictions. In this paper, however, we demonstrate the diversity of predictions on the training set does not nece…
While adversarial training can improve robust accuracy (against an adversary), it sometimes hurts standard accuracy (when there is no adversary). Previous work has studied this tradeoff between standard and robust accuracy, but only in the setting where no predictor performs well on both objectives in the infinite data…
The ability of many powerful machine learning algorithms to deal with large data sets without compromise is often hampered by computationally expensive linear algebra tasks, of which calculating the log determinant is a canonical example. In this paper we demonstrate the optimality of Maximum Entropy methods in approxi…
Large corporate credit models may be adapted for small business risk assessment.
Improved noise estimation in latent neural SDEs enhances model accuracy.
New MIP methods improve training of integer-valued neural networks.
Study improves AI's handling of uncertainty.
Machine learning predicts liquid water properties from cluster data.
Maximum Likelihood Estimation (MLE) is the bread and butter of system inference for stochastic systems. In some generality, MLE will converge to the correct model in the infinite data limit. In the context of physical approaches to system inference, such as Boltzmann machines, MLE requires the arduous computation of pa…
In many mobile health interventions, treatments should only be delivered in a particular context, for example when a user is currently stressed, walking or sedentary. Even in an optimal context, concerns about user burden can restrict which treatments are sent. To diffuse the treatment delivery over times when a user i…
Estimators computed from adaptively collected data do not behave like their non-adaptive brethren. Rather, the sequential dependence of the collection policy can lead to severe distributional biases that persist even in the infinite data limit. We develop a general method -- -decorrelation -- for transformi…
Recent years have seen rapid advances in the data-driven analysis of dynamical systems based on Koopman operator theory and related approaches. On the other hand, low-rank tensor product approximations -- in particular the tensor train (TT) format -- have become a valuable tool for the solution of large-scale problems …
With the rapid increase of available data for complex systems, there is great interest in the extraction of physically relevant information from massive datasets. Recently, a framework called Sparse Identification of Nonlinear Dynamics (SINDy) has been introduced to identify the governing equations of dynamical systems…
IntelligentPooling learns personalized mHealth policies from limited data.
Invariances to translation, rotation and other spatial transformations are a hallmark of the laws of motion, and have widespread use in the natural sciences to reduce the dimensionality of systems of equations. In supervised learning, such as in image classification tasks, rotation, translation and scale invariances ar…
Generative model learns wireless channel distributions efficiently.
Study compares exponential and power-law kernels in modeling high-frequency trading data.
This work characterizes reward function partial identifiability and its impact on policy optimization.
Develops CLDS models to model neural activity with nonlinear dynamics.
Theoretical study shows AI models can recover from contaminated training data.
A new criterion selects models in overparameterized settings.
Study on optimal ReLU networks with weight decay for interpolation.
Gradient descent converges geometrically to optimal self-attention parameters.
Image classification system identifies bumble bee species from images.
Study integrates deep learning with financial data for improved trading strategies.
Bayesian approach learns linear operators from noisy data.
Identifying small subsets of features that are relevant for prediction and/or classification tasks is a central problem in machine learning and statistics. The feature selection task is especially important, and computationally difficult, for modern datasets where the number of features can be comparable to, or even ex…
Samplets and multiwavelets constructed from scattered data converge to specific densities in the limit.
Diffusion models can memorize training data, limiting their creativity and privacy.
Bayesian Neural Networks are robust to gradient-based attacks in the large-data limit.
Deep Neural Networks (DNNs) are susceptible to model stealing attacks, which allows a data-limited adversary with no knowledge of the training dataset to clone the functionality of a target model, just by using black-box query access. Such attacks are typically carried out by querying the target model using inputs that…
Predicting bioactivity and physical properties of small molecules is a central challenge in drug discovery. Deep learning is becoming the method of choice but studies to date focus on mean accuracy as the main metric. However, to replace costly and mission-critical experiments by models, a high mean accuracy is not eno…
CalNF models rare failures with limited data, improving safety in autonomous systems.