Formalizes weak and strong verification for LLMs, controlling errors without assumptions.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
The study explores the strengths and weaknesses of models that generalize from weak to strong supervision.
Improved machine learning models outperform their simpler counterparts by using imperfect labels.
New theory explains how strong models can learn from weak ones.
CB-SLICE identifies concept-based error slices in deep learning models.
A new adaptive splitting method improves accuracy for Cox-Ingersoll-Ross model.
We establish the first nonasymptotic error bounds for Kaplan-Meier-based nearest neighbor and kernel survival probability estimators where feature vectors reside in metric spaces. Our bounds imply rates of strong consistency for these nonparametric estimators and, up to a log factor, match an existing lower bound for c…
We consider the approximation of stochastic differential equations (SDEs) with non-Lipschitz drift or diffusion coefficients. We present a modified explicit Euler-Maruyama discretisation scheme that allows us to prove strong convergence, with a rate. Under some regularity and integrability conditions, we obtain the opt…
Although kernel methods are widely used in many learning problems, they have poor scalability to large datasets. To address this problem, sketching and stochastic gradient methods are the most commonly used techniques to derive efficient large-scale learning algorithms. In this study, we consider solving a binary class…
New tests for identifying the number of latent factors in short panels with small time dimensions.
In spite of the accomplishments of deep learning based algorithms in numerous applications and very broad corresponding research interest, at the moment there is still no rigorous understanding of the reasons why such algorithms produce useful results in certain situations. A thorough mathematical analysis of deep lear…
Boosting improves accuracy by combining weak learners into a voting classifier.
Finding biologically plausible alternatives to back-propagation of errors is a fundamentally important challenge in artificial neural network research. In this paper, we propose a learning algorithm called error-driven Local Representation Alignment (LRA-E), which has strong connections to predictive coding, a theory t…
New method improves optimization and DP in FL.
Self-training improves weak classifiers in mixture models.
While active learning offers potential cost savings, the actual data efficiency---the reduction in amount of labeled data needed to obtain the same error rate---observed in practice is mixed. This paper poses a basic question: when is active learning actually helpful? We provide an answer for logistic regression with t…
New algorithm optimally evaluates policies with linear approximations.
A contraction analysis improves model-based RL's error recovery.
Improves test set performance and reduces out-of-sample disappointment for unstable models.
Paper proposes deep neural networks for nonparametric regression from dependent data.
Bayesian sequence prediction is a simple technique for predicting future symbols sampled from an unknown measure on infinite sequences over a countable alphabet. While strong bounds on the expected cumulative error are known, there are only limited results on the distribution of this error. We prove tight high-probabil…
W2S FT often outperforms weak teachers due to low intrinsic dimensionality.
New ensemble SVM model reduces prediction error without choosing best kernel.
Bayesian Additive Regression Trees (BART) is a fully Bayesian approach to modeling with ensembles of trees. BART can uncover complex regression functions with high dimensional regressors in a fairly automatic way and provide Bayesian quantification of the uncertainty through the posterior. However, BART assumes IID nor…
SGD-trained deep nets often generalize well due to a strong inductive bias towards low-error, low-complexity functions.
New algorithms improve community detection in network data with strong consistency.
This article analyzes the weak error of SGD optimization schemes.
New findings show privacy affects generalization error in a non-monotonic way.
We derive error estimates for multinomial approximations of American options in a multidimensional jump--diffusion Merton's model. We assume that the payoffs are Markovian and satisfy Lipschitz type conditions. Error estimates for such type of approximations were not obtained before. Our main tool is the strong approxi…
Bayesian framework tackles measurement error in covariates.
We justify and give error estimates for binomial approximations of game (Israeli) options in the Black--Scholes market with Lipschitz continuous path dependent payoffs which are new also for usual American style options. We show also that rational (optimal) exercise times and hedging self-financing portfolios of binomi…
End-to-end ASR error detection using audio-transcript entailment.
This paper investigates tradeoffs among optimization errors, statistical rates of convergence and the effect of heavy-tailed errors for high-dimensional robust regression with nonconvex regularization. When the additive errors in linear models have only bounded second moment, we show that iteratively reweighted $\ell_1…
K-fold cross-validation (CV) with squared error loss is widely used for evaluating predictive models, especially when strong distributional assumptions cannot be taken. However, CV with squared error loss is not free from distributional assumptions, in particular in cases involving non-i.i.d. data. This paper analyzes …
Spectral feature learning improves IV regression for causal effect estimation.
In this paper, we are interested in the strong convergence properties of the Ninomiya-Victoir scheme which is known to exhibit weak convergence with order 2. We prove strong convergence with order . This study is aimed at analysing the use of this scheme either at each level or only at the finest level of a multil…
The paper reveals three mechanisms for weak-to-strong generalization.
An active learner is given a hypothesis class, a large set of unlabeled examples and the ability to interactively query labels to an oracle of a subset of these examples; the goal of the learner is to learn a hypothesis in the class that fits the data well by making as few label queries as possible. This work addresses…
Strong inductive biases prevent harmless interpolation in overparameterized models.
We study high-dimensional asymptotic performance limits of binary supervised classification problems where the class conditional densities are Gaussian with unknown means and covariances and the number of signal dimensions scales faster than the number of labeled training samples. We show that the Bayes error, namely t…
Study provides error estimates for approximating game options with diffusion asset prices.
New algorithm solves saddle point problems in Banach spaces.
In high dimensions, most machine learning methods are brittle to even a small fraction of structured outliers. To address this, we introduce a new meta-algorithm that can take in a base learner such as least squares or stochastic gradient descent, and harden the learner to be resistant to outliers. Our method, Sever, p…
Develops algorithms for multi-class Neyman-Pearson classification with cost sensitivity.
ECN framework improves training on noisy structured labels.
Study linear regression with missing or corrupted data, showing error bounds.
Model-based reinforcement learning is an appealing framework for creating agents that learn, plan, and act in sequential environments. Model-based algorithms typically involve learning a transition model that takes a state and an action and outputs the next state---a one-step model. This model can be composed with itse…
We show that the sets in a family with finite VC dimension can be uniformly approximated within a given error by a finite partition. Immediate corollaries include the fact that VC classes have finite bracketing numbers, satisfy uniform laws of averages under strong dependence, and exhibit uniform mixing. Our results ar…