Batch normalization biases linear models towards uniform margins, improving performance in binary classification.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
New findings on PAC learning and marginal distribution estimation.
Bayesian network structure learning is often performed in a Bayesian setting, evaluating candidate structures using their posterior probabilities for a given data set. Score-based algorithms then use those posterior probabilities as an objective function and return the maximum a posteriori network as the learned model.…
Bayesian network structure learning is often performed in a Bayesian setting, by evaluating candidate structures using their posterior probabilities for a given data set. Score-based algorithms then use those posterior probabilities as an objective function and return the maximum a posteriori network as the learned mod…
Margin enlargement over training data has been an important strategy since perceptrons in machine learning for the purpose of boosting the robustness of classifiers toward a good generalization ability. Yet Breiman (1999) showed a dilemma that a uniform improvement on margin distribution does NOT necessarily reduces ge…
Study improves curvature estimate for stable marginally outer trapped hypersurfaces with a free boundary.
Study shows uniform-time chaos propagation in mean field Langevin dynamics.
Default-ERM shortcut learning persists even without additional information.
Stability result for a popular algorithm in optimal transport.
A fast method for training linear classifiers maximizes margins.
Active learning can't improve over passive in certain settings.
Proposes MFSWB for marginal fairness in SWB, improving efficiency and performance.
We derive and analyze a new, efficient, pool-based active learning algorithm for halfspaces, called ALuMA. Most previous algorithms show exponential improvement in the label complexity assuming that the distribution over the instance space is close to uniform. This assumption rarely holds in practical applications. Ins…
New method controls error in low-dimensional marginals of spatial models.
We design and mathematically analyze sampling-based algorithms for regularized loss minimization problems that are implementable in popular computational models for large data, in which the access to the data is restricted in some way. Our main result is that if the regularizer's effect does not become negligible as th…
We introduce a new category of multivariate conditional generative models and demonstrate its performance and versatility in probabilistic time series forecasting and simulation. Specifically, the output of quantile regression networks is expanded from a set of fixed quantiles to the whole Quantile Function by a univar…
Identifying components and estimating mixing weights in unlabeled finite mixtures under marginal independence.
Constructs supermartingale couplings with full marginals constraints.
Study on computational aspects of replicable learning, bridging statistical and algorithmic perspectives.
Adversarial training is by far the most successful strategy for improving robustness of neural networks to adversarial attacks. Despite its success as a defense mechanism, adversarial training fails to generalize well to unperturbed test set. We hypothesize that this poor generalization is a consequence of adversarial …
A new method models volatile financial time series using v-transforms and copulas.
Entropy regularization is used to get improved optimization performance in reinforcement learning tasks. A common form of regularization is to maximize policy entropy to avoid premature convergence and lead to more stochastic policies for exploration through action space. However, this does not ensure exploration in th…
Deep generative networks such as GANs and normalizing flows flourish in the context of high-dimensional tasks such as image generation. However, so far exact modeling or extrapolation of distributional properties such as the tail asymptotics generated by a generative network is not available. In this paper, we address …
New insights into how linear classifiers and leaky ReLU networks can overfit without harming generalization.
We propose a novel training procedure for improving the performance of generative adversarial networks (GANs), especially to bidirectional GANs. First, we enforce that the empirical distribution of the inverse inference network matches the prior distribution, which favors the generator network reproducibility on the se…
New method finds closest martingale to Brownian motion.
Flexible copula model using implicit generative neural networks.
This paper studies convergence properties of multivariate distributions constructed by endowing empirical margins with a copula. This setting includes Latin Hypercube Sampling with dependence, also known as the Iman--Conover method. The primary question addressed here is the convergence of the component sum, which is r…
The paper develops p-values for outlier detection using conformal inference.
We develop quantile regression models in order to derive risk margin and to evaluate capital in non-life insurance applications. By utilizing the entire range of conditional quantile functions, especially higher quantile levels, we detail how quantile regression is capable of providing an accurate estimation of risk ma…
This paper considers a new family of variational distributions motivated by Sklar's theorem. This family is based on new copula-like densities on the hypercube with non-uniform marginals which can be sampled efficiently, i.e. with a complexity linear in the dimension of state space. Then, the proposed variational densi…
Exploration is critical to a reinforcement learning agent's performance in its given environment. Prior exploration methods are often based on using heuristic auxiliary predictions to guide policy behavior, lacking a mathematically-grounded objective with clear properties. In contrast, we recast exploration as a proble…
Conformal Test Martingales can be 'blind' to significant changes in data distribution.
We present a simple noise-robust margin-based active learning algorithm to find homogeneous (passing the origin) linear separators and analyze its error convergence when labels are corrupted by noise. We show that when the imposed noise satisfies the Tsybakov low noise condition (Mammen, Tsybakov, and others 1999; Tsyb…
Deep ReLU networks generalize well with few parameters.
This study optimizes offline reinforcement learning methods for various tasks without rewards.
New methods for estimating and inferring nonparametric structural functions and elasticities.
UDM reparameterization improves language model generation.
A new copula, the checkerboard copula, maximizes entropy and preserves dependence.
We provide a set of copulas that can be interpreted as having the negative extreme dependence. This set of copulas is interesting because it coincides with countermonotonic copula for a bivariate case, and more importantly, is shown to be minimal in concordance ordering in the sense that no copula exists which is stric…
The univariate piecing-together approach (PT) fits a univariate generalized Pareto distribution (GPD) to the upper tail of a given distribution function in a continuous manner. We propose a multivariate extension. First it is shown that an arbitrary copula is in the domain of attraction of a multivariate extreme value …
A preference order or ranking aggregated from pairwise comparison data is commonly understood as a strict total order. However, in real-world scenarios, some items are intrinsically ambiguous in comparisons, which may very well be an inherent uncertainty of the data. In this case, the conventional total order ranking c…
The paper tackles learning from non-irreducible Markov chains, proving learnability and generalization bounds.
Proposes MvTPMSVM to improve multiview learning with reduced computational complexity.
The study models insurance dependence using Bernstein copulas.
New findings show score matching's accuracy doesn't ensure numerical stability in diffusion sampling.
This paper establishes a precise high-dimensional asymptotic theory for boosting on separable data, taking statistical and computational perspectives. We consider a high-dimensional setting where the number of features (weak learners) scales with the sample size , in an overparametrized regime. Under a class of …
Multiple kernel learning (MKL), structured sparsity, and multi-task learning have recently received considerable attention. In this paper, we show how different MKL algorithms can be understood as applications of either regularization on the kernel weights or block-norm-based regularization, which is more common in str…