Bayesian method improves quantile estimation and subset selection.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We consider binary classification problems with positive definite kernels and square loss, and study the convergence rates of stochastic gradient methods. We show that while the excess testing loss (squared loss) converges slowly to zero as the number of observations (and thus iterations) goes to infinity, the testing …
This work investigates square loss in overparametrized neural networks, revealing its advantages in robustness and calibration.
Artificial neural network training with stochastic gradient descent can be destabilized by "bad batches" with high losses. This is often problematic for training with small batch sizes, high order loss functions or unstably high learning rates. To stabilize learning, we have developed adaptive learning rate clipping (A…
The paper analyzes error bounds and KL properties for noisy matrix recovery problems.
The paper improves Kaczmarz algorithm with momentum for linear least squares.
Paper unifies bias and variance models for classification.
The paper studies the loss landscape of regularized deep matrix factorization, revealing unique and sharp minimizers.
New loss function improves accuracy of MRI parameter estimation.
Paper establishes a universal growth rate for smooth surrogate losses in classification.
The logcosh loss function helps neural networks learn set-valued functions better.
Transformer-based models overfit financial time series data, leading to increased prediction variance.
K-fold cross-validation (CV) with squared error loss is widely used for evaluating predictive models, especially when strong distributional assumptions cannot be taken. However, CV with squared error loss is not free from distributional assumptions, in particular in cases involving non-i.i.d. data. This paper analyzes …
Improved speech enhancement using diffusion models with MSE loss.
Proposes a new probabilistic framework for domain generalization.
In this paper, we consider the nonparametric least square regression in a Reproducing Kernel Hilbert Space (RKHS). We propose a new randomized algorithm that has optimal generalization error bounds with respect to the square loss, closing a long-standing gap between upper and lower bounds. Moreover, we show that our al…
Principal Component Analysis (PCA) is a very successful dimensionality reduction technique, widely used in predictive modeling. A key factor in its widespread use in this domain is the fact that the projection of a dataset onto its first principal components minimizes the sum of squared errors between the original …
Cross-validation pitfalls in change-point regression are addressed with new approaches.
This study explains gradient flow dynamics in neural networks for small initialisation.
In this work we propose an adversarial learning approach to generate high resolution MRI scans from low resolution images. The architecture, based on the SRGAN model, adopts 3D convolutions to exploit volumetric information. For the discriminator, the adversarial loss uses least squares in order to stabilize the traini…
The Nyström method improves learning efficiency for convex losses.
Paper develops estimators for unbounded density ratios with applications in error control.
Previous studies have shown that deep neural networks (DNNs) with common settings often capture target functions from low to high frequency, which is called Frequency Principle (F-Principle). It has also been shown that F-Principle can provide an understanding to the often observed good generalization ability of DNNs. …
This paper presents a learning method for convolutional autoencoders (CAEs) for extracting features from images. CAEs can be obtained by utilizing convolutional neural networks to learn an approximation to the identity function in an unsupervised manner. The loss function based on the pixel loss (PL) that is the mean s…
We introduce the implicitly constrained least squares (ICLS) classifier, a novel semi-supervised version of the least squares classifier. This classifier minimizes the squared loss on the labeled data among the set of parameters implied by all possible labelings of the unlabeled data. Unlike other discriminative semi-s…
Super learner with Huber loss improves cost prediction and causal effect estimation in healthcare expenditure data.
Analyzes double descent in binary classification models with different losses.
We study tensor completion in the agnostic setting. In the classical tensor completion problem, we receive entries of an unknown rank- tensor and wish to exactly complete the remaining entries. In agnostic tensor completion, we make no assumption on the rank of the unknown tensor, but attempt to predict unknown …
We derive the mapping between two of the most pervasive utility functions, the mean square error () and the concordance correlation coefficient (CCC, ). Despite its drawbacks, is one of the most popular performance metrics (and a loss function); along with lately in many of the sequence prediction…
In this paper, we explore ordinal classification (in the context of deep neural networks) through a simple modification of the squared error loss which not only allows it to not only be sensitive to class ordering, but also allows the possibility of having a discrete probability distribution over the classes. Our formu…
We study the error landscape of deep linear and nonlinear neural networks with the squared error loss. Minimizing the loss of a deep linear neural network is a nonconvex problem, and despite recent progress, our understanding of this loss surface is still incomplete. For deep linear networks, we present necessary and s…
DEQGAN uses GANs to solve differential equations without supervision.
We prove a new generalization bound that shows for any class of linear predictors in Gaussian space, the Rademacher complexity of the class and the training error under any continuous loss can control the test error under all Moreau envelopes of the loss . We use our finite-sample bound to directly recover…
This article provides, through theoretical analysis, an in-depth understanding of the classification performance of the empirical risk minimization framework, in both ridge-regularized and unregularized cases, when high dimensional data are considered. Focusing on the fundamental problem of separating a two-class Gauss…
Two new algorithms improve Q* approximation in batch RL with linear error propagation.
The paper explores MAE as a loss function for DNN vector-to-vector regression, proving its advantages over MSE.
With the recent advancement in the deep learning technologies such as CNNs and GANs, there is significant improvement in the quality of the images reconstructed by deep learning based super-resolution (SR) techniques. In this work, we propose a robust loss function based on the preservation of edges obtained by the Can…
Cryptocurrency prices predicted using LSTM, SVM, and polynomial regression.
Mack's estimator improves chain ladder prediction for large exposure insurance models.
Hard to learn ReLU with Gaussian data, but can approximate efficiently.
Quantification of the stationary points and the associated basins of attraction of neural network loss surfaces is an important step towards a better understanding of neural network loss surfaces at large. This work proposes a novel method to visualise basins of attraction together with the associated stationary points…
Biased mean regression estimates factors exceeding expected loss or radiation release severity.
Given a task of predicting from , a loss function , and a set of probability distributions on , what is the optimal decision rule minimizing the worst-case expected loss over ? In this paper, we address this question by introducing a generalization of the principle of maximum entropy. Applying t…
Least Squares Estimators are suboptimal for 5D convex functions.
We introduce a novel semi-supervised version of the least squares classifier. This implicitly constrained least squares (ICLS) classifier minimizes the squared loss on the labeled data among the set of parameters implied by all possible labelings of the unlabeled data. Unlike other discriminative semi-supervised method…
Chain-ladder reserving is sensitive to outliers, leading to unreliable estimates.
Extends matrix factorization for deviance-based losses with GLM theory.
Efficiently removes specific data subsets without retraining for GDPR compliance.