Most of machine learning approaches have stemmed from the application of minimizing the mean squared distance principle, based on the computationally efficient quadratic optimization methods. However, when faced with high-dimensional and noisy data, the quadratic error functionals demonstrated many weaknesses including…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Bayes-optimal learning of a neural network with quadratic activations is achieved with GAMP-RIE.
New bound matches exact generalization error for quadratic Gaussian problem.
New findings on kernel regression in the quadratic regime, improving understanding of machine learning models.
We consider the tensor completion problem of predicting the missing entries of a tensor. The commonly used CP model has a triple product form, but an alternate family of quadratic models, which are the sum of pairwise products instead of a triple product, have emerged from applications such as recommendation systems. N…
The minimum error entropy (MEE) criterion has been verified as a powerful approach for non-Gaussian signal processing and robust machine learning. However, the implementation of MEE on robust classification is rather a vacancy in the literature. The original MEE only focuses on minimizing the Renyi's quadratic entropy …
We consider the problem of solving a large-scale Quadratically Constrained Quadratic Program. Such problems occur naturally in many scientific and web applications. Although there are efficient methods which tackle this problem, they are mostly not scalable. In this paper, we develop a method that transforms the quadra…
Study how generalization scales with model size and data in quadratic neural networks.
WildCat efficiently compresses neural network attention mechanisms.
Estimates surface count with prescribed foliations.
We study the performance of the certainty equivalent controller on Linear Quadratic (LQ) control problems with unknown transition dynamics. We show that for both the fully and partially observed settings, the sub-optimality gap between the cost incurred by playing the certainty equivalent controller on the true system …
We examine optimal quadratic hedging of barrier options in a discretely sampled exponential Lévy model that has been realistically calibrated to reflect the leptokurtic nature of equity returns. Our main finding is that the impact of hedging errors on prices is several times higher than the impact of other pricing bias…
New method reduces training time for deep hedging networks.
Optimizes prediction error method for time-varying models.
Recently, deep learning has achieved huge successes in many important applications. In our previous studies, we proposed quadratic/second-order neurons and deep quadratic neural networks. In a quadratic neuron, the inner product of a vector of data and the corresponding weights in a conventional neuron is replaced with…
Compress++ speeds up distribution compression to near-linear time.
Study on PG learning for LQ MFC problems with common noise, proving convergence and sample complexity.
Improved HGF networks avoid negative precision errors in volatility updates.
Sharp asymptotics reveal how network width controls learnability in quadratic neural networks.
We characterize the asymptotic performance of nonparametric one- and two-sample testing. The exponential decay rate or error exponent of the type-II error probability is used as the asymptotic performance metric, and an optimal test achieves the maximum rate subject to a constant level constraint on the type-I error pr…
Robust GQDA improves classification accuracy in non-Normal data.
Despite their practical success, a theoretical understanding of the loss landscape of neural networks has proven challenging due to the high-dimensional, non-convex, and highly nonlinear structure of such models. In this paper, we characterize the training landscape of the mean squared error loss for neural networks wi…
HA-SME models SGD dynamics with Hessian info for better escaping behaviors.
The paper debiases mini-batch approximations in deep learning for more accurate optimization and uncertainty quantification.
We consider high-dimensional quadratic classifiers in non-sparse settings. The target of classification rules is not Bayes error rates in the context. The classifier based on the Mahalanobis distance does not always give a preferable performance even if the populations are normal distributions having known covariance m…
The runtime for Kernel Partial Least Squares (KPLS) to compute the fit is quadratic in the number of examples. However, the necessity of obtaining sensitivity measures as degrees of freedom for model selection or confidence intervals for more detailed analysis requires cubic runtime, and thus constitutes a computationa…
Sharp 2-Wasserstein bounds for DDPMs derived from Föllmer process.
Quantification of the stationary points and the associated basins of attraction of neural network loss surfaces is an important step towards a better understanding of neural network loss surfaces at large. This work proposes a novel method to visualise basins of attraction together with the associated stationary points…
We analyze the errors arising from discrete readjustment of the hedging portfolio when hedging options in exponential Levy models, and establish the rate at which the expected squared error goes to zero when the readjustment frequency increases. We compare the quadratic hedging strategy with the common market practice …
Deep neural networks enforce non-crossing quantile regression curves.
Itô processes are the most common form of continuous semimartingales, and include diffusion processes. This paper is concerned with the nonparametric regression relationship between two such Itô processes. We are interested in the quadratic variation (integrated volatility) of the residual in this regression, over a un…
Gradient descent dynamics in quadratic regression models are analyzed, revealing five phases: monotonic, catapult, periodic, chaotic, and divergent.
Logarithmic regret achieved in continuous-time linear-quadratic reinforcement learning.
This paper addresses the optimal control problem known as the Linear Quadratic Regulator in the case when the dynamics are unknown. We propose a multi-stage procedure, called Coarse-ID control, that estimates a model from a few experimental trials, estimates the error in that model with respect to the truth, and then d…
New bounds for adaptive control in high dimensions without fixed state space.
We characterize the asymptotic performance of nonparametric goodness of fit testing. The exponential decay rate of the type-II error probability is used as the asymptotic performance metric, and a test is optimal if it achieves the maximum rate subject to a constant level constraint on the type-I error probability. We …
Neural operators correct PDE residuals to improve BIP solutions.
New bounds show linear predictors rarely overfit with certain optimization methods.
This paper aims at refined error analysis for binary classification using support vector machine (SVM) with Gaussian kernel and convex loss. Our first result shows that for some loss functions such as the truncated quadratic loss and quadratic loss, SVM with Gaussian kernel can reach the almost optimal learning rate, p…
Proposes QDF to improve multi-step time-series forecasting.
Quantum algorithm speeds up Lasso regression by quadratically faster per iteration.
New algorithm for robust regression with subgaussian error bound.
A contraction analysis improves model-based RL's error recovery.
Two new algorithms improve Q* approximation in batch RL with linear error propagation.
Paper analyzes holdout cross-validation for large non-Gaussian covariance estimation.
New method solves constrained stochastic optimization problems efficiently.
This paper solves quadratic systems with sparse or generative priors.
This paper tightens information-theoretic bounds on generalization errors.