New guarantees for SGD in non-convex optimization without strict noise bounds.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Comparing with traditional learning criteria, such as mean square error (MSE), the minimum error entropy (MEE) criterion is superior in nonlinear and non-Gaussian signal processing and machine learning. The argument of the logarithm in Renyis entropy estimator, called information potential (IP), is a popular MEE cost i…
New results show flat minima in neural networks suffer from high dimensionality.
We propose two sparsity-aware normalized subband adaptive filter (NSAF) algorithms by using the gradient descent method to minimize a combination of the original NSAF cost function and the l1-norm penalty function on the filter coefficients. This l1-norm penalty exploits the sparsity of a system in the coefficients upd…
Double Q-learning has the same mean-squared error as Q-learning under certain conditions.
The kernel least mean squares (KLMS) algorithm is a computationally efficient nonlinear adaptive filtering method that "kernelizes" the celebrated (linear) least mean squares algorithm. We demonstrate that the least mean squares algorithm is closely related to the Kalman filtering, and thus, the KLMS can be interpreted…
Estimates latent inner products from an anisotropic Gaussian graph with improved spectral method.
The paper provides mean-square error bounds for stochastic approximation algorithms.
Constrained adaptive filtering algorithms inculding constrained least mean square (CLMS), constrained affine projection (CAP) and constrained recursive least squares (CRLS) have been extensively studied in many applications. Most existing constrained adaptive filtering algorithms are developed under mean square error (…
This paper analyzes sampling from heavy-tailed distributions using discretized Itô diffusions.
This work analyzes -learning with adaptive stepsizes for finite-time convergence.
This paper presents a stochastic behavior analysis of a kernel-based stochastic restricted-gradient descent method. The restricted gradient gives a steepest ascent direction within the so-called dictionary subspace. The analysis provides the transient and steady state performance in the mean squared error criterion. It…
This study calculates the maximum error of a famous estimation method.
New CH covariance class improves spatial statistics by balancing differentiability and tail behavior.
We consider adaptive system identification problems with convex constraints and propose a family of regularized Least-Mean-Square (LMS) algorithms. We show that with a properly selected regularization parameter the regularized LMS provably dominates its conventional counterpart in terms of mean square deviations. We es…
TACAM improves argument mining by integrating topic and external context.
New method optimizes tail dependence coefficient estimation.
Nonparametric modeling approaches show very promising results in the area of system identification and control. A naturally provided model confidence is highly relevant for system-theoretical considerations to provide guarantees for application scenarios. Gaussian process regression represents one approach which provid…
Cryptocurrency prices predicted using LSTM, SVM, and polynomial regression.
We formulate the problem of neural network optimization as Bayesian filtering, where the observations are the backpropagated gradients. While neural network optimization has previously been studied using natural gradient methods which are closely related to Bayesian inference, they were unable to recover standard optim…
New method finds best arm minimizing mean-squared error in correlated bandits.
Study on LMMSE estimation with model mismatch, quantifying MSE trade-offs.
We define two new notions of projection of a stochastic differential equation (SDE) onto a submanifold: the Ito-vector and Ito-jet projections. This allows one to systematically develop low dimensional approximations to high dimensional SDEs using differential geometric techniques. The approach generalizes the notion o…
A fast method for LOOCV in k-NN regression reduces computation time.
Develops optimal low-dimensional approximations to high-dimensional SDEs.
Kernel adaptive filters (KAF) are a class of powerful nonlinear filters developed in Reproducing Kernel Hilbert Space (RKHS). The Gaussian kernel is usually the default kernel in KAF algorithms, but selecting the proper kernel size (bandwidth) is still an open important issue especially for learning with small sample s…
New methods for estimating complex causal effects in econometrics.
We study the relationship between online Gaussian process (GP) regression and kernel least mean squares (KLMS) algorithms. While the latter have no capacity of storing the entire posterior distribution during online learning, we discover that their operation corresponds to the assumption of a fixed posterior covariance…
Unified framework for finite-sample RL algorithms using Lyapunov theory.
The purpose of this paper is to indicate that the recently proposed Momentum fractional least mean squares (mFLMS) algorithm has some serious flaws in its design and analysis. Our apprehensions are based on the evidence we found in the derivation and analysis in the paper titled: \textquotedblleft \textit{Momentum frac…
Paper presents a Siamese network for identifying more convincing evidence.
The most important aspect of any classifier is its error rate, because this quantifies its predictive capacity. Thus, the accuracy of error estimation is critical. Error estimation is problematic in small-sample classifier design because the error must be estimated using the same data from which the classifier has been…
This letter presents an improved version of diffusion least mean ppower (LMP) algorithm for distributed estimation. Instead of sum of mean square errors, a weighted sum of mean square error is defined as the cost function for global and local cost functions of a network of sensors. The weight coefficients are updated b…
Paper explores arbitrage and CAPM in continuous time.
Enhances RL for jump processes using MSBVE algorithm.
Paper optimizes diffusion models for denoising tasks with theoretical guarantees.
Improved multi-task averaging reduces mean squared error in high-dimensional data.
ABae efficiently computes subset means with expensive predicates using stratified sampling.
Simplified argument for second order estimate in quaternionic Calabi-Yau problem.
Study non-asymptotic estimation bounds for LTI models with Gaussian noise.
This is a continuation of our first paper in [WY16]. There are two purposes of this paper: One is to give a proof of the main result in [WY16] without going through the argument depending on numerical effectiveness. The other one is to provide a proof of our conjecture, mentioned in [TY], where the assumption of negati…
MIC improves VAR order selection accuracy.
RMSNorm simplifies LayerNorm, reducing computational cost.
Unified framework for robust A/B testing under model misspecification.
Improved stock volume prediction using Kalman Filters with various hidden states.
For -holomorphic mappings for a strongly pseudo-convex manifold, we prove elliptic regularity by the argument of boots-strapping.
Transformer model with mixed-frequency data improves stock volatility prediction.
The aim of this paper is to propose distributed strategies for adaptive learning of signals defined over graphs. Assuming the graph signal to be bandlimited, the method enables distributed reconstruction, with guaranteed performance in terms of mean-square error, and tracking from a limited number of sampled observatio…