Study shows deterministic equivalent for neural network kernel convergence.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We find a deterministic equivalent for random feature regression's test error, independent of feature map dimension.
New theory extends LQ control to non-exponential discount scenarios.
Deterministic method for certifying neural network robustness.
Study extends bounds on sample covariance matrices with general dependence.
We study an equivalence of (i) deterministic pathwise statements appearing in the online learning literature (termed \emph{regret bounds}), (ii) high-probability tail bounds for the supremum of a collection of martingales (of a specific form arising from uniform laws of large numbers for martingales), and (iii) in-expe…
Unified derivation of high-dimensional linear models using stochastic gradient descent.
We derive deterministic criteria for the existence and non-existence of equivalent (local) martingale measures for financial markets driven by multi-dimensional time-inhomogeneous diffusions. Our conditions can be used to construct financial markets in which the \emph{no unbounded profit with bounded risk} condition ho…
The paper proves a non-asymptotic test error approximation for KRR.
Study of eigenvalues in nonlinear kernels for classification of separable data.
SVM generalizes well even with many support vectors in high dimensions.
Analyzes SGD dynamics in high-dimensional settings for GLMs and multi-index models.
For random matrix models, the parameter estimation based on the traditional likelihood functions is not straightforward in particular when we have only one sample matrix. We introduce a new parameter optimization method for random matrix models which works even in such a case. The method is based on the spectral distri…
In this paper, we analyze the fundamental conditions for low-rank tensor completion given the separation or tensor-train (TT) rank, i.e., ranks of unfoldings. We exploit the algebraic structure of the TT decomposition to obtain the deterministic necessary and sufficient conditions on the locations of the samples to ens…
In deterministic optimization, line searches are a standard tool ensuring stability and efficiency. Where only stochastic gradients are available, no direct equivalent has so far been formulated, because uncertain gradients do not allow for a strict sequence of decisions collapsing the search space. We construct a prob…
In deterministic optimization, line searches are a standard tool ensuring stability and efficiency. Where only stochastic gradients are available, no direct equivalent has so far been formulated, because uncertain gradients do not allow for a strict sequence of decisions collapsing the search space. We construct a prob…
In this paper we propose a look at the capital risk problem inspired by deterministic, known from classical mechanics, problem of juggling. We propose capital equivalents to the Newton's laws of motion and on this basis we determine the most secure form of credit repayment with regard to maximisation of profit. Then we…
The paper sets criteria for no arbitrage in complex financial models.
Develops a reinforcement learning algorithm for learning deterministic equilibrium policies in time-inconsistent control problems.
In this work we construct an optimal linear shrinkage estimator for the covariance matrix in high dimensions. The recent results from the random matrix theory allow us to find the asymptotic deterministic equivalents of the optimal shrinkage intensities and estimate them consistently. The developed distribution-free es…
Two new deterministic offspring selection methods reduce statistical distance in SMC and pMCMC.
We consider the problem of low canonical polyadic (CP) rank tensor completion. A completion is a tensor whose entries agree with the observed entries and its rank matches the given CP rank. We analyze the manifold structure corresponding to the tensors with the given rank and define a set of polynomials based on the sa…
The study proves Gaussian universality of deep random features learning.
Study online learning with set-valued feedback, showing differences between deterministic and randomized approaches.
Study the tradeoffs of bandit feedback in multiclass classification.
Recently, a number of mostly -norm regularized least squares type deterministic algorithms have been proposed to address the problem of \emph{sparse} adaptive signal estimation and system identification. From a Bayesian perspective, this task is equivalent to maximum a posteriori probability estimation under a …
Solves Merton's investment-consumption problem with certainty equivalent approach.
In this paper, we continue our study on a general time-inconsistent stochastic linear--quadratic (LQ) control problem originally formulated in [6]. We derive a necessary and sufficient condition for equilibrium controls via a flow of forward--backward stochastic differential equations. When the state is one dimensional…
Diffusion models' sampling paths lie in a low-dimensional subspace, resembling boomerangs.
In this paper we study arbitrage theory of financial markets in the absence of a numéraire both in discrete and continuous time. In our main results, we provide a generalization of the classical equivalence between no unbounded profits with bounded risk (NUPBR) and the existence of a supermartingale deflator. To obtain…
Kernel methods form a powerful, versatile, and theoretically-grounded unifying framework to solve nonlinear problems in signal processing and machine learning. The standard approach relies on the kernel trick to perform pairwise evaluations of a kernel function, which leads to scalability issues for large datasets due …
We present new algorithms for computing and approximating bisimulation metrics in Markov Decision Processes (MDPs). Bisimulation metrics are an elegant formalism that capture behavioral equivalence between states and provide strong theoretical guarantees on differences in optimal behaviour. Unfortunately, their computa…
Dropout is a simple yet effective algorithm for regularizing neural networks by randomly dropping out units through Bernoulli multiplicative noise, and for some restricted problem classes, such as linear or logistic regression, several theoretical studies have demonstrated the equivalence between dropout and a fully de…
Categorical d-separation criterion simplifies probability graph analysis.
Graph-EFM models weather uncertainty with graph-based ensembles.
Random matrix theory explains how neural networks adapt to data.
Improved machine learning models outperform their simpler counterparts by using imperfect labels.
We present a new family of models that is based on graphs that may have undirected, directed and bidirected edges. We name these new models marginal AMP (MAMP) chain graphs because each of them is Markov equivalent to some AMP chain graph under marginalization of some of its nodes. However, MAMP chain graphs do not onl…
This work exploits action equivariance for representation learning in reinforcement learning. Equivariance under actions states that transitions in the input space are mirrored by equivalent transitions in latent space, while the map and transition functions should also commute. We introduce a contrastive loss function…
This paper considers the problem of high dimensional signal detection in a large distributed network whose nodes can collaborate with their one-hop neighboring nodes (spatial collaboration). We assume that only a small subset of nodes communicate with the Fusion Center (FC). We design optimal collaboration strategies w…
Optimal SD improves ridge regression performance strictly and precisely.
Paper investigates separating times for general diffusions, providing new insights.
Study resolvent convergence for random matrices with general covariance profiles.
The article examines in some detail the convergence rate and mean-square-error performance of momentum stochastic gradient methods in the constant step-size and slow adaptation regime. The results establish that momentum methods are equivalent to the standard stochastic gradient method with a re-scaled (larger) step-si…
In this work we construct an optimal shrinkage estimator for the precision matrix in high dimensions. We consider the general asymptotics when the number of variables and the sample size so that . The precision matrix is estimated directly, wit…
PS-IG improves feature attribution by reducing noise and variance.
This work uses a scalable approach to identify partially observed nonlinear systems.
The study shows that generative models can be effectively used to understand neural network performance.