In this paper, we propose a novel sufficient decrease technique for stochastic variance reduced gradient descent methods such as SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of stochastic varianc…
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose a novel sufficient decrease technique for variance reduced stochastic gradient descent methods such as SAG, SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of variance redu…
The paper analyzes the convergence of CART under a SID condition, improving previous results.
Study connects curvature bounds to map existence and flow solutions.
Numerical observations on martingale couplings are confirmed under certain conditions.
Topological entropy decreases strictly along Ricci flow near hyperbolic metrics.
New conditions for weighted composition operators in group homomorphisms.
Novel evolutionary strategy solves stochastic constrained optimization problems.
Given a geometrically finite hyperbolic cone-manifold, with the cone singularity sufficiently short, we construct a one parameter family of cone-manifolds decreasing the cone angle to zero. We also control the geometry of this one parameter family via the Schwarzian derivative of the projective boundary and the length …
In this paper, we study the relation of the monotonicity of Hawking Mass and geometric flow problems. We show that along the Hamilton-DeTurck flow with bounded curvature coupled with the modified mean curvature flow, the Hawking mass of the hypersphere with a sufficiently large radius in Schwarzschild spaces is monoton…
TREGO improves EGO for global optimization of high-dimensional problems.
The object of our investigation is a point that gives the maximum value of a potential with a strictly decreasing radially symmetric kernel. It defines a center of a body in Rm. When we choose the Riesz kernel or the Poisson kernel as the kernel, such centers are called a radial center or an illuminating center, respec…
Selective regression allows abstention to improve fairness criteria.
Many optimization algorithms converge to stationary points. When the underlying problem is nonconvex, they may get trapped at local minimizers and occasionally stagnate near saddle points. We propose the Run-and-Inspect Method, which adds an "inspect" phase to existing algorithms that helps escape from non-global stati…
We study the use of knowledge distillation to compress the U-net architecture. We show that, while standard distillation is not sufficient to reliably train a compressed U-net, introducing other regularization methods, such as batch normalization and class re-weighting, in knowledge distillation significantly improves …
We propose a method for recognizing moving vehicles, using data from roadside audio sensors. This problem has applications ranging widely, from traffic analysis to surveillance. We extract a frequency signature from the audio signal using a short-time Fourier transform, and treat each time window as an individual data …
ATSM are widely applied for pricing of bonds and interest rate derivatives but the consistency of ATSM when the short rate, r, is unbounded from below remains essentially an open question. First, the standard approach to ATSM uses the Feynman-Kac theorem which is easily applicable only when r is bounded from below. Sec…
Decentralized Bayesian learning reduces KL-divergence exponentially.
This paper optimizes high-dimensional oblique splits for decision trees, enhancing performance and computational efficiency.
Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporatin…
Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information matrix (FIM), which als…
We consider training over-parameterized two-layer neural networks with Rectified Linear Unit (ReLU) using gradient descent (GD) method. Inspired by a recent line of work, we study the evolutions of network prediction errors across GD iterations, which can be neatly described in a matrix form. When the network is suffic…
In this article, we derive a Bayesian model to learning the sparse and low rank PARAFAC decomposition for the observed tensor with missing values via the elastic net, with property to find the true rank and sparse factor matrix which is robust to the noise. We formulate efficient block coordinate descent algorithm and …
It is known that evolution strategies in continuous domains might not converge in the presence of noise. It is also known that, under mild assumptions, and using an increasing number of resamplings, one can mitigate the effect of additive noise and recover convergence. We show new sufficient conditions for the converge…
Batch normalization (batch norm) is often used in an attempt to stabilize and accelerate training in deep neural networks. In many cases it indeed decreases the number of parameter updates required to achieve low training error. However, it also reduces robustness to small adversarial input perturbations and noise by d…
Poly-view contrastive learning improves image representation learning.
Let be a reduced alternating diagram of a non-split link and be the link whose diagram is obtained from by a crossing change. If is alternating, then . In this paper we explore when holds and obtain a simple sufficient and necessary cond…
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
This work reveals how label noise can cause a final ascent in neural network performance curves.
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
Investment decisions shift earlier as patience decreases, with implications for pasting conditions.
Characterizes a general range decreasing group homomorphism.
Study shows HFT benefits large traders under certain conditions.
Margulis space-times with parabolic holonomy elements are stable under sufficiently small deformations.
Most random ReLU networks are vulnerable to small, Euclidean adversarial perturbations.
The study examines how investor protection and past information affect stock returns and interest rates.
D2SRM solves complex PDEs using deep learning.
We expect manifolds obtained by Dehn filling to inherit properties from the knot manifold. To what extent does that hold true for the Heegaard structure? We study four changes to the Heegaard structure that may occur after filling: (1) Heegaard genus decreases, (2) a new Heegaard surface is created, (3) a non-stabilize…
In general, homeowners refinance in response to a decrease in interest rates, as their borrowing costs are lowered. However, it is worth investigating the effects of refinancing after taking the underlying costs into consideration. Here we develop a synthetic mortgage calculator that sufficiently accounts for such cost…
We propose a method to decrease the number of hidden units of the restricted Boltzmann machine while avoiding decrease of the performance measured by the Kullback-Leibler divergence. Then, we demonstrate our algorithm by using numerical simulations.
New method improves convergence for smooth games.
LMC loss barrier decreases to zero with large network width.
Informed traders strategically reveal noisier signals, making prices less responsive to public information.
Sharp estimate for flow in any dimension.
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
Paper proves genus of surfaces decreases in mean curvature flow.
GD monotonically decreases GFS sharpness in neural networks and scalar models.
In this paper, we study the problem of learning a mixture of Gaussians with streaming data: given a stream of points in dimensions generated by an unknown mixture of spherical Gaussians, the goal is to estimate the model parameters using a single pass over the data stream. We analyze a streaming version of …