The paper analyzes the convergence of CART under a SID condition, improving previous results.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In this paper, we propose a novel sufficient decrease technique for stochastic variance reduced gradient descent methods such as SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of stochastic varianc…
In this paper, we propose a novel sufficient decrease technique for variance reduced stochastic gradient descent methods such as SAG, SVRG and SAGA. In order to make sufficient decrease for stochastic optimization, we design a new sufficient decrease criterion, which yields sufficient decrease versions of variance redu…
Study connects curvature bounds to map existence and flow solutions.
Numerical observations on martingale couplings are confirmed under certain conditions.
New conditions for weighted composition operators in group homomorphisms.
Selective regression allows abstention to improve fairness criteria.
TREGO improves EGO for global optimization of high-dimensional problems.
The object of our investigation is a point that gives the maximum value of a potential with a strictly decreasing radially symmetric kernel. It defines a center of a body in Rm. When we choose the Riesz kernel or the Poisson kernel as the kernel, such centers are called a radial center or an illuminating center, respec…
ATSM are widely applied for pricing of bonds and interest rate derivatives but the consistency of ATSM when the short rate, r, is unbounded from below remains essentially an open question. First, the standard approach to ATSM uses the Feynman-Kac theorem which is easily applicable only when r is bounded from below. Sec…
Adam and RMSProp are two of the most influential adaptive stochastic algorithms for training deep neural networks, which have been pointed out to be divergent even in the convex setting via a few simple counterexamples. Many attempts, such as decreasing an adaptive learning rate, adopting a big batch size, incorporatin…
Normalization methods play an important role in enhancing the performance of deep learning while their theoretical understandings have been limited. To theoretically elucidate the effectiveness of normalization, we quantify the geometry of the parameter space determined by the Fisher information matrix (FIM), which als…
Investment decisions shift earlier as patience decreases, with implications for pasting conditions.
It is known that evolution strategies in continuous domains might not converge in the presence of noise. It is also known that, under mild assumptions, and using an increasing number of resamplings, one can mitigate the effect of additive noise and recover convergence. We show new sufficient conditions for the converge…
The -1 norm based optimization is widely used in signal processing, especially in recent compressed sensing theory. This paper studies the solution path of the -1 norm penalized least-square problem, whose constrained form is known as Least Absolute Shrinkage and Selection Operator (LASSO). A solution path …
Decentralized Bayesian learning reduces KL-divergence exponentially.
Let be a reduced alternating diagram of a non-split link and be the link whose diagram is obtained from by a crossing change. If is alternating, then . In this paper we explore when holds and obtain a simple sufficient and necessary cond…
How many samples are sufficient to guarantee that the eigenvectors and eigenvalues of the sample covariance matrix are close to those of the actual covariance matrix? For a wide family of distributions, including distributions with finite second moment and distributions supported in a centered Euclidean ball, we prove …
Paper generalizes Schwarz Lemma for VT harmonic maps with conditions.
Study optimal stopping times for multi-dimensional processes with non-exponential discounting.
Topological entropy decreases strictly along Ricci flow near hyperbolic metrics.
Study shows HFT benefits large traders under certain conditions.
Game-theoretic analysis of mining gaps in blockchain systems.
We expect manifolds obtained by Dehn filling to inherit properties from the knot manifold. To what extent does that hold true for the Heegaard structure? We study four changes to the Heegaard structure that may occur after filling: (1) Heegaard genus decreases, (2) a new Heegaard surface is created, (3) a non-stabilize…
Novel evolutionary strategy solves stochastic constrained optimization problems.
Given a geometrically finite hyperbolic cone-manifold, with the cone singularity sufficiently short, we construct a one parameter family of cone-manifolds decreasing the cone angle to zero. We also control the geometry of this one parameter family via the Schwarzian derivative of the projective boundary and the length …
Adaptive step-size improves optimization in complex geometries.
In this paper, we study the relation of the monotonicity of Hawking Mass and geometric flow problems. We show that along the Hamilton-DeTurck flow with bounded curvature coupled with the modified mean curvature flow, the Hawking mass of the hypersphere with a sufficiently large radius in Schwarzschild spaces is monoton…
New method improves convergence for smooth games.
In 1963, Polyak proposed a simple condition that is sufficient to show a global linear convergence rate for gradient descent. This condition is a special case of the Łojasiewicz inequality proposed in the same year, and it does not require strong convexity (or even convexity). In this work, we show that this much-older…
Many optimization algorithms converge to stationary points. When the underlying problem is nonconvex, they may get trapped at local minimizers and occasionally stagnate near saddle points. We propose the Run-and-Inspect Method, which adds an "inspect" phase to existing algorithms that helps escape from non-global stati…
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
Study on continuity of solutions for complex Monge-Ampère equations with movable singularities.
DG algorithms often fail to generalize well in limited domains, highlighting necessary vs. sufficient conditions.
We study the use of knowledge distillation to compress the U-net architecture. We show that, while standard distillation is not sufficient to reliably train a compressed U-net, introducing other regularization methods, such as batch normalization and class re-weighting, in knowledge distillation significantly improves …
We propose a method for recognizing moving vehicles, using data from roadside audio sensors. This problem has applications ranging widely, from traffic analysis to surveillance. We extract a frequency signature from the audio signal using a short-time Fourier transform, and treat each time window as an individual data …
The study defines new surfaces with specific cut locus properties and provides conditions for their existence.
New conditions prevent gaps in optimal control problems.
The paper studies entropy and free energy for harmonic metrics on cyclic Higgs bundles.
In this note, we show that the solution to the Dirichlet problem for the minimal surface system in any codimension is unique in the space of distance-decreasing maps. This follows as a corollary of the following stability theorem: if a minimal submanifold is the graph of a (strictly) distance-decreasing map, then $…
The optimal capital structure model with endogenous bankruptcy was first studied by Leland (1994) and Leland and Toft (1996), and was later extended to the spectrally negative Levy model by Hilberink and Rogers (2002) and Kyprianou and Surya (2007). This paper incorporates the scale effects by allowing the values of ba…
Establishes a condition for multiclass classification-calibration of Gamma-Phi losses.
This paper optimizes high-dimensional oblique splits for decision trees, enhancing performance and computational efficiency.
Paper investigates monotonicity issues in AI preference learning.
Study abelianization of Lie algebroids and groupoids, providing conditions for existence.
New findings on flatness for specific driftless systems.
Optimal dividend payout strategies with drawdown constraint identified.
We consider training over-parameterized two-layer neural networks with Rectified Linear Unit (ReLU) using gradient descent (GD) method. Inspired by a recent line of work, we study the evolutions of network prediction errors across GD iterations, which can be neatly described in a matrix form. When the network is suffic…