Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

1223 · Aug 202519922001200920172026
48 results for first-moment

The paper analyzes Adam and SGD in nonstationary optimization, revealing tradeoffs between noise and drift.

problem Analyzing Adam and SGD in nonstationary optimization problems.
method Theoretical analysis of Adam and SGD under non-stationary stochastic objectives, separating two regimes.
result Characterizes the tradeoff between noise and drift in Adam and SGD, revealing when adaptive step-sizing is beneficial or harmful.

Adaptive gradient methods such as Adam have been shown to be very effective for training deep neural networks (DNNs) by tracking the second moment of gradients to compute the individual learning rates. Differently from existing methods, we make use of the most recent first moment of gradients to compute the individual …

2019-02-24abs ↗pdf ↗

In this paper we will study the statistics of the unit geodesic flow normal to the boundary of a hyperbolic manifold with non-empty totally geodesic boundary. Viewing the time it takes this flow to hit the boundary as a random variable, we derive a formula for its moments in terms of the orthospectrum. The first moment…

2013-03-26abs ↗pdf ↗

We propose SWA-Gaussian (SWAG), a simple, scalable, and general purpose approach for uncertainty representation and calibration in deep learning. Stochastic Weight Averaging (SWA), which computes the first moment of stochastic gradient descent (SGD) iterates with a modified learning rate schedule, has recently been sho…

2019-02-07abs ↗pdf ↗

We consider harmonic measures that arise from random walks on the mapping class group determined by probability distributions that have finite first moment with respect to the Teichmuller metric, and whose supports generate non-elementary subgroups. We prove that Teichmuller space with the Teichmuller metric is statist…

2019-09-30abs ↗pdf ↗

Random walks on hyperbolic spaces show linear growth in translation lengths.

problem Investigate the growth of translation lengths in random walks on hyperbolic spaces.
method Prove linear growth without moment conditions and apply to Teichmüller spaces.
result Linear growth of translation lengths in random walks on hyperbolic spaces.

Study resolvent convergence for random matrices with general covariance profiles.

problem Analyzing resolvent convergence for random matrices with non-identically distributed columns.
method Using moments of quadratic forms and deterministic equivalents, the study provides bounds on the trace of matrix products.
result The trace of matrix products is close to the trace of a deterministic equivalent, controlled by matrix norms.

We consider random walks on the mapping class group that have finite first moment with respect to the word metric, whose support generates a non-elementary subgroup and contains a pseudo-Anosov map whose invariant Teichmuller geodesic is in the principal stratum of quadratic differentials. We show that a Teichmuller ge…

2017-06-06abs ↗pdf ↗

Paper shows robust generative learning with minimal assumptions on target distributions.

problem Learning generative models with minimal assumptions on target distributions.
method Lipschitz-regularized αα-divergences with minimal assumptions.
result Stable learning across various target distributions with minimal assumptions.

We model non-stationary volume-price distributions with a log-normal distribution and collect the time series of its two parameters. The time series of the two parameters are shown to be stationary and Markov-like and consequently can be modelled with Langevin equations, which are derived directly from their series of …

2017-04-30abs ↗pdf ↗

This article presents a new model for demographic simulation which can be used to forecast and estimate the number of people in pension funds (contributors and retirees) as well as workers in a public institution. Furthermore, the model introduces opportunities to quantify the financial ows coming from future populatio…

2017-12-12abs ↗pdf ↗

Solves Christoffel-Minkowski problem and Hessian equations with radial symmetry.

problem Christoffel-Minkowski problem and Hessian equations under rotational symmetries.
method Constructing explicit convex solutions to mixed Monge-Ampère equations on \(\mathbb{R}^n\) under radial symmetry.
result Explicit representation formula for the support function of the resulting convex body.

This article presents an empirical study of thirteen derivative markets for commodity and financial assets. It compares the statistical properties of futures contracts's daily returns at different maturities, from 1998 to 2010 and for delivery dates up to 120 months. The analysis of the fourth first moments of the dist…

2010-10-28abs ↗pdf ↗

A pivotal problem in Bayesian nonparametrics is the construction of prior distributions on the space M(V) of probability measures on a given domain V. In principle, such distributions on the infinite-dimensional space M(V) can be constructed from their finite-dimensional marginals---the most prominent example being the…

2011-01-24abs ↗pdf ↗

Study shows singularity of stationary measure on Furstenberg boundary for certain random walks.

problem Singularity of stationary measure on Furstenberg boundary for random walks.
method Analysis of random walks on semisimple Lie groups with specific properties.
result Stationary measure is singular to Lebesgue measure in certain cases.

Robust estimation under Huber's εε-contamination model has become an important topic in statistics and theoretical computer science. Statistically optimal procedures such as Tukey's median and other estimators based on depth functions are impractical because of their computational intractability. In this paper, we est…

2018-10-04abs ↗pdf ↗

New EM algorithm improves deep generative network training.

problem Training deep generative networks with complex posterior and likelihood distributions.
method Derive analytical posterior and marginal distributions using CPA property, derive analytical EM algorithm.
result EM training yields higher likelihood than Variational Autoencoders (VAEs).

Study uniform learnability of binary classification networks with communication.

problem Learning a network with communication between vertices from uniform ergodic Random Graph Process.
method Introduced structural Rademacher complexity and used martingale method and Marton's coupling.
result Uniform learnability as worst-case theoretical limits for binary classification problems.

The paper develops methods for novelty detection on path space using signature-based statistics.

problem Novelty detection on path space as a hypothesis testing problem.
method Signature-based test statistics, transportation-cost inequalities, CVaR, one-class SVM algorithms.
result Established lower bounds on type-II\mathrm{II} error and general power bounds.

Randomly biased data makes complex models as easy to learn as simple ones.

problem Learning complex models like multi-index and sparse Boolean functions.
method Introducing a small random shift in the first moment of the data distribution.
result Randomly biased data makes Gaussian single index models and sparse Boolean functions as easy to learn as linear functions.

This paper refines bounds on random walk speed in Teichmüller space.

problem Understanding the speed of random walks on Teichmüller space.
method Analyzing Jenkins-Strebel directions and Lebesgue geodesics.
result The drift of random walks grows exponentially for typical geodesics and oscillates between linear and exponential for some geodesics.

Scale-free distributions and correlation functions found in financial data are reminiscent of the scale invariance of physical observables in the vicinity of a critical point. Here, we present empirical evidence for a transition phenomenon, accompanied by a symmetry breaking, in the investors' demand for stocks. We stu…

2001-11-19abs ↗pdf ↗

New methods target conditional demographic parity using optimal transport distances.

problem Auditing and enforcing conditional demographic parity (CDP) in models with complex conditioning variables.
method Developed novel measures of conditional demographic disparity (CDD) based on optimal transport distances and regularization-based approaches.
result Validated methods airbit{} and airlp{} effectively target CDP in real-world datasets with continuous model outputs.

The paper offers a framework to analyze machine learning problems using concentration of measure.

problem Analyzing machine learning algorithms defined by implicit equations.
method Develops a concentration of measure framework to solve convex problems and implicit formulations.
result Provides precise estimations for the first moments of the solution, describing the behavior and performance of machine learning classifiers.

New summary measures reveal geometric structure in weighted measures on manifolds.

problem Lack of geometric information in standard weight-only summaries.
method Heat-kernel entropy profiles, tracking nonuniformity across scales.
result Geometric effective sample size discounts nearby or duplicate particles.

Paper proposes robust risk measures for non-negative risks with partial information.

problem Tackles robustness of distortion risk measures under distributional uncertainty.
method Introduces new uncertainty sets and derives closed-form expressions for risk maximization.
result Derives closed-form expressions for risk maximization over uncertainty sets.

MTAdam optimizes multiple loss terms in neural models, balancing gradients dynamically.

problem Balancing multiple loss terms in neural model training is challenging and computationally demanding.
method Generalized Adam algorithm that computes separate derivatives and balances gradients across layers dynamically.
result Training with MTAdam leads to faster recovery from suboptimal initial loss weighting and matches conventional training outcomes.

Paper proposes an efficient algorithm to handle high-order portfolio moments.

problem Designing portfolios with high-order moments (skewness and kurtosis) is computationally challenging.
method Proposes a SCA algorithm framework for solving high-order portfolios efficiently.
result Demonstrates the efficiency of the proposed algorithm through numerical experiments.

Optimal transport theory characterizes convex order between probability measures.

problem Characterizing convex order between probability measures using optimal transport.
method Quantitative bounds on optimal transport, infimum of functionals over 1-Lipschitz functions.
result Two measures are in convex order if and only if a specific cost functional inequality holds.

Study problem-dependent rates in statistical learning theory, achieving optimal generalization error bounds.

problem Generalization error in statistical learning theory.
method Uniform localized convergence framework.
result Optimal generalization error bounds for various learning problems.

We relate ergodic-theoretic properties of a very small tree or lamination to the behavior of folding and unfolding paths in Outer space that approximate it, and we obtain a criterion for unique ergodicity in both cases. Our main result is that non-unique ergodicity gives rise to a transverse decomposition of the foldin…

2014-10-31abs ↗pdf ↗

Modified EAT method improves Poisson gradient estimation.

problem Challenging differentiation through Poisson-distributed latent variables.
method Exponential Arrival Time (EAT) simulation with modifications and Gumbel-SoftMax relaxation.
result Modified EAT method provides unbiased first moment and reduced second-moment bias.

The paper relaxes the stability condition to boost confidence in generalization for randomized learning algorithms.

problem The tension between uniform stability and L2L_2-stability in generalization bounds.
method Establishes in-expectation first moment generalization error bounds for L2L_2-stable randomized learning algorithms and uses subbagging to achieve near-tight exponential bounds.
result Improves generalization bounds for convex and non-convex optimization problems with SGD.

In this paper we study iterative procedures for stationary equilibria in games with large number of players. Most of learning algorithms for games with continuous action spaces are limited to strict contraction best reply maps in which the Banach-Picard iteration converges with geometrical convergence rate. When the be…

2012-10-17abs ↗pdf ↗

Iteratively reweighted least squares (IRLS) is a widely-used method in machine learning to estimate the parameters in the generalised linear models. In particular, IRLS for L1 minimisation under the linear model provides a closed-form solution in each step, which is a simple multiplication between the inverse of the we…

2016-05-24abs ↗pdf ↗

We study the column subset selection problem with respect to the entrywise 1\ell_1-norm loss. It is known that in the worst case, to obtain a good rank-kk approximation to a matrix, one needs an arbitrarily large nΩ(1)n^{Ω(1)} number of columns to obtain a (1+ε)(1+ε)-approximation to the best entrywise 1\ell_1-norm low ra…

2020-04-16abs ↗pdf ↗