Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

2.1%4.2%6.3%8.4% · May 202619922001200920172026
48 results for variance explosion

Noise injection before gradient steps helps in regularization for neural networks.

problem Improving generalization in overparametrized neural networks.
method Injecting small noise perturbations before computing gradient steps, especially in layer-wise fashion.
result Small noise perturbations can explicitly regularize neural networks without variance explosion.

We conduct mathematical analysis on the effect of batch normalization (BN) on gradient backpropogation in residual network training, which is believed to play a critical role in addressing the gradient vanishing/explosion problem, in this work. By analyzing the mean and variance behavior of the input and the gradient i…

2018-12-02abs ↗pdf ↗

We propose a randomised version of the Heston model-a widely used stochastic volatility model in mathematical finance-assuming that the starting point of the variance process is a random variable. In such a system, we study the small-and large-time behaviours of the implied volatility, and show that the proposed random…

2016-08-25abs ↗pdf ↗

We show that the moment explosion time in the rough Heston model [El Euch, Rosenbaum 2016, arxiv:1609.02108] is finite if and only if it is finite for the classical Heston model. Upper and lower bounds for the explosion time are established, as well as an algorithm to compute the explosion time (under some restrictions…

2018-01-29abs ↗pdf ↗

We present transductive Boltzmann machines (TBMs), which firstly achieve transductive learning of the Gibbs distribution. While exact learning of the Gibbs distribution is impossible by the family of existing Boltzmann machines due to combinatorial explosion of the sample space, TBMs overcome the problem by adaptively …

2018-05-21abs ↗pdf ↗

Study on martingale property and moment explosions in signature volatility models.

problem Analyzing the martingale property and moment explosions in signature volatility models.
method Fine analysis of the explosion time of a signature stochastic differential equation.
result The price process is a true martingale if and only if the order of the linear form is odd and a correlation parameter is negative.

Classical (Itô diffusions) stochastic volatility models are not able to capture the steepness of small-maturity implied volatility smiles. Jumps, in particular exponential Lévy and affine models, which exhibit small-maturity exploding smiles, have historically been proposed to remedy this (see \cite{Tank} for an overvi…

2015-03-27abs ↗pdf ↗

We study the explosion of the solutions of the SDE in the quasi-Gaussian HJM model with a CEV-type volatility. The quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation which simplifies their numerical implementation …

2019-08-19abs ↗pdf ↗

New RL algorithm GDPO improves DLM reasoning efficiency.

problem Adapting RL to DLMs for efficient, unbiased likelihood estimation.
method Group Diffusion Policy Optimization (GDPO) using semi-deterministic Monte Carlo.
result GDPO outperforms existing methods on math, reasoning, and coding benchmarks.

SCOPE-FE improves feature engineering efficiency for high-dimensional datasets.

problem Expanding and reducing feature space in tabular learning becomes computationally expensive with increased dimensionality.
method SCOPE-FE controls the search space by regulating operator and feature-pair spaces, using OperatorProbing and FeatureClustering.
result SCOPE-FE reduces feature engineering time while maintaining competitive predictive performance.

Bayesian model improves categorization of explosions from sparse data.

problem Challenges in categorizing explosions from limited data.
method Bayesian update to Event Categorization Matrix model with Bayesian Decision Theory.
result Consistent gains in overall accuracy and lower false negative rates.

In this paper we investigate the asymptotics of forward-start options and the forward implied volatility smile in the Heston model as the maturity approaches zero. We prove that the forward smile for out-of-the-money options explodes and compute a closed-form high-order expansion detailing the rate of the explosion. Fu…

2013-03-18abs ↗pdf ↗

Recently, there is an explosive growth of activities to understand stringy properties of orbifolds. In this article, we survey some of recent developments.

2002-01-15abs ↗pdf ↗

Quasi-Gaussian HJM models are a popular approach for modeling the dynamics of the yield curve. This is due to their low dimensional Markovian representation, which greatly simplifies their numerical implementation. We present a qualitative study of the solutions of the quasi-Gaussian log-normal HJM model. Using a small…

2019-08-19abs ↗pdf ↗

Study on VIX options pricing in SABR model, showing infinite prices due to volatility explosion.

problem Infinite VIX futures and call prices due to volatility explosion in SABR model.
method Analyzing SABR model, showing vtv_t as unique solution to diffusion process, proving explosion using Feller test, proposing capped volatility process.
result VIX futures and call prices are infinite for any maturity due to volatility explosion, but capped volatility process mitigates this issue.

RestoreAI predicts landmine risk from patterns, improving clearance efficiency.

problem Predicting landmine risk from spatial patterns to enhance clearance efficiency.
method RestoreAI uses landmine patterns for risk prediction, implementing three deminers: linear, curved, and Bayesian.
result RestoreAI significantly boosts clearance efficiency, achieving a 14.37 percentage point increase in cleared landmines per timestep.

Graph Convolutional Networks (GCNs) are powerful models for learning representations of attributed graphs. To scale GCNs to large graphs, state-of-the-art methods use various layer sampling techniques to alleviate the "neighbor explosion" problem during minibatch training. We propose GraphSAINT, a graph sampling based …

2019-07-10abs ↗pdf ↗

Wide-band Electromagnetic Induction Sensors (WEMI) have been used for a number of years in subsurface detection of explosive hazards. While WEMI sensors have proven effective at localizing objects exhibiting large magnetic responses, detecting objects lacking or containing very low amounts of conductive materials can b…

2019-03-22abs ↗pdf ↗

In the LIBOR market model, forward interest rates are log-normal under their respective forward measures. This note shows that their distributions under the other forward measures of the tenor structure have approximately log-normal tails.

2010-08-12abs ↗pdf ↗

Paper addresses theoretical risks in neural MCCFR, proposing Robust Deep MCCFR for improved performance.

problem Theoretical risks in neural MCCFR, especially in large games.
method Adaptive framework with selective component deployment, including target networks, exploration, and variance-aware training.
result Robust Deep MCCFR achieves significant exploitability improvements in both Kuhn and Leduc Poker.

We exhibit sufficient conditions such that components of a multidimensional SDE giving rise to a local martingale MM are strict local martingales or martingales. We assume that the equations have diffusion coefficients of the form σ(Mt,vt),σ(M_t,v_t), with vtv_t being a stochastic volatility term.

2019-03-06abs ↗pdf ↗

At present, there is an explosion of practical interest in the pricing of interest rate (IR) derivatives. Textbook pricing methods do not take into account the leptokurticity of the underlying IR process. In this paper, such a leptokurtic behaviour is illustrated using LIBOR data, and a possible martingale pricing sche…

2004-01-23abs ↗pdf ↗

New neural networks model complex phenomena with fewer parameters.

problem Challenges in studying higher-order interactions in neural networks.
method Introducing curved neural networks using the maximum entropy principle.
result Curved neural networks accelerate memory retrieval and exhibit explosive phase transitions.

We consider an interest rate model with log-normally distributed rates in the terminal measure in discrete time. Such models are used in financial practice as parametric versions of the Markov functional model, or as approximations to the log-normal Libor market model. We show that the model has two distinct regimes, a…

2011-04-02abs ↗pdf ↗

Lipschitz normalization boosts deep attention models, especially for graph neural networks.

problem Gradient explosion in deep graph attention networks leads to poor performance.
method Enforcing Lipschitz continuity by normalizing attention scores.
result Deep GAT models with LipschitzNorm achieve state-of-the-art results for tasks with long-range dependencies.

We consider the stochastic volatility model dSt=σtStdWt,dσt=ωσtdZtdS_t = σ_t S_t dW_t,dσ_t = ωσ_t dZ_t, with (Wt,Zt)(W_t,Z_t) uncorrelated standard Brownian motions. This is a special case of the Hull-White and the β=1β=1 (log-normal) SABR model, which are widely used in financial practice. We study the properties of this model, discretized in …

2017-07-04abs ↗pdf ↗

Detects anomalies in astronomical time series data.

problem Identifying new and interesting transients in large astronomical surveys.
method Two novel methods: a probabilistic neural network and a Bayesian parametric model.
result Neural networks are less suitable for anomaly detection in time series data compared to parametric models.

Early training phase affects deep neural network optimization and generalization.

problem The choice of learning rate influences generalization in deep learning models.
method Showed that SGD implicitly penalizes the trace of the Fisher Information Matrix (FIM) from the start of training, and explicitly penalizing the trace of FIM improves generalization.
result Catastrophic Fisher explosion (large trace of FIM early in training) is linked to poor generalization.

Ripple Walk Training tackles graph neural network training issues for large and deep graphs.

problem Neighbors explosion, node dependence, and oversmoothing in large and deep GNNs.
method Subgraph-based training framework with Ripple Walk Sampler for high-quality subgraph sampling.
result RWT improves training efficiency and reduces space complexity for deep and large GNNs.