Markov regime switching models have been used in numerous empirical studies in economics and finance. However, the asymptotic distribution of the likelihood ratio test statistic for testing the number of regimes in Markov regime switching models has been an unresolved problem. This paper derives the asymptotic distribu…
We propose a new procedure for inference on optimal treatment regimes in the model-free setting, which does not require to specify an outcome regression model. Existing model-free estimators for optimal treatment regimes are usually not suitable for the purpose of inference, because they either have nonstandard asympto…
Dynamic treatment regimes are of growing interest across the clinical sciences as these regimes provide one way to operationalize and thus inform sequential personalized clinical decision making. A dynamic treatment regime is a sequence of decision rules, with a decision rule per stage of clinical intervention; each de…
Motivated by the study of Fano type varieties we define a new class of log pairs that we call asymptotically log Fano varieties and strongly asymptotically log Fano varieties. We study their properties in dimension two under an additional assumption of log smoothness, and give a complete classification of two dimension…
New tools reveal simple structure in complex hyperparameter loss surfaces near optima.
problem Understanding the behavior of hyperparameter loss surfaces near model optima.
method Developed a novel technique based on random search to uncover asymptotic features of the loss surface.
result Random search within the asymptotic regime yields a new distribution with parameters defining the loss surface.
Study of Lévy flights on Zoll surfaces, revealing geometric information.
problem Understanding the mean first capture time of Lévy flights on Zoll surfaces.
method Analysis of geodesic Lévy processes on Zoll surfaces, focusing on the first correction term.
result The first correction term encodes geometric information, specifically the degree of the conjugate point.
When can reliable inference be drawn in the "Big Data" context? This paper presents a framework for answering this fundamental question in the context of correlation mining, with implications for general large scale inference. In large scale data applications like genomics, connectomics, and eco-informatics the dataset…
New findings on k-NN algorithm's robustness under random data corruption.
problem Impact of random data corruption on k-NN algorithm performance. method Theoretical analysis of k-NN algorithm under random perturbation scenarios. result Phase transition phenomenon in asymptotic regret: small-ω regime vs large-ω regime. This note provides an error bound for the Hartman-Watson integral's leading term.
problem Bounding the error of the leading term of the Hartman-Watson integral.
method Asymptotic expansion analysis focusing on the regime rt=ρ constant. result The error term is bounded uniformly as ∣ϑ(t,ρ)∣≤701t. The study provides precise asymptotic theory for in-context learning by Transformers.
problem Understanding the sample complexity, pretraining task diversity, and context length for successful in-context learning.
method An exactly solvable model of linear regression task by linear attention, deriving sharp asymptotics.
result Double-descent learning curve with increasing pretraining examples, phase transition between low and high task diversity regimes.
New findings on kernel regression in the quadratic regime, improving understanding of machine learning models.
problem Understanding kernel ridge regression in the quadratic asymptotic regime.
method Extended study of kernel regression to the quadratic regime, establishing approximation bounds and spectral distributions.
result Broad class of inner-product kernels exhibit behavior similar to a quadratic kernel, with precise asymptotic training and test errors characterized.
We develop a local theory for the construction of singular spacetimes in all spacetime dimensions which become asymptotically self-similar as the singularity is approached. The techniques developed also allow us to construct and classify exact self-similar solutions which correspond to the formal asymptotic expansions …
This paper optimizes importance sampling for rare-event options pricing under the Heston model.
problem Efficiently pricing European call options with short maturity and deep out-of-the-money strikes.
method Asymptotic importance sampling schemes leveraging the large deviation principle and state-dependent change of measure.
result Proposed IS methods achieve logarithmic efficiency in short-maturity and deep OTM regimes, significantly reducing variance.
GNA optimally identifies the best arm with small gaps.
problem Best arm identification in fixed-budget settings.
method Generalized Neyman Allocation (GNA) for asymptotically locally minimax optimal BAI.
result GNA's worst-case bounds match the lower and upper bounds in the small-gap regime.
Comment on entropy learning for dynamic treatment regimes.
problem Evaluating dynamic treatment regimes using entropy loss.
method Optimization-based alternative to IPW estimate.
result Suggests optimization-based approach for evaluation.
Proves convergence of neural networks in a two-timescale regime.
problem Training dynamics of shallow neural networks.
method Two-timescale regime analysis of gradient flow.
result Gradient flow converges to global optimum in non-convex optimization.
Study on implied volatility of an affine jump-diffusion model.
problem Characterize implied volatility of an affine jump-diffusion model.
method Explicit moment generating function derived from solving ODEs; large deviation principle applied.
result Asymptotic behaviors of implied volatility in large-maturity and large-strike regimes characterized.
Optimal strategies are found for a repeated betting game using diffusion approximation.
problem Finding optimal strategies for a repeated betting game with i.i.d. outcomes.
method Constructing a diffusion approximation of the repeated game and analyzing the wealth share process.
result Necessary and sufficient conditions for the wealth share process to be transient or recurrent are derived.
Optimal graph classification uses message-passing neural networks.
problem Node classification on sparse graphs with fixed feature dimensions.
method Asymptotic local Bayes optimality, message-passing graph neural networks.
result Optimal message-passing architecture interpolates between MLP and convolution.
We consider a non-trapping n-dimensional Lorentzian manifold endowed with an end structure modeled on the radial compactification of Minkowski space. We find a full asymptotic expansion for tempered forward solutions of the wave equation in all asymptotic regimes. The rates of decay seen in the asymptotic expansion a…
We propose a novel time discretization for the log-normal SABR model and derive its asymptotic properties.
problem Analyzing the log-normal SABR model's time-discretized behavior and implied volatility surface.
method We use the Euler-Maruyama scheme for time discretization and derive asymptotic properties in the limit of large number of time steps.
result We derive an exact representation of the implied volatility surface for arbitrary maturity and strike in the asymptotic regime.
Study optimal stopping times under regime-switching models with constraints.
problem Optimal stopping times for discounted payoffs on a regime-switching geometric Brownian motion.
method Solve variational inequality to find value functions and optimal thresholds.
result Existence and expressions of optimal stopping times under specific conditions.
Derives TAP approximation for Bayesian linear regression.
problem Log-normalizing constant of posterior distribution in high-dimensional linear regression.
method Variational representation and Thouless-Anderson-Palmer approximation.
result Proves TAP approximation for spherical prior in proportional asymptotic regime.
Study leading-order asymptotics for VIX option prices in Bergomi models.
problem Understanding VIX option pricing in Bergomi models.
method Analytical approach to derive leading-order asymptotics for VIX option prices in Bergomi models.
result Closed-form solutions for VIX option prices in Bergomi models are derived.
The paper compares Bayesian uncertainty to MAP estimator in random features regression.
problem Comparing Bayesian uncertainty to MAP estimator in random features regression.
method Analyzing the variance of the posterior predictive distribution and comparing it to the risk of the MAP estimator.
result Asymptotic agreement between Bayesian uncertainty and MAP estimator under specific signal-to-noise ratios and sample sizes.
The paper analyzes learning curves for kernel ridge regression with dot-product kernels.
problem Understanding the learning curves for different scaling regimes of data and model.
method Precise formulas for mean test error, bias, and variance in the mo∞ with m/dr constant regime. result A peak in the learning curve at m≈dr/r! for any integer r. Stochastic gradient descent (SGD) is almost ubiquitously used for training non-convex optimization tasks. Recently, a hypothesis proposed by Keskar et al. [2017] that large batch methods tend to converge to sharp minimizers has received increasing attention. We theoretically justify this hypothesis by providing new pro…
New algorithms improve sampling from complex distributions.
problem Sampling from complex probability distributions efficiently.
method Regime-switching Langevin dynamics and Monte Carlo algorithms.
result Convergence guarantees and iteration complexities provided.
In a market with a rough or Markovian mean-reverting stochastic volatility there is no perfect hedge. Here it is shown how various delta-type hedging strategies perform and can be evaluated in such markets in the case of European options. A precise characterization of the hedging cost, the replication cost caused by th…
Study reveals three limiting regimes for neural network functionals.
problem Understanding the behavior of functionals of random neural networks.
method Central and non-central limit theorems, Hermite expansions, Diagram Formula, Stein-Malliavin techniques.
result Three distinct limiting regimes based on fixed points of covariance function.
We analyze waiting times for price changes in a foreign currency exchange rate. Recent empirical studies of high frequency financial data support that trades in financial markets do not follow a Poisson process and the waiting times between trades are not exponentially distributed. Here we show that our data is well ap…
We provide explicit conditions on the distribution of risk-neutral log-returns which yield sharp asymptotic estimates on the implied volatility smile. We allow for a variety of asymptotic regimes, including both small maturity (with arbitrary strike) and extreme strike (with arbitrary bounded maturity), extending previ…
We extend the Lyapunov-Schmidt analysis of outlying stable CMC spheres in the work of S. Brendle and the second-named author to the "far-off-center" regime and to include general Schwarzschild asymptotics. We obtain sharp existence and non-existence results for large stable CMC spheres that depend very delicately on th…
Study resolves conjecture on overparameterized linear models' generalization.
problem Asymptotic generalization of multiclass classification with overparameterized models.
method Gaussian covariates bi-level model, Hanson-Wright inequality variant.
result Min-norm interpolating classifier can be suboptimal compared to noninterpolating classifiers.
Log-normal continuous random cascades form a class of multifractal processes that has already been successfully used in various fields. Several statistical issues related to this model are studied. We first make a quick but extensive review of their main properties and show that most of these properties can be analytic…
Large batch sizes reduce gradient variance in DP-SGD, improving privacy.
problem Understanding why large batch sizes work in DP-SGD.
method Decomposed total gradient variance into subsampling and noise-induced variances, proving batch size independence in the limit.
result Large batch sizes reduce effective total gradient variance, improving privacy in DP-SGD.
Study short-maturity VIX and European option prices with jumps.
problem Analyzing VIX and European options with jumps in short-maturity models.
method Local-stochastic volatility models with compound Poisson jumps, leading-order asymptotics in closed-form.
result Closed-form solutions for VIX and European option prices in short-maturity models.
Study examines how two-layer networks learn features after one gradient step.
problem Understanding feature learning in neural networks after a single gradient descent step.
method Modeling the trained network as a spiked Random Features (sRF) model and leveraging Gaussian universality.
result Exact asymptotic description of the generalization error of the sRF in high-dimensional limit.
Optimal tuning for estimating ECC in proportional asymptotics.
problem Estimating Expected Conditional Covariance (ECC) under proportional asymptotics.
method Debiased ridge regression estimators for nuisance functions, sample splitting strategies, and asymptotic variance analysis.
result Prediction-optimal tuning parameters may not minimize asymptotic variance of ECC estimator.
In this paper, we study stochastic volatility models in regimes where the maturity is small, but large compared to the mean-reversion time of the stochastic volatility factor. The problem falls in the class of averaging/homogenization problems for nonlinear HJB-type equations where the "fast variable" lives in a noncom…
Optimizes stochastic linear bandits with efficient, asymptotically optimal algorithm.
problem Optimizing stochastic linear bandits with multiple actions.
method Frequentist information-directed sampling (IDS) with a surrogate for information gain.
result Asymptotically optimal and nearly worst-case optimal in finite time.
We study a portfolio selection problem in a continuous-time Itô-Markov additive market with prices of financial assets described by Markov additive processes which combine Lévy processes and regime switching models. Thus the model takes into account two sources of risk: the jump diffusion risk and the regime switching …
Study on generalisation in random feature learning and hidden manifold models.
problem Generalisation in high-dimensional learning problems.
method Replica method from statistical physics for asymptotic generalisation performance.
result Closed-form expression for generalisation performance in various high-dimensional settings.
A fast regime-split Black-Scholes implied volatility solver
problem Fast computation of implied volatility
method Analytical and numerical expansions
result Achieves near-machine precision with minimal iterations
The paper studies Hawkes processes under mean-field limits and criticality conditions.
problem Analyzing nearly unstable Hawkes processes in a mean-field regime.
method Extending the method by Jaisson and Rosenbaum, establishing scaling limits and propagation of chaos.
result Scaling limits of Hawkes processes are stochastic Volterra diffusions of affine type, with three distinct limiting regimes.
This work shows linear convergence for two-layer neural networks in mean-field regime.
problem Optimizing two-layer neural networks in the mean-field regime.
method Mean-field analysis and continuous-time noisy gradient descent.
result Establishes linear convergence rate for two-layer neural networks.
The paper studies neural networks with wide layers and finds a deformed semicircle law.
problem Investigating spectral distributions of neural networks in the ultra-wide regime.
method Analyzes empirical kernel matrices, proves deformed semicircle law, provides nonlinear Hanson-Wright inequality.
result Emergence of a deformed semicircle law in the ultra-wide neural network regime.
This work analyzes self-attention matrices using random matrix theory.
problem Understanding the theoretical behavior of self-attention layers in neural networks.
method Asymptotic spectral analysis of the attention matrix, Gaussian equivalence, and linearization.
result The singular value distribution of the attention matrix is asymptotically characterized by a linear model.