New framework estimates treatment effects in extreme data.
problem Hindered by unavailability of counterfactual outcomes and rarity of extreme data.
method Proposes a new framework based on extreme value theory.
result Quantifies treatment effects using tail decay rates of potential outcomes.
The paper identifies five extreme learning regimes for large linear autoencoders.
problem Understanding the learning dynamics of large weight-tied linear autoencoders.
method Formal loss-expansion hierarchy and analysis of gradient flow.
result Five extreme regimes associated with faces of a triangular prism.
Cryptocurrency markets show higher spreads during extreme fear and greed phases.
problem Understanding and predicting liquidity withdrawal in cryptocurrency markets.
method Analysis of Crypto Fear & Greed Index and Bitcoin daily data.
result Extreme fear and greed regimes exhibit significantly higher spreads than neutral periods.
Inference over tails is usually performed by fitting an appropriate limiting distribution over observations that exceed a fixed threshold. However, the choice of such threshold is critical and can affect the inferential results. Extreme value mixture models have been defined to estimate the threshold using the full dat…
The study identifies and analyzes different market regimes in equity markets using advanced signal processing techniques.
problem Understanding and quantifying the dynamics of different market regimes in equity markets.
method Data-driven Hilbert--Huang Transform for regime identification, Holo--Hilbert Spectral Analysis for profiling, and Variable-Length Markov Chains for return dynamics modeling.
result Developed markets normalize more effectively as stress subsides, while developing markets retain residual tail dependence and downside persistence.
We analyze architectural features of Deep Neural Networks (DNNs) using the so-called Neural Tangent Kernel (NTK), which describes the training and generalization of DNNs in the infinite-width setting. In this setting, we show that for fully-connected DNNs, as the depth grows, two regimes appear: "order", where the (sca…
Model predicts global financial market risks and asset allocation.
problem Predicting downside risk and market regime shifts.
method Dynamic regime switching model based on GARCH-DCC-Copula.
result Significantly improves risk and alpha-based asset allocation strategies.
A new approach is presented to describe the change in the statistics of the log return distribution of financial data as a function of the timescale. To this purpose a measure is introduced, which quantifies the distance of a considered distribution to a reference distribution. The existence of a small timescale regime…
Estimates treatment effects in rare extreme events using EVT.
problem Estimating treatment effects in rare, impactful events like extreme climate events.
method Introduces a novel framework using EVT and multivariate regular variation for consistent treatment effect estimation.
result Developed a consistent estimator for extreme treatment effects with rigorous non-asymptotic analysis.
Two BO methods improve reliability optimization for rare failures.
problem Maximizing reliability of designs subject to random perturbations.
method Bayesian optimization with Thompson sampling and knowledge gradient.
result Proposed methods outperform existing techniques in extreme failure probability scenarios.
This paper presents the construction of a particle filter, which incorporates elements inspired by genetic algorithms, in order to achieve accelerated adaptation of the estimated posterior distribution to changes in model parameters. Specifically, the filter is designed for the situation where the subsequent data in on…
This research shows loss weighting remains effective in last layer retraining despite model overparameterization.
problem Overcoming biases in machine learning models at scale.
method Theoretical and practical exploration of last layer retraining in an overparameterized setting.
result Loss weighting is still effective in last layer retraining, but weights must account for model overparameterization.
Novel SVM approach for extreme quantile regression with heavy tailed inputs.
problem Learning from extreme values in quantile regression.
method Support Vector Machine framework for handling high-dimensional and nonlinear settings.
result Established finite-sample learning guarantees under mild regularity assumptions.
The problem of learning tree-structured Gaussian graphical models from independent and identically distributed (i.i.d.) samples is considered. The influence of the tree structure and the parameters of the Gaussian distribution on the learning rate as the number of samples increases is discussed. Specifically, the error…
We here present a model of the dynamics of extremism based on opinion dynamics in order to understand the circumstances which favour its emergence and development in large fractions of the general public. Our model is based on the bounded confidence hypothesis and on the evolution of initially anti-conformist agents to…
A robust implementation of a Dupire type local volatility model is an important issue for every option trading floor. Typically, this (inverse) problem is solved in a two step procedure : (i) a smooth parametrization of the implied volatility surface; (ii) computation of the local volatility based on the resulting call…
Comment on entropy learning for dynamic treatment regimes.
problem Evaluating dynamic treatment regimes using entropy loss.
method Optimization-based alternative to IPW estimate.
result Suggests optimization-based approach for evaluation.
Study free energy in spherical spin glasses, proving universality dichotomy.
problem Analyzing free energy in spherical spin glass models with different tail exponents.
method Introduced a tail-adapted normalization and used universality dichotomy.
result Sharp universality dichotomy for free energy across different tail exponents.
We assess cluster stability by trimming extreme points and tracking data range reduction.
problem Assessing stability of one-dimensional clusters.
method Probabilistic method using diameter-shrinkage ratio to track data range reduction.
result Our method achieves higher accuracy than classical tests in small or noisy samples.
Develops a new framework to analyze gradient flow regimes and derive explicit solutions.
problem Analyzing scaling regimes and deriving explicit analytic solutions for gradient flow in large learning problems.
method Formal power series expansion of the loss evolution with coefficients encoded by diagrams.
result Reveals different learning phases and obtains explicit solutions in some cases.
The distribution of recurrence times or return intervals between extreme events is important to characterize and understand the behavior of physical systems and phenomena in many disciplines. It is well known that many physical processes in nature and society display long range correlations. Hence, in the last few year…
Deep ensembles don't necessarily improve calibration in low data regimes.
problem Calibration issues in deep learning models, especially in low data regimes.
method Examination of data-augmentation, ensembling, and post-processing calibration methods.
result Standard ensembling techniques can lead to less calibrated models in low data regimes.
A robust algorithm for non-negative matrix factorization (NMF) is presented in this paper with the purpose of dealing with large-scale data, where the separability assumption is satisfied. In particular, we modify the Linear Programming (LP) algorithm of [9] by introducing a reduced set of constraints for exact NMF. In…
We provide explicit conditions on the distribution of risk-neutral log-returns which yield sharp asymptotic estimates on the implied volatility smile. We allow for a variety of asymptotic regimes, including both small maturity (with arbitrary strike) and extreme strike (with arbitrary bounded maturity), extending previ…
Deep learning models learn chaotic system dynamics from real and simulated data.
problem Training deep learning models for chaotic systems requires big data.
method Jointly train deep neural networks on real and simulated data, enforcing physical laws.
result Proposes knowledge-based deep learning (KDL) for accurate forecasting of chaotic systems.
The paper analyzes Kernel Density Estimation in high dimensions with varying data and dimensionality.
problem High-dimensional Kernel Density Estimation with growing data and dimensionality.
method Examines the behavior of Kernel Density Estimators in the regime where both data points and dimensionality grow with a fixed ratio.
result Three distinct statistical regimes are identified for Kernel-based density estimates, each with different statistical properties.
Paper analyzes and predicts Covid19 in Romania using neural networks and regime switching.
problem Inaccurate reported numbers and multiple influencing factors in pandemic prediction.
method Three-stage analysis using SIR model refined with neural networks and regime switching.
result Daily estimation of parameters and identification of regime turning points for predictions.
More data can actually hurt linear regression performance in certain conditions.
problem The test risk of linear regression estimators increases with additional samples in overparameterized settings.
method An analysis of linear regression with isotropic Gaussian covariates using gradient descent.
result The bias decreases with more samples, but variance increases, leading to a surprising increase in test risk.
Paper improves SVaR estimation for stress testing under macro scenarios using a hybrid GPR-HS framework.
problem Numerical instability in traditional SVaR estimation under extreme shocks.
method Extends GPR-HS framework to forward-looking stress scenarios with SACS for stable covariance.
result Stable SVaR ranges from -2.1020% to -2.2231%, preserving coherence property.
Backtests of structured strategies lose much of their predictive power in live trading.
problem Uncertainty in how marketed backtests predict live performance of structured strategies.
method Analysis of 1,726 structured strategies from ten global institutions.
result Raw backtests have limited portability into live trading and deteriorate sharply.
Bayesian VAR and Elliptical Black-Litterman models improve portfolio optimization during regime changes and heavy-tailed returns.
problem Portfolio optimization under market regime changes and heavy-tailed returns.
method BAVAR-BLED algorithm combining BAVAR and Black-Litterman models with Elliptical Distributions.
result Significant outperformance of state-of-the-art methods in Sharpe, Sortino ratios, and total returns.
New method estimates extreme outcomes in heavy-tailed data, breaking circular dependence.
problem Estimating outcomes for extreme events in heavy-tailed data.
method Proposes an ADRF estimator that includes a structured tail-shape output and a diagnostic to evaluate tail shape.
result Successfully reduces MAE in deep-tail and conditional-shortfall predictions.
New method reduces bias in learning from large action spaces using selective importance sampling.
problem Learning from large-scale recommendation systems with bandit feedback and supervised labels.
method Selective Importance Sampling (sIS) and Policy Optimization for eXtreme Models (POXM) algorithm.
result POXM method significantly outperforms existing methods in learning from bandit feedback on XMC tasks.
We consider a stochastic volatility model which captures relevant stylized facts of financial series, including the multi-scaling of moments. The volatility evolves according to a generalized Ornstein-Uhlenbeck processes with super-linear mean reversion. Using large deviations techniques, we determine the asymptotic sh…
Investors in stock market are usually greedy during bull markets and scared during bear markets. The greed or fear spreads across investors quickly. This is known as the herding effect, and often leads to a fast movement of stock prices. During such market regimes, stock prices change at a super-exponential rate and ar…
We study a nonparametric contextual bandit problem where the expected reward functions belong to a Hölder class with smoothness parameter β. We show how this interpolates between two extremes that were previously studied in isolation: non-differentiable bandits (β≤1), where rate-optimal regret is achieved by run…
Quant strategies lost during market selloff.
problem Why equities Statistical Arbitrage strategies failed during the COVID-19 market downturn.
method Explains strategies' normal performance and market regimes, discusses their limitations.
result Quant strategies are vulnerable to extreme market events.
Minimal DAMs can recognize patterns in high noise, even with minimal data.
problem Pattern recognition in high noise conditions with limited data.
method Interpolating between DAMs and spin glasses, using minimal dense associative networks and extremizing quenched free-energy.
result Minimal DAMs can correctly recognize patterns even when the signal is very weak and noise is high.
Theoretical and empirical taxonomy of imbalance in binary classification.
problem Class imbalance degrades binary classification performance.
method Proposed a principled framework based on three scales: imbalance coefficient, sample-dimension ratio, and intrinsic separability. Derived closed-form Bayes errors and analyzed degradation across models.
result The triplet (η, κ, Δ) provides a model-agnostic explanation of imbalance-induced deterioration.
Gradient descent and SGD achieve low test error in specific network weight regimes.
problem Optimizing two-layer ReLU networks with standard initialization.
method Gradient flow and stochastic gradient descent, analyzing margins and weight norms.
result Gradient descent and SGD can achieve globally maximal margins under certain constraints.
Knowledge transfer speeds up neural classifier training.
problem Lack of theoretical analysis of knowledge transfer in neural networks.
method Regularization of fit between teacher and student networks using privileged information.
result Wide two-layer networks can interpolate between privileged information and data, improving generalization.
Stablecoin system improves resilience to extreme market events.
problem Vulnerability of stablecoins to extreme volatility and adversarial attacks.
method MVF-Composer uses multi-agent simulations to stress-test and down-weight manipulative signals.
result Reduces peak peg deviation by 57% and mean recovery time by 3.1x under adversarial conditions.
RDLI integrates domain logic and context grounding to detect crypto anomalies under scarce labels.
problem Extreme label scarcity and evasion strategies in crypto networks.
method Relational Domain Logic Integration (RDLI) with Retrieval Grounded Context (RGC).
result RDLI outperforms GNN baselines by 28.9% in F1 score under 0.01% label scarcity.
New limits found for training deep learning models efficiently.
problem Optimizing the training speed of deep learning models without sacrificing accuracy.
method Applied stochastic thermodynamics to set speed limits for neural network training.
result Training neural networks is optimal within certain scaling assumptions.
Proposes QEP to mitigate quantization error propagation in layer-wise post-training quantization.
problem Growth of quantization errors across layers degrades performance, especially in low-bit regimes.
method Quantization Error Propagation (QEP) framework that explicitly propagates and compensates for quantization errors.
result QEP-enhanced layer-wise PTQ achieves substantially higher accuracy, especially in low-bit regimes.
We propose a novel time discretization for the log-normal SABR model and derive its asymptotic properties.
problem Analyzing the log-normal SABR model's time-discretized behavior and implied volatility surface.
method We use the Euler-Maruyama scheme for time discretization and derive asymptotic properties in the limit of large number of time steps.
result We derive an exact representation of the implied volatility surface for arbitrary maturity and strike in the asymptotic regime.
Paper introduces PHI to identify structurally distinct payment patterns in UK municipal procurement.
problem Vulnerability of public procurement to error, fraud, and corruption in high-volume transactions.
method Introduces Payment Heterogeneity Index (PHI) using Gaussian Mixture Model (GMM) and non-parametric statistics.
result Identifies a significant cohort with structurally distinct payment patterns, improving procurement oversight.
Exact tail probability bounds for bounded kurtosis.
problem Determining worst-case tail probabilities under kurtosis constraints.
method AI-guided search and certifying pipeline to compute bounds.
result A four-regime map of tail probabilities with explicit formulas.