A simple log-transform fixes heavy-tailed data for generative models.
problem Standard generative models struggle with heavy-tailed data.
method Apply the soft-log transform to data before training and exponentiate samples after generation.
result Log-FM outperforms specialized baselines on multivariate benchmarks.
New measures capture tail dependence and non-exchangeability in financial data.
problem Underestimation of tail dependence and inability to capture non-exchangeable tail dependence.
method Tail copulas and novel tail dependence measures (MTCM, ATCM) are proposed.
result Captures non-exchangeable tail dependence and provides analytical forms for various copulas.
PH-VAE models heavy-tailed data with flexible Phase-Type distributions.
problem Standard VAEs fail to capture heavy-tailed behavior in real-world data.
method PH-VAE uses Phase-Type distributions defined by continuous-time Markov chains to adaptively model tail behavior.
result PH-VAE significantly outperforms existing heavy-tail-aware VAEs in approximating diverse heavy-tailed distributions.
The paper improves machine learning for heavy-tailed panel data.
problem Improving estimates for financial and economic data with fat tails.
method Sparse-group LASSO regularization and Fuk-Nagaev concentration inequality.
result Oracle inequalities for panel data estimators.
SS-GEN simulates rare events in heavy and light-tailed data.
problem Estimating probabilities of extreme events in multivariate data.
method Self-Similar Generative Estimation (SS-GEN) decomposes tail distribution into radial and angular components.
result SS-GEN generates representative extreme scenarios and estimates rare-event probabilities beyond observed data.
Study improves ERM for heavy-tailed data with dependent inputs.
problem Empirical Risk Minimization with dependent and heavy-tailed data.
method Extending risk bounds for ERM with heavy-tailed, dependent data.
result Established risk bounds for ERM with dependent and heavy-tailed data.
HTFM improves mode coverage and tail-statistic recovery for heavy-tailed data.
problem Tackles heavy-tailed data in various domains with rare events.
method Proposes a framework using clock-conditioned Gaussian sources and truncated logsignature features.
result Improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and baselines.
The book chapter discusses tail risk analysis for financial data using extreme value statistics.
problem Serial dependence in financial time series complicates tail risk assessment.
method The approach involves unconditional and conditional quantile forecasting.
result Serial dependence impacts multivariate tail dependence.
A Gaussian mixture model improves generalization for long-tailed data.
problem Optimizing generalization for rare data in long-tailed distributions.
method Suggested Gaussian mixture model and comparison of linear vs. nonlinear classifiers.
result Nonlinear classifiers outperform linear ones for long-tailed data.
Paper improves ETF tail-risk monitoring reliability.
problem Unreliable ETF risk monitoring under degraded data.
method Combines quality checks, prediction, scoring, and adjustment.
result Improves tail-risk monitoring, especially during stressed periods.
The paper investigates heavy-tailed behavior in offline SGD, showing it approximates power-law tails.
problem Understanding heavy-tailed behavior in offline (multi-pass) SGD with finite data.
method Proves nonasymptotic Wasserstein convergence bounds for offline SGD to online SGD.
result Offline SGD exhibits approximate power-law tails as the number of data points increases.
Markov chain decoders improve generative models' ability to produce heavy-tailed data.
problem Generative models struggle with heavy-tailed distributions.
method Replaced Gaussian decoder with Markov chain-based Phase-Type distributions.
result Significantly reduced tail Kolmogorov-Smirnov distance and extreme quantile error.
New framework controls generalization for heavy-tailed data in RLHF and SGLD.
problem Heavy-tailed data in modern learning pipelines.
method Tail-dependent information-theoretic framework for sub-Weibull data.
result Sharp generalization bounds for heavy-tailed data.
The paper uses EVT to improve tail risk measures under ambiguity sets.
problem Misspecification of tail risk measures leads to inflated risk estimates.
method Applies Extreme Value Theory to derive worst-case tail risk under ambiguity sets.
result Proposes a tail-calibrated ambiguity design that preserves nominal tail asymptotic scaling.
Paper improves normalizing flows to better capture distribution tails.
problem Difficult to learn tail behavior of distributions.
method Develops a new type of flows using flexible base distributions and data-driven linear layers.
result Improves accuracy, especially on distribution tails, and generates heavy-tailed data.
Study optimizes sampling to avoid extreme tail risks in unknown heavy-tailed distributions.
problem Identify optimal alternative with minimal extreme tail risk from unknown heavy-tailed distributions.
method Data-driven sequential sampling policies to maximize likelihood of selecting the optimal alternative.
result Proposed methods outperform existing approaches in identifying the optimal alternative.
New method estimates extreme outcomes in heavy-tailed data, breaking circular dependence.
problem Estimating outcomes for extreme events in heavy-tailed data.
method Proposes an ADRF estimator that includes a structured tail-shape output and a diagnostic to evaluate tail shape.
result Successfully reduces MAE in deep-tail and conditional-shortfall predictions.
Proposes a method to model financial returns with extreme shocks using flexible tail transformations.
problem Capturing extreme shocks in financial return data.
method Introduces a transformation layer in normalizing flows to model heavy-tailed distributions.
result Trained models can generate synthetic sets of extreme returns.
New method approximates CVaR with less data for heavy-tailed risks.
problem Lack of data for accurate CVaR approximation in heavy-tailed distributions.
method Importance sampling based extrapolation for heavy-tailed distributions.
result Statistically consistent approximations with reduced data requirements.
Generative Adversarial Network (GAN) simulates realistic multi-asset scenarios for tail risk.
problem Simulating realistic joint dynamics of multi-asset portfolios for tail risk estimation.
method Designing a GAN that preserves Value-at-Risk (VaR) and Expected Shortfall (ES) tail risk features.
result Correctly captures tail risk for a broad class of trading strategies and demonstrates strong generalization.
Framework for handling long-tailed multi-modal data.
problem Class imbalance and long-tailed distributions in multi-modal data.
method Multi-expert architecture with modality-specific networks and dynamic fusion weights.
result Framework outperforms existing methods in long-tailed, class-imbalanced scenarios.
DE-SGD shows heavy-tailed behavior in decentralized settings.
problem Heavy-tailed behavior in decentralized SGD.
method Analyzes the emergence of heavy-tails in DE-SGD, considering both quadratic and twice continuously differentiable strongly convex loss functions.
result DE-SGD exhibits heavier tails than centralized SGD, and tail behavior depends on network parameters.
DRAGON improves learning for rare classes in unbalanced datasets using class descriptions.
problem Learning rare classes in unbalanced datasets with deep models.
method DRAGON is a late-fusion architecture that corrects bias towards frequent classes and fuses class-descriptions to improve tail-class accuracy.
result DRAGON outperforms state-of-the-art models on new benchmarks for long-tail learning with class descriptors.
DSI improves tail-risk estimation in generative models by averaging checkpoints.
problem Generative models' instability in rare adverse scenarios.
method Diachronic Sample Integration (DSI) ensembles generated samples across checkpoints.
result DSI reduces tail-estimation error compared to single-checkpoint baselines.
Efficiently estimates sparse linear regression with heavy-tailed data and outliers.
problem Sparse estimation of linear regression coefficients with heavy-tailed covariates and noises, including outliers.
method Efficient computation of robust estimator with nearly optimal error bound.
result Nearly optimal error bound for robust sparse estimation.
Generative models learn better with data-adaptive noise.
problem Learning heavy-tailed distributions in flow-based models.
method Data-adaptive latent noise using 1D quantile functions optimized via Wasserstein distance.
result Flexibility and effectiveness in learning heavy-tailed and compactly supported distributions.
This work analyzes CVaR under heavy-tailed data, providing generalization and robustness bounds.
problem Understanding CVaR's behavior under heavy-tailed data and rare high-impact losses.
method Learning-theoretic analysis of CVaR-based empirical risk minimization.
result Sharp, high-probability generalization and excess risk bounds under minimal moment assumptions.
I report a new statistical distribution formulated to confront the infamous, long-standing, computational/modeling challenge presented by highly skewed and/or leptokurtic ("fat- or heavy-tailed") data. The distribution is straightforward, flexible and effective. Even when working with far fewer data points than are rou…
New approach tackles class imbalance in long-tailed datasets using domain adaptation techniques.
problem Class imbalance in long-tailed datasets leading to poor model performance.
method Proposes a meta-learning approach to estimate differences between class-conditioned distributions.
result Validated approach on six benchmark datasets and three loss functions.
Paper introduces Balanced Meta-Softmax for better long-tailed visual recognition.
problem Long-tailed distribution mismatch between training and testing data.
method Balanced Meta-Softmax, an unbiased extension of Softmax, using a Meta Sampler.
result Balanced Meta-Softmax outperforms state-of-the-art solutions on visual recognition and instance segmentation.
New algorithm detects changes in heavy-tailed data streams.
problem Detecting changes in heavy-tailed data streams.
method Clipped Stochastic Gradient Descent (SGD) combined with union bound.
result First algorithm with finite-sample false-positive rate guarantees for heavy-tailed data.
Efficiently estimates sparse linear regression with heavy-tailed and outlier-contaminated data.
problem Estimating sparse linear regression coefficients with heavy-tailed and outlier-contaminated data.
method Efficient computation of estimators with sharp error bounds.
result Sharp error bounds for efficient estimators.
TailGAN uses GANs to detect anomalies near data distribution tails.
problem Anomaly detection near data distribution tails with current GAN limitations.
method TailGAN leverages GANs with maximum entropy regularization to generate and detect anomalies near data distribution tails.
result TailGAN achieves competitive performance on various datasets compared to existing methods.
Paper develops heavy-tailed embeddings for better text classification and augmentation.
problem Improving text classification, especially for extreme values.
method Develops heavy-tailed embeddings using multivariate extreme value theory and introduces a scale-invariant classifier.
result The classifier outperforms baselines and generates meaningful augmented text.
New method allocates capital based on tail central moments for financial risk assessment.
problem Inability of CTE-based capital allocation to reflect tail behavior of losses.
method Developed TCM-based capital allocation for normal mean-variance mixture distributions.
result TCM-based method captures tail risk contributions not detected by CTE.
This work extends diffusion models to handle heavy-tailed targets, improving score estimation and sampling guarantees.
problem Score estimation and sampling guarantees for heavy-tailed targets in diffusion models.
method Kernel density estimation and minimax rates analysis for score estimation and sampling guarantees.
result Sharp minimax rates for score estimation and sampling guarantees for heavy-tailed targets, revealing qualitative differences between exponential and polynomial tails.
Heavy-tailed outliers are more resilient to robust estimation than adversarial ones.
problem Developing robust estimators for data with outliers.
method Analyzing the relationship between adversarial and heavy-tailed outlier models.
result Optimal estimators for heavy-tailed outliers are also optimal for adversarial settings, but not vice versa.
Improved VAE for heavy-tailed data using Student's t-distributions.
problem Over-regularization in VAEs with Gaussian priors.
method Proposed t3VAE framework with Student's t-distributions for prior, encoder, and decoder. result Significantly outperforms other models on heavy-tailed datasets.
Improved scalable machine learning under heavy-tailed data.
problem Machine learning scalability under heavy-tailed data without strong convexity.
method Simple robust validation sub-routine to boost confidence in gradient-based sub-processes.
result Substantial improvement in dimension dependence without strong convexity.
Improved privacy-preserving methods for convex optimization with heavy-tailed data.
problem Privacy-preserving optimization of convex functions with heavy-tailed data.
method Developed algorithms for private mean estimation and convex optimization under concentrated differential privacy constraints.
result Achieved improved upper bounds on excess population risk for convex and strongly convex loss functions.
Improved estimation of hedge fund tail risks using a novel model.
problem Estimation inefficiencies and need for manual threshold selection in extreme value regression models.
method Extended tail regression model with automatic threshold selection and artificial censoring.
result Significant link between tail risks and factors like equity momentum and financial stability index.
Study on error probability for classification of heavy-tailed renewal processes.
problem Error probability in classification of heavy-tailed renewal processes.
method Asymptotic expressions for Bhattacharyya bound on misclassification error probabilities.
result Obtained asymptotic expressions for misclassification error probabilities.
Study reveals heavy-tailed behavior in training ReLU gates.
problem Understanding heavy-tailed distribution in stochastic deep learning.
method Experimental study of heavy-tail index for S.G.D. and a variant.
result Two algorithms exhibit similar heavy-tail behavior on ReLU data.
Constructs tail-specific prediction intervals for financial applications
problem Financial applications require strict control on the left tail
method Extends classical conformal frameworks to provide explicit tail-specific guarantees
result Improved directional calibration in skewed data
A new method improves posterior approximation for complex distributions.
problem Difficulty in capturing multimodal and heavy-tailed posteriors with standard normalizing flows.
method StiCTAF: stick-breaking mixture base with component-wise tail adaptation.
result Improved tail recovery and better mode coverage compared to benchmarks.
Tail-GNNs improve protein function prediction using relational reinforcement.
problem Predicting hierarchical protein functions from sequence data.
method Combining Tail-GNNs with dilated convolutional networks for multi-task learning.
result Significant improvement in F_1 score for protein function prediction.
Kurtosis is seen as a measure of the discrepancy between the observed data and a Gaussian distribution and is defined when the 4th moment is finite. In this work an empirical study is conducted to investigate the behaviour of the sample estimate of kurtosis with respect to sample size and the tail index when applied to…
Novel framework for reliable long-tailed classification.
problem Challenges of long-tailed imbalance and specific error risks.
method Bayesian Decision Theory and variational optimization.
result Demonstrates reliability and flexibility in diverse tasks.