Learning shrinks hard tail, improving inference performance.
problem Improving inference performance in neural networks.
method Latent Instance Difficulty (LID) model analyzing fine-tuning of neural networks.
result Training-dependent inference scaling, with βexteff growing with sample size before saturating. ELF improves long-tailed classification by focusing on hard examples.
problem Overfitting to majority classes in long-tailed data distributions.
method EARLY-exiting Framework with auxiliary branches.
result Improves accuracy by more than 3 percent on ImageNet LT and iNaturalist'18.
We consider the problem of sparsity-constrained M-estimation when both explanatory and response variables have heavy tails (bounded 4-th moments), or a fraction of arbitrary corruptions. We focus on the k-sparse, high-dimensional regime where the number of variables d and the sample size n are related through $…
Privacy affects how much data is needed for CVaR optimization.
problem Privacy constraints impact the effective sample size for CVaR optimization.
method Analyzes the privacy-relevant sample size and decomposes CVaR excess risk.
result The effective private tail sample size is εnτ, affecting CVaR learning rates.
Algorithm distinguishes light-tailed from non-light-tailed distributions.
problem Characterize the tail of a distribution using hazard rate.
method Careful bucketing scheme based on hazard rate.
result Polynomial number of samples required for success.
Private learning is hard when data is long-tailed.
problem Achieving both privacy and fairness in machine learning with long-tailed data.
method Theoretical analysis and experimental validation on various datasets and algorithms.
result Relaxing overall accuracy can lead to good fairness even with strict privacy requirements.
Proposes a new sampling policy for ranking and selection problems.
problem Improving ranking and selection in adaptive sampling policies.
method Annealed entropic allocation, using soft-min weights and saddlepoint corrections.
result Consistently competitive performance in various settings.
Studied how heavy-tailed behavior affects SGD's generalization in quadratic optimization.
problem Link between heavy-tailed behavior and generalization in SGD.
method Used heavy-tailed stochastic differential equation and proved stability bounds.
result Stability of SGD depends on the loss function's tail behavior.
Study shows flash crashes in finance are self-organized criticality events.
problem Understanding and predicting anomalous price events in high-frequency finance.
method Investigated volume distributions during flash crashes and linked them to self-organized criticality.
result Volume distributions during flash crashes indicate a diverging second moment, suggesting self-organized criticality.
AIHT improves online high-dimensional quantile regression by separating support discovery and refinement.
problem Online high-dimensional quantile regression with structural sparsity.
method Adaptive Iterative Hard Thresholding (AIHT) alternates stochastic updates with adaptive hard-thresholding steps.
result AIHT achieves logarithmic regret for the sliding-window objective in high-dimensional settings.
New algorithm resists contamination in high-dimensional regression with optimal performance.
problem Adversarial and measurement errors in high-dimensional data.
method Adversarial Contamination-resistant Iterative Hard Thresholding (AC-IHT) algorithm.
result Achieves minimax near-optimal estimation and signal-adaptive support recovery.
Model predicts methane emissions from oil sands tailing ponds, suggesting significant environmental impact.
problem Estimating methane emissions from inactive oil sands tailing ponds.
method Physics constrained machine learning model using real-time weather data and laboratory experiments.
result Active oil sands tailing ponds emit between 950 to 1500 tonnes of methane per year, equivalent to 6000 gasoline vehicles.
TuneUp improves GNN training by focusing on hard-to-learn nodes.
problem Sub-optimal training of GNNs on all nodes equally.
method Two-stage training: base GNN + tail node improvement.
result Significant improvement in tail node prediction performance.
The paper optimizes stock portfolios with constraints based on performance attribution.
problem Optimizing stock portfolios with performance attribution constraints.
method Minimizes expected tail loss, constrains asset allocation and selection effect, tests on Dow Jones stocks.
result Imposing constraints on asset allocation and selection effect improves portfolio performance.
Tail-Safe hedging uses reinforcement learning with a safety layer to manage financial risks.
problem Managing financial risks in derivatives trading with robustness and explainability.
method Combines distributional reinforcement learning with a CBF-QP safety layer to enforce financial constraints.
result Improves risk management without degrading central performance and avoids hard constraint violations.
We study the problem of robust linear regression with response variable corruptions. We consider the oblivious adversary model, where the adversary corrupts a fraction of the responses in complete ignorance of the data. We provide a nearly linear time estimator which consistently estimates the true regression vector, e…
The condition for stationary increments, not scaling, detemines long time pair autocorrelations. An incorrect assumption of stationary increments generates spurious stylized facts, fat tails and a Hurst exponent H_s=1/2, when the increments are nonstationary, as they are in FX markets. The nonstationarity arises from s…
Annealed Entropic Allocation improves ranking and selection by mitigating hard switching and improving finite-budget discrimination.
problem Sequential budget allocation in ranking and selection
method Annealed weighted soft-min framework
result Surrogate converges uniformly to the hard minimum, soft-min weights concentrate on active challengers, and target allocation map is continuous.
Local Gaussian correlation struggles in tails but a new method improves it.
problem Local Gaussian correlation's limitations in tail dependence.
method A new adaptive bandwidth method for LGC, optimizing for local effective sample size.
result Adaptive bandwidths outperform global ones in moderate dependence, but not in strong or weak dependence.
Soft diamond regularizers improve deep learning performance and sparsity.
problem Improving deep learning performance and sparsity of trained weights.
method New soft diamond synaptic weight priors based on thick-tailed symmetric alpha stable probability curves.
result Soft diamond regularizers outperform state-of-the-art methods in deep learning tasks.
We use daily data on bilateral interbank exposures and monthly bank balance sheets to study network characteristics of the Russian interbank market over Aug 1998 - Oct 2004. Specifically, we examine the distributions of (un)directed (un)weighted degree, nodal attributes (bank assets, capital and capital-to-assets ratio…
New model approximates sparse mean-CVaR portfolio optimization efficiently.
problem NP-hard ℓ0-constrained mean-CVaR optimization. method Proximal alternating linearized minimization algorithm with nested fixed-point proximity.
result The model offers a guaranteed approximation of the ℓ0-constrained mean-CVaR model. Proposes a new method to assess Wrong-Way Risk in cross-currency swaps.
problem Addressing Wrong-Way Risk (WWR) in cross-currency swaps with stochastic correlation modeling.
method Proposes a stochastic correlation approach to model the dependency between exposure and counterparty credit risk, capturing tail dependence.
result The impact of stochastic correlation on calculated CVA is substantial, providing a promising method to model WWR.
GAttNHP predicts future events in temporal knowledge graphs by encoding long-range dependencies and handling mutual excitation.
problem Forecasting future events in temporal knowledge graphs due to long-range dependencies, mutual excitation, and heavy-tailed inter-arrival times.
method GAttNHP uses a self-attention encoder, semantic soft-grouping, and NCQ regression to address these issues.
result GAttNHP improves entity and time prediction on six benchmark TKG datasets compared to state-of-the-art baselines.
Improved regret bounds for linear bandits with heavy-tailed rewards.
problem Stochastic linear bandits with heavy-tailed rewards.
method Elimination-based algorithm guided by experimental design.
result Regret bound of \(\tilde{\mathcal{O}}(d^\frac{1+3ε}{2(1+ε)} T^\frac{1}{1+ε})\) for \(ε\in (0,1)\).
Paper proposes a new method to stabilize noisy gradient algorithms.
problem Stochastic-gradient Langevin algorithms can introduce bias when taming denominators depend on stochastic-gradient realizations.
method Proposes a structure-preserving framework for designing tamed denominators that avoid unnecessary taming and maintain the stabilizing effect of taming.
result The method avoids stationary bias and explains the stationary error split into bias and remaining error.
In this paper, we generalize Huber's criterion to multichannel sparse recovery problem of complex-valued measurements where the objective is to find good recovery of jointly sparse unknown signal vectors from the given multiple measurement vectors which are different linear combinations of the same known elementary vec…
New method models fat-tailed distributions with anisotropic tail-adaptive flows.
problem Gaussian-based variational inference fails to accurately capture tail decay in fat-tailed distributions.
method Improved theory on tails of flows, developed anisotropic tail-adaptive flows (ATAF).
result ATAF models tail-anisotropy, outperforming prior work on synthetic and real-world targets.
A neural network estimates sampling distributions for hard problems where classical methods fail.
problem Bootstrap failure in estimating sampling distributions for specific statistics.
method Neural network trained on simulated datasets using pinball loss.
result Neural network attains 95% nominal coverage and 97% improvement over classical methods on four bootstrap-failure problems.
Paper develops a model-based RL framework for portfolio optimization in financial markets.
problem Complex, non-Gaussian environment dynamics in financial markets.
method Heavy-tailed preserving normalizing flows for environment simulation; model-based reinforcement learning framework.
result Proposed method outperforms in various financial markets, especially during the pandemic.
New measures capture tail dependence and non-exchangeability in financial data.
problem Underestimation of tail dependence and inability to capture non-exchangeable tail dependence.
method Tail copulas and novel tail dependence measures (MTCM, ATCM) are proposed.
result Captures non-exchangeable tail dependence and provides analytical forms for various copulas.
The paper examines how heavy-tailed risks behave under Gaussian copula models.
problem Understanding tail risk probabilities with heavy-tailed marginal risks and Gaussian dependence.
method Modeling heavy-tailed risks using regular variation and analyzing tail probabilities under Gaussian copula.
result The rate of decay of tail set probabilities varies with the type of tail sets and Gaussian correlation matrix.
New tail dependence measures for stock indices.
problem Measuring tail dependence between financial variables.
method Introducing a new stochastic order and studying monotone tail dependence measures.
result Advantage of new tail dependence measures over classical ones.
Study tail behavior of sum of heavy-tailed risks with copulas.
problem Analyzing the tail behavior of sums of heavy-tailed risks with dependence modeled by copulas.
method Modeling dependence with copulas and analyzing tail asymptotics of sums of heavy-tailed risks.
result Obtained asymptotic expansions for Value-at-Risk of aggregate risk.
A new method for density estimation using mixture discrepancy and moments.
problem Generalizing histogram statistics to higher dimensions.
method Density estimation via mixture discrepancy and moments (DSP-mix and MSP).
result DSP-mix and MSP are computationally tractable and maintain accuracy with increased speed.
Paper provides tail bounds for stochastic mirror descent in heavy-tailed noise.
problem Optimizing convex and Lipschitz functions with heavy-tailed noise.
method Develops tail bounds for optimization error of Stochastic Mirror Descent.
result Tail bounds extend to heavier-tailed noise regimes without diameter constraints.
A simple log-transform fixes heavy-tailed data for generative models.
problem Standard generative models struggle with heavy-tailed data.
method Apply the soft-log transform to data before training and exponentiate samples after generation.
result Log-FM outperforms specialized baselines on multivariate benchmarks.
SS-GEN simulates rare events in heavy and light-tailed data.
problem Estimating probabilities of extreme events in multivariate data.
method Self-Similar Generative Estimation (SS-GEN) decomposes tail distribution into radial and angular components.
result SS-GEN generates representative extreme scenarios and estimates rare-event probabilities beyond observed data.
This work extends diffusion models to handle heavy-tailed targets, improving score estimation and sampling guarantees.
problem Score estimation and sampling guarantees for heavy-tailed targets in diffusion models.
method Kernel density estimation and minimax rates analysis for score estimation and sampling guarantees.
result Sharp minimax rates for score estimation and sampling guarantees for heavy-tailed targets, revealing qualitative differences between exponential and polynomial tails.
The literature of heavy tails (typically) starts with a random walk and finds mechanisms that lead to fat tails under aggregation. We follow the inverse route and show how starting with fat tails we get to thin-tails when deriving the probability distribution of the response to a random variable. We introduce a general…
Efficiently computes optimal policies for Entropic Risk Measures.
problem Optimizing risk-sensitive metrics in MDPs is computationally expensive.
method Uses Entropic Risk Measures and novel structural analysis for efficient computation.
result Achieves strong performance in various decision-making scenarios.
This paper improves tail dependence analysis by introducing a path-based approach.
problem The classical tail dependence coefficient fails to capture non-exchangeable features of tail dependence.
method The paper introduces a path-based maximal tail dependence approach to capture the most pronounced feature of dependence over all possible paths.
result The paper proves the existence and provides an explicit characterization of the path-based maximal TDC, improving analytical and computational tractability.
HTFM improves mode coverage and tail-statistic recovery for heavy-tailed data.
problem Tackles heavy-tailed data in various domains with rare events.
method Proposes a framework using clock-conditioned Gaussian sources and truncated logsignature features.
result Improves mode coverage, sample quality, and tail-statistic recovery over Gaussian flow matching and baselines.
Robust CG methods avoid data corruption and solve structured statistical estimation problems.
problem Data corruption and heavy-tailed data in structured statistical estimation.
method Robustification of Conditional Gradient (CG) type methods using Huber's corruption model and robust mean gradient estimation.
result Robust CG methods converge linearly with correct sample complexity, even for high-dimensional problems.
The paper explores tail diversification in financial markets using entropy and mutual information.
problem Tail diversification in financial time series.
method Statistical independence through differential entropy and mutual information, using moments as contrast functions.
result Tail covariance matrix is a key driver of tail diversification.
The paper uses EVT to improve tail risk measures under ambiguity sets.
problem Misspecification of tail risk measures leads to inflated risk estimates.
method Applies Extreme Value Theory to derive worst-case tail risk under ambiguity sets.
result Proposes a tail-calibrated ambiguity design that preserves nominal tail asymptotic scaling.
Study on U-statistics with heavy-tailed samples, providing tail bounds and LDP.
problem Deviation of U-statistics with heavy-tailed samples.
method Exponential tail bounds and Large Deviation Principle (LDP) for U-statistics.
result Obtained an exponential upper bound for U-statistics tail decay, showing two regions of decay.
TTF improves performance of normalizing flows for heavy-tailed distributions.
problem Improving performance of normalizing flows for heavy-tailed distributions.
method Uses a Gaussian base distribution and a final transformation layer to produce heavy tails.
result Experimental results show TTF outperforms current methods, especially in high-dimensional or heavy-tailed scenarios.