Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

4.3%8.6%12.9%17.3% · May 202619922001200920172026
48 results for Extreme Scale

We consider strictly stationary heavy tailed time series whose finite-dimensional exponent measures are concentrated on axes, and hence their extremal properties cannot be tackled using classical multivariate regular variation that is suitable for time series with extremal dependence. We recover relevant information ab…

2013-07-05abs ↗pdf ↗

HAXMLNet tackles extreme multi-label text classification with hierarchical attention.

problem Tagging each text with relevant labels from an extreme-scale label set.
method Proposes a hierarchical structure with multi-label attention for efficient and effective XMTC.
result HAXMLNet achieves competitive performance compared to state-of-the-art methods.

New method simulates multivariate extreme events using GANs and Aitchison coordinates.

problem Simulating multivariate extreme events for economic risk assessment.
method Wasserstein-Aitchison GAN approach combining tail dependence and marginal tail modeling.
result Strong performance in capturing tail dependence and generating accurate extreme observations.

Sharp inequalities and extremizers for J functional on Kähler manifolds.

problem Understanding the large scale asymptotic of the J functional on Kähler metrics.
method Proving sharpness of inequalities and studying extremizing potentials/rays on toric Kähler manifolds, and existence of radial extremizers on general Kähler manifolds.
result Sharpness of inequalities and existence of extremizing potentials/rays on toric Kähler manifolds, and equivalence with plurisupported currents on general Kähler manifolds.

Study shows one-dimensional location-scale-shape models are flat in Wasserstein geometry.

problem Investigating curvature in location-scale-shape models under Wasserstein metric.
method Introduced location-scale-shape model and investigated its geometry.
result Location-scale-shape model is intrinsically flat but extrinsically curved in Wasserstein geometry.

Recently, large-scale cascading failures in complex systems have garnered substantial attention. Such extreme events have been treated as an integral part of the self-organized criticality (SOC). Recent empirical work has suggested that some extreme events systematically deviate from the SOC paradigm, requiring a diffe…

2015-02-24abs ↗pdf ↗

New neural network models extreme value distributions with preserved shape constraints.

problem Modeling multivariate extreme value distributions with preserved shape constraints.
method d-max-decreasing neural network architecture for non-parametric calibration and generation of MEVs.
result The proposed architecture approximates the dependence structure of MEVs at parametric rate and preserves essential shape constraints.

Deep learning models complex multivariate extremes using geometric shapes.

problem Modeling complex extremal dependencies in high-dimensional data.
method Geometric representation and deep learning for flexible semi-parametric models.
result First approach to modeling limit sets using deep learning for high-dimensional data.

Anomaly-aware forecast improves accuracy for extreme events.

problem Challenges in automatically detecting and learning from extreme events and anomalies in large-scale datasets.
method Proposes an anomaly-aware forecast framework that automatically detects and incorporates anomalies using an attention mechanism and dynamic uncertainty optimization.
result Demonstrated superior accuracy and reduced uncertainty on three datasets with different types of anomalies.

Changing initialization scale affects deep model generalization, leading to memorization or improved performance.

problem Understanding how initialization scale impacts deep model generalization and memorization.
method Experimental setup with varying initialization scales, analysis of activation and loss functions, and development of an alignment measure.
result Increasing initialization scale leads to memorization, and decreasing it improves generalization, depending on activation and loss functions.

The study identifies extremal dependence in financial markets using a bootstrap-based testing procedure.

problem Accurately identifying extremal dependence in multivariate heavy-tailed financial data.
method Bootstrap-based testing procedure applied to U.S. and Chinese stock returns.
result The U.S. exhibits more isolated clustering of dependent assets compared to China.

Paper proposes efficient GCN learning method for limited data.

problem Learning GCNs from data with extremely limited annotations.
method Adaptive sampling strategy and model compression.
result Cut down annotation requirement by 90% and compress parameters 6x.

Study on price fluctuations and persistence in European electricity spot markets.

problem Analyzing variability and persistence of electricity prices in European spot markets.
method Analysis of hourly, intraday, and 15-min intraday market prices; quantification of fluctuations, correlations, and extreme events; classification into circulation weather types.
result Different time scales in market dynamics; multifractal behavior below 12 hours; anti-correlation and mean reversion above 12 hours; long-term behavior influenced by four-day weather patterns; qq-Gaussian distributions as best fit.

Extends geometric approach to model non-stationary extremal dependence.

problem Capturing evolving extremal dependence in multivariate data.
method Geometric framework for non-stationary multivariate extreme value modelling.
result Framework can capture various dependence forms and is robust to different model formulations.

This paper presents a novel scaling method for unbiased risk estimation.

problem Challenges in risk assessment due to limited data, non-stationarity, and heavy tails.
method Develops a statistical framework for efficient risk scaling, extending beyond the square-root-of-time rule.
result Ensures robust and conservative risk estimation, applicable to small sample settings.

The paper uses EVT to improve tail risk measures under ambiguity sets.

problem Misspecification of tail risk measures leads to inflated risk estimates.
method Applies Extreme Value Theory to derive worst-case tail risk under ambiguity sets.
result Proposes a tail-calibrated ambiguity design that preserves nominal tail asymptotic scaling.

Many modern clustering methods scale well to a large number of data items, N, but not to a large number of clusters, K. This paper introduces PERCH, a new non-greedy algorithm for online hierarchical clustering that scales to both massive N and K--a problem setting we term extreme clustering. Our algorithm efficiently …

2017-04-06abs ↗pdf ↗

We define regularity scales to study the behavior of the Calabi flow. Based on estimates of the regularity scales, we obtain convergence theorems of the Calabi flow on extremal Kahler surfaces, under the assumption of global existence of the Calabi flow solutions. Our results partially confirm Donaldson's conjectural p…

2015-01-08abs ↗pdf ↗

Extreme classification problems are multiclass and multilabel classification problems where the number of outputs is so large that straightforward strategies are neither statistically nor computationally viable. One strategy for dealing with the computational burden is via a tree decomposition of the output space. Whil…

2015-11-10abs ↗pdf ↗

DEFRAG accelerates extreme classification by reducing feature dimensions.

problem High precision and scalability in assigning labels from a vast label space.
method Adaptive feature agglomeration to reduce feature dimensions.
result Significant reduction in training and prediction times (up to 40%) for extreme classification algorithms.

Novel algorithm speeds up log-determinant estimation for large matrices.

problem Efficiently estimating log-determinants of large positive definite matrices under memory constraints.
method Hierarchical algorithm based on block-wise computation of LDL decomposition.
result Accurate estimation of NTK log-determinants from a tiny fraction of the full dataset.

In this paper, we first demonstrate that b-bit minwise hashing, whose estimators are positive definite kernels, can be naturally integrated with learning algorithms such as SVM and logistic regression. We adopt a simple scheme to transform the nonlinear (resemblance) kernel into linear (inner product) kernel; and hence…

2011-06-06abs ↗pdf ↗

Classification of G2-structures on Lie groups with Ricci pinched conditions.

problem Classifying G2-structures on Lie groups under specific geometric conditions.
method Complete classification of left-invariant closed G2-structures on Lie groups, extremally Ricci pinched, up to equivalence and scaling.
result Five distinct G2-structures on five different completely solvable Lie groups, with one unimodular case being exact.

New method uses neural networks to predict extreme wildfires, improving accuracy over traditional models.

problem Predicting extreme wildfires using complex, non-linear relationships.
method Partially-interpretable neural networks for extreme quantile regression.
result Significant improvement in predictive performance over traditional methods.

Accurate forecasting of risk is the key to successful risk management techniques. Using the largest stock index futures from twelve European bourses, this paper presents VaR measures based on their unconditional and conditional distributions for single and multi-period settings. These measures underpinned by extreme va…

2011-03-29abs ↗pdf ↗

The goal in extreme multi-label classification is to learn a classifier which can assign a small subset of relevant labels to an instance from an extremely large set of target labels. Datasets in extreme classification exhibit a long tail of labels which have small number of positive training instances. In this work, w…

2018-03-05abs ↗pdf ↗

The objective in extreme multi-label learning is to train a classifier that can automatically tag a novel data point with the most relevant subset of labels from an extremely large label set. Embedding based approaches make training and prediction tractable by assuming that the training label matrix is low-rank and hen…

2015-07-09abs ↗pdf ↗

This research shows loss weighting remains effective in last layer retraining despite model overparameterization.

problem Overcoming biases in machine learning models at scale.
method Theoretical and practical exploration of last layer retraining in an overparameterized setting.
result Loss weighting is still effective in last layer retraining, but weights must account for model overparameterization.

Efficient neural Bayes estimators for censored peaks-over-threshold models improve inference speed and accuracy.

problem Computational burden in inference with spatial extremal dependence models due to intractable or censored likelihoods.
method Developed neural Bayes estimators using data augmentation techniques to encode censoring information.
result Significant gains in computational and statistical efficiency compared to traditional methods.

Labeled Latent Dirichlet Allocation (LLDA) is an extension of the standard unsupervised Latent Dirichlet Allocation (LDA) algorithm, to address multi-label learning tasks. Previous work has shown it to perform in par with other state-of-the-art multi-label methods. Nonetheless, with increasing label sets sizes LLDA enc…

2017-09-16abs ↗pdf ↗

New loss functions improve extreme classification with missing labels.

problem Large number of infrequent labels and missing labels in XMC.
method Derive unbiased loss functions for XMC, incorporating them into existing algorithms.
result Significant improvement in extreme classification performance (up to 20%) over existing methods.

Accurately predicting customer churn using large scale time-series data is a common problem facing many business domains. The creation of model features across various time windows for training and testing can be particularly challenging due to temporal issues common to time-series data. In this paper, we will explore …

2018-02-09abs ↗pdf ↗