Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

3336659981,330 · Jun 202019922001200920172026
48 results for high data regime

Data pruning algorithms struggle in high compression regimes, as shown by theoretical and empirical studies.

problem Limitations of score-based data pruning algorithms in high compression regimes.
method Theoretical and empirical analysis of score-based data pruning algorithms.
result Score-based data pruning algorithms fail in high compression regimes due to 'No Free Lunch' theorems.

Analyzes SGD dynamics in two-layer networks, bridging different regimes.

problem Understanding SGD dynamics in high-dimensional and mean-field settings.
method Rigorous analysis via deterministic low-dimensional description of sufficient statistics.
result Infinite-width dynamics remains close to a low-dimensional subspace.

New bounds for high-dimensional sparse linear bandits, balancing information and regret.

problem Stochastic linear bandits with high-dimensional sparse features.
method Derivation of minimax regret lower and upper bounds for explore-then-commit algorithm.
result Optimal rate of Θ(n2/3)Θ(n^{2/3}) for data-poor regime, complemented by O(n)O(\sqrt{n}) under signal magnitude assumption.

The study identifies and analyzes different market regimes in equity markets using advanced signal processing techniques.

problem Understanding and quantifying the dynamics of different market regimes in equity markets.
method Data-driven Hilbert--Huang Transform for regime identification, Holo--Hilbert Spectral Analysis for profiling, and Variable-Length Markov Chains for return dynamics modeling.
result Developed markets normalize more effectively as stress subsides, while developing markets retain residual tail dependence and downside persistence.

New method detects and clusters market regimes in multidimensional data.

problem Detecting and clustering market regimes in complex data structures.
method Non-parametric online market regime detection and clustering using path-wise two-sample tests and maximum mean discrepancy.
result Successfully detected and clustered market regimes in various data structures.

New method identifies nonstationary causal structures in time series data.

problem Identifying causal relationships in time series data that change over time.
method High-order Markov Switching Models for regime-dependent causal discovery.
result Scalable approach for estimating high-order regime-dependent causal structures.

Volatility forecasting and return prediction in high-frequency Chinese equity markets.

problem Improving statistical forecasting performance and economic strategy outcomes in equity markets.
method Developing a sequential two-stage framework combining realized volatility modeling and XGBoost return prediction.
result Regime-aware volatility forecasting outperforms baseline models.

Paper analyzes Gibbs and Langevin Monte Carlo for interpolation regime, showing generalization from low errors.

problem Analyzing Gibbs and Langevin Monte Carlo in overparameterized interpolation regime.
method Data-dependent bounds and stability under approximation with Langevin Monte Carlo.
result Generalization is signaled by small training errors in noisy regime, with bounds stable under approximation.

Method predicts which high-dimensional correlation signs will change in the future.

problem Predicting which correlation matrix coefficients will change signs in high-dimensional data.
method Stability of correlation signs depends on three-by-three relationships, inspired by Heider social cohesion theory.
result The method accurately predicts the stability of correlation signs in high-dimensional data.

Local averaging accurately distills manifold structure from noisy data.

problem Tackles the challenge of uncovering manifold structure from noisy data.
method Two-round mini-batch local averaging method applied to noisy samples.
result Achieves accuracy bound of $d(\hat{\mathbf q}, \mathcal M) \leq σ\sqrt{d\left(1+\frac{κ\mathrm{diam}(\mathcal {M})}{\log(D)} ight)}$.

Proposes a new robust expectile regression method for high-dimensional data.

problem Heterogeneity in high-dimensional data with heteroscedastic variance or inhomogeneous covariate effects.
method Iteratively reweighted ℓ1-penalization for robust expectile regression (retire).
result Oracle convergence rate after log(log d) iterations in high-dimensional settings.

The paper analyzes PLS-SVD in high-dimensional data integration, revealing its strengths and limitations.

problem Understanding the behavior of PLS-SVD in high-dimensional data integration.
method Analysis using random matrix theory and singular value decomposition.
result PLS-SVD exhibits counter-intuitive or limiting behavior in certain regimes and outperforms PCA when detecting common latent subspace.

The paper finds non-Gaussian directions in high-dimensional data using Wasserstein distance.

problem Locating interesting non-Gaussian features in high-dimensional data.
method Projection pursuit using 2-Wasserstein distance to maximize the difference from Gaussian.
result Statistical guarantees for accurately approximating an unknown low-dimensional non-Gaussian subspace.

The paper analyzes Kernel Density Estimation in high dimensions with varying data and dimensionality.

problem High-dimensional Kernel Density Estimation with growing data and dimensionality.
method Examines the behavior of Kernel Density Estimators in the regime where both data points and dimensionality grow with a fixed ratio.
result Three distinct statistical regimes are identified for Kernel-based density estimates, each with different statistical properties.

Corrected whitening restores orthogonality in high-dimensional spherical Gaussian mixtures.

problem In high-dimensional data, standard whitening fails to preserve orthogonality of mixture means.
method Derived exact limits for whitened means dot products using random matrix theory, constructed a corrected whitening matrix.
result Corrected whitening allows for improved estimation of spherical Gaussian mixtures in the large-dimensional regime.

DIVI clusters noisy high-dimensional data with stable feature gating.

problem Challenging clustering in high-dimensional noisy data.
method Data-informed variational clustering framework combining global feature gating and adaptive structure growth.
result DIVI performs competitively under severe feature noise and remains computationally feasible.

sWk-means clusters multidimensional financial time series into distinct market regimes.

problem Classifying distinct market regimes in multidimensional financial time series.
method Approximated multidimensional Wasserstein distance as sliced Wasserstein distance for clustering.
result sWk-means successfully identifies distinct market regimes in real financial data.

Investigates how SGD behaves in high-dimensional neural networks, distinguishing between global convergence and local minima.

problem Understanding the behavior of SGD in high-dimensional shallow neural networks.
method Extends statistical physics analysis to study SGD dynamics, focusing on mean-field/hydrodynamic regime and learning rate.
result Identifies the critical number of hidden units and learning rate for SGD to avoid local minima.

A large number of online services provide automated recommendations to help users to navigate through a large collection of items. New items (products, videos, songs, advertisements) are suggested on the basis of the user's past history and --when available-- her demographic profile. Recommendations have to satisfy the…

2013-01-08abs ↗pdf ↗

Paper proposes a 1-bit quantization scheme for high-dimensional statistical estimation.

problem High-dimensional statistical estimation with limited data.
method Uniformly dithered 1-bit quantization for sparse covariance matrix estimation, sparse linear regression, and matrix completion.
result Near minimax rates in sub-Gaussian regime and improved rates in heavy-tailed regime.

ReCAP adapts to dynamic financial markets by segmenting and combining policy vectors.

problem Inefficient traditional PM approaches in non-stationary financial markets.
method Integrates continual learning into PM, segmenting regimes and adapting policies.
result Consistently outperforms baselines in real-world financial datasets.

Proposes a method to estimate personalized treatments from high-dimensional data.

problem Estimating individualized treatment regimes (ITRs) from high-dimensional covariates.
method Directly targets the contrast between potential outcomes, using dimension-reduced outcome-weighted learning.
result Achieves universal consistency, converging to the Bayes risk under mild conditions.

We study the problem of finding the best linear model that can minimize least-squares loss given a data-set. While this problem is trivial in the low dimensional regime, it becomes more interesting in high dimensions where the population minimizer is assumed to lie on a manifold such as sparse vectors. We propose proje…

2019-07-03abs ↗pdf ↗

LCD improves causal discovery in high-dimensional gene data.

problem Predicting causal effects in large-scale gene expression data.
method Local Causal Discovery (LCD) with practical estimators, ICP algorithm inspiration, preselection method, and statistical tests.
result LCD estimator closely matches ICP's accuracy but is simpler and faster.

Kernel methods and MLPs perform similarly to linear models in high dimensions.

problem Understanding the performance of kernel methods and MLPs in high-dimensional settings.
method Analysis of kernel methods and MLPs in a high-dimensional regime with proportional asymptotics.
result Linear models are optimal in high-dimensional settings when data is generated by kernel models with nonlinear relationships.

Study examines how two-layer networks learn features after one gradient step.

problem Understanding feature learning in neural networks after a single gradient descent step.
method Modeling the trained network as a spiked Random Features (sRF) model and leveraging Gaussian universality.
result Exact asymptotic description of the generalization error of the sRF in high-dimensional limit.

AJL framework detects dynamic patterns in high-dimensional time-varying models.

problem Complex time-varying associations and abrupt regime shifts in longitudinal processes.
method Hierarchical regularization framework integrating functional variable selection with structural changepoint detection.
result The refined estimator achieves the oracle property in ultra-high-dimensional settings.

Semi-supervised learning improves classification in high dimensions.

problem Combining labeled and unlabeled data for high-dimensional classification.
method Information theoretic and computational lower bounds analysis for feature selection.
result Semi-supervised learning is advantageous for classification in high dimensions.

Modular pipeline improves stock portfolio prediction robustness under regime changes.

problem Overfitting in deep learning models for non-stationary datasets.
method Modular machine learning pipeline with GBDT models and online learning techniques.
result GBDT models with dropout show high performance, robustness, and generalisability.

The paper analyzes SGD in high-dimensional networks, revealing new scaling limits.

problem Understanding SGD dynamics in high-dimensional networks.
method Analyzing the effective dynamics of SGD using recent work on the subject.
result A new correction term emerges at the critical scaling regime, changing the phase diagram.

GRIP2 improves deep learning feature selection robustness in correlated and noisy data.

problem Identifying predictive features in correlated and noisy data.
method Integrates first-layer feature activity over a two-dimensional regularization surface to control sparsity and geometry, using efficient block-stochastic sampling.
result Demonstrates improved robustness and power in high correlation and low signal-to-noise ratio regimes.

PCA++ improves robustness to background noise in contrastive learning.

problem Recovering shared signal subspaces from positive pairs in high-dimensional data with structured background noise.
method PCA++ uses hard uniformity-constrained contrastive learning to enforce identity covariance on projected features.
result PCA++ outperforms standard PCA and alignment-only PCA+ in simulations and real-world datasets.

This paper improves adversarial robustness of deep learning models.

problem Vulnerability of machine learning models to adversarial perturbations.
method Analyzes adversarial training for linear regression and neural networks, incorporating L1 penalty.
result Incorporating L1 penalty leads to consistent adversarially robust estimation in high-dimensional settings.

Model captures external influences through random parameters and regime switching.

problem Capturing external influences in asset dynamics with uncertainty and regime changes.
method Developed a stochastic model with random parameters and regime switching, mathematically consistent and interpretable.
result Demonstrated the model's versatility through local volatility models and characteristic functions.

New bounds for SGD in high dimensions improve inference efficiency.

problem Quantifying uncertainty in high-dimensional SGD.
method Established non-asymptotic Berry--Esseen bounds for online least-squares SGD.
result Gaussian Central Limit Theorem holds for td1+δt \gtrsim d^{1+δ}, extending dimensional scaling.

New method detects changes in high-dimensional data from small samples.

problem Detecting changes in high-dimensional data with limited samples.
method Angular kernel scan framework for detecting marginal distributional shifts.
result Exact population mean factorization and asymptotically distribution-free test.

Study on SGD dynamics and scaling laws for training quadratic neural networks in high dimensions.

problem Optimizing and understanding the training dynamics of quadratic neural networks in high-dimensional settings.
method Sharp analysis of SGD dynamics, combining matrix Riccati differential equations and matrix monotonicity arguments.
result Derivation of scaling laws for prediction risk, highlighting power-law dependencies on optimization time, sample size, and model width.