Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,932 papers · 148 categories

Trend · papers per month

105210314419 · Jun 202019922001200920172026
48 results for Feature concentration

Conformal prediction fails under severe feature turnover in COVID-19 supply chain tasks.

problem Dealing with distribution shift in conformal prediction models.
method Using COVID-19 as a natural experiment across 8 supply chain tasks, analyzing SHAP explanations.
result Coverage drops vary widely (0% to 86.7%) and correlate with single-feature dependence.

Random feature matrices' singular values concentrate near their full expectation in high dimensions.

problem Characterizing the spectra of random feature matrices for regression problems.
method Analyzing two settings of input variables (random or well-separated) with conditions on dimension, complexity ratio, and sampling variance.
result The singular values of random feature matrices concentrate near their full expectation and near one with high probability.

Paper addresses concentration of distances for fractional quasi p-norms, identifying conditions for concentration and anti-concentration.

problem Understanding concentration of distances for fractional quasi p-norms in high dimensions.
method Analyzes conditions for concentration and anti-concentration of distances for fractional quasi p-norms.
result Identifies conditions for concentration and anti-concentration of fractional quasi p-norms, ruling out some approaches and specifying conditions for control.

Bayesian neural networks explore rare fluctuations for better feature learning.

problem Understanding rare but dominant fluctuations in Bayesian neural networks.
method Large-deviation theory and joint optimization over predictors and internal kernels.
result Posterior rate function optimization reveals data-dependent kernel selection.

We find a deterministic equivalent for random feature regression's test error, independent of feature map dimension.

problem Understanding the generalization performance of random feature ridge regression.
method We derive a deterministic equivalent for the test error of RFRR under a concentration property, showing it can be approximated by a closed-form expression dependent on feature map eigenvalues.
result Our approximation guarantee is non-asymptotic, multiplicative, and independent of the feature map dimension, providing a tight result for the smallest number of features achieving optimal minimax error rate.

This work explores the characteristics of financial contagion in networks whose links distributions approaches a power law, using a model that defines banks balance sheets from information of network connectivity. By varying the parameters for the creation of the network, several interbank networks are built, in which …

2014-10-09abs ↗pdf ↗

A new algorithm detects out-of-distribution samples by concentrating them in feature space.

problem Building safe AI systems requires effective out-of-distribution detection.
method The paper proposes a novel algorithm based on the observation that OoD samples concentrate in feature space.
result The algorithm achieves state-of-the-art performance on various OoD detection benchmarks.

Free boundary minimal submanifolds with boundaries on concentric spheres

problem Finding minimal submanifolds with boundaries on concentric spheres in Euclidean space
method Using a Steklov problem with an indefinite weight
result Exact Morse index of an mm-dimensional flat annulus in an nn-dimensional spherical shell

A limaçon-like curve, allowing 2π-transition with monotone curvature between concentric curvature elements, is presented. The curve is 4th degree algebraic, 4th degree rational, and shares other common features with Pascal's limaçon.

2013-09-22abs ↗pdf ↗

Quantum kernel methods can lead to trivial models due to exponential concentration of kernel values.

problem Exponential concentration of quantum kernel values can lead to trivial models in QML.
method Analyzing the resources needed to accurately estimate quantum kernel values and identifying four sources of concentration.
result Quantum kernel values can be exponentially concentrated, leading to trivial models.

Random feature maps are ubiquitous in modern statistical machine learning, where they generalize random projections by means of powerful, yet often difficult to analyze nonlinear operators. In this paper, we leverage the "concentration" phenomenon induced by random matrix theory to perform a spectral analysis on the Gr…

2018-05-30abs ↗pdf ↗

New loss function improves adversarial robustness without sacrificing standard accuracy.

problem Adversarial robustness requires more samples than standard accuracy, and new data collection is costly.
method Proposed Max-Mahalanobis center (MMC) loss to concentrate feature points in the feature space.
result Empirical results show MMC loss significantly improves robustness under strong adaptive attacks.

Develops a fast variational approximation for high-dimensional empirical Bayes posteriors.

problem Optimal posterior computation in high-dimensional settings with prior tails effect.
method Variational approximation of empirical Bayes posterior with data-driven centers and thin-tailed conjugate priors.
result Retains optimal concentration rate properties and superior performance compared to existing methods.

New analysis proves sketching operators' RIP guarantees for mixture models without importance sampling.

problem Proving sketching operators' Restricted Isometry Property (RIP) for mixture models without assuming importance sampling.
method Proposed alternative analysis based on new deterministic bounds and concentration inequalities.
result Theoretical guarantees for sketching operators without importance sampling.

High-dimensional spectroscopy data makes ML models achieve near-perfect accuracy, even when chemical distinctions are absent.

problem Why machine learning models achieve near-perfect accuracy in spectroscopic classification tasks without chemically meaningful features.
method Theoretical analysis grounded in the Feldman-Hajek theorem and concentration of measure, combined with specific experiments on synthetic and real fluorescence spectra.
result Infinitesimal distributional differences in high-dimensional spaces can lead to perfect separability, making models achieve near-perfect accuracy in spectroscopy.

New method approximates complex kernel norms with random features, making learning tractable.

problem Complexity of learning with kernel methods in high dimensions.
method Random features approximations to Fp\mathcal{F}_p norms, focusing on p>1p>1.
result For p>1p>1, the number of random features required is polynomial in the sample size, making learning tractable.

The study examines how investor protection and past information affect stock returns and interest rates.

problem Empirical regularities related to investor protection and past information in asset pricing models.
method Developed a dynamic asset pricing model with a controlling shareholder and good/bad memory in budget dynamics.
result Good/bad memory of investors on historical market information affects stock returns and interest rates, strengthening investor protection in high ownership concentration.

Analyzes neural networks using linear models to understand their behavior.

problem Understanding multi-layer neural networks through linear models.
method Recalls and reviews four models: linear regression with concentrated features, kernel ridge regression, random feature model, and neural tangent model.
result Highlights limitations of linear theory and discusses approaches to overcome them.

TabPFN model shows strong robustness to noisy data.

problem TabPFN tackles robustness to noisy and imperfect tabular data.
method Empirical robustness analysis of TabPFN's attention mechanisms under various perturbations.
result TabPFN maintains high predictive performance and coherent internal behavior under noisy and imperfect data.

Devoted to multi-task learning and structured output learning, operator-valued kernels provide a flexible tool to build vector-valued functions in the context of Reproducing Kernel Hilbert Spaces. To scale up these methods, we extend the celebrated Random Fourier Feature methodology to get an approximation of operator-…

2016-05-09abs ↗pdf ↗

Random features and KRR generalize similarly when N is large enough.

problem Understanding the generalization error of random features and KRR methods.
method Analyzing spectral conditions and hypercontractivity on kernel eigenfunctions.
result The test error of random features is larger than KRR when N is small, but they achieve the same error when N is large.

In this paper we study the concentration properties for the eigenvalues of kernel matrices, which are central objects in a wide range of kernel methods and, more recently, in network analysis. We present a set of concentration inequalities tailored for each individual eigenvalue of the kernel matrix with respect to its…

2018-12-05abs ↗pdf ↗

We propose a new online algorithm for cumulative regret minimization in a stochastic linear bandit. The algorithm pulls the arm with the highest estimated reward in a linear model trained on its perturbed history. Therefore, we call it perturbed-history exploration in a linear bandit (LinPHE). The perturbed history is …

2019-03-21abs ↗pdf ↗

Interactive news recommendation has been launched and attracted much attention recently. In this scenario, user's behavior evolves from single click behavior to multiple behaviors including like, comment, share etc. However, most of the existing methods still use single click behavior as the unique criterion of judging…

2018-11-30abs ↗pdf ↗

A novel algorithm predicts customized allergy seasons using multi-variate triple-regression.

problem Predicting customized allergy seasons for individual patients.
method Triple-regression algorithm with pre-processing and three-stage regressions.
result Improved forecasting accuracy and reduced uncertainty.

Study Finsler metric measure manifolds' concentration properties.

problem Understanding concentration properties in Finsler metric measure manifolds.
method Established relationships with observable diameter, isoperimetric inequalities, and first eigenvalue.
result Derived a Cheng type upper bound estimate for the first closed eigenvalue.

MapLUR uses deep learning on map images to estimate NO2 pollution, outperforming traditional methods.

problem Limited availability of data for traditional LUR models makes them hard to adapt to new areas.
method Data-driven, open-source approach using convolutional neural networks trained on map data.
result MapLUR significantly outperforms traditional LUR models, including those with manually engineered features.

Paper develops a novel approach to identify clusters of features in multivariate extremes.

problem Understanding the complex structure of multivariate extremes in various fields.
method Optimization-based approach to assess the dependence structure of extremes.
result Estimating clusters of features that best capture the support of extremes.

We survey recent results related to the concentration of eigenfunctions. We also prove some new results concerning ball-concentration, as well as showing that eigenfunctions saturating lower bounds for L1L^1-norms must also, in a measure theoretical sense, have extreme concentration near a geodesic.

2015-10-26abs ↗pdf ↗