Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

52105157209 · Jun 202019922001200920172026
48 results for moderate dimension

New lower bounds show challenges in clustering in moderate dimensions.

problem Clustering points from mixtures of isotropic Gaussians in moderate dimensions.
method Established low-degree polynomial lower bounds and developed a novel non-spectral algorithm.
result New lower bounds reveal a 'non-parametric rate' in moderate dimensions.

A new graph-based clustering method for moderate-dimensional data.

problem Performance degradation of existing graph-based clustering methods in high dimensions.
method Introduces UN-CCDs using NND-based MC-SRT for covering radii determination.
result UN-CCDs provide stable and competitive performance in moderate-sized datasets.

Algorithm identifies fractal system's scaling exponents in high dimensions.

problem Statistical identification of Hurst distribution in high-dimensional fractal systems.
method Wavelet random matrices, modified spectral clustering, model selection.
result Algorithm consistently estimates Hurst distribution in moderately high dimensions.

Proposes a method for interpreting time-varying causal effect moderation in high-dimensional data.

problem Interpreting causal effect moderation in high-dimensional data with interpretability and avoiding false positives.
method Two-step method: 1) Selects a smaller model for linear causal effect moderation using Gaussian randomization, 2) Conditions on selection to construct a pivot for uniformly asymptotic semi-parametric inference.
result Consistently achieves valid coverage rates and shorter, bounded intervals in time-varying causal effect moderation.

For spherically symmetric distributions, efficient quantisation can be achieved with moderate sample sizes.

problem Optimal quantisation in high dimensions requires large sample sizes, making it impractical.
method Uniformly distributed random quantisers on a sphere of suitable radius achieve exceptional performance.
result For moderate sample sizes, quantisation error can be efficiently computed and approximated.

Study on benign overfitting in leaky ReLUs with moderate input dimensions.

problem Understanding when overfitting is beneficial in neural networks.
method Two-layer leaky ReLU networks trained with hinge loss, considering signal-to-noise ratio.
result Characterization of conditions for benign overfitting based on signal-to-noise ratio.

PQMass assesses generative model quality using chi-squared tests.

problem Assessing the quality of generative models without density assumptions.
method Divides sample space into regions, applies chi-squared tests to p-values.
result Effectively assesses generative model quality, novelty, and diversity.

Study shows limits of certain normalizing flows in higher dimensions.

problem Understanding the representation power of normalizing flows in different dimensions.
method Rigorously established bounds on expressive power of basic normalizing flows.
result Limited representation power in higher dimensions, especially with moderate depth.

Understanding how features interact with each other is of paramount importance in many scientific discoveries and contemporary applications. Yet interaction identification becomes challenging even for a moderate number of covariates. In this paper, we suggest an efficient and flexible procedure, called the interaction …

2016-05-28abs ↗pdf ↗

Importance sampling has become an important tool for the computation of tail-based risk measures. Since such quantities are often determined mainly by rare events standard Monte Carlo can be inefficient and importance sampling provides a way to speed up computations. This paper considers moderate deviations for the wei…

2013-06-27abs ↗pdf ↗

We develop an approach for feature elimination in statistical learning with kernel machines, based on recursive elimination of features.We present theoretical properties of this method and show that it is uniformly consistent in finding the correct feature space under certain generalized assumptions.We present four cas…

2013-04-18abs ↗pdf ↗

Paper analyzes ensemble Kalman updates for effective dimension and localization.

problem Why small ensemble sizes work well in inverse problems and data assimilation.
method Non-asymptotic analysis of ensemble Kalman updates, focusing on effective dimension and localization.
result Rigorously explains why a small ensemble size is sufficient when prior covariance has moderate effective dimension.

Computes invariants distinguishing between immersions and embeddings of doodles and blobs on surfaces.

problem Distinguishing between immersions and embeddings of doodles and blobs on surfaces.
method Regular embeddings, bordisms, and exact sequences of abelian groups.
result Exact sequence describing bordisms of immersions and embeddings of doodles on A=RimesIA = \mathbb R imes I.

New method interprets deep learning for causal effects, separating prognostic and moderating covariates.

problem Estimating individual causal/treatment effects under confounders.
method Deep counterfactual learning architecture for estimating CATE with interpretable score functions.
result Demonstrated improved interpretability and quantification of uncertainty in CATE estimation.

We provide a unifying treatment of pathwise moderate deviations for models commonly used in financial applications, and for related integrated functionals. Suitable scaling allows us to transfer these results into small-time, large-time and tail asymptotics for diffusions, as well as for option prices and realised vari…

2018-03-12abs ↗pdf ↗

Bayesian optimization techniques have been successfully applied to robotics, planning, sensor placement, recommendation, advertising, intelligent user interfaces and automatic algorithm configuration. Despite these successes, the approach is restricted to problems of moderate dimension, and several workshops on Bayesia…

2013-01-09abs ↗pdf ↗

Stochastic Gradient Descent shows directional bias with moderate learning rates, impacting optimization outcomes.

problem Understanding the bias of SGD with moderate learning rates in practical scenarios.
method Analyzing SGD and GD on an overparameterized linear regression problem.
result SGD converges along large eigenvalue directions, GD along small ones, affecting early stopping outcomes.

We consider call option prices in diffusion models close to expiry, in an asymptotic regime ("moderately out of the money") that interpolates between the well-studied cases of at-the-money options and out-of-the-money fixed-strike options. First and higher order small-time moderate deviation estimates of call prices an…

2016-04-05abs ↗pdf ↗

Multidimensional scaling is an important dimension reduction tool in statistics and machine learning. Yet few theoretical results characterizing its statistical performance exist, not to mention any in high dimensions. By considering a unified framework that includes low, moderate and high dimensions, we study multidim…

2018-10-24abs ↗pdf ↗

Random feature matrices' singular values concentrate near their full expectation in high dimensions.

problem Characterizing the spectra of random feature matrices for regression problems.
method Analyzing two settings of input variables (random or well-separated) with conditions on dimension, complexity ratio, and sampling variance.
result The singular values of random feature matrices concentrate near their full expectation and near one with high probability.

Inspired by the Bruhat-Tits building of SLn_n(Qp\mathbb Q_p), we construct a complete metric space X with an action of the tame automorphism group of the affine space Tame(KnK^n). The points in X are certain monomial valuations, and X admits a natural structure of Euclidean CW-complex of dimension n-1. When n = 3, and…

2018-02-01abs ↗pdf ↗

The paper analyzes the dynamics of tokens in transformer models at moderate interaction levels.

problem Understanding the evolution of tokens in transformer models at moderate interaction levels.
method Modeling transformer models as a system of particles interacting in a mean-field way and studying the corresponding dynamics.
result Characterization and convergence of the limiting dynamics in different phases of the system.

Method estimates multivariate counterfactual distributions efficiently and accurately.

problem Estimating multivariate counterfactual distributions in causal models with correlation structures.
method Proposes a method leveraging a one-dimensional subspace to capture correlation structures and efficiently estimate multivariate counterfactual distributions.
result Demonstrates superior performance over existing methods on synthetic and real-world data.

Paper optimizes change-point detection using learned distributions from training sequences.

problem Optimal change-point detection with unknown pre- and post-change distributions.
method Designs a change-point estimator using training sequences and test sequences.
result Optimal confidence width characterized as a function of undetected error.

A method for high-dimensional Bayesian optimization reduces dimensionality using EDR and Gaussian process.

problem Extending Bayesian optimization to high-dimensional settings.
method Two-step framework: EDR subspace identification followed by Gaussian process optimization.
result Algorithm converges in high-dimensional contexts, validated by numerical experiments.

Efficient EP algorithm improves smoothing distribution inference in financial models.

problem Computational intractability of smoothing distribution in high dimensions.
method Adapted expectation propagation (EP) algorithms for the unified skew-normal family.
result Accuracy gains in financial illustrations over existing approximate algorithms.

SMAVE optimizes SDR by projecting onto a low-dimensional subspace on a Riemannian manifold.

problem High-dimensional regression challenges due to the curse of dimensionality.
method SMAVE combines nearest-neighbor localization and Riemannian stochastic gradient ascent.
result SMAVE achieves almost-sure convergence and matches RMAVE's synthetic subspace recovery rate.

This work improves tensor decomposition methods, especially for large datasets.

problem Lack of efficient methods for estimating Tucker decompositions.
method Applies Johnson-Lindenstrauss type guarantees to Tucker decompositions with random embeddings.
result Effective dimension reduction with minimal error for large tensors.

In this paper, we propose the uncertain volatility models with stochastic bounds. Like the regular uncertain volatility models, we know only that the true model lies in a family of progressively measurable and bounded processes, but instead of using two deterministic bounds, the uncertain volatility fluctuates between …

2017-02-16abs ↗pdf ↗

Practitioners of Bayesian statistics have long depended on Markov chain Monte Carlo (MCMC) to obtain samples from intractable posterior distributions. Unfortunately, MCMC algorithms are typically serial, and do not scale to the large datasets typical of modern machine learning. The recently proposed consensus Monte Car…

2015-06-09abs ↗pdf ↗

In this article, we show that solving the system of linear equations by manipulating the kernel and the range space is equivalent to solving the problem of least squares error approximation. This establishes the ground for a gradient-free learning search when the system can be expressed in the form of a linear matrix e…

2018-10-27abs ↗pdf ↗