Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

25.0%50.0%75.0%100.0% · Feb 199419922001200920172026
48 results for Finite data

Estimates causal effects in Gaussian Linear SCMs with finite data.

problem Estimating causal effects from observational data with latent confounders.
method Centralized Gaussian Linear SCMs (CGL-SCMs) and EM-based estimation algorithm.
result Learned CGL-SCM parameters accurately recover causal distributions from finite observational samples.

The paper develops finite knot theory using ropelength-filtered Reidemeister graphs.

problem Understanding knot types in bounded ropelength sublevel spaces.
method Study thick representatives in bounded ropelength sublevel spaces through lifted Reidemeister graphs.
result Define characteristic Reidemeister patterns and finite recognition length.

New guarantees for uniquely identifying transport maps and vector fields from finite measure-valued data.

problem Unique recovery of transport maps and vector fields from finite measure-valued data.
method Use of Whitney and Takens embedding theorems to establish conditions for unique identification.
result New metric for comparing diffeomorphisms and analogous results in infinitesimal settings.

The paper studies how more data affects prediction risk in high-dimensional models.

problem The impact of increasing data on prediction risk in high-dimensional models.
method Derives central limit theorem and provides finite-sample distribution and confidence interval for prediction risk.
result Demonstrates 'more data hurt' phenomenon in high-dimensional least squares estimation.

New method reduces over-parametrization in neural networks, ensuring sparsity and finite network size.

problem Over-parametrization leads to too many active neurons in neural networks, especially with large data.
method Investigates a nonconvex regularization method for shallow ReLU networks.
result Locally optimal networks are finite even with infinite data, maintaining approximation guarantees and network size bounds.

We identify linear models from nonlinear systems with initialization constraints.

problem Identifying linear models from nonlinear systems with initialization constraints.
method Multiple trajectories-based deterministic data acquisition algorithm followed by regularized least squares.
result We provide a finite sample error bound on the learned linearized dynamics.

In this paper we prove that a complete, embedded minimal surface MM in R3\mathbb{R}^3 with finite topology and compact boundary (possibly empty) is conformally a compact Riemann surface M\overline{M} with boundary punctured in a finite number of interior points and that MM can be represented in terms of meromorphic …

2015-06-25abs ↗pdf ↗

Finite resources limit false discovery rate control in structured hypothesis spaces.

problem Controlling false discovery rate in hypothesis testing with finite data and structured hypothesis spaces.
method Framework for exact FDR control and adaptive power maximization.
result Exact FDR control and adaptive power maximization.

End-to-end algorithm for controlling bilinear systems with probabilistic noise.

problem Controlling bilinear systems with noisy data.
method Proposes an end-to-end algorithm using statistical learning theory and robust controller design.
result Derived finite sample identification error bounds and structurally suitable for control.

FMM fails to accurately determine the number of components even with consistent posterior.

problem Determining the number of subpopulations in a data set using FMM.
method Analysis of FMM component-count posterior under model misspecification.
result FMM component-count posterior diverges under model misspecification, contrary to intuition.

We learn linear models from nonlinear systems using multiple trajectories and regularization.

problem Identifying linear models from data when the underlying dynamics are nonlinear.
method Multiple trajectories data acquisition followed by regularized least squares.
result Learn linearized dynamics with arbitrarily small error given enough samples.

Graph Laplacians and machine learning predict properties of finite graphs.

problem Understanding properties of finite graphs using spectral and topological methods.
method Combining graph Laplacians, spectral inequalities, machine learning, and topological data analysis.
result Neural networks can accurately predict graph properties like Ricci-flatness and spectral gaps.

Framework detects and mitigates data-poisoning attacks in causal effect estimation.

problem Vulnerability to append-only attacks in observational causal analyses.
method Develops a data-poisoning audit for augmented inverse-probability-weighted estimation.
result Proposes a greedy scan to compute exact worst-case movement at every append budget.

The paper sets sample complexity bounds for identifying LTI systems from a finite set.

problem Identifying an LTI system from a finite set of possible systems using trajectory data.
method Maximum likelihood estimator and information theory tools.
result Upper and lower bounds for sample complexity are derived, independent of stability assumption.

Paper provides unbiased spectral moment estimates from finite data.

problem Challenges in estimating spectral moments from limited data.
method Dynamic programming approach to estimate spectral moments of kernel integral operator.
result Demonstrates consistency with theoretical spectra and practical utility in neural networks.

We quantify Peter Scott's Theorem that surface groups are locally extended residually finite (LERF) in terms of geometric data. In the process, we will quantify another result by Scott that any closed geodesic in a surface lifts to an embedded loop in a finite cover.

2012-04-23abs ↗pdf ↗

Investigates the impact of finite VC dimension on neural network approximation and learning.

problem The influence of VC dimension on neural network approximation and learning from samples.
method Analysis of high-dimensional geometry and statistical learning theory, focusing on VC dimension.
result Finite VC dimension is beneficial for uniform convergence of empirical errors but not for approximation of functions from a probability distribution.

A neural network model predicts the critical point of the Ising phase transition.

problem Predicting the critical point of the Ising phase transition using supervised learning.
method Proposed a minimal one-free-parameter neural network model to describe the supervised learning problem for the Ising model.
result Just one free parameter is enough to describe the universal finite-size-scaling function in the network output.

The paper analyzes how the one-dimensional Wasserstein distance captures pointwise density differences in finite samples.

problem Uncertainty in identifying density differences when supports overlap and densities have substantial pointwise differences.
method Analysis using the Poisson process and neural spike train decoding.
result The one-dimensional Wasserstein distance highlights meaningful density differences related to both rate and support.

Study on Dirichlet process mixtures for clustering consistency.

problem Consistency of clustering with Dirichlet process mixtures.
method Analysis of posterior distribution as sample size increases, focusing on consistency for the number of clusters.
result Consistency for the number of clusters can be achieved with a properly adapted concentration parameter in a Bayesian setting.

We study minimal annuli in S2×R\mathbb{S}^2 \times \mathbb{R} of finite type by relating them to harmonic maps CS2\mathbb{C} \to \mathbb{S}^2 of finite type. We rephrase an iteration by Pinkall-Sterling in terms of polynomial Killing fields. We discuss spectral curves, spectral data and the geometry of the isospectral set…

2012-10-20abs ↗pdf ↗

Paper introduces FNM framework for learning finite-dimensional parametrized models.

problem Efficiently learning finite-dimensional parametrized models from limited data.
method Fourier Neural Mappings (FNMs) framework for operator learning.
result End-to-end learning of PtO maps can be less data-efficient than learning the solution operator first.

EbC learns equivariant embeddings from unlabeled group actions.

problem Learning equivariant embeddings from unlabeled group actions.
method Equivariance by Contrast (EbC) method to learn equivariant embeddings from observation pairs (y,gy)(\mathbf{y}, g \cdot \mathbf{y}).
result High-fidelity equivariance in latent space for diverse groups.

Privacy concerns have led to the development of privacy-preserving approaches for learning models from sensitive data. Yet, in practice, even models learned with privacy guarantees can inadvertently memorize unique training examples or leak sensitive features. To identify such privacy violations, existing model auditin…

2019-11-08abs ↗pdf ↗

We study the sample complexity of private synthetic data generation over an unbounded sized class of statistical queries, and show that any class that is privately proper PAC learnable admits a private synthetic data generator (perhaps non-efficient). Previous work on synthetic data generators focused on the case that …

2019-02-09abs ↗pdf ↗

We demonstrate how a 3-manifold, a Heegaard diagram, and a group presentation can each be interpreted as a pair of signed permutations in the symmetric group Sd.S_d. We demonstrate the power of permutation data in programming and discuss an algorithm we have developed that takes the permutation data as input and determi…

2011-08-19abs ↗pdf ↗