Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

76152228304 · Jun 202019922001200920172026
48 results for statistical fitting

Proposes φφ-table for statistical SHAP explanations in regression models.

problem Lack of clear directional summaries, uncertainty, and fidelity in SHAP feature importance.
method SHAP importance selection, fitting a standardized linear surrogate, reporting coefficients, uncertainty, fidelity, and stability.
result Extends SHAP into a statistical global explanation with direction, uncertainty, fidelity, and stability.

We propose a nonparametric statistical test for goodness-of-fit: given a set of samples, the test determines how likely it is that these were generated from a target density function. The measure of goodness-of-fit is a divergence constructed via Stein's method using functions from a Reproducing Kernel Hilbert Space. O…

2016-02-09abs ↗pdf ↗

Study evaluates different mathematical models for three case studies using statistical fitting.

problem Estimating outcomes in population dynamics, temperature variations, and market equilibrium.
method Applied various statistical equations (e.g., fractional exponential, sinusoidal) to three case studies.
result Optimal models differ by case study (fractional exponential for population dynamics, sinusoidal for temperature and market equilibrium).

Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to identify absence of fit through finding a classification rule that can efficientl…

2007-12-02abs ↗pdf ↗

Two methods monitor high-dimensional processes via manifold fitting or learning.

problem Monitoring high-dimensional, dynamic industrial processes.
method Manifold fitting and learning approaches for online SPC.
result Manifold-fitting approach achieves performance competitive with classical methods.

A new method improves fitting neural data with spiking network models.

problem Fitting spiking network models to neural activity does not produce realistic data.
method Augment log-likelihood with dissimilarity terms measured by summary statistics and optimized via back-propagation.
result The new method generates more realistic neural activity statistics and improves network connectivity inference.

Develops a goodness-of-fit test for self-exciting processes.

problem Quantifying how well generative models capture self-exciting point processes.
method Connects to Quasi-maximum-likelihood estimator (QMLE) theory and develops a non-parametric self-normalizing statistic, the Generalized Score (GS) statistics.
result Validates the proposed GS test's good performance through numerical simulation and real-data experiments.

These are the written discussions of the paper "Bayesian measures of model complexity and fit" by D. Spiegelhalter et al. (2002), following the discussions given at the Annual Meeting of the Royal Statistical Society in Newcastle-upon-Tyne on September 3rd, 2013.

2013-10-10abs ↗pdf ↗

T-Rex uses EM to fit robust factor models in noisy data.

problem Robustly fitting factor models in high-dimensional data with heavy tails and outliers.
method Expectation-Maximization (EM) algorithm based on Tyler's M-estimator for elliptical distributions.
result Demonstrates robustness in direction-of-arrival estimation and subspace recovery.

The paper fits a seven-parameter GTS distribution to financial data.

problem Nonexistence of GTS probability density function makes MLE inadequate.
method Used fractional Fourier transform to circumvent MLE and provide good parameter estimation.
result The GTS distribution fits financial data significantly better than other models.

Latent block models are used for probabilistic biclustering, which is shown to be an effective method for analyzing various relational data sets. However, there has been no statistical test method for determining the row and column cluster numbers of latent block models. Recent studies have constructed statistical-test…

2019-06-10abs ↗pdf ↗

The paper advances U-statistics in dependent settings, improving spectral estimation and goodness-of-fit tests.

problem Non-asymptotic analysis of U-statistics in dependent Markov chain settings.
method Proved new concentration and exponential inequalities for U-statistics, applied to spectral estimation, online algorithms, and goodness-of-fit tests.
result Established new results for spectral estimation, online algorithms, and goodness-of-fit tests in Markov chain settings.

A new distributed algorithm for fitting sparse additive models with feature division and decorrelation.

problem Fitting high-dimensional sparse additive models efficiently and accurately.
method Divide, decorrelate, and conquer approach.
result Effective and efficient recovery of sparsity patterns and statistical inference for each component.

We propose two nonparametric statistical tests of goodness of fit for conditional distributions: given a conditional probability density function p(yx)p(y|x) and a joint sample, decide whether the sample is drawn from p(yx)rx(x)p(y|x)r_x(x) for some density rxr_x. Our tests, formulated with a Stein operator, can be applied to any…

2020-02-24abs ↗pdf ↗

The paper proposes a test to determine the number of latent classes in ordinal categorical data.

problem Determining the correct number of latent classes in latent class models with ordinal categorical data.
method The test statistic centers the largest singular value of a normalized residual matrix by a simple sample-size adjustment.
result The test statistic converges to zero under the null hypothesis and exceeds a fixed positive constant under an under-fitted alternative.

This work improves fair tensor decomposition using a kernel criterion.

problem Learning fair low-rank tensor decompositions with statistical parity.
method Regularizes Canonical Polyadic Decomposition with KHSIC to ensure approximate statistical parity.
result The proposed algorithm achieves better fairness and fit than state-of-the-art FATR.

This article describes a multivariate polynomial regression method where the uncertainty of the input parameters are approximated with Gaussian distributions, derived from the central limit theorem for large weighted sums, directly from the training sample. The estimated uncertainties can be propagated into the optimal…

2013-10-03abs ↗pdf ↗

We show that univariate and symmetric multivariate Hawkes processes are only weakly causal: the true log-likelihoods of real and reversed event time vectors are almost equal, thus parameter estimation via maximum likelihood only weakly depends on the direction of the arrow of time. In ideal (synthetic) conditions, test…

2017-09-25abs ↗pdf ↗

Improved modeling of persistence diagrams for data analysis.

problem Determining significant outliers in persistence diagrams.
method Modification of the RST (Replicating Statistical Topology) model using MCMC Metropolis-Hastings algorithm.
result The modified RST model improves the goodness of fit in persistence diagram analysis.

The two key issues of modern Bayesian statistics are: (i) establishing principled approach for distilling statistical prior that is consistent with the given data from an initial believable scientific prior; and (ii) development of a Bayes-frequentist consolidated data analysis workflow that is more effective than eith…

2018-02-01abs ↗pdf ↗

The paper calculates ruin probabilities for insurers with phase-type distributed claims.

problem Calculating ruin probabilities for insurers with specific claim distributions.
method Change-of-measure technique applied to phase-type distributed claim amounts.
result The mixture of Erlangs best fits real-world loss data, improving risk assessment.

Bidirectional attention is shown to be equivalent to a continuous bag of words model with mixture-of-experts.

problem Understanding the statistical underpinnings of bidirectional attention.
method Exploring bidirectional attention as a mixture-of-experts model and reparameterizing it.
result Bidirectional attention can be viewed as a continuous bag of words model with mixture-of-experts weights.

PFNs pre-train models on simulated data to predict class probabilities.

problem Training machine learning models on large datasets.
method Pre-train a fixed model on small simulated datasets and use it to infer class probabilities in-context.
result PFNs achieve state-of-the-art performance and improve with larger inference data.

Given two candidate models, and a set of target observations, we address the problem of measuring the relative goodness of fit of the two models. We propose two new statistical tests which are nonparametric, computationally efficient (runtime complexity is linear in the sample size), and interpretable. As a unique adva…

2018-10-27abs ↗pdf ↗

The κκ-generalised distribution fits daily stock returns well.

problem Stock returns are often heavy-tailed, not normally distributed.
method Used the κκ-generalised distribution with a Monte-Carlo goodness of fit test.
result The κκ-generalised distribution fits historic daily stock returns well for a significant proportion of analyzed stocks.

Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.

problem Statistical inference challenges with missing at random labels and spatial dependence.
method Doubly robust estimator with cross-fit nuisances and jackknife spatial HAC variance correction.
result Asymptotically valid confidence intervals with improved finite-sample calibration.

CPCR mitigates bias in PCR for overparameterized models.

problem Bias in Principal Component Regression (PCR) for overparameterized models.
method Calibrated Principal Component Regression (CPCR) learns a low-variance prior in the PC subspace and calibrates the model in the original feature space.
result CPCR outperforms standard PCR in overparameterized settings, improving prediction across multiple problems.

The balance property is crucial for insurance pricing, ensuring total actuarial price equals loss. Maximum likelihood GLMs fulfill it, but Lindholm-Wüthrich suggests three methods, with constrained GLM being superior.

problem Ensuring the balance property in insurance pricing models
method Using constrained GLM fitting
result Constrained GLM fitting is superior to the two previously discussed balance correction methods