Proposes -table for statistical SHAP explanations in regression models.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
We propose a nonparametric statistical test for goodness-of-fit: given a set of samples, the test determines how likely it is that these were generated from a target density function. The measure of goodness-of-fit is a divergence constructed via Stein's method using functions from a Reproducing Kernel Hilbert Space. O…
Study evaluates different mathematical models for three case studies using statistical fitting.
Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to identify absence of fit through finding a classification rule that can efficientl…
Many fits of Hawkes processes to financial data look rather good but most of them are not statistically significant. This raises the question of what part of market dynamics this model is able to account for exactly. We document the accuracy of such processes as one varies the time interval of calibration and compare t…
Two methods monitor high-dimensional processes via manifold fitting or learning.
A new method improves fitting neural data with spiking network models.
Develops a goodness-of-fit test for self-exciting processes.
Complex phenomena in engineering and the sciences are often modeled with computationally intensive feed-forward simulations for which a tractable analytic likelihood does not exist. In these cases, it is sometimes necessary to estimate an approximate likelihood or fit a fast emulator model for efficient statistical inf…
CEDA improves understanding of data fit to models.
Hawkes processes have seen a number of applications in finance, due to their ability to capture event clustering behaviour typically observed in financial systems. Given a calibrated Hawkes process, of concern is the statistical fit to empirical data, particularly for the accurate quantification of self- and mutual-exc…
These are the written discussions of the paper "Bayesian measures of model complexity and fit" by D. Spiegelhalter et al. (2002), following the discussions given at the Annual Meeting of the Royal Statistical Society in Newcastle-upon-Tyne on September 3rd, 2013.
The statistical description and modeling of volatility plays a prominent role in econometrics, risk management and finance. GARCH and stochastic volatility models have been extensively studied and are routinely fitted to market data, albeit providing a phenomenological description only. In contrast, the field of econop…
T-Rex uses EM to fit robust factor models in noisy data.
In large-scale statistical learning, data collection and model fitting are moving increasingly toward peripheral devices---phones, watches, fitness trackers---away from centralized data collection. Concomitant with this rise in decentralized data are increasing challenges of maintaining privacy while allowing enough in…
The paper fits a seven-parameter GTS distribution to financial data.
Latent block models are used for probabilistic biclustering, which is shown to be an effective method for analyzing various relational data sets. However, there has been no statistical test method for determining the row and column cluster numbers of latent block models. Recent studies have constructed statistical-test…
Boosting improves data fitting while maintaining fairness guarantees.
The paper advances U-statistics in dependent settings, improving spectral estimation and goodness-of-fit tests.
A new test method improves goodness-of-fit tests for copulas.
Maximum likelihood estimation and a test of fit based on the Anderson-Darling statistic is presented for the case of the power law distribution when the parameters are estimated from a left-censored sample. Expressions for the maximum likelihood estimators and tables of asymptotic percentage points for the A^2 statisti…
A new distributed algorithm for fitting sparse additive models with feature division and decorrelation.
We derive a new discrepancy statistic for measuring differences between two probability distributions based on combining Stein's identity with the reproducing kernel Hilbert space theory. We apply our result to test how well a probabilistic model fits a set of observations, and derive a new class of powerful goodness-o…
We propose two nonparametric statistical tests of goodness of fit for conditional distributions: given a conditional probability density function and a joint sample, decide whether the sample is drawn from for some density . Our tests, formulated with a Stein operator, can be applied to any…
The paper proposes a test to determine the number of latent classes in ordinal categorical data.
We seek to utilize the nonextensive statistics to the microscopic modeling of the interacting many-investor dynamics that drive the price changes in a market. The statistics of price changes are known to be fit well by the Students-T and power-law distributions of the nonextensive statistics. We therefore derive models…
This work improves fair tensor decomposition using a kernel criterion.
New autoencoder uses goodness-of-fit tests for better model performance.
Latent space models are effective tools for statistical modeling and exploration of network data. These models can effectively model real world network characteristics such as degree heterogeneity, transitivity, homophily, etc. Due to their close connection to generalized linear models, it is also natural to incorporat…
Paper introduces statistical learning for point processes.
This chapter reviews classic regression methods and their evolution to physics-informed approaches.
This article describes a multivariate polynomial regression method where the uncertainty of the input parameters are approximated with Gaussian distributions, derived from the central limit theorem for large weighted sums, directly from the training sample. The estimated uncertainties can be propagated into the optimal…
We show that univariate and symmetric multivariate Hawkes processes are only weakly causal: the true log-likelihoods of real and reversed event time vectors are almost equal, thus parameter estimation via maximum likelihood only weakly depends on the direction of the arrow of time. In ideal (synthetic) conditions, test…
Improved modeling of persistence diagrams for data analysis.
The two key issues of modern Bayesian statistics are: (i) establishing principled approach for distilling statistical prior that is consistent with the given data from an initial believable scientific prior; and (ii) development of a Bayes-frequentist consolidated data analysis workflow that is more effective than eith…
The paper calculates ruin probabilities for insurers with phase-type distributed claims.
Bidirectional attention is shown to be equivalent to a continuous bag of words model with mixture-of-experts.
Study presents MMC model for better fitting multiple choice data.
PFNs pre-train models on simulated data to predict class probabilities.
New test assesses probabilistic model calibration without expensive approximations.
Given two candidate models, and a set of target observations, we address the problem of measuring the relative goodness of fit of the two models. We propose two new statistical tests which are nonparametric, computationally efficient (runtime complexity is linear in the sample size), and interpretable. As a unique adva…
The -generalised distribution fits daily stock returns well.
Unified framework for various probability distribution distances.
Proposes a method to make statistical inferences robust in spatially dependent settings with missing at random labels.
CPCR mitigates bias in PCR for overparameterized models.
A new test assesses how well observed networks fit a specified ERGM model.
The balance property is crucial for insurance pricing, ensuring total actuarial price equals loss. Maximum likelihood GLMs fulfill it, but Lindholm-Wüthrich suggests three methods, with constrained GLM being superior.
We present a new variable selection method based on model-based gradient boosting and randomly permuted variables. Model-based boosting is a tool to fit a statistical model while performing variable selection at the same time. A drawback of the fitting lies in the need of multiple model fits on slightly altered data (e…