Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,695 papers · 148 categories

Trend · papers per month

12.5%25.0%37.5%50.0% · Nov 199319922001200920172026
48 results for distribution fitting

New method extends fitted Q-evaluation for distributional off-policy reinforcement learning.

problem Estimating return distribution in reinforcement learning using offline data.
method Developed a set of guiding principles and new FDE methods with theoretical justification.
result FDE methods outperform existing approaches in simulations and real-world games.

This paper is a step-by-step tutorial for fitting a mixture distribution to data. It merely assumes the reader has the background of calculus and linear algebra. Other required background is briefly reviewed before explaining the main algorithm. In explaining the main algorithm, first, fitting a mixture of two distribu…

2019-01-20abs ↗pdf ↗

Paper explores how text generation quality and diversity metrics relate to distribution fitting.

problem Unclear relation between text generation quality and diversity metrics and distribution fitting.
method Theoretical approach to prove a linear combination of quality and diversity metrics can be a divergence metric.
result CR/NRR proposed as a better substitute for BLEU/Self-BLEU metrics.

We study a resource utilization scenario characterized by intrinsic fitness. To describe the growth and organization of different cities, we consider a model for resource utilization where many restaurants compete, as in a game, to attract customers using an iterative learning process. Results for the case of restauran…

2014-03-07abs ↗pdf ↗

This paper illustrates a procedure for fitting financial data with αα-stable distributions. After using all the available methods to evaluate the distribution parameters, one can qualitatively select the best estimate and run some goodness-of-fit tests on this estimate, in order to quantitatively assess its quality. I…

2006-08-23abs ↗pdf ↗

The κκ-generalised distribution fits daily stock returns well.

problem Stock returns are often heavy-tailed, not normally distributed.
method Used the κκ-generalised distribution with a Monte-Carlo goodness of fit test.
result The κκ-generalised distribution fits historic daily stock returns well for a significant proportion of analyzed stocks.

The paper fits a seven-parameter GTS distribution to financial data.

problem Nonexistence of GTS probability density function makes MLE inadequate.
method Used fractional Fourier transform to circumvent MLE and provide good parameter estimation.
result The GTS distribution fits financial data significantly better than other models.

Improves decision-making in models fit with AEVB by using distinct approximate posteriors.

problem Bias in expected risk estimates due to variational distribution use.
method Use multiple approximate posteriors, including those distinct from variational, for decision-making.
result Proposed approach outperforms state-of-the-art methods in single-cell RNA sequencing.

A new method for modeling insurance claim frequencies using random proportions.

problem Inaccurate fitting of classical distributions to insurance claim frequency data.
method Modeling claim frequencies using random proportions of insurance contracts and applying goodness-of-fit tests.
result A new statistical approach for better modeling insurance claim frequencies.

In the spirit of the emergent field of econophysics, a goodness-of-fit test for the Power-Law distribution, based on the Empirical Distribution Function (EDF) is presented, and related problems are discussed. An analysis of the tail behaviour of the daily logarithmic variation of the Mexican Stock Market Index (IPC), s…

2003-03-27abs ↗pdf ↗

Ancestral graph models, introduced by Richardson and Spirtes (2002), generalize both Markov random fields and Bayesian networks to a class of graphs with a global Markov property that is closed under conditioning and marginalization. By design, ancestral graphs encode precisely the conditional independence structures t…

2012-07-11abs ↗pdf ↗

Autonomous agents that must exhibit flexible and broad capabilities will need to be equipped with large repertoires of skills. Defining each skill with a manually-designed reward function limits this repertoire and imposes a manual engineering burden. Self-supervised agents that set their own goals can automate this pr…

2019-03-08abs ↗pdf ↗

FIT evaluates time series model feature importance quantifying distributional shift.

problem Lack of explanations for time series models in high-stakes applications.
method FIT framework quantifies feature importance based on distributional shift using KL-divergence.
result FIT identifies important time points and observations superiorly compared to baselines.

We prove that Student's t-distribution provides one of the better fits to returns of S&P component stocks and the generalized inverse gamma distribution best fits VIX and VXO volatility data. We further argue that a more accurate measure of the volatility may be possible based on the fact that stock returns can be unde…

2013-05-17abs ↗pdf ↗

Algorithm infers sampling distribution from i.i.d. samples without supervision.

problem Learning probability distributions from unlabeled data.
method Unsupervised tree boosting using additive tree ensembles and new distributional operations.
result Algorithm outperforms deep learning in multivariate density estimation.

Using tax and census data, we demonstrate that the distribution of individual income in the USA is exponential. Our calculated Lorenz curve without fitting parameters and Gini coefficient 1/2 agree well with the data. From the individual income distribution, we derive the distribution function of income for families wi…

2000-08-21abs ↗pdf ↗

Employing profits data of Japanese companies in 2002 and 2003, we identify the non-Gibrat's law which holds in the middle profits region. From the law of detailed balance in all regions, Gibrat's law in the high region and the non-Gibrat's law in the middle region, we kinematically derive the profits distribution funct…

2005-08-24abs ↗pdf ↗

The paper calculates ruin probabilities for insurers with phase-type distributed claims.

problem Calculating ruin probabilities for insurers with specific claim distributions.
method Change-of-measure technique applied to phase-type distributed claim amounts.
result The mixture of Erlangs best fits real-world loss data, improving risk assessment.

We provide an analytically treatable model that describes in a unified manner income distribution for all income categories. The approach is based on a master equation with growth and reset terms. The model assumptions on the growth and reset rates are tested on an exhaustive database with incomes on individual level s…

2019-11-06abs ↗pdf ↗

GRASP tests goodness-of-fit for binary classifiers without parametric assumptions.

problem Assessing the fit of a binary classifier to the underlying conditional law of labels given features.
method Formulates a tolerance hypothesis testing problem and proposes a novel test called GRASP.
result Proposes GRASP and Model-X GRASP tests for assessing goodness-of-fit in finite sample settings.

A new framework improves kernel Stein discrepancy tests for validating distributions.

problem Improving goodness-of-fit testing for non-normal distributions.
method Introducing Sf-KSD, a unifying framework for studying Stein operators in KSD-based tests.
result Sf-KSD guides the development of new tests and outperforms existing methods.

In regression modelling approach, the main step is to fit the regression line as close as possible to the target variable. In this process most algorithms try to fit all of the data in a single line and hence fitting all parts of target variable in one go. It was observed that the error between predicted and target var…

2018-05-04abs ↗pdf ↗

Fitting models for non-Poisson point processes is complicated by the lack of tractable models for much of the data. By using large samples of independent and identically distributed realizations and statistical learning, it is possible to identify absence of fit through finding a classification rule that can efficientl…

2007-12-02abs ↗pdf ↗

Typically, operational risk losses are reported above a threshold. Fitting data reported above a constant threshold is a well known and studied problem. However, in practice, the losses are scaled for business and other factors before the fitting and thus the threshold is varying across the scaled data sample. A report…

2009-04-27abs ↗pdf ↗

Improved KSD test for better detection of differences in distributions.

problem Low power of KSD test when distributions have same modes but different mixing proportions.
method Perturb the observed sample using Markov transition kernels to improve KSD test power.
result Perturbed KSD test can lead to substantially higher power than the original KSD test.

How to forecast next year's portfolio-wide credit default rate based on last year's default observations and the current score distribution? A classical approach to this problem consists of fitting a mixture of the conditional score distributions observed last year to the current score distribution. This is a special (…

2014-06-23abs ↗pdf ↗

The Black-Scholes theory of option pricing has been considered for many years as an important but very approximate zeroth-order description of actual market behavior. We generalize the functional form of the diffusion of these systems and also consider multi-factor models including stochastic volatility. Daily Eurodoll…

2000-01-23abs ↗pdf ↗

A homogeneously saturated equation for the time development of the price of a financial asset is presented and investigated for the pricing of European call options using noise that is distributed as a Student's t-distribution. In the limit that the saturation parameter of the equation equals zero, the standard model o…

2013-01-24abs ↗pdf ↗

Study analyzes Airbnb lead-time distributions for Nights Booked and Gross Booking Value, finding divergent shapes and tail behavior.

problem Analyzing lead-time distributions for Airbnb demand metrics.
method Compositional analysis of daily lead-time vectors, fitting Gamma, Weibull, and Lognormal distributions, using generalized Pareto for tail inference.
result Lead-time distributions for Nights Booked and Gross Booking Value diverge, with GBV concentrating more in mid-range horizons.

The accumulation of individual fitness or wealth is modelled as a population game in which pairs of individuals are recurrently and randomly matched to play a game over a resource. In addition, all individuals have random access to a constant background resource, and their fitness or wealth depreciates over time. For b…

2017-07-04abs ↗pdf ↗

The multivariate version of the Mixed Tempered Stable is proposed. It is a generalization of the Normal Variance Mean Mixtures. Characteristics of this new distribution and its capacity in fitting tails and capturing dependence structure between components are investigated. We discuss a random number generating procedu…

2016-09-04abs ↗pdf ↗

Unified neural network model for astro-particle physics predictions with coverage, systematics, and goodness-of-fit.

problem Lack of statistical uncertainties, coverage, systematic uncertainties, and goodness-of-fit in neural network predictions.
method KL-divergence objective for joint distribution of data and labels, conditional normalizing flows, amortized with neural networks.
result Unified supervised learning and VAEs under stochastic variational inference for event property predictions.