Unified representation of density-power-based divergences simplifies estimation to M-estimation.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Study explores relationship between Hölder and FDPD divergences.
New method minimizes robust density power-based divergences for general parametric densities.
This paper improves active learning by using robust divergences for committee disagreement.
Study compares statistical properties and power of divergence measures for credit risk monitoring.
Improved Bayesian inference using power priors with historical data.
The paper introduces a new method for estimating optimal policies in dynamic treatment regimes using information geometry.
Improved VAE for heavy-tailed data using Student's t-distributions.
A new robust PCA estimator combining M-estimators and minimum divergence estimators.
This paper generalizes beta divergence beyond its classical form associated with power variance functions of Tweedie models. Generalized form is represented by a compact definite integral as a function of variance function of the exponential dispersion model. This compact integral form simplifies derivations of many pr…
For distributions and with different supports or undefined densities, the divergence may not exist. We define a Spread Divergence on modified and and describe sufficient conditions for t…
Unified technique for sequential estimation of convex divergences.
A minimal model of a market of myopic non-cooperative agents who trade bilaterally with random bids reproduces qualitative features of short-term electric power markets, such as those in California and New England. Each agent knows its own budget and preferences but not those of any other agent. The near-equilibrium pr…
Proposes practical kernel tests for -divergences with theoretical guarantees.
New metrics defined on SPD matrices link to divergences and curvature.
We extend CS divergence to conditional distributions and show its advantages in time series data and sequential decision making.
Differential privacy is a de facto standard in data privacy, with applications in the public and private sectors. A way to explain differential privacy, which is particularly appealing to statistician and social scientists is by means of its statistical hypothesis testing interpretation. Informally, one cannot effectiv…
The t-distributed Stochastic Neighbor Embedding (t-SNE) is a powerful and popular method for visualizing high-dimensional data. It minimizes the Kullback-Leibler (KL) divergence between the original and embedded data distributions. In this work, we propose extending this method to other f-divergences. We analytically a…
In high-dimensional data, many sparse regression methods have been proposed. However, they may not be robust against outliers. Recently, the use of density power weight has been studied for robust parameter estimation and the corresponding divergences have been discussed. One of such divergences is the -divergence a…
EM optimizes tensor density estimation by relaxing -divergence to KL-divergence.
Robust VB framework handles contamination using min-max median aggregation.
New Stein operator improves robustness in model inference.
While Generative Adversarial Networks (GANs) have empirically produced impressive results on learning complex real-world distributions, recent works have shown that they suffer from lack of diversity or mode collapse. The theoretical work of Arora et al. suggests a dilemma about GANs' statistical properties: powerful d…
Locally private mechanisms' output divergence bounds derived.
A new distribution addresses scalability and numerical stability issues of the vMF.
Exponential dispersion model is a useful framework in machine learning and statistics. Primarily, thanks to the additive structure of the model, it can be achieved without difficulty to estimate parameters including mean. However, tight conditions on cumulant function, such as analyticity, strict convexity, and steepne…
New bounds for sequential tests under power-one error levels.
Paper develops a method to compare generative models using KL divergence.
We study the risk criterion for investments based on the drawdown from the maximal value of the capital in the past. Depending on investor's risk attitude, thus his risk exposure, we find that the distribution of these drawdowns follows a general power law. In particular, if the risk exposure is Kelly-optimal, the expo…
Generative adversarial network (GAN) is a minimax game between a generator mimicking the true model and a discriminator distinguishing the samples produced by the generator from the real training samples. Given an unconstrained discriminator able to approximate any function, this game reduces to finding the generative …
We study the geometry of probability distributions with respect to a generalized family of Csiszár -divergences. A member of this family is the relative -entropy which is also a Rényi analog of relative entropy in information theory and known as logarithmic or projective power divergence in statistics. We apply E…
In the presence of model risk, it is well-established to replace classical expected values by worst-case expectations over all models within a fixed radius from a given reference model. This is the "robustness" approach. We show that previous methods for measuring this radius, e.g. relative entropy or polynomial diverg…
We propose a modified -divergence, give some of its properties, and show that this leads to the definition of a generalized Fisher information. We give generalized Cramér-Rao inequalities, involving this Fisher information, an extension of the Fisher information matrix, and arbitrary norms and power of the estimat…
Proposes a new method to improve Bayesian computation accuracy using flexible classification.
Local mass perspective on Bayesian inference
Paper introduces SDM for detecting LLM hallucinations, improving on entropy tests.
Study finds no evidence dual-class stocks are effective predictors.
Improved hypothesis testing and change-point detection using diffusion-based methods.
In this note, we point out a basic link between generative adversarial (GA) training and binary classification -- any powerful discriminator essentially computes an (f-)divergence between real and generated samples. The result, repeatedly re-derived in decision theory, has implications for GA Networks (GANs), providing…
We introduce a new test for detection of power-law cross-correlations among a pair of time series - the rescaled covariance test. The test is based on a power-law divergence of the covariance of the partial sums of the long-range cross-correlated processes. Utilizing a heteroskedasticity and auto-correlation robust est…
In this paper we study a homological version of the higher-dimensional divergence invariants defined by Brady and Farb. We show that they are quasi-isometry invariants in the class of proper cocompact Hadamard spaces in the sense of Alexandrov and that they can moreover be used to detect the Euclidean rank of such spac…
We propose a direct estimation method for Rényi and f-divergence measures based on a new graph theoretical interpretation. Suppose that we are given two sample sets and , respectively with and samples, where is a constant value. Considering the -nearest neighbor (-NN) graph of in the j…
Paper tackles non-convex constrained DRO with a stochastic algorithm for large-scale applications.
Paper resolves bias in ALFT training using generalized alignment games.
Measuring divergence between two distributions is essential in machine learning and statistics and has various applications including binary classification, change point detection, and two-sample test. Furthermore, in the era of big data, designing divergence measure that is interpretable and can handle high-dimensiona…
Paper analyzes inclusive KL inference using Wasserstein gradient flows.
The variational autoencoder (VAE) is a powerful generative model that can estimate the probability of a data point by using latent variables. In the VAE, the posterior of the latent variable given the data point is regularized by the prior of the latent variable using Kullback Leibler (KL) divergence. Although the stan…
Variational Autoencoder (VAE), a simple and effective deep generative model, has led to a number of impressive empirical successes and spawned many advanced variants and theoretical investigations. However, recent studies demonstrate that, when equipped with expressive generative distributions (aka. decoders), VAE suff…