Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

17355269 · May 202619922001200920172026
48 results for JS divergence

The paper introduces a new divergence measure for variational autoencoders to improve reconstruction and generation.

problem Balancing reconstruction and generalizability in latent space of variational autoencoders.
method Presented a regularisation mechanism based on skew-geometric Jensen-Shannon divergence.
result The skew-geometric Jensen-Shannon divergence leads to better reconstruction and generation in variational autoencoders.

Improved GANs estimate convergence rate for density estimation.

problem Improving the accuracy of density estimation with GANs.
method Proved an oracle inequality for JS divergence between GAN estimate and true density.
result JS-divergence rate of convergence is (logn/n)2β/(2β+d)(\log{n}/n)^{2β/(2β+ d)}.

Implicit generative models are difficult to train as no explicit density functions are defined. Generative adversarial nets (GANs) present a minimax framework to train such models, which however can suffer from mode collapse due to the nature of the JS-divergence. This paper presents a learning by teaching (LBT) approa…

2018-07-10abs ↗pdf ↗

Generative adversarial network (GAN) is a minimax game between a generator mimicking the true model and a discriminator distinguishing the samples produced by the generator from the real training samples. Given an unconstrained discriminator able to approximate any function, this game reduces to finding the generative …

2018-10-28abs ↗pdf ↗

A new objective function using Jensen-Shannon divergence improves generative learning from multiple data types.

problem Learning from multiple data types efficiently and accurately.
method Proposes a novel objective function using Jensen-Shannon divergence to approximate multimodal posteriors directly.
result The mmJSD objective optimizes an ELBO and improves generative learning tasks.

WDAIL uses Wasserstein distance for more effective reward shaping in IL.

problem Fixed reward functions in GAIL limit performance on complex tasks.
method Introduces Wasserstein distance and PPO for improved reward shaping and stability.
result Significant performance improvement in complex MuJoCo tasks.

We study risk-sensitive imitation learning where the agent's goal is to perform at least as well as the expert in terms of a risk profile. We first formulate our risk-sensitive imitation learning setting. We consider the generative adversarial approach to imitation learning (GAIL) and derive an optimization problem for…

2018-08-13abs ↗pdf ↗

Complex computer simulators are increasingly used across fields of science as generative models tying parameters of an underlying theory to experimental observations. Inference in this setup is often difficult, as simulators rarely admit a tractable density or likelihood function. We introduce Adversarial Variational O…

2017-07-22abs ↗pdf ↗

PolySwarm uses a swarm of LLMs to predict and arbitrage prediction markets.

problem Real-time prediction market trading and latency arbitrage inefficiencies.
method PolySwarm employs a swarm of 50 diverse LLMs, Bayesian combination, and risk-controlled execution.
result Swarm aggregation outperforms single-model baselines in prediction tasks.

Quantum Clustering is a powerful method to detect clusters in data with mixed density. However, it is very sensitive to a length parameter that is inherent to the Schrödinger equation. In addition, linking data points into clusters requires local estimates of covariance that are also controlled by length parameters. Th…

2019-02-14abs ↗pdf ↗

This paper considers the problem of estimating a high-dimensional vector of parameters θRn\boldsymbolθ \in \mathbb{R}^n from a noisy observation. The noise vector is i.i.d. Gaussian with known variance. For a squared-error loss function, the James-Stein (JS) estimator is known to dominate the simple maximum-likelihood (…

2016-02-01abs ↗pdf ↗

In safety-critical applications of machine learning, it is often important to abstain from making predictions on low confidence examples. Standard abstention methods tend to be focused on optimizing top-k accuracy, but in many applications, accuracy is not the metric of interest. Further, label shift (a shift in class …

2018-02-20abs ↗pdf ↗

C-SURE improves complex-valued deep learning models by shrinking estimates, outperforming MLE and SurReal.

problem Improving accuracy and robustness of complex-valued deep learning models.
method Proposes a Stein's unbiased risk estimate (SURE) for complex-valued data and integrates it into a prototype CNN classifier.
result C-SURE outperforms SurReal and MLE in accuracy and robustness on complex-valued datasets.

The public package registry npm is one of the biggest software registry. With its 216 911 software packages, it forms a big network of software dependencies. In this paper we evaluate various methods for finding similar packages in the npm network, using only the structure of the graph. Namely, we want to find a way of…

2016-02-11abs ↗pdf ↗

Divergence functions play a key role as to measure the discrepancy between two points in the field of machine learning, statistics and signal processing. Well-known divergences are the Bregman divergences, the Jensen divergences and the f-divergences. In this paper, we show that the symmetric Bregman divergence can be …

2018-10-03abs ↗pdf ↗

Study explores relationship between Hölder and FDPD divergences.

problem Understanding the relationship between Hölder and FDPD divergences.
method Intersection and generalization of divergence families, proving nonnegativity, deriving inequalities.
result Established ξξ-Hölder divergence and derived inequalities.

Unified representation of density-power-based divergences simplifies estimation to M-estimation.

problem Outliers in density estimation.
method Define a norm-based Bregman density power divergence (NB-DPD) that reduces to M-estimation.
result NB-DPD connects and generalizes existing divergences, highlighting robustness properties.

This paper improves active learning by using robust divergences for committee disagreement.

problem Active learning with high measurement costs.
method Query by committee with Bregman divergence (including Kullback-Leibler divergence as a special case).
result The proposed method is more robust and performs as well as or better than conventional methods.

The paper improves semi-supervised learning using ff-divergences and αα-Rényi divergences.

problem Improving semi-supervised learning with noisy pseudo-labels.
method Inspired by ff-divergences and αα-Rényi divergences, the paper develops new empirical risk functions and regularization techniques.
result The new methods show better performance than traditional self-training methods, especially in noisy pseudo-label scenarios.

ff-divergences are a general class of divergences between probability measures which include as special cases many commonly used divergences in probability, mathematical statistics and information theory such as Kullback-Leibler divergence, chi-squared divergence, squared Hellinger distance, total variation distance e…

2013-02-02abs ↗pdf ↗

We introduce a new quasi-isometry invariant, called the divergence spectrum, to study finitely generated groups. We compare the concept of divergence spectrum with the other classical notions of divergence and we examine the divergence spectra of relatively hyperbolic groups. We show the existence of an infinite collec…

2016-11-15abs ↗pdf ↗

We study the logarithmic L(α)L^{(α)}-divergence which extrapolates the Bregman divergence and corresponds to solutions to novel optimal transport problems. We show that this logarithmic divergence is equivalent to a conformal transformation of the Bregman divergence, and, via an explicit affine immersion, is equivalent t…

2019-06-17abs ↗pdf ↗

The study defines divergence for multivector fields on infinite-dimensional manifolds.

problem Defining divergence for multivector fields on infinite-dimensional manifolds.
method Definition of divergence consistent with finite-dimensional geometry, properties transferred from finite to infinite dimensions.
result Natural properties of divergence are preserved in infinite dimensions.

The paper evaluates biased methods for alpha-divergence minimization.

problem The impact of bias on solutions found for alpha-divergence minimization.
method Empirical evaluation of biased methods for alpha-divergence minimization, focusing on bias effects and dimensionality.
result Solutions are biased towards KL-divergence minimizers and require impractical computation in high dimensions to minimize alpha-divergence.

Develops a new divergence framework that combines ff-divergences and IPMs.

problem Comparing distributions that are not absolutely continuous.
method Introduces (f,Γ)(f,Γ)-divergences as a two-stage mass-redistribution/mass-transport process.
result Improves estimation, learning, and uncertainty quantification in GANs for heavy-tailed distributions.

Study compares statistical properties and power of divergence measures for credit risk monitoring.

problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.

New α\alpha-divergence loss function improves neural density ratio estimation.

problem Optimization challenges in existing DRE methods, especially overfitting and high sample requirements.
method Derived α\alpha-divergence loss function (α\alpha-Div) for neural density ratio estimation.
result The α\alpha-divergence loss function (α\alpha-Div) offers stable and effective optimization for DRE.

Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.

problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.

The paper explores how information geometry impacts classical CR inequalities.

problem Deriving and generalizing CR inequalities using information geometry.
method Examining Eguchi's theory and applying Amari-Nagoaka's theory to KL-divergence, and then extending to other divergences.
result Generalized CR inequalities derived from various divergences.

We extend CS divergence to conditional distributions and show its advantages in time series data and sequential decision making.

problem Quantifying the closeness between conditional distributions.
method Developed and estimated a conditional Cauchy-Schwarz divergence using kernel density estimation.
result Conditional CS divergence outperforms previous methods in time series clustering and sequential decision making.

Rényi divergence is related to Rényi entropy much like Kullback-Leibler divergence is related to Shannon's entropy, and comes up in many settings. It was introduced by Rényi as a measure of information that satisfies almost the same axioms as Kullback-Leibler divergence, and depends on a parameter that is called its or…

2012-06-12abs ↗pdf ↗

Classifies divergence and thickness in right-angled Coxeter groups.

problem Characterizing the divergence and thickness of right-angled Coxeter groups.
method Completely classifies divergence functions and proves conditions for thickness using the hypergraph index.
result Exact divergence functions of RACGs can be computed from their defining graphs.

The paper explores statistical and topological properties of sliced probability divergences.

problem Understanding the topological, statistical, and computational consequences of slicing divergences.
method Deriving theoretical properties of sliced probability divergences, including metric axioms preservation and weak continuity.
result Sliced divergences share similar topological properties and have stable sample complexity.

New framework using Jensen-Shannon divergence improves domain adaptation theory.

problem Incoherence between empirical domain adversarial training and theoretical H\mathcal{H}-divergence.
method Established new theoretical framework based on Jensen-Shannon divergence, derived bi-directional upper bounds.
result Framework exhibits flexibilities for various transfer learning problems.