New bound on machine learning model performance using Jensen-Shannon information.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
A new objective function using Jensen-Shannon divergence improves generative learning from multiple data types.
New method improves understanding of machine learning model performance.
Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
Proposes a new divergence measure for probability distributions.
New framework using Jensen-Shannon divergence improves domain adaptation theory.
In this paper, we present a new wrapper feature selection approach based on Jensen-Shannon (JS) divergence, termed feature selection with maximum JS-divergence (FSMJ), for text categorization. Unlike most existing feature selection approaches, the proposed FSMJ approach is based on real-valued features which provide mo…
Proposes a new loss function for learning with noisy labels.
Computing approximate nearest neighbors in high dimensional spaces is a central problem in large-scale data mining with a wide range of applications in machine learning and data science. A popular and effective technique in computing nearest neighbors approximately is the locality-sensitive hashing (LSH) scheme. In thi…
In this report, we derive a non-negative series expansion for the Jensen-Shannon divergence (JSD) between two probability distributions. This series expansion is shown to be useful for numerical calculations of the JSD, when the probability distributions are nearly equal, and for which, consequently, small numerical er…
New insights link RLHF and contrastive learning for better model alignment.
Study compares statistical properties and power of divergence measures for credit risk monitoring.
We introduce the Mutual Information Machine (MIM), a novel formulation of representation learning, using a joint distribution over the observations and latent state in an encoder/decoder framework. Our key principles are symmetry and mutual information, where symmetry encourages the encoder and decoder to learn differe…
We introduce the Mutual Information Machine (MIM), a probabilistic auto-encoder for learning joint distributions over observations and latent variables. MIM reflects three design principles: 1) low divergence, to encourage the encoder and decoder to learn consistent factorizations of the same underlying distribution; 2…
In this report, we present an unsupervised machine learning method for determining groups of molecular systems according to similarity in their dynamics or structures using Ward's minimum variance objective function. We first apply the minimum variance clustering to a set of simulated tripeptides using the information …
The paper introduces a new divergence measure for variational autoencoders to improve reconstruction and generation.
This paper raises an implicit manifold learning perspective in Generative Adversarial Networks (GANs), by studying how the support of the learned distribution, modelled as a submanifold , perfectly match with , the support of the real data distribution. We show that optimizing Jensen-Sha…
The coefficient of determination, known as , is commonly used as a goodness-of-fit criterion for fitting linear models. is somewhat controversial when fitting nonlinear models, although it may be generalised on a case-by-case basis to deal with specific models such as the logistic model. Assume we are fittin…
SMT trains generative models by estimating mixture scores, outperforming existing methods.
Formula derived for sample complexity in binary hypothesis testing.
The success of popular algorithms for deep reinforcement learning, such as policy-gradients and Q-learning, relies heavily on the availability of an informative reward signal at each timestep of the sequential decision-making process. When rewards are only sparsely available during an episode, or a rewarding feedback i…
Generative Adversarial Networks (GANs) have become a widely popular framework for generative modelling of high-dimensional datasets. However their training is well-known to be difficult. This work presents a rigorous statistical analysis of GANs providing straight-forward explanations for common training pathologies su…
Paper introduces metrics to evaluate missing data imputation without ground truth.
The study reveals the hierarchical structure of the international FOREX market using currency fluctuation distribution similarities.
Two methods factor out prior knowledge from low-dimensional embeddings.
Proposes a scalable method for counterfactual prediction using machine learning.
Paper introduces SDM for detecting LLM hallucinations, improving on entropy tests.
A new framework for offline RL improves policy flexibility and regularity.
Paper trains two models to improve synthesis quality and unsupervised learning.
PolyGraph Discrepancy improves graph generative model evaluation.
DeepInversion generates images from trained networks without additional data.
This study considers the multivariate segmentation procedure under the assumption of the multivariate Gaussian mixture. Jensen-Shannon divergence between two multivariate Gaussian distributions is employed as a discriminator and a recursive segmentation procedure is proposed. The daily log-return time series for 30 cur…
Paper proposes f-DPG for aligning language models with preferences.
The von Neumann graph entropy (VNGE) facilitates measurement of information divergence and distance between graphs in a graph sequence. It has been successfully applied to various learning tasks driven by network-based data. While effective, VNGE is computationally demanding as it requires the full eigenspectrum of the…
Parametric adversarial divergences, which are a generalization of the losses used to train generative adversarial networks (GANs), have often been described as being approximations of their nonparametric counterparts, such as the Jensen-Shannon divergence, which can be derived under the so-called optimal discriminator …
Proposes a new method for fairness in machine learning with multiple protected attributes.
In the field of statistics, many kind of divergence functions have been studied as an amount which measures the discrepancy between two probability distributions. In the differential geometrical approach in statistics (information geometry), dually flat spaces play a key role. In a dually flat space, there exist dual a…
New GAN loss functions improve image generation quality and stability.
PolySwarm uses a swarm of LLMs to predict and arbitrage prediction markets.
Clust-PSI-PFL uses PSI to improve accuracy and fairness in federated learning.
We propose JECL, a method for clustering image-caption pairs by training parallel encoders with regularized clustering and alignment objectives, simultaneously learning both representations and cluster assignments. These image-caption pairs arise frequently in high-value applications where structured training data is e…
We study risk-sensitive imitation learning where the agent's goal is to perform at least as well as the expert in terms of a risk profile. We first formulate our risk-sensitive imitation learning setting. We consider the generative adversarial approach to imitation learning (GAIL) and derive an optimization problem for…
We present two related methods for deriving connectivity-based brain atlases from individual connectomes. The proposed methods exploit a previously proposed dense connectivity representation, termed continuous connectivity, by first performing graph-based hierarchical clustering of individual brains, and subsequently a…
Improved GANs estimate convergence rate for density estimation.
High-frequency financial data of the foreign exchange market (EUR/CHF, EUR/GBP, EUR/JPY, EUR/NOK, EUR/SEK, EUR/USD, NZD/USD, USD/CAD, USD/CHF, USD/JPY, USD/NOK, and USD/SEK) are analyzed by utilizing the Kullback-Leibler divergence between two normalized spectrograms of the tick frequency and the generalized Jensen-Sha…
Empirical analysis of the foreign exchange market is conducted based on methods to quantify similarities among multi-dimensional time series with spectral distances introduced in [A.-H. Sato, Physica A, 382 (2007) 258--270]. As a result it is found that the similarities among currency pairs fluctuate with the rotation …
LoRA-Curve connects independent LoRA optima through continuous low-loss valleys, improving Bayesian model averaging.
We address the problem of imitation learning with multi-modal demonstrations. Instead of attempting to learn all modes, we argue that in many tasks it is sufficient to imitate any one of them. We show that the state-of-the-art methods such as GAIL and behavior cloning, due to their choice of loss function, often incorr…