Study on geometric Jensen-Shannon divergence for Gaussian measures in Hilbert space.
problem Computing divergence between Gaussian measures in infinite-dimensional Hilbert space.
method Closed form expression and regularization for divergence calculation.
result Closed form expression and regularization for Geometric Jensen-Shannon divergence.
In this paper, we introduce new classes of divergences by extending the definitions of the Bregman divergence and the skew Jensen divergence. These new divergence classes (g-Bregman divergence and skew g-Jensen divergence) satisfy some properties similar to the Bregman or skew Jensen divergence. We show these g-diverge…
Transformers for binary decisions are sensitive to evidence order, leading to unreliable outcomes.
problem Order sensitivity in Transformers for binary decisions leads to unreliable outcomes.
method Formalized an expectation-realization gap and developed QMV and EDFL bounds.
result Uniform permutation mixtures reduce dispersion and improve reliability.
In this report, we present an unsupervised machine learning method for determining groups of molecular systems according to similarity in their dynamics or structures using Ward's minimum variance objective function. We first apply the minimum variance clustering to a set of simulated tripeptides using the information …
Proposes a new divergence measure for probability distributions.
problem Challenges in estimating divergences from empirical samples.
method Embeds data into RKHS, computes Jensen-Shannon divergence between covariance operators.
result Establishes RJSD as a lower bound on Jensen-Shannon divergence, enabling variational estimation.
New bound on machine learning model performance using Jensen-Shannon information.
problem Understanding the performance of machine learning models.
method Proposes a new information-theoretic bound on generalization error.
result Shows that the new bound can be tighter than mutual information-based bounds under certain conditions.
New framework using Jensen-Shannon divergence improves domain adaptation theory.
problem Incoherence between empirical domain adversarial training and theoretical H-divergence. method Established new theoretical framework based on Jensen-Shannon divergence, derived bi-directional upper bounds.
result Framework exhibits flexibilities for various transfer learning problems.
A new objective function using Jensen-Shannon divergence improves generative learning from multiple data types.
problem Learning from multiple data types efficiently and accurately.
method Proposes a novel objective function using Jensen-Shannon divergence to approximate multimodal posteriors directly.
result The mmJSD objective optimizes an ELBO and improves generative learning tasks.
This paper introduces Jensen, an easily extensible and scalable toolkit for production-level machine learning and convex optimization. Jensen implements a framework of convex (or loss) functions, convex optimization algorithms (including Gradient Descent, L-BFGS, Stochastic Gradient Descent, Conjugate Gradient, etc.), …
Proposes a new loss function for learning with noisy labels.
problem Improving model learnability with noisy labels.
method Uses generalized Jensen-Shannon divergence as a noise-robust loss function.
result Shows state-of-the-art results on noisy data.
In this report, we derive a non-negative series expansion for the Jensen-Shannon divergence (JSD) between two probability distributions. This series expansion is shown to be useful for numerical calculations of the JSD, when the probability distributions are nearly equal, and for which, consequently, small numerical er…
Study compares statistical properties and power of divergence measures for credit risk monitoring.
problem Detecting distributional shifts in credit risk models.
method Derives statistical properties and chi-square benchmark values for Jensen-Shannon Divergence and Kullback-Leibler Divergence, demonstrating their applicability in credit risk monitoring.
result Jensen-Shannon Divergence and Kullback-Leibler Divergence follow chi-square distributions and reveal practical trade-offs in minimizing false positives vs. detecting changes.
New method improves understanding of machine learning model performance.
problem Understanding how well machine learning models generalize from training data to unseen data.
method Auxiliary Distribution Method to derive new generalization error bounds.
result Upper bounds on generalization errors are tighter and more applicable.
Divergence functions play a key role as to measure the discrepancy between two points in the field of machine learning, statistics and signal processing. Well-known divergences are the Bregman divergences, the Jensen divergences and the f-divergences. In this paper, we show that the symmetric Bregman divergence can be …
The paper proves a Jensen's inequality in spaces with lower bounded curvature.
problem Proving Jensen's inequality in geodesic spaces with curvature constraints.
method Using properties of tangent cones and gradients for semi-concave functions in spaces with lower bounded curvature.
result The inequality holds for geodesically convex functions in spaces with curvature lower bounded.
In the paper we give necessary and sufficient conditions for the Jensen inequality to hold for the generalized Choquet integral with respect to a pair of capacities. Next, we apply obtained result to the theory of risk aversion by providing the assumptions on utility function and capacities under which an agent is risk…
Mathematical study of excess growth rate connects info theory with finance.
problem Understanding the excess growth rate in portfolio theory.
method Axiomatic characterization theorems of excess growth rate in terms of relative entropy, Jensen's inequality gap, and logarithmic divergence.
result Established rich connections between information theory and finance.
Paper improves particle variational inference by optimizing generalization error bound.
problem Improving the diversity of models in particle variational inference to enhance generalization.
method Develops a new second-order Jensen inequality with a repulsion term based on the loss function, leading to a tighter generalization error bound.
result The proposed PVI optimizes the generalization error bound directly, improving performance compared to existing methods.
Extends online learning to metric spaces using exponential weights.
problem Online learning in metric spaces.
method Exponentially weighted average forecaster, barycenters, Jensen's inequality, measure contraction property.
result Results in a statistical learning framework.
The paper introduces a new divergence measure for variational autoencoders to improve reconstruction and generation.
problem Balancing reconstruction and generalizability in latent space of variational autoencoders.
method Presented a regularisation mechanism based on skew-geometric Jensen-Shannon divergence.
result The skew-geometric Jensen-Shannon divergence leads to better reconstruction and generation in variational autoencoders.
We consider invariant Einstein metrics on the Stiefel manifold $V_q\bb{R} ^n$ of all orthonormal q-frames in $\bb{R}^n$. This manifold is diffeomorphic to the homogeneous space $\SO(n)/\SO(n-q)$ and its isotropy representation contains equivalent summands. %This causes difficulty in the description of all $\SO(n)$-in…
Holomorphic automorphisms on hyperkähler manifolds with high entropy are Kummer examples.
problem Characterizing holomorphic automorphisms with high entropy on hyperkähler manifolds.
method Using Jensen's inequality and properties of stable and unstable distributions, the authors show uniform contraction and expansion, leading to the conclusion that the manifold is birational to a torus quotient.
result Holomorphic automorphisms with high entropy on hyperkähler manifolds are Kummer examples.
This paper raises an implicit manifold learning perspective in Generative Adversarial Networks (GANs), by studying how the support of the learned distribution, modelled as a submanifold Mθ, perfectly match with Mr, the support of the real data distribution. We show that optimizing Jensen-Sha…
The coefficient of determination, known as R2, is commonly used as a goodness-of-fit criterion for fitting linear models. R2 is somewhat controversial when fitting nonlinear models, although it may be generalised on a case-by-case basis to deal with specific models such as the logistic model. Assume we are fittin…
In this paper, we present a new wrapper feature selection approach based on Jensen-Shannon (JS) divergence, termed feature selection with maximum JS-divergence (FSMJ), for text categorization. Unlike most existing feature selection approaches, the proposed FSMJ approach is based on real-valued features which provide mo…
SMT trains generative models by estimating mixture scores, outperforming existing methods.
problem Training one-step generative models efficiently and effectively.
method Score-of-Mixture Training (SMT) estimates the score of mixture distributions between real and fake samples.
result SMT/SMD outperform existing methods on CIFAR-10 and ImageNet 64x64 datasets.
Generative Adversarial Networks (GANs) have become a widely popular framework for generative modelling of high-dimensional datasets. However their training is well-known to be difficult. This work presents a rigorous statistical analysis of GANs providing straight-forward explanations for common training pathologies su…
The submanifold quantum mechanics was opened by Jensen and Koppe (Ann. Phys. {\bf 63} (1971) 586-591) and has been studied for these three decades. This article gives its more algebraic definition and show what is the essential of the submanifold quantum mechanics from an algebraic viewpoint.
This expository paper presents elementary proofs of four basic results concerning derivatives of quasi-convex functions. They are combined into a fifth theorem which is simple to apply and adequate in many cases. Along the way we establish the equivalence of the basic lemmas of Jensen and Slodkowski.
We propose a general framework to learn deep generative models via \textbf{V}ariational \textbf{Gr}adient Fl\textbf{ow} (VGrow) on probability spaces. The evolving distribution that asymptotically converges to the target distribution is governed by a vector field, which is the negative gradient of the first variation o…
Paper introduces metrics to evaluate missing data imputation without ground truth.
problem Handling missing data in time series without ground truth.
method Introduces Wasserstein distance (WD) and Jensen-Shannon divergence (JSD) as metrics to evaluate imputation quality.
result WD and JSD are effective metrics for assessing missing data imputation quality.
The study reveals the hierarchical structure of the international FOREX market using currency fluctuation distribution similarities.
problem Understanding the hierarchical structure of the international FOREX market.
method Using Jensen-Shannon divergence to quantify the similarity between normalized logarithmic return distributions of currencies.
result Clusters of currencies are consistent with the nature of underlying economies but diverge during crises.
New taxonomy reveals different detection limits for various types of fraud.
problem Existing fraud detection treats all fraud as the same, ignoring its diverse forms.
method Introduced an observation-mechanism taxonomy with five fraud classes.
result Separate estimation by fraud class outperforms pooled estimation.
New bounds on homological eigenvalues relate to Weil-Petersson length.
problem Bounding growth of homological eigenvalues for pseudo-Anosov automorphisms.
method Established inequality linking homological Jensen square sum to Weil-Petersson translation length.
result Homological Jensen square sum grows at most linearly with covering degree compared to Weil-Petersson translation length.
We introduce a new approximation of f-divergences for machine learning.
problem Variational representations of f-divergences for machine learning. method Definition and analysis of Moreau-Yosida approximation of f-divergences with the Wasserstein-1 metric. result Generalization and relaxation of hard Lipschitz constraints in f-divergences. The paper proposes a new framework to generate synthetic data with human-like imperfections to prevent model collapse.
problem Model collapse due to statistical optimization of synthetic data.
method Introduces Prompt-driven Cognitive Computing Framework (PMCSF) with Cognitive State Decoder (CSD) and Cognitive Text Encoder (CTE).
result The framework generates text with cognitive imperfections, reducing maximum drawdown and delivering defensive alpha.
Using Jeff Holman's comments in Quantitative Finance to illustrate 4 critical errors students should learn to avoid: 1) Mistaking tails (4th moment) for volatility (2nd moment), 2) Missing Jensen's Inequality, 3) Analyzing the hedging wihout the underlying, 4) The necessity of a numeraire in finance.
Two methods factor out prior knowledge from low-dimensional embeddings.
problem Visualizing data without considering background knowledge.
method JEDI for tSNE and CONFETTI for any embedding.
result Embeddings reveal meaningful structure hidden by prior knowledge.
Proposes a scalable method for counterfactual prediction using machine learning.
problem De-bias causal estimators with high-dimensional data in observational studies.
method Uses entropy balancing to learn weights minimizing Jensen-Shannon divergence, leading to robust counterfactual predictions.
result Consistent causal estimation if either propensity score or outcome model is correctly specified.
A new framework for offline RL improves policy flexibility and regularity.
problem Lack of environmental interactions in offline RL leads to poor policy performance.
method Proposes a behavior-regularized implicit policy framework with modified policy-matching methods.
result The framework improves policy effectiveness and robustness beyond static datasets.
Improved lower bound for first Dirichlet eigenvalue using variance refinement.
problem Finding a more precise lower bound for the first Dirichlet eigenvalue.
method Refined Jensen-Hölder averaging using variance term.
result Explicit closed-form in-diameter bound strictly stronger than previous estimates.
Study bounds for European basket call options in a discrete-time market model with price jumps.
problem Bounding the prices of European basket call options in a market model with price jumps.
method Computed bounds using a binomial model and proved that the lower bound coincides with Jensen's bound.
result The upper bound of the price interval of European basket call options can be computed by restricting to a binomial model.
Computing approximate nearest neighbors in high dimensional spaces is a central problem in large-scale data mining with a wide range of applications in machine learning and data science. A popular and effective technique in computing nearest neighbors approximately is the locality-sensitive hashing (LSH) scheme. In thi…
New insights link RLHF and contrastive learning for better model alignment.
problem Aligning large language models with human values.
method Interpreting RLHF and DPO as contrastive learning methods based on mutual information.
result Proposed Mutual Information Optimization (MIO) improves model performance.
Paper resolves bias in ALFT training using generalized alignment games.
problem Systematic bias in estimating logarithmic rewards from small batches.
method Generalized Distributional Alignment Games, U-statistics, minimax polynomial estimators, Variance-Optimal Augmented Polynomial Optimization Program (AQP) Estimator.
result Proves optimal bias and accelerated convergence in ALFT training.
Bayesian Attention Networks compress data by focusing on key training samples.
problem Lossless data compression for efficiency.
method Bayesian Attention Networks with attention factors and latent space.
result Efficient prediction using a few correlated training samples.
PolyGraph Discrepancy improves graph generative model evaluation.
problem Inability of existing metrics to provide an absolute performance measure and comparability across different graph descriptors.
method Approximates Jensen-Shannon distance using binary classifiers trained to distinguish between real and generated graphs.
result PGD provides a more robust and insightful evaluation compared to MMD metrics.
To achieve a high learning accuracy, generative adversarial networks (GANs) must be fed by large datasets that adequately represent the data space. However, in many scenarios, the available datasets may be limited and distributed across multiple agents, each of which is seeking to learn the distribution of the data on …