Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

306090120 · May 202619922001200920182026
48 results for φ-divergence penalties

Bayesian priors and penalties are equivalent in variational inference.

problem Understanding the relationship between Bayesian priors and penalties in variational inference.
method Characterizing the regularizers that can arise in variational inference and providing a systematic way to compute the prior corresponding to a given penalty.
result Equivalence between Bayesian priors and penalties in variational inference.

Paper proposes f-DPG for aligning language models with preferences.

problem Aligning language models with user preferences.
method Uses f-divergence to approximate target distributions and minimizes a forward KL from it using DPG.
result Jensen-Shannon divergence often outperforms forward KL divergence, leading to significant improvements.

Unified framework for reinforcement learning using entropic regularization.

problem Learning optimal controllers for unknown MDPs without diverging to dangerous regions.
method Entropic proximal policy optimization with αα-divergences.
result Unified perspective on actor-critic architectures and asymptotic analysis of solutions.

DPEs use KL divergence to approximate BNNs, improving uncertainty estimates for active learning.

problem Improving uncertainty estimates in active learning for visual classification.
method Regularized ensemble approach with KL divergence penalty for variational inference.
result DPEs steadily improve active learning performance with increased annotation budgets.

A new method for non-negative matrix factorization using generalized dual divergence.

problem Non-negative matrix factorization for various noise structures.
method Theoretical framework based on generalized dual Kullback-Leibler divergence, with algorithms developed and proven convergence using Expectation-Maximization.
result Generalizes existing methods and provides an alternative for non-negative matrix factorizations.

Optimal transport with ff-divergence regularization using generalized Sinkhorn algorithm.

problem Optimal transport with ff-divergence regularization.
method Generalized Sinkhorn algorithm for solving optimal transport problems with various ff-divergences.
result Strong duality holds, optimums are attained, and convergence to an optimal solution is guaranteed under certain conditions.

There has been a growing interest in mutual information measures due to their wide range of applications in Machine Learning and Computer Vision. In this paper, we present a generalized structured regression framework based on Shama-Mittal divergence, a relative entropy measure, which is introduced to the Machine Learn…

2014-09-26abs ↗pdf ↗

Study extends DRO with IPMs, linking robustness to regularization and GANs.

problem Addressing robustness of deep neural networks to adversarial attacks.
method Distributionally Robust Optimization (DRO) with Integral Probability Metrics (IPMs).
result DRO under any IPM corresponds to a family of regularization penalties.

This paper explores policy improvement using various f-divergences, enhancing stability in reinforcement learning.

problem Ensuring stability in reinforcement learning algorithms through policy improvement with trust regions.
method The paper considers a general class of f-divergences and derives policy update rules, including the KL divergence as a special case.
result The study reveals different policy updates and evaluations for various f-divergences, including Pearson χ2χ^2-divergence and KL divergence.

Develops a statistical learning framework for personalized asset allocation.

problem Continuous-action decision-making with a large number of characteristics.
method Discretization approach with generalized penalties for penalized regression.
result Improves financial well-being with individualized optimal asset allocation.

The paper proposes a method to select tuning parameters for high-dimensional data analysis.

problem Selecting the tuning parameter in penalized likelihood methods for high-dimensional data.
method Optimizing the generalized information criterion (GIC) with an appropriate model complexity penalty.
result The proposed model complexity penalty should diverge at the rate of some power of log p.

The paper introduces MU for NMF with ββ-divergences and disjoint constraints.

problem Nonnegative matrix factorization with constraints.
method Design multiplicative updates for NMF based on ββ-divergences with disjoint constraints.
result Multiplicative updates satisfy constraints and decrease the objective function.

New algorithms improve robust estimation in contaminated Gaussian models.

problem Simultaneous estimation of location and variance matrix in contaminated Gaussian models.
method Tractable adversarial algorithms with spline discriminators for robust estimation.
result Achieve minimax optimal rates or near-optimal rates under Huber's contamination model.

LSGANs improve GANs by using least squares loss, leading to better image quality and stability.

problem Vanishing gradients in GANs during training.
method Introducing LSGANs with least squares loss for both discriminator and generator.
result LSGANs generate higher quality images and are more stable during training.

New RL algorithm learns good actions from offline data, reducing uncertainty and divergence.

problem Limited applicability of current RL algorithms in real-world settings due to high costs of exploration.
method Proposes an algorithm for batch RL using a fixed offline dataset, with penalties for policy and value constraints.
result Compared favorably to state-of-the-art methods on 32 continuous-action benchmarks.

Study EM and GD for clustering with penalties for misspecification and high dimensions.

problem Clustering with misspecification and high-dimensional data.
method Model-based Gaussian Mixture Models, EM algorithm, GD optimization with AD, penalized likelihood.
result GD outperforms EM on high-dimensional data but both have poor cluster interpretation.

Generative models often misrepresent class frequencies; this paper calibrates them.

problem Miscalibration of class frequencies in generative models.
method Formulated as constrained optimization, using surrogate objectives to approximate constraints.
result Significant reduction in calibration error across various models and applications.

The paper calibrates robust optimization models to reduce sensitivity to model errors.

problem Reducing sensitivity of expected reward to model errors in empirical optimization.
method Develops a theory for data-driven calibration of robustness parameter δ using resampling methods.
result Substantial variance reduction is possible at little cost if δ is properly calibrated.

VED framework learns low-dimensional latent representations of physical systems.

problem Learning latent representations of complex physical systems.
method Variational Encoder-Decoder (VED) framework with KL divergence and covariance regularization.
result VED achieves lower-dimensional latent representations with improved feature disentanglement.

Momentum SGD fails to track nonstationary optima due to drift amplification.

problem Tracking nonstationary optima in stochastic optimization.
method Theoretical analysis of SGD and momentum variants under strong convexity and smoothness.
result Momentum incurs a drift-amplification penalty that diverges as the momentum parameter approaches 1, leading to systematic lag.

SAIL-RevKL improves SAIL's convergence by regularizing the objective function.

problem Convergence of self-improving online LLM alignment algorithms.
method Proposed SAIL-RevKL, a regularized objective function to improve optimization landscape.
result Proved SAIL-RevKL satisfies the Polyak-Lojasiewicz (PL) condition with near-linear sample complexity.

Binacox detects multiple cut-points in high-dimensional Cox models for genetic cancer data.

problem Detecting multiple cut-points in high-dimensional Cox models with many continuous features.
method Combines one-hot encoding with binarsity penalty for feature selection and regularization.
result Significantly outperforms state-of-the-art survival models in terms of C-index and computational speed.

Bayesian learning is often hampered by large computational expense. As a powerful generalization of popular belief propagation, expectation propagation (EP) efficiently approximates the exact Bayesian computation. Nevertheless, EP can be sensitive to outliers and suffer from divergence for difficult cases. To address t…

2012-04-18abs ↗pdf ↗

This paper shows RL with KL penalties is equivalent to Bayesian inference for fine-tuning LMs.

problem Fine-tuning large language models to avoid undesirable features.
method Analyzed KL-regularized RL and showed it's equivalent to variational inference.
result KL-regularized RL avoids distribution collapse and is more insightful as Bayesian inference.

We propose a new method of learning a sparse nonnegative-definite target matrix. Our primary example of the target matrix is the inverse of a population covariance or correlation matrix. The algorithm first estimates each column of the target matrix by the scaled Lasso and then adjusts the matrix estimator to be symmet…

2012-02-13abs ↗pdf ↗

Paper proposes algorithms for robust 1-bit compressive sensing with nonconvex penalties.

problem Recovering sparse signals from one-bit measurements.
method Develops algorithms based on convex and nonconvex penalties, providing analytical solutions.
result Analytical solutions for several nonconvex penalties are found, making the recovery process faster and more efficient.

Gradient penalty improves GAN performance by inducing a large-margin classifier.

problem Improving GAN performance and addressing vanishing gradients.
method A unifying framework of expected margin maximization, showing gradient penalties induce large-margin classifiers.
result Gradient penalties reduce vanishing gradients and produce better generated outputs.

Study examines insider trading with penalties, finding optimal penalties increase quickly for small orders.

problem Analyzing the impact of penalties on insider trading behavior and market efficiency.
method Formal economic model with penalty functions, existence and uniqueness theorems, and optimization.
result Optimal penalties increase quickly for small orders, signaling extreme events and incorporating information into prices.

A new approach to risk-sensitive reinforcement learning tackles computational challenges.

problem Computational challenges in estimating risk-sensitive policies for MDPs with finite state and action spaces.
method Proposes a new risk measure called 'caution' and uses a stochastic primal-dual method with KL divergence.
result Demonstrates improved reliability in reward accumulation without additional computational costs.

The paper studies robust risk measures with linear penalties under uncertain distributions.

problem Risk measurement under distributional uncertainty.
method Robust distortion risk measures with linear penalty function under distributional constraints.
result Explicit characterization of optimal quantile distribution and value function.