Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,341 papers · 148 categories

Trend · papers per month

285684112 · May 202619922001200920182026
48 results for Entropy penalty

Insider trading is reduced when penalized, affecting expected penalties in a non-monotone way.

problem Reducing insider trading behavior when insiders face legal penalties.
method Characterized via a backward stochastic differential equation (BSDE) with a non-linear operator.
result The insider's expected penalties are non-monotone in the fee structure and determined by relative entropy.

In an incomplete Brownian-motion market setting, we propose a convex monotonic pricing functional for nonattainable bounded contingent claims which is compatible with prices for attainable claims. The pricing functional is defined as the convex conjugate of a generalized entropy penalty functional and an interpretation…

2008-04-01abs ↗pdf ↗

This paper reformulates FβF_β for better model performance and interpretation.

problem Optimizing model performance and interpretation using FβF_β metric.
method Reformulate FβF_β metric to facilitate statistical distributions and dynamic penalty weights.
result Better and interpretable results with a 14% boost in F1F_1 score for IMDB data.

Proposes a neural generator network for efficient energy-based model learning.

problem Challenges in maximum likelihood estimation of energy-based models.
method Uses a neural generator network to approximate the log-likelihood gradient and maximizes entropy of generated samples.
result Generates sharp images with competitive Inception and FID scores, and is robust to mode collapse.

ProSelfLC improves robustness of deep neural networks by automatically deciding trust in predictions.

problem Training robust deep neural networks requires addressing issues like label noise and low entropy predictions.
method ProSelfLC progressively increases trust in predicted labels over time, considering entropy and learning time.
result ProSelfLC demonstrates improved robustness in both clean and noisy settings through empirical validation.

New scalable algorithm for non-negative linear regression with entropy-regularized OT loss.

problem Generalizing task-specific linear models to broader applications.
method Sinkhorn-like scaling iterations for convex penalty and datafit terms.
result Simple multiplicative updates for various penalty and datafit terms.

Study risk-sensitive market making with entropy regularization for better quote control.

problem Risk-sensitive market making with exponential utility and penalties.
method Entropy-regularized certainty-equivalent Bellman policies for discrete-time market dynamics.
result Proves convergence and performance bounds for entropy-regularized policies.

Study optimal reinsurance pricing under model uncertainty for multiple insurers.

problem Optimal reinsurance pricing in the presence of multiple sources of model uncertainty.
method Solves a continuous-time Stackelberg game for general reinsurance contracts, considering entropy penalties and ambiguity in insurers' models.
result Reinsurer prices under a distortion of the barycentre of insurers' models, maximizing expected wealth with an entropy penalty.

Generative AI connects to Schrödinger bridge problems with soft constraints for stability.

problem Stability issues in generative AI due to hard terminal constraints.
method Soft-constrained Schrödinger bridge formulation and convergence analysis.
result Existence and convergence of optimal solutions as penalty grows.

Bayesian optimization method predicts high costs for unstable robot controllers.

problem Time-consuming and challenging learning robot controllers with unknown penalties.
method Proposes a Bayesian model that predicts high costs in unstable regions.
result Improves robot learning by guiding exploration toward stable regions.

Recent contributions have framed linear system identification as a nonparametric regularized inverse problem. Relying on 2\ell_2-type regularization which accounts for the stability and smoothness of the impulse response to be estimated, these approaches have been shown to be competitive w.r.t classical parametric met…

2015-08-12abs ↗pdf ↗

This paper introduces a method to incorporate risk sensitivity in RL using quadratic variation penalties.

problem Risk-sensitive reinforcement learning under entropy regularization.
method Equivalent martingale property and quadratic variation penalty for value process.
result The proposed method improves finite-sample performance in linear-quadratic control problems.

The paper sets lower bounds for adversarial robustness in multiclass classification.

problem Adversarial robustness in multiclass classification with arbitrary loss functions.
method Dual and barycentric reformulations for robust risk minimization.
result Sharp lower bounds for adversarial risks are computed efficiently.

Paper introduces ENZ to measure significant coefficients in sparse recovery, improving over classical methods.

problem Numerical noise creates long tails of negligible coefficients in sparse recovery.
method Entropy-based notion of effective sparsity (ENZ) to measure significant coefficients, proving stability under restricted isometry condition.
result ENZ decomposes into support cardinality and efficiency factor, providing a precise measure of sparsity.

Develops a new duality between entropy martingale optimal transport and nonlinear pricing-hedging.

problem Entropy Martingale Optimal Transport problem and its associated optimization problem.
method Combines Entropy Optimal Transport and Martingale Optimal Transport theories, with novel penalization terms and constraints.
result Establishes a nonlinear robust pricing-hedging duality, covering various known robust results.

Tree tensor networks balance model complexity and empirical risk for high-dimensional function approximation.

problem Selecting optimal tree structure and ranks for high-dimensional function approximation.
method Proposes a complexity-based model selection method for tree tensor networks in empirical risk minimization.
result Demonstrates near-minimax adaptive performance across various smoothness classes.

Detecting and recovering labels in binomial logistic mixtures is challenging due to an information gap.

problem Detecting and recovering labels in binomial logistic mixtures
method Propose two feasibility-aware inference procedures
result Avoid misleading component selections and improve label probability calibration

The paper derives oracle inequalities for estimators with fast and slow rates.

problem Developing fast and slow oracle inequalities for estimators.
method Direct study of analysis estimator and adaptation of Dalalyan, Hebiri and Lederer's arguments.
result Constant-friendly rates for (square root) total variation regularized estimators over graphs.

There has been a growing interest in mutual information measures due to their wide range of applications in Machine Learning and Computer Vision. In this paper, we present a generalized structured regression framework based on Shama-Mittal divergence, a relative entropy measure, which is introduced to the Machine Learn…

2014-09-26abs ↗pdf ↗

Paper introduces a new method for risk-sensitive investment management using RL.

problem Risk-sensitive portfolio management with unknown model parameters.
method Combines RL and risk-sensitive stochastic control with Gaussian perturbations for exploration.
result Endogenous relative-entropy regularization and optimal investment strategy derived.

LSGANs improve GANs by using least squares loss, leading to better image quality and stability.

problem Vanishing gradients in GANs during training.
method Introducing LSGANs with least squares loss for both discriminator and generator.
result LSGANs generate higher quality images and are more stable during training.

Paper proposes algorithms for robust 1-bit compressive sensing with nonconvex penalties.

problem Recovering sparse signals from one-bit measurements.
method Develops algorithms based on convex and nonconvex penalties, providing analytical solutions.
result Analytical solutions for several nonconvex penalties are found, making the recovery process faster and more efficient.

Gradient penalty improves GAN performance by inducing a large-margin classifier.

problem Improving GAN performance and addressing vanishing gradients.
method A unifying framework of expected margin maximization, showing gradient penalties induce large-margin classifiers.
result Gradient penalties reduce vanishing gradients and produce better generated outputs.

Study examines insider trading with penalties, finding optimal penalties increase quickly for small orders.

problem Analyzing the impact of penalties on insider trading behavior and market efficiency.
method Formal economic model with penalty functions, existence and uniqueness theorems, and optimization.
result Optimal penalties increase quickly for small orders, signaling extreme events and incorporating information into prices.

The paper studies robust risk measures with linear penalties under uncertain distributions.

problem Risk measurement under distributional uncertainty.
method Robust distortion risk measures with linear penalty function under distributional constraints.
result Explicit characterization of optimal quantile distribution and value function.

This work extends entropic optimal transport to non-product reference couplings, focusing on Gaussian cases.

problem Finding a diffuse coupling between two measures with non-product reference couplings.
method Reduction of the entropic optimal transport problem to a matrix optimization problem.
result Complete description of the solution for non-product reference couplings, including primal and dual variables.

New approach avoids excess empirical risk in domain generalization.

problem Learning models that generalize to unseen distributions from diverse data sets.
method Minimizes penalty under constraint of optimal empirical risk, leveraging rate-distortion theory.
result Significant improvements in domain generalization performance across multiple methods.

UCPO improves diversity in reinforcement learning models, maintaining high accuracy.

problem RLVR objectives often lead to diversity collapse, reducing coverage of correct solutions.
method UCPO adds a conditional uniformity penalty to GRPO, redistributing probability mass.
result UCPO improves Pass@K and diversity while maintaining competitive Pass@1 accuracy.

Curvature penalties improve interpretability of KANs without sacrificing accuracy.

problem Pathologically high-curvature oscillations in KANs activations make them hard to interpret.
method Derived a curvature penalty and proved an upper bound on model curvature.
result KANs with curvature penalties achieve substantially smoother activations while maintaining accuracy.