A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
In real-world decision-making problems, for instance in the fields of finance, robotics or autonomous driving, keeping uncertainty under control is as important as maximizing expected returns. Risk aversion has been addressed in the reinforcement learning literature through risk measures related to the variance of retu…
We present a fully nonparametric method to estimate the value function, via simulation, in the context of expected infinite-horizon discounted rewards for Markov chains. Estimating such value functions plays an important role in approximate dynamic programming and applied probability in general. We incorporate "soft in…
We study the problem of stopping a Brownian motion at a given distribution ν while optimizing a reward function that depends on the (possibly randomized) stopping time and the Brownian motion. Our first result establishes that the set T(ν) of stopping times embedding ν is weakly dense in the set $\mathc…
We study the optimal transport between two probability measures on the real line, where the transport plans are laws of one-step martingales. A quasi-sure formulation of the dual problem is introduced and shown to yield a complete duality theory for general marginals and measurable reward (cost) functions: absence of a…
In [S. Basu, A. Gabrielov, N. Vorobjov, Semi-monotone sets. arXiv:1004.5047v2 (2011)] we defined semi-monotone sets, as open bounded sets, definable in an o-minimal structure over the reals, and having connected intersections with all translated coordinate cones in R^n. In this paper we develop this theory further by d…
Policy Gradient (PG) algorithms are among the best candidates for the much-anticipated applications of reinforcement learning to real-world control tasks, such as robotics. However, the trial-and-error nature of these methods poses safety issues whenever the learning process itself must be performed on a physical syste…
Entropy regularization improves policy optimization in reinforcement learning.
problem Improving policy optimization in reinforcement learning.
method Entropy regularization is introduced to soften the greedy policy towards a more diverse softmax policy, leading to a continuously parameterized algorithm that interpolates between policy gradient and Q-learning.
result An intermediate algorithm can improve performance in reinforcement learning.
We present results on simulations of a stock market with heterogeneous, cumulative information setup. We find a non-monotonic behaviour of traders' returns as a function of their information level. Particularly, the average informed agents underperform random traders; only the most informed agents are able to beat the …
In the paper, the author studies properties of three functions relating to the exponential function and the existence of partitions of unity, including accurate and explicit computation of their derivatives, analyticity, complete monotonicity, logarithmically complete monotonicity, absolute monotonicity, and the like.
In this paper, we study monotonicity formulas of eigenvalues and entropies along the rescaled List's extended Ricci flow. We derive some monotonicity formulas of eigenvalues of Laplacian which generalize those of Li in [8] and Cao-Hou-Ling in [3]. Moreover, we also consider monotonicity formulas of Fk-func…
In this paper we generalize the monotonicity formulas of [C] for manifolds with nonnegative Ricci curvature. Monotone quantities play a key role in analysis and geometry; see, e.g., [A], [CM1] and [GL] for applications of monotonicity to uniqueness. Among the applications here is that level sets of Green's function on …
Higher conservative training increases reward-hacking in reasoning models.
problem Reward hacking during online adaptation in reasoning models.
method Conservative offline training with varying levels of conservatism (β) was applied to a Qwen3-14B policy, and online adaptation was measured against a reward ensemble.
result Higher conservatism (β) increases reward-hacking damage, measured by the Goodhart gap and AUGC.
We study, to the best of our knowledge, the first Bayesian algorithm for unimodal Multi-Armed Bandit (MAB) problems with graph structure. In this setting, each arm corresponds to a node of a graph and each edge provides a relationship, unknown to the learner, between two nodes in terms of expected reward. Furthermore, …
The paper is motivated by a problem concerning the monotonicity of insurance premiums with respect to their loading parameter: the larger the parameter, the larger the insurance premium is expected to be. This property, usually called loading monotonicity, is satisfied by premiums that appear in the literature. The inc…
We propose learning deep models that are monotonic with respect to a user-specified set of inputs by alternating layers of linear embeddings, ensembles of lattices, and calibrators (piecewise linear functions), with appropriate constraints for monotonicity, and jointly training the resulting network. We implement the l…