The paper tackles rested bandits with non-decreasing and concave rewards, deriving lower bounds and an efficient algorithm.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Numerical observations on martingale couplings are confirmed under certain conditions.
Optimal insurance minimizes ruin probability with non-decreasing functions.
Study personalizes user experience to maximize rewards with patience budget.
New inequalities for convex hypersurfaces in various spaces.
It is a well-known fact that on a bounded spectral interval the Dirac spectrum can be described locally by a non-decreasing sequence of continuous functions of the Riemannian metric. In the present article we extend this result to a global version. We think of the spectrum of a Dirac operator as a function from the int…
We define an explicit quasi-local mass functional which is non-decreasing along all foliations (satisfying a convexity assumption) of null cones. We use this new functional to prove the null Penrose conjecture under fairly generic conditions.
We assume that an agent's rate of consumption is {\it ratcheted}; that is, it forms a non-decreasing process. Given the rate of consumption, we act as financial advisers and find the optimal investment strategy for the agent who wishes to minimize his probability of ruin.
We propose a global invariant for contact manifolds which admit a strictly pseudoconvex CR structure, analogous to the Yamabe invariant . We prove that this invariant is non-decreasing under handle attaching and under connected sum. We then give a lower bound on in a particular case.
Let be a complete Riemannian manifold possessing a strictly convex Lipschitz continuous exhaustion function. We show that the isoperimetric profile of is a continuous and non-decreasing function. Particular cases are Hadamard manifolds and complete non-compact manifolds with strictly positive sectional curvatur…
In this paper, we study the relation of the monotonicity of Hawking Mass and geometric flow problems. We show that along the Hamilton-DeTurck flow with bounded curvature coupled with the modified mean curvature flow, the Hawking mass of the hypersphere with a sufficiently large radius in Schwarzschild spaces is monoton…
Let be a compact Kähler manifold of dimension and fix . We prove that the total mass of the complex Hessian measure of --subharmonic functions is non-decreasing with respect to the singularity type. We then solve complex Hessian equations with prescribed singularity, and prove a Hodge i…
We consider an arbitrage-free, discrete time and frictionless market. We prove that an investor maximising the expected utility of her terminal wealth can always find an optimal investment strategy provided that her dissatisfaction of infinite losses is infinite and her utility function is non-decreasing, continuous an…
We define an invariant of contact structures in dimension three from Heegaard Floer homology. This invariant takes values in the set . It is zero for overtwisted contact structures, for Stein fillable contact structures, non-decreasing under Legendrian surgery, and computable …
We establish interior Lipschitz regularity for continuous viscosity solutions of fully nonlinear, conformally invariant, degenerate elliptic equations. As a by-product of our method, we also prove a weak form of the strong comparison principle, which we refer to as the principle of propagation of touching points, for o…
In this paper, we use the distance comparison principle, first been developed by G. Huisken, to study the spatial curve shortening flow. We have got the result that if the initial curve is the helix, then the local minimum of the ratio of the extrinsic and intrinsic distance is non-decreasing. And we have proved a Gray…
Paper analyzes convergence rates for multi-agent learning in games.
We study the restless bandit associated with an extremely simple scalar Kalman filter model in discrete time. Under certain assumptions, we prove that the problem is indexable in the sense that the Whittle index is a non-decreasing function of the relevant belief state. In spite of the long history of this problem, thi…
We show that every non-decreasing function bounded from above by for some can be realized (up to a natural equivalence) as the conjugacy growth function of a finitely generated group. We also construct a finitely generated group and a subgroup of index 2 such…
In this note, we study Liouville type theorem for conformal Gaussian curvature equation (also called the mean field equation) where is a smooth function on . When is a sign-changing smooth function in the real line , we have a non-existence result for the finite to…
We develop a new approach to the existence of time functions on Lorentzian manifolds, based on Conley's work regarding Lyapunov functions for dynamical systems. We recover Hawking's result that a stably causal admits a time function through a more general result giving the existence of a continuous function that is non…
Bartnik mass is positive and non-decreasing for black holes
We study a simple problem that arises from the study of Lorentz surfaces and Anosov flows. For a non decreasing map of degree one , we are interested in groups of circle diffeomorphisms that act on the complement of the graph of in by preserving a vo…
The Cheeger constant increases under Ricci flow on spheres.
It is conjectured that the full (spacetime) Bartnik mass of a surface is realised as the ADM mass of some stationary asymptotically flat manifold with boundary data prescribed by . Assuming this holds true for a 1-parameter family of surfaces evolving in an initial data set {with the dominant energy condit…
Paper proves a generalized Alexandrov-Fenchel inequality for convex hypersurfaces with capillary boundary.
In this paper we obtain a splitting theorem for the symmetric diffusion operator and a non-constant function in a complete Riemannian manifold , under the assumptions that the Ricci curvature associated with satisfies , that $|…
Assuming that agents' preferences satisfy first-order stochastic dominance, we show how the Expected Utility paradigm can rationalize all optimal investment choices: the optimal investment strategy in any behavioral law-invariant (state-independent) setting corresponds to the optimum for an expected utility maximizer w…
We describe a novel family of models of multi- layer feedforward neural networks in which the activation functions are encoded via penalties in the training problem. Our approach is based on representing a non-decreasing activation function as the argmin of an appropriate convex optimiza- tion problem. The new framewor…
We introduce a new entropy functional for nonnegative solutions of the heat equation on a manifold with time-dependent Riemannian metric. Under certain integral assumptions, we show that this entropy is non-decreasing, and moreover convex if the metric evolves under super Ricci flow (which includes Ricci flow and fixed…
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
This work analyzes the value of future reward information in RL.
Self-supervised reward prediction improves RL in sparse reward settings.
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
Reward models need more than just accuracy for effective RLHF.
Proposes a method to boost deep reinforcement learning with sparse rewards.
Action guidance helps agents learn true objectives in games with sparse rewards.
New RL method uses distance between states instead of rewards for sparse reward environments.
Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.
Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …
Enhances reward specification in RL with a novel language-based approach.
Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.
We propose a generic, Bayesian, information geometric approach to the exploration--exploitation trade-off in multi-armed bandit problems. Our approach, BelMan, uniformly supports pure exploration, exploration--exploitation, and two-phase bandit problems. The knowledge on bandit arms and their reward distributions is su…
Many reinforcement-learning researchers treat the reward function as a part of the environment, meaning that the agent can only know the reward of a state if it encounters that state in a trial run. However, we argue that this is an unnecessary limitation and instead, the reward function should be provided to the learn…
Extends reinforcement learning alignment to scalar rewards, improving math reasoning.
This work characterizes reward function partial identifiability and its impact on policy optimization.