Optimizes algorithms for non-concave bandit problems.
problem Optimizing algorithms for non-concave bandit problems.
method Unified zeroth-order optimization paradigm.
result Minimax-optimal algorithms in the dimension for low-rank generalized linear bandit problems.
The paper analyzes portfolio selection with non-concave utility and transaction costs.
problem Non-concave utility maximization with proportional transaction costs.
method Two-step procedure: asymptotic terminal behavior analysis and discontinuous viscosity solution.
result Optimal portfolio strategies can differ significantly from the frictionless case due to transaction costs.
We treat a discrete-time asset allocation problem in an arbitrage-free, generically incomplete financial market, where the investor has a possibly non-concave utility function and wealth is restricted to remain non-negative. Under easily verifiable conditions, we establish the existence of optimal portfolios.
Optimizes investment under uncertain time horizons with non-concave utility.
problem Optimizing investment decisions with non-concave utility and uncertain time horizons.
method Established necessary and sufficient conditions for optimality, suggested recursive procedure for non-concave utility.
result Optimal investment strategies under uncertain time horizons exhibit multimodal distribution, indicating flexibility in switching between local maximizers.
We consider non-concave and non-smooth random utility functions with do- main of definition equal to the non-negative half-line. We use a dynamic pro- gramming framework together with measurable selection arguments to establish both the no-arbitrage condition characterization and the existence of an optimal portfolio i…
We study a non-concave optimization problem in which a financial company maximizes the expected utility of the surplus under a risk-based regulatory constraint. For this problem, we consider four different prevalent risk constraints (Expected Shortfall, Expected Discounted Shortfall, Value-at-Risk, and Average Value-at…
This paper investigates the problem of maximizing expected terminal utility in a (generically incomplete) discrete-time financial market model with finite time horizon. In contrast to the standard setting, a possibly non-concave utility function U is considered, with domain of definition R. Simple conditio…
Paper analyzes adversarial dynamics in neural networks.
problem Vulnerability of neural networks to adversarial perturbations.
method Analyzed the dynamics of maximization step in adversarial training.
result Projected gradient ascent finds a local maximum in polynomial iterations.
A convex surface contracting by a strictly monotone, homogeneous degree one function of curvature remains smooth until it contracts to a point in finite time, and is asymptotically spherical in shape. No assumptions are made on the concavity of the speed as a function of principal curvatures.
We solve S-shaped utility portfolio selection with SD constraints using algorithms and neural networks.
problem Optimizing portfolios with S-shaped utility functions under SD constraints.
method First-order SD constraint solution, numerical algorithm for SSD, neural network approach.
result Effective numerical and neural network solutions for SSD constrained problems.
We study a wide class of non-convex non-concave min-max games that generalizes over standard bilinear zero-sum games. In this class, players control the inputs of a smooth function whose output is being applied to a bilinear zero-sum game. This class of games is motivated by the indirect nature of the competition in Ge…
New framework for ranking distributions using variable fractional parameters.
problem Ordering distributions with varying steepness and local non-concavities.
method Introducing a function γ:Ro[0,1] to replace the fixed parameter in fractional SD. result Enables ranking of a broader range of distributions and incorporates dynamic greediness.
In this paper, we focus on solving a class of constrained non-convex non-concave saddle point problems in a decentralized manner by a group of nodes in a network. Specifically, we assume that each node has access to a summand of a global objective function and nodes are allowed to exchange information only with their n…
PAPAL algorithm finds mixed Nash equilibria in continuous games.
problem Finding mixed Nash equilibria in non-convex, non-concave games.
method Particle-based Primal-Dual Algorithm (PAPAL) for weakly entropy-regularized min-max optimization.
result PAPAL offers non-asymptotic convergence guarantees for ε-mixed Nash equilibrium. Novel framework for portfolio selection considering utility and risk.
problem Maximizing utility subject to risk constraints with various utility and risk functionals.
method General framework accommodating non-concave utilities and non-convex risk measures. Characterization of well-posedness using a simple either-or criterion.
result Minimal condition for well-posedness: either utility or risk must be sensitive to large losses.
The paper analyzes how optimization algorithms affect the generalization of minimax models.
problem The generalization performance of minimax models trained with different optimization algorithms.
method Analysis of gradient descent ascent (GDA) and proximal point method (PPM) algorithms under convex concave and non-convex non-concave settings.
result The PPM algorithm ensures a bounded excess risk in convex concave problems, while GDA's generalization depends on solving subproblems simultaneously.
Gradient Descent Ascent converges to von-Neumann solution in hidden zero-sum games.
problem Understanding dynamics of zero-sum games with hidden structure.
method Gradient Descent Ascent applied to hidden zero-sum games with specific convex-concave structure.
result Gradient Descent Ascent converges to von-Neumann solution in strictly convex-concave hidden games.
Study optimal control strategy for hedge funds managers with PSAHARA utility family.
problem Optimizing risk and reward in incomplete markets with non-monotone risk aversion and convex compensation.
method Introduced PSAHARA utility family to model non-monotone risk aversion and convex compensation. Proved concavification techniques for non-concave utility functions. Derived explicit optimal control strategy.
result PSAHARA utility induces risk-taking behavior even with convex compensation, leading to high returns and volatility.
We consider the problem of maximizing a non-concave Lipschitz multivariate function over a compact domain by sequentially querying its (possibly perturbed) values. We study a natural algorithm designed originally by Piyavskii and Shubert in 1972, for which we prove new bounds on the number of evaluations of the functio…
Framework for robust control under model uncertainty, improving financial derivatives hedging.
problem Model uncertainty in financial derivatives hedging.
method Dynamic programming principle for solving one-step optimization problems.
result Robust hedging strategy outperforms model-based strategies during adverse scenarios.
Develops a new parabolic equation for surfaces, proving long-time existence and convergence.
problem Extending elliptic equations to parabolic settings for surfaces.
method Introduces a parabolic analogue of the elliptic split-type Monge-Ampère equation.
result Proves long-time existence and convergence conditions for the new equation.
In this paper, we first investigate the flow of convex surfaces in the space form R3(κ) (κ=0,1,−1) expanding by F−α, where F is a smooth, symmetric, increasing and homogeneous of degree one function of the principal curvatures of the surfaces and the power α∈(0,1] for κ=0,−1 and α=1 for κ=1…
This paper studies the optimal risk-averse timing to sell a risky asset. The investor's risk preference is described by the exponential, power, or log utility. Two stochastic models are considered for the asset price -- the geometric Brownian motion and exponential Ornstein-Uhlenbeck models -- to account for, respectiv…
Study nonconcave portfolio choice with smooth ambiguity and Bayesian learning.
problem Nonconcave portfolio choice under smooth ambiguity and Bayesian learning.
method Developed a general framework for dynamic, non-concave asset allocation.
result Dynamic consistency achieved through a robust representation.
In this paper, we propose a novel reinforcement- learning algorithm consisting in a stochastic variance-reduced version of policy gradient for solving Markov Decision Processes (MDPs). Stochastic variance-reduced gradient (SVRG) methods have proven to be very successful in supervised learning. However, their adaptation…
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.
This work analyzes the value of future reward information in RL.
problem Analyzing the impact of knowing future rewards in reinforcement learning.
method Competitive analysis and worst-case reward distribution.
result Exact ratios between standard RL agents and those with future-reward lookahead.
Self-supervised reward prediction improves RL in sparse reward settings.
problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.
Develops deep learning methods for solving S-shaped utility maximisation problems.
problem Optimizing portfolios with S-shaped utility and random benchmarks.
method Uses deep learning and duality methods to solve the Hamilton-Jacobi-Bellman equation and adjoint equation.
result Demonstrates the accuracy of deep learning methods for non-concave utility maximisation problems.
This study examines how earnings announcements affect option volatility and pricing.
problem The impact of earnings announcements on option volatility and pricing.
method Analysis of extremely short-term options data to study bimodality and concavity in IV curves.
result Investors pay a premium to hedge against extreme volatility during earnings announcements in the presence of concave IV smiles.
Reward models need more than just accuracy for effective RLHF.
problem The effectiveness of reward models in RLHF is not fully understood.
method An optimization perspective to evaluate reward models.
result Reward models with low reward variance can lead to a flat optimization landscape, hindering performance.
Action guidance helps agents learn true objectives in games with sparse rewards.
problem Training agents in games with sparse rewards requires significant exploration.
method Action guidance, a novel technique that combines exploration with reward shaping.
result Action guidance enables agents to optimize true objectives efficiently.
Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.
problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.
Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …
Enhances reward specification in RL with a novel language-based approach.
problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.
Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.
problem Optimizing behavior for long-term goals in reinforcement learning with unstable long planning horizons.
method Reward tweaking learns a surrogate reward function that induces optimal behavior for the original task.
result Reward tweaking guides agents towards better long-term returns while planning for short horizons.
We propose a generic, Bayesian, information geometric approach to the exploration--exploitation trade-off in multi-armed bandit problems. Our approach, BelMan, uniformly supports pure exploration, exploration--exploitation, and two-phase bandit problems. The knowledge on bandit arms and their reward distributions is su…
Many reinforcement-learning researchers treat the reward function as a part of the environment, meaning that the agent can only know the reward of a state if it encounters that state in a trial run. However, we argue that this is an unnecessary limitation and instead, the reward function should be provided to the learn…
Gaussian random fields are a powerful tool for modeling environmental processes. For high dimensional samples, classical approaches for estimating the covariance parameters require highly challenging and massive computations, such as the evaluation of the Cholesky factorization or solving linear systems. Recently, Anit…
Extends reinforcement learning alignment to scalar rewards, improving math reasoning.
problem Designing reinforcement learning algorithms for general LLM alignment.
method Introduces f-GRPO and f-HAL, estimating f-divergences between reward-aligned and unaligned distributions.
result Improves math reasoning RLVR tasks and mitigates reward hacking.
Paper improves adversarial training using a learned optimizer.
problem Improving robustness of deep learning models against adversarial attacks.
method Empirically identified PGD attack's limitations and used a learning-to-learn framework to train an adaptive inner optimizer.
result The proposed framework consistently improves model robustness over traditional adversarial training methods.
This work characterizes reward function partial identifiability and its impact on policy optimization.
problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.
This paper introduces a new reward shaping method for average-reward reinforcement learning.
problem Speeding up convergence to an optimal policy in average-reward reinforcement learning tasks.
method Developed a temporal logic-based approach to automatically generate reward shaping functions.
result The optimal policy can be recovered using the proposed reward shaping framework.
The paper examines how background risk affects portfolio selection and optimal reinsurance design.
problem Maximizing the probability of reaching a financial goal in the presence of background risk.
method Quantile formulation method to derive optimal solutions explicitly.
result The presence of background risk does not change the solution shape but alters the parameter values.
While recent progress in deep reinforcement learning has enabled robots to learn complex behaviors, tasks with long horizons and sparse rewards remain an ongoing challenge. In this work, we propose an effective reward shaping method through predictive coding to tackle sparse reward problems. By learning predictive repr…