New algorithm for fair ranking in contextual bandits with concave rewards.
problem Fair ranking in recommendation systems.
method Geometric interpretation of CBCR as optimization, Frank-Wolfe analyses.
result First algorithm with provably vanishing regret for CBCR.
Optimizes algorithms for non-concave bandit problems.
problem Optimizing algorithms for non-concave bandit problems.
method Unified zeroth-order optimization paradigm.
result Minimax-optimal algorithms in the dimension for low-rank generalized linear bandit problems.
A new algorithm uses concavity in Gaussian processes to optimize decisions in bandit problems.
problem Optimizing decisions in sequential problems with context-dependent rewards.
method Proposes a UCB algorithm using a shape-constrained reward function estimator based on a Gaussian Process model with concavity constraints.
result Derives regret bounds for the proposed UCB algorithm.
The paper tackles rested bandits with non-decreasing and concave rewards, deriving lower bounds and an efficient algorithm.
problem Studying the sample complexity and optimal strategies for rested bandits with specific reward properties.
method Deriving regret lower bounds and designing an efficient algorithm R-ed-UCB with theoretical and empirical analysis.
result An efficient algorithm R-ed-UCB with a regret bound of O ~ ( T 2 3 ) \widetilde{\mathcal{O}}(T^{\frac{2}{3}}) O ( T 3 2 ) under certain conditions. Algorithm tackles constrained reinforcement learning with concave-convex and knapsack constraints.
problem Constrained episodic reinforcement learning with concave rewards and convex constraints.
method Modular analysis with strong theoretical guarantees for concave-convex and knapsack settings.
result Significantly outperforms existing approaches in constrained episodic environments.
This paper studies GAIL's global convergence for general MDP and nonlinear rewards.
problem Understanding when GAIL algorithms achieve global convergence for general MDP and nonlinear rewards.
method Characterization of global convergence for various policy gradient algorithms applied to GAIL.
result First systematic theoretical study of GAIL for global convergence.
The paper provides tight bounds for improving multi-armed bandits problem.
problem Improving multi-armed bandits problem with concave reward functions.
method Upper and lower bounds for randomized online algorithms, providing an O ( k log k ) O(\sqrt{k} \log k) O ( k log k ) approximation. result Achieved nearly-tight approximation guarantees for the improving multi-armed bandits problem.
New algorithm for optimizing statistical utilities in bandits.
problem Optimizing statistical functionals of long-run reward distributions.
method Influence-function calculus for stochastic gradient estimation, entropic mirror-ascent algorithm.
result Regret bounds that separate optimization and estimation errors.
We consider the classical problem of sequential resource allocation where a decision maker must repeatedly divide a budget between several resources, each with diminishing returns. This can be recast as a specific stochastic optimization problem where the objective is to maximize the cumulative reward, or equivalently …
This work overcomes bias in concave multi-objective reinforcement learning.
problem Gradient bias in policy gradient methods for concave scalarized multi-objective reinforcement learning.
method Developed a Natural Policy Gradient (NPG) algorithm with a multi-level Monte Carlo (MLMC) estimator.
result Achieved optimal O ~ ( ε − 2 ) \widetilde{\mathcal{O}}(ε^{-2}) O ( ε − 2 ) sample complexity for computing an ε ε ε -optimal policy. New star-shaped acceptability indexes generalize existing methods.
problem Generalizing existing acceptability measures.
method Characterizing acceptability indexes through star-shaped risk measures and sets.
result Introducing concrete examples linked to various financial measures.
We study the global convergence of generative adversarial imitation learning for linear quadratic regulators, which is posed as minimax optimization. To address the challenges arising from non-convex-concave geometry, we analyze the alternating gradient algorithm and establish its Q-linear rate of convergence to a uniq…
Despite the success of single-agent reinforcement learning, multi-agent reinforcement learning (MARL) remains challenging due to complex interactions between agents. Motivated by decentralized applications such as sensor networks, swarm robotics, and power grids, we study policy evaluation in MARL, where agents with jo…
We consider an agent who is involved in a Markov decision process and receives a vector of outcomes every round. Her objective is to maximize a global concave reward function on the average vectorial outcome. The problem models applications such as multi-objective optimization, maximum entropy exploration, and constrai…
Finding optimal policies which maximize long term rewards of Markov Decision Processes requires the use of dynamic programming and backward induction to solve the Bellman optimality equation. However, many real-world problems require optimization of an objective that is non-linear in cumulative rewards for which dynami…
Generative Adversarial Imitation Learning (GAIL) is a powerful and practical approach for learning sequential decision-making policies. Different from Reinforcement Learning (RL), GAIL takes advantage of demonstration data by experts (e.g., human), and learns both the policy and reward function of the unknown environme…
This paper tackles robust control of noisy systems with uncertain distributions.
problem Optimal control of sampled-data stochastic systems with multiplicative noise and distributional ambiguity.
method Develops a convex relaxation to handle the ``concave-max'' geometry and derives a probabilistic performance guarantee.
result Derives an explicit, non-asymptotic bound on the duality gap and proves robust viability conditions.
Algorithm finds near-optimal VaR portfolios using MILP, improving risk management.
problem Computing optimal VaR portfolios is hard due to non-convexity and combinatorial nature.
method Formulates VaR portfolio problem as MILP, uses alternate formulations for guarantees.
result Near-optimal VaR portfolios with near-optimality guarantees.
We study the out-of-sample properties of robust empirical optimization problems with smooth φ φ φ -divergence penalties and smooth concave objective functions, and develop a theory for data-driven calibration of the non-negative "robustness parameter" δ δ δ that controls the size of the deviations from the nominal model. Bu…
We consider a contextual version of multi-armed bandit problem with global knapsack constraints. In each round, the outcome of pulling an arm is a scalar reward and a resource consumption vector, both dependent on the context, and the global knapsack constraints require the total consumption for each resource to be bel…
Safe exploration in RF-RL doesn't increase sample complexity.
problem Achieving optimal policies with safety constraints in reward-free RL.
method Proposed SWEET framework for tabular and low-rank MDP settings, leveraging truncated value functions.
result Sample complexities match or outperform constraint-free counterparts, proving safety constraints have little impact.
In reinforcement learning (RL) , one of the key components is policy evaluation, which aims to estimate the value function (i.e., expected long-term accumulated reward) of a policy. With a good policy evaluation method, the RL algorithms will estimate the value function more accurately and find a better policy. When th…
The paper establishes conditions for strict power concavity in convolutions.
problem Conditions for strict power concavity in convolutions.
method Analyzes sufficient conditions for strict parabolic power concavity of convolutions.
result Establishes sufficient conditions for strict power concavity of convolutions.
Minimal graph level sets are concave if boundary is concave.
problem Understanding curvature of minimal graph level sets.
method Proved an inequality and showed geometric properties.
result Level sets of minimal graphs are concave if boundary is concave.
Proves log-concavity of cluster algebra coefficients for type A n A_n A n .
problem Log-concavity of cluster algebra coefficients.
method Introduced atomic theta basis and proved log-concavity for type A n A_n A n . result Proved log-concavity of coefficients for cluster algebra variables of type A n A_n A n . Study improves sampling from non-log-concave distributions using Fisher information.
problem Sampling from non-log-concave distributions with high Fisher information guarantees.
method Proximal sampler with RGO implementation, leveraging log-concave sampling results.
result Improved complexity guarantee in relative Fisher information for non-log-concave sampling.
Established concavity principle for curved spaces.
problem Solving equations on curved spaces with nonnegative curvature.
method Applied concavity principle to elliptic and parabolic equations on locally symmetric spaces with nonnegative curvature.
result First general concavity principle on spaces with non-constant sectional curvature.
Establishes log-concavity estimates for convex domains' first Dirichlet eigenfunctions.
problem Quantifying the Hessian of log-concave eigenfunctions on convex domains.
method Analyzes log-concavity properties of the first Dirichlet eigenfunction on convex domains.
result Obtains quantitative estimates for the Hessian of log u \log u log u . New saddle network architectures preserve convex-concave geometry in optimization problems.
problem Optimization models with convex x and concave y components.
method Structured separable decomposition and saddle network architectures.
result Proven one-dimensional approximation theorem and high accuracy on various test functions.
Heat flow fails to preserve concavity in curved spaces.
problem Non-preservation of concavity properties in curved spaces.
method Analysis of Dirichlet heat flow on Riemannian manifolds.
result No concavity properties are preserved unless curvature is zero.
We present a simple connection between differential Harnack inequalities for hypersurface flows and natural concavity properties of their time-of-arrival functions. We prove these concavity properties directly for a large class of flows by applying a concavity maximum principle argument to the corresponding level set f…
Investigates concavity of spacetimes, showing conditions for local concavity.
problem Understanding the concavity of spacetimes in Finsler geometry.
method Analyzes flag curvature and future capsules to characterize concavity.
result Berwald spacetimes are locally concave if and only if their flag curvature is nonnegative in timelike directions.
We define a class of L-convex-concave subsets of R P n \Bbb{R}P^n R P n , where L is a projective subspace of dimension l in R P n \Bbb{R}P^n R P n . These are sets whose sections by any (l+1)-dimensional space L' containing L are convex and concavely depend on L'. We introduce an L-duality for these sets, and prove that the L-dual to an L-…
Geodesic concavity and hypersymplectic structures in G 2 G2 G 2 -structures space.
problem Analyzing the geodesic concavity and hypersymplectic structures in the space of closed G 2 G2 G 2 -structures. method Utilising the geodesic constructed in the previous article, we show geodesic concavity and decrease in length of G 2 G2 G 2 Laplacian flow. result Hitchin's volume functional is geodesically concave and the G 2 G2 G 2 Laplacian flow decreases the length. Gradient methods converge exponentially in concave network games.
problem Finding Nash equilibria in concave network zero-sum games.
method Gradient Ascent and Optimistic Gradient Ascent analyses.
result Exponential convergence rates in various game settings.
This study examines how earnings announcements affect option volatility and pricing.
problem The impact of earnings announcements on option volatility and pricing.
method Analysis of extremely short-term options data to study bimodality and concavity in IV curves.
result Investors pay a premium to hedge against extreme volatility during earnings announcements in the presence of concave IV smiles.
Log-concavity of eigenfunctions on curved surfaces is proven, leading to fundamental gap estimates.
problem Proving log-concavity of eigenfunctions on curved surfaces.
method Analyzing the Laplacian eigenfunctions on positively curved surfaces.
result Strong log-concavity of the first eigenfunction on positively curved surfaces.
Improved sampling guarantees for weakly log-concave distributions.
problem Sampling from distributions that are not strongly log-concave.
method Proximal sampler with convergence guarantees under weaker assumptions.
result New state-of-the-art sampling guarantees for various target distributions.
We explain a general construction through which concave elliptic operators on complex manifolds give rise to concave functions on cohomology. In particular, this leads to generalized versions of the Khovanskii-Teissier inequalities.
Establishes a concavity property for positive Hessian quotient operators.
problem Analyzing positive Hessian quotient operators on Riemannian manifolds.
method Proves a special concavity property and a Jacobi inequality.
result Proves a Jacobi inequality for symmetric tensors.
Unified routing and arbitrage with concave continuation.
problem Combining routing and arbitrage in financial markets.
method Extending AMM trade functions to negative inputs via concave continuation.
result Unified approach unifies routing and arbitrage.
The study proves non-existence of concave functions on specific metric spaces.
problem Proving the non-existence of concave functions on certain metric spaces.
method Analogue theorems for Alexandrov spaces and C α C^α C α -Hölder Riemannian manifolds. result Proves non-existence of concave functions on complete manifolds with finite volume and specific metric spaces.
A scalable gradient-based framework for sparse portfolio selection.
problem Sparse minimum-variance portfolio selection with cardinality constraint.
method Gradient-based optimization with Boolean relaxation and tunable parameter.
result Matches commercial solvers in most instances, differing by a few assets with negligible error in portfolio variance.
Log-concavity proven for multinomial likelihoods under specific constraints.
problem Log-concavity of multinomial likelihoods under interval censoring constraints.
method Proved log-concavity by showing M-convex subsets of the discrete simplex.
result Likelihood function is completely log-concave.
The Links-Gould polynomial of alternating knots is shown to be log-concave and positive.
problem Verifying the positivity and log-concavity of the Links-Gould polynomial for alternating knots.
method Formulated a conjecture and verified it computationally for all 51.3 million knots with up to 19 crossings.
result All but 544 knots satisfy a stronger log-concavity condition.
The paper transforms a convex hull into a concave surface around a point cloud.
problem Creating a concave surface that encloses all points in a point cloud.
method Iterative facet replacement and expansion of the convex hull.
result A method to evolve a convex hull into a concave surface that fits the point cloud.
Introduces CSLC models to bridge deep generative models and classical algorithms.
problem Mode collapse and memorization issues in deep generative models and restrictive assumptions in classical algorithms.
method Introduces conditionally strongly log-concave (CSLC) models, factorizing data distribution into strongly log-concave conditional distributions.
result Efficient parameter estimation and sampling algorithms with theoretical guarantees for non-log-concave data distributions.
Paper finds convexity in translating solitons for concave flows.
problem Understanding convexity in translating solitons for concave extrinsic flows.
method Analyzes convexity estimates for translating solitons evolving under concave functions in R n + 1 \mathbb{R}^{n+1} R n + 1 . result Establishes convexity estimates for translating solitons of concave extrinsic geometric flows.