Efficient RL algorithms for linear function approximation with limited adaptivity constraints.
problem Limited adaptivity in reinforcement learning with linear function approximation.
method Proposed two efficient online RL algorithms for episodic linear Markov decision processes under batch learning and rare policy switch models.
result Achieved efficient regret bounds for both batch learning and rare policy switch models, with substantial reduction in adaptivity.
New 2-cocycles introduced on Lie algebra of S^3H, leading to central extensions and root space decompositions.
problem Introducing new 2-cocycles on Lie algebra of S^3H and their implications.
method Extending 2-cocycles to Lie algebra S 3 g l ( n , H ) S^3gl(n,H) S 3 g l ( n , H ) , defining central extensions and root spaces. result Obtained root space decomposition and Chevalley generators for the Lie algebra s l ^ ( n , H ) \hat{sl}(n,H) s l ^ ( n , H ) . Researchers found a special basis for cycles on a K3 surface.
problem Understanding the structure of two-cycles on K3 surfaces.
method Constructed a canonical basis of two-cycles using formal sums of smooth submanifolds.
result The intersection form of the basis takes a specific canonical form.
Study complex reflections in 3D hyperbolic geometry, finding new representations.
problem Deforming groups in 3D complex hyperbolic geometry.
method Representations of abstract groups in PU(3,1), using Ford domains as guides.
result First nontrivial example of a discrete and faithful representation of a subgroup in PU(3,1).
Study on error rates for approximating rough volatility models.
problem Simulation of rough volatility models with fractional Brownian motion.
method Analysis of weak error rates for numerical schemes, focusing on fBm and cubic test functions.
result Convergence rates for approximations are ( 3 H + 1 2 ) ∧ 1 (3H+ \frac{1}{2}) \wedge 1 ( 3 H + 2 1 ) ∧ 1 for exact left-point discretization and H + 1 2 H+\frac{1}{2} H + 2 1 for hybrid schemes. New RL algorithm achieves sublinear regret and constraint violation without simulators.
problem Maximizing reward under utility constraints in large-scale systems.
method Model-free, simulator-free algorithm using LSVI-UCB with primal-dual optimization and soft-max policy.
result Achieves i l d e O ( d 3 H 3 T ) ilde{\mathcal{O}}(\sqrt{d^3H^3T}) i l d e O ( d 3 H 3 T ) regret and i l d e O ( d 3 H 3 T ) ilde{\mathcal{O}}(\sqrt{d^3H^3T}) i l d e O ( d 3 H 3 T ) constraint violation bounds. Elementary proof shows no specific torus to sphere cover with certain branching points.
problem Existence of specific branched covers between torus and sphere.
method Elementary topological proof using properties of the torus.
result No such branched cover exists with specified branching points.
Paper develops Euler scheme for fractional delay diff. eqs with additive noise.
problem Developing a consistent Euler-Maruyama scheme for fractional stochastic delay diff. eqs.
method Euler-Maruyama scheme for fractional Brownian motion with additive noise.
result Achieved convergence rate of H+1/2 for smooth delays when H>1/2.
Study rough volatility models using path-dependent PDEs and fractional Brownian motions.
problem Modeling and analyzing rough volatility in financial markets.
method Showed conditional expectations are unique classical solutions to path-dependent PDEs derived from functional Itô formula. Leverage these to study weak rates of convergence for discretized stochastic integrals.
result Obtained optimal weak error rates for approximating log-stock prices in rough volatility models.
Study improves weak error estimates for rough volatility models.
problem Efficient numerical schemes for non-Markovian stochastic processes with rough volatility.
method Analyzes weak rates for a class of stochastic processes with rough stochastic volatility.
result Weak rate is of order min{3H+0.5, 1} for a large class of test functions.
New algorithm explores reinforcement learning with noisy data.
problem Exploration in reinforcement learning with complex value functions.
method Randomized exploration with i.i.d. scalar noises and optimistic reward sampling.
result Achieves worst-case regret bound of O ~ ( p o l y ( d E H ) T ) \widetilde{O}(\mathrm{poly}(d_EH)\sqrt{T}) O ( poly ( d E H ) T ) . Safe reinforcement learning tackles safety constraints with linear approximations.
problem Ensuring safety in reinforcement learning without violating constraints.
method Modeling safety as a linear cost function, developing SLUCB-QVI and RSLUCB-QVI algorithms for MDPs with linear function approximation.
result Achieved a nearly optimal regret bound for safe reinforcement learning, matching state-of-the-art unsafe algorithms.
Paper presents efficient RL algorithm for linear dynamics without simulator assumptions.
problem Designing efficient RL algorithms with function approximation for linear settings.
method Optimistic modification of Least-Squares Value Iteration (LSVI).
result Achieves i l d e O ( d 3 H 3 T ) ilde{\mathcal{O}}(\sqrt{d^3H^3T}) i l d e O ( d 3 H 3 T ) regret, independent of states and actions. Paper presents an efficient algorithm for linear MDP with low switching cost.
problem Large state space reinforcement learning problems with low switching cost.
method First algorithm for linear MDP with low switching cost, achieving near-optimal regret and switching cost.
result Regret bound of $\widetilde{O}\left(\sqrt{d^3H^4K}
ight)$ and near-optimal switching cost of $O\left(d H\log K
ight)$ .
New RL algorithm handles delayed feedback with posterior sampling.
problem Challenges of delayed feedback in reinforcement learning with linear function approximation.
method Posterior sampling with delayed feedback for value-based RL.
result Achieves optimal regret guarantee with improved computational efficiency.
New RLHF algorithm identifies optimal policies from human feedback without explicit reward inference.
problem Training large language models with human feedback without reward inference.
method Model-free RLHF algorithm B S A D \mathsf{BSAD} BSAD that identifies optimal policies directly from human preference. result Provable, instance-dependent sample complexity i l d e O ( c M S A 3 H 3 M log 1 δ ) ilde{\mathcal{O}}(c_{\mathcal{M}}SA^3H^3M\log\frac{1}δ) i l d e O ( c M S A 3 H 3 M log δ 1 ) . The study improves geometric estimates for surfaces of constant mean curvature in 3-manifolds.
problem Estimating geometric properties of surfaces with constant mean curvature in 3-manifolds.
method Application of the Hierarchy Structure Theorem to understand global properties of surfaces, including area and diameter.
result Improved geometric estimates for the area and diameter of surfaces with constant mean curvature, including genus bounds.