Paper solves pendulum swing-up problem using RL.
problem Solving the classic pendulum swing-up problem.
method Deep Deterministic Policy Gradient algorithm applied to continuous action domain.
result Optimal pendulum achieved with increasing average return and decreasing loss.
RL algorithms compare in controlling a complex dynamical system.
problem Optimal control of complex, nonlinear systems.
method Temporal-difference, policy gradient actor-critic, value function approximation compared.
result RL algorithms outperform standard LQR in controlling the cartpole system.
Photonic quantum reinforcement learning for control problems.
problem Solving continuous control problems with noisy quantum computers.
method Proximal policy optimization for photonic variational quantum agents.
result Photonic policy learning achieves comparable performance to classical neural networks.
Adapts model-based advice to stabilize black-box policies for nonlinear control.
problem Stabilizing machine-learned policies for nonlinear control with limited model information.
method Proposes an adaptive λ-confident policy to combine black-box and model-based advice. result Proves the stability of the adaptive λ-confident policy and its competitive ratio. The paper introduces MDP homomorphic networks for faster reinforcement learning.
problem Current reinforcement learning approaches do not exploit symmetries in the joint state-action space.
method Equivariant neural networks with group-structured symmetries (reflections, rotations).
result MDP homomorphic networks converge faster than unstructured baselines on various tasks.
A lightweight FPGA-based reinforcement learning approach for edge devices.
problem Resource constraints and inefficiency of DQN on edge devices.
method OS-ELM based training algorithm and L2 regularization for stability.
result 29.77x and 89.40x faster than conventional DQN-based approach for CartPole-v0 task.
We adapt the ideas underlying the success of Deep Q-Learning to the continuous action domain. We present an actor-critic, model-free algorithm based on the deterministic policy gradient that can operate over continuous action spaces. Using the same learning algorithm, network architecture and hyper-parameters, our algo…
A simple DQN-based multi-agent RL system for binary actions.
problem Complexity and training issues in multi-agent reinforcement learning.
method Shared state and rewards, agent-specific actions, experience replay pool.
result Better performance and faster convergence compared to conventional methods.
This paper applies CPI to deep RL, improving stability and performance.
problem Improving stability and performance in deep reinforcement learning.
method Combines Conservative Policy Iteration with deep neural networks and adaptive mixture rates.
result Demonstrates improved stability and performance in deep RL algorithms.
Deep reinforcement learning has led to several recent breakthroughs, though the learned policies are often based on black-box neural networks. This makes them difficult to interpret and to impose desired specification constraints during learning. We present an iterative framework, MORL, for improving the learned polici…
This paper detects Markov violations in RL with noise, improving policy development.
problem Partial observability and sensor/actuator noise invalidate Markovian assumptions in RL.
method Combines PCMCI causal discovery with Markov Violation score (MVS).
result Even substantial noise doesn't always disrupt multi-step dependencies.
WAPPO optimizes feature distributions for better visual transfer in RL.
problem Improving visual transfer in reinforcement learning.
method WAPPO uses Wasserstein Confusion to minimize feature distribution distance.
result WAPPO outperforms previous methods in visual transfer across different environments.
We present a data-efficient reinforcement learning algorithm resistant to observation noise. Our method extends the highly data-efficient PILCO algorithm (Deisenroth & Rasmussen, 2011) into partially observed Markov decision processes (POMDPs) by considering the filtering process during policy evaluation. PILCO conduct…
Watermarks DRL policies with minimal performance impact.
problem Detecting and managing unauthorized use of proprietary DRL policies.
method Integrates a unique identifier into DRL policies' responses to a sequence of state transitions.
result Demonstrates feasibility with DQN policy in Cartpole environment.
Quantum variational circuits improve reinforcement learning efficiency.
problem Improving reinforcement learning algorithms using quantum computing.
method Investigation of quantum variational circuits for DQN and Double DQN, encoding classical data for quantum circuits.
result Quantum variational circuits can solve reinforcement learning tasks with a smaller parameter space.
Paper benchmarks DRL policies' resilience to state transitions.
problem Measuring DRL policies' resilience to state perturbations.
method Disentangled representation learning and RL-based techniques.
result Demonstrated feasibility of resilience benchmarking in DQN, A2C, and PPO2.
New method in Bayesian optimization finds optimal inputs knowing the optimal outputs.
problem Finding optimal inputs when the optimal outputs are known in advance.
method Transform Gaussian process surrogate using known optimum output; propose two acquisition functions.
result Our approaches give quantitatively better performance than standard BO methods.
WMPG reduces policy gradient variance using world models.
problem Reducing variance in policy gradient estimates.
method Trains a world model online to estimate policy gradients and uses imagined trajectories as a baseline.
result WMPG achieves better sample efficiency compared to AC and MAC.
LM optimization outperforms other methods in deep learning tasks but at high computational cost.
problem Finding efficient optimization methods for deep learning models.
method Comparing first-order (CG, SGD, LM, L-BFGS) and higher-order optimization functions.
result Levemberg-Marquardt (LM) optimization significantly improves convergence but at a high computational cost.
Differentiable MPC improves reinforcement learning efficiency.
problem Combining model-free and model-based reinforcement learning approaches.
method Differentiating through KKT conditions of a convex approximation of MPC.
result MPC policies are more data-efficient and superior to traditional system identification.
D2D-SPL uses discrete states and a classifier to train RL faster.
problem Training neural networks in RL due to correlated samples.
method Discretizes state space, uses actor-critic, selects input/target pairs, trains classifier.
result Trains faster than state-of-the-art methods.
This project proposes using reinforcement learning to train spiking neural networks.
problem Training spiking neural networks using traditional methods is challenging due to the discrete nature of spikes.
method The project investigates two approaches: 1) treating each neuron as an RL agent, 2) applying the reparameterization trick.
result The project demonstrates that reinforcement learning can be applied to train spiking neural networks.
New STDP rule for spiking neurons solves discrete action reinforcement learning tasks.
problem Applying standard STDP to discrete action reinforcement learning tasks.
method Feedback-modulated TD-STDP learning rule for spiking neuron networks.
result Feedback modulation improves credit assignment in reinforcement learning.
Bayesian optimization outperforms other methods in hyperparameter tuning for reinforcement learning.
problem Finding optimal hyperparameters that generalize across random seeds in reinforcement learning.
method Benchmarked Successive Halving, Random Search, and Bayesian Optimization with and without repetitions on PPO2 algorithms for Cartpole and Inverted Pendulum tasks.
result Bayesian optimization with noise robust acquisition function is the best choice.
AdaRL adapts quickly to new environments with minimal data.
problem Quickly adapting to new environments in reinforcement learning.
method AdaRL uses a parsimonious graphical representation to encode changes across domains.
result AdaRL can efficiently adapt policies to target domains with few samples.
CARL safely adapts RL agents for safety-critical tasks.
problem Safety hazards in RL for safety-critical tasks.
method CARL combines model-based RL and cautious adaptation.
result CARL achieves higher rewards with fewer failures in safety-critical tasks.
BCPO optimizes offline RL policies by converting uncertainty into conservative bounds.
problem Offline RL's fragility under distribution shifts and model errors.
method Bayesian approach with credible lower bounds and KL regularization.
result BCPO yields an uncertainty-calibrated policy that avoids exploiting model errors.
Paper introduces MVS to detect non-Markovian observations in reinforcement learning.
problem Real-world sensors violate Markov property, leading to suboptimal reinforcement learning performance.
method Uses prediction-based Markov Violation Score (MVS) combining random forest and ridge regression.
result MVS detects non-Markovian structure in observation trajectories, quantifying its impact.
We survey the status of some decision problems for 3-manifolds and their fundamental groups. This includes the classical decision problems for finitely presented groups (Word Problem, Conjugacy Problem, Isomorphism Problem), and also the Homeomorphism Problem for 3-manifolds and the Membership Problem for 3-manifold gr…
Optimal transport reformulates multiple quantile hedging problem.
problem Multiple quantile hedging problem in incomplete markets.
method Reformulated as Monge optimal transport problem, introduced Kantorovitch version, proved no duality gap.
result Multiple quantile hedging problem can be seen as semi-discrete optimal transport problem.
Solves four problems related to circle families in the plane.
problem Four basic problems of circle families in the plane.
method Solves all four basic problems of circle families in the plane.
result All four basic problems are solved.
Solves four problems related to sphere families in 3D space.
problem Four basic problems of sphere families in Euclidean 3-space.
method Solves all four basic problems of sphere families in Euclidean 3-space.
result All four basic problems are solved.
The paper solves optimal control problems for various convex sets using convex trigonometry.
problem Optimal control problems with 2D convex compact sets.
method Using convex trigonometry to derive extremals for various problems.
result Geodesics in multiple sub-Finsler problems are derived.
Explains eigenvalue and generalized eigenvalue problems with examples.
problem Eigenvalue and generalized eigenvalue problems.
method Introduction and examples from machine learning.
result Solutions to eigenvalue and generalized eigenvalue problems.
This paper solves the Christoffel problem in hyperbolic space and its equivalent on spheres.
problem Prescribing curvatures for convex hypersurfaces in hyperbolic space.
method Proving a full rank theorem to establish the existence of solutions.
result Existence of solutions to the Christoffel problem and its equivalent Nirenberg-Kazdan-Warner problem on spheres.
In the present paper, the primal-dual problem consisting of the investment risk minimization problem and the expected return maximization problem in the mean-variance model is discussed using replica analysis. As a natural extension of the investment risk minimization problem under only a budget constraint that we anal…
Study proves only origin-centered spheres solve certain curvature problems.
problem Proving uniqueness of solutions to curvature problems.
method Using the Heintze-Karcher inequality, the study proves the uniqueness of smooth, strictly convex solutions to a class of Minkowski type problems.
result Only origin-centered spheres solve isotropic and Lp-Gaussian-Minkowski problems. MathChat uses LLM agents to solve challenging math problems through conversational problem-solving.
problem Solving math problems expressed in natural language.
method MathChat is a conversational framework combining an LLM agent and a user proxy agent for collaborative problem-solving.
result MathChat improves tool-using prompting methods by 6% on difficult math problems.
New algorithm solves non-convex min-max problems in signal processing.
problem Non-convex min-max problems in signal processing and communication.
method Hybrid Block Successive Approximation (HiBSA) algorithm alternating gradient descent and ascent steps.
result HiBSA converges to first-order stationary solutions with global rates.
This article reviews ranking problems and their solutions.
problem Ranking problems in statistical learning.
method Systematic review of ranking problems and optimization techniques.
result Unified notation for optimization problems and identification of strengths and limitations of algorithms.
The paper solves a generalized Christoffel-Minkowski problem using a curvature flow.
problem Solving the (p,q)-Christoffel-Minkowski problem.
method Investigating the problem via an expanding curvature flow.
result Existence and uniqueness of smooth solutions to the (p,q)-Christoffel-Minkowski problem.
Paper solves Gromov-Wasserstein for point clouds efficiently.
problem Quantifying similarity between two formations or shapes.
method Reformulates QAP as low-rank concave quadratic optimization problem.
result Global solution for large-scale problems with thousands of points.
The paper explains how microlocal analysis solves geometric inverse problems.
problem Recovering geometric information from boundary measurements.
method Microlocal analysis applied to three inverse problems.
result Microlocal techniques solve specific inverse problems in Riemannian geometry.
Proves NP and co-NP status for knot core recognition in solid torus.
problem Determining if a knot is the core of a solid torus.
method Alternate proof and corollary of Hopf link recognition problem.
result Proves NP and co-NP status for solid torus core recognition problem.
A new method solves complex control problems with random coefficients.
problem Solving LQ McKean-Vlasov control problems with random coefficients.
method Decomposes the problem into two decoupled stochastic optimal control problems.
result The sum of optimal controls of auxiliary problems equals the original problem's optimal control.
This is a survey of some problems in geometric group theory which I find interesting. The problems are from different areas of group theory. Each section is devoted to problems in one area. It contains an introduction where I give some necessary definitions and motivations, problems and some discussions of them. For ea…
We present updates to the problems on Hirzebruch's 1954 problem list focussing on open problems, and on those where substantial progress has been made in recent years. We discuss some purely topological problems, as well as geometric problems about (almost) complex structures, both algebraic and non-algebraic, about co…
27 problems identified in automating movie/TV subtitle translation.
problem Challenges in translating movie/TV subtitles.
method Categorized problems into three categories and evaluated translation quality.
result Frontier NLP systems struggle with subtitles and require post-processing.