New algorithm achieves data-dependent regret bounds in MDPs with unknown transitions.
problem Achieving best-of-both-worlds guarantees with data-dependent regret bounds in MDPs with unknown transitions.
method Optimistic follow-the-regularized-leader algorithm with new optimistic Q-function estimators and transition bonus.
result First-order, second-order, and path-length bounds with polylog(T) regret in the stochastic regime.
Optimizes bus schedules to improve on-time performance.
problem Improving on-time performance of public transit systems.
method Formulated as a single-objective optimization task, solved using greedy algorithm, GA, and PSO.
result Enhanced bus timetables leading to better on-time performance.
Study shows optimal RL with transition look-ahead is NP-hard for ℓ ≥ 2 \ell \geq 2 ℓ ≥ 2 .
problem Optimal reinforcement learning with transition look-ahead is computationally hard.
method Proved NP-hardness for ℓ ≥ 2 \ell \geq 2 ℓ ≥ 2 using linear programming. result There is a precise boundary between tractable and intractable cases for RL with look-ahead.
Characterizing the phase transitions of convex optimizations in recovering structured signals or data is of central importance in compressed sensing, machine learning and statistics. The phase transitions of many convex optimization signal recovery methods such as ℓ 1 \ell_1 ℓ 1 minimization and nuclear norm minimization are…
Bayesian method estimates dynamics from near-optimal trajectories.
problem Estimating dynamics from near-optimal expert trajectories in reinforcement learning.
method Constraint-based Bayesian approach integrating expert near-optimality.
result Significant improvements in decision-making and transfer success.
Study optimal holomorphic extensions on complex manifolds with transitivity property.
problem Optimal holomorphic extensions on complex manifolds with transitivity property.
method Use Toeplitz operators and transitivity property for optimal holomorphic extensions.
result Transitivity property of optimal holomorphic extensions with small defect.
Optimal spectral method found for inhomogeneous spiked Wigner model.
problem Structured noise in learning scenarios.
method Random matrix theory and spectral analysis.
result Optimal threshold for phase transition in block-structured Wigner model.
Deep reinforcement learning finds optimal learning policies for adaptive systems.
problem Finding individualized learning plans for learners with unknown latent traits.
method Formulated as a Markov decision process, applied deep Q-learning with a transition model estimator.
result The algorithm efficiently discovers optimal learning policies with small data sets.
New RL algorithm optimizes policies with bandit feedback, matching previous bounds.
problem Optimizing policies with unknown transitions and bandit feedback.
method Optimistic Trust Region Policy Optimization (TRPO) algorithm.
result Sub-linear regret bounds for both stochastic and adversarial rewards.
A new algorithm for deep Q-learning with robustness to state transition uncertainty.
problem Model uncertainty in state transitions for non-tabular, continuous state spaces.
method Distributionally robust approach using worst-case transition ball and dualized Bellman operator with Sinkhorn distance.
result Optimal policy found through solving non-linear Bellman equation with neural network parameterization.
High-dimensional random geometry shows phase transitions in various problems.
problem Phase transitions in high-dimensional random geometry.
method Analysis of various financial, optimization, and ecological problems.
result Links between seemingly distant fields and further ramifications.
Framework optimizes transit routes based on crowd movements using demand prediction and supply optimization.
problem Dynamic optimization of transit routes in areas of crowd movements.
method Combines demand prediction (Quantile Regression) and supply optimization (Linear Programming) to dynamically redesign routes.
result Framework often obtains optimal solutions and outperforms conventional methods.
Develops CLTs for Markov chain transition probabilities and policies.
problem Estimating transition probabilities and policies in controlled Markov chains.
method Non-parametric estimator for transition matrices; CLTs for value, Q-, and advantage functions; goodness-of-fit tests.
result Asymptotic normality of estimators under specific logging policies.
New algorithm reduces dynamic regret for MDPs with unknown transition and adversarial rewards.
problem Episodic linear mixture MDPs with unknown transition and adversarial rewards.
method Combines occupancy-measure-based global optimization and policy-based variance-aware value-targeted regression.
result Achieves near-optimal dynamic regret of O ~ ( d H 3 K + H K ( H + P ˉ K ) ) \widetilde{\mathcal{O}}(d \sqrt{H^3 K} + \sqrt{HK(H + \bar{P}_K)}) O ( d H 3 K + H K ( H + P ˉ K ) ) . The paper tackles robust control for insurance contracts under uncertain transition rates.
problem Maximizing utility in insurance contracts with uncertain transition rates.
method Novel robust utility maximization problem under bounded cumulative transition rate uncertainty, using worst-case scenario analysis.
result Existence and uniqueness of worst-case and best-case reserves for insurance contracts.
New algorithm reduces suboptimality in imitation learning to nearly optimal levels.
problem Statistical limits of imitation learning in MDPs with known transitions.
method Mimic-MD algorithm and reduction to value estimation problem.
result Upper bound of O ( ∣ S ∣ H 3 / 2 / N ) O(|\mathcal{S}|H^{3/2}/N) O ( ∣ S ∣ H 3/2 / N ) for suboptimality, with efficient computation. This paper presents several models addressing optimal portfolio choice, optimal portfolio liquidation, and optimal portfolio transition issues, in which the expected returns of risky assets are unknown. Our approach is based on a coupling between Bayesian learning and dynamic programming techniques that leads to partia…
Study on bandit problems with switching constraints, revealing phase transitions in regret.
problem Stochastic multi-armed bandit problem with switching cost constraints.
method Proved matching upper and lower bounds on optimal regret, provided efficient algorithms.
result Phase transitions in optimal regret rate with respect to switching budget.
Paper introduces a new value function for state transitions and optimal policy learning.
problem Learning optimal policies from state transitions and actions.
method Develops a forward dynamics model to maximize a novel value function Q ( s , s ′ ) Q(s, s') Q ( s , s ′ ) . result Demonstrates benefits in value function transfer, redundant action spaces, and off-policy learning.
Study optimal timing to divest from assets with uncertain future scenarios.
problem Optimal timing to divest from assets with uncertain future scenarios.
method Smooth model of decision making under ambiguity aversion, optimal stopping problem with learning.
result Proves a minimax result reducing the problem to standard optimal stopping problems with learning.
TempLe learns transition templates for efficient multi-task RL.
problem Efficiently transferring knowledge across different RL tasks with varying state/action spaces.
method Generates transition dynamics templates to abstract similarities between tasks.
result Achieves significantly lower sample complexity than single-task or multi-task methods.
We study the problem of approximate ranking from observations of pairwise interactions. The goal is to estimate the underlying ranks of n n n objects from data through interactions of comparison or collaboration. Under a general framework of approximate ranking models, we characterize the exact optimal statistical error …
We study optimal solutions to an abstract optimization problem for measures, which is a generalization of classical variational problems in information theory and statistical physics. In the classical problems, information and relative entropy are defined using the Kullback-Leibler divergence, and for this reason optim…
SCPO learns robust policies without modeling disturbance, improving real-world task performance.
problem Poor performance of reinforcement learning in real-world tasks due to disturbance in transition dynamics.
method State-conservative policy optimization (SCPO) that reduces disturbance to state space and approximates it with a gradient-based regularizer.
result SCPO learns robust policies without prior knowledge of disturbance or simulators, improving performance in robot control tasks.
NetOTC compares and aligns directed or undirected networks via random walk transitions.
problem Comparing and aligning networks of different types and sizes.
method NetOTC uses a transport-based approach to find optimal transition couplings of random walks.
result NetOTC quantifies network differences and provides vertex and edge alignments.
Proposes MIVI for efficient posterior estimation and design of MCMC transitions.
problem Efficiently estimating posterior distributions in constrained time.
method Combines variational inference and MCMC with a variational distribution and optimized Markov chain.
result Optimized Markov chain improves variational distribution and vice versa, leading to more accurate posteriors.
New method reduces discrete flow transitions, improving perplexity estimation.
problem Stochasticity in discrete paths makes rectification strategies ineffective.
method Dynamic-optimal-transport-like minimization objective with minibatch strategies.
result 32 times reduction in transitions for same perplexity.
This work improves transferability of rewards inferred from expert demonstrations.
problem Transferability of rewards inferred from expert demonstrations under limited access to the expert's policy.
method Proposed principal angles as a measure of similarity and dissimilarity between transition laws. Established sufficient conditions for transferability under limited access.
result Two key results on sufficient conditions for transferability to any and local changes in transition laws.
WRAAC uses Wasserstein distance for robust reinforcement learning.
problem Lack of quantified robustness to system dynamics in existing reinforcement learning algorithms.
method Leverages Wasserstein distance to connect state disturbance to transition kernel disturbance, reducing infinite-dimensional optimization to a finite-dimensional problem.
result Designs a novel algorithm, WRAAC, that achieves robust reinforcement learning.
RRPI improves offline RL by optimizing policies against worst-case dynamics.
problem Offline RL's performance degrades under distribution shift and transition uncertainty.
method Formulates offline RL as robust policy optimization, treating transition kernel as decision variable.
result RRPI achieves strong average performance on D4RL benchmarks, outperforming recent baselines.
Study reveals phase transition in neural networks near interpolation.
problem Understanding generalization and learning transitions in neural networks.
method Effective theory for approximating Bayes-optimal generalisation error.
result Unveils a discontinuous phase transition between universal and specialisation phases.
Optimizes electric aircraft deployment for Canadian aviation to reduce emissions.
problem Limited fleet capacity and operational structure hinder electric aircraft transition.
method Multi-period mixed-integer linear programming (MILP) framework.
result Electric aircraft can reduce emissions by over 70% within five years.
New model for pairwise comparisons without stochastic transitivity.
problem Suboptimal performance of models assuming stochastic transitivity in real-world scenarios.
method Proposes a general family of statistical models using a skew-symmetric matrix.
result Achieves minimax-rate optimality and adapts to data sparsity.
We consider reinforcement learning in changing Markov Decision Processes where both the state-transition probabilities and the reward functions may vary over time. For this problem setting, we propose an algorithm using a sliding window approach and provide performance guarantees for the regret evaluated against the op…
We address the problem of portfolio optimization under the simplest coherent risk measure, i.e. the expected shortfall. As it is well known, one can map this problem into a linear programming setting. For some values of the external parameters, when the available time series is too short, the portfolio optimization is …
Enhances machine learning for complex systems by embedding transition manifolds.
problem Identifying low-dimensional dynamics in high-dimensional multiscale systems.
method Kernel embeddings of transition manifolds in reproducing kernel Hilbert spaces.
result Robust and more efficient algorithm for identifying reaction coordinates.
Estimates bisimulation metrics from sample streams, not full transition models.
problem Estimating Markov chain metrics from limited sample data.
method Stochastic optimization using linear programming and primal-dual method.
result Validated through empirical evaluations, providing sample complexity guarantees.
A new method learns robust policies from offline data with latent structures.
problem Conservative policies under unrealistic dynamics shifts.
method d-RRMDP framework with f f f -divergence regularization and R2PVI algorithm. result R2PVI learns robust policies with superior computational efficiency.
We establish minimax optimal rates of convergence for estimation in a high dimensional additive model assuming that it is approximately sparse. Our results reveal an interesting phase transition behavior universal to this class of high dimensional problems. In the {\it sparse regime} when the components are sufficientl…
Optimal sequential testing for Markovian data with lower and upper bounds.
problem Sequential hypothesis testing for Markovian data.
method Non-asymptotic lower bounds and optimal test design.
result Optimal test matches lower bound asymptotically.
Paper analyzes LPSA algorithm for constrained optimization, revealing phase transitions and bias-variance trade-offs.
problem Optimization problems with linear constraints.
method Loopless projection stochastic approximation (LPSA) with jump diffusion approximation.
result LPSA trajectories converge to SDEs, revealing asymptotic behaviors and phase transitions.
New protocols show 1-bit mean estimation can be order-optimal without interaction.
problem Can 1-bit mean estimation be optimal without interaction?
method Adaptive and non-adaptive threshold and interval queries, with one adaptive transition.
result Arbitrary non-adaptive quantizers can match the adaptive rate, suggesting interaction is not necessary.
Deep reinforcement learning method finds rare events in complex systems.
problem Computing transition pathways in high-dimensional systems.
method Formulated as a cost minimization problem, solved using DDPG with physical properties.
result Efficiently samples and computes globally optimal transition pathways.
Model-based Bayesian Reinforcement Learning (BRL) allows a found formalization of the problem of acting optimally while facing an unknown environment, i.e., avoiding the exploration-exploitation dilemma. However, algorithms explicitly addressing BRL suffer from such a combinatorial explosion that a large body of work r…
In recent years, non-parametric methods utilizing random walks on graphs have been used to solve a wide range of machine learning problems, but in their simplest form they do not scale well due to the quadratic complexity. In this paper, a new dual-tree based variational approach for approximating the transition matrix…
AdaSGD combines SGD and Adam benefits, eliminating the need for transition.
problem Understanding when to transition from Adam to SGD for optimal performance.
method Adapting a single global learning rate for SGD (AdaSGD).
result AdaSGD combines the benefits of both SGD and Adam, improving convergence and generalization.
Study uses MTD model to optimize portfolios by capturing complex financial asset relationships.
problem Capturing nonlinear and directional relationships in financial markets.
method Directed and weighted financial networks using Mixture Transition Distribution (MTD) model.
result Portfolio optimization with network-based assortativity measures outperforms classical methods.
New method for optimizing risk in financial models using Fourier transforms.
problem Optimizing risk in financial models with multi-period mean-CVaR.
method Strictly monotone 2D integration scheme via Fourier-trained transition kernels.
result Established robust and accurate optimization method for financial models.