Bayesian approach to optimal transport with stochastic costs.
problem Inferring optimal transport plans with uncertain costs.
method Bayesian framework and Hamiltonian Monte Carlo (HMC) sampling.
result Inference of optimal transport plans under stochastic cost functions.
New approach handles stochastic and partially-observable environments using discrete autoencoders and Monte Carlo tree search.
problem Challenges in planning for stochastic and partially-observable environments.
method Uses discrete autoencoders and a stochastic variant of Monte Carlo tree search.
result Significantly outperforms MuZero on stochastic chess and scales to DeepMind Lab.
The hierarchical structure of production planning has the advantage of assigning different decision variables to their respective time horizons and therefore ensures their manageability. However, the restrictive structure of this top-down approach implying that upper level decisions are the constraints for lower level …
Investment strategies in occupational pension plans are optimized for non-tradable income risk.
problem Optimizing investment strategies for occupational pension plans in the presence of non-tradable income risk.
method Formulated as a stochastic optimization problem, analyzed in both constant and stochastic volatility environments.
result Random contributions induce the optimal glide path structure, influenced by initial wealth, contributions, and risk aversion.
We designed a grid world task to study human planning and re-planning behavior in an unknown stochastic environment. In our grid world, participants were asked to travel from a random starting point to a random goal position while maximizing their reward. Because they were not familiar with the environment, they needed…
Study integrates reliability constraints into generation planning models.
problem Challenges in integrating reliability constraints with generation planning models.
method Leverages a weighted oblique decision tree (WODT) technique to embed reliability verification constraints.
result Demonstrates effectiveness in achieving reliable and optimal planning solutions.
Optimal withdrawal strategy for DC pension plans maximizes total withdrawals while managing risk.
problem Maximizing withdrawals from DC pension plans while managing risk.
method Optimal stochastic control approach with constraints on withdrawal and asset allocation.
result Optimal strategy yields higher average withdrawals with minimal increase in risk.
A neural network approach solves optimal decumulation problems for pension plans.
problem Optimal asset allocation and withdrawal strategies for DC pension holders.
method Data-driven neural network optimization with customized activation functions.
result The neural network approach learns near-optimal solutions comparable to HJB PDE methods.
New algorithm minimizes worst-case regret in uncertain, time-varying dynamics.
problem Model-based policy learning in uncertain, time-varying dynamics.
method Planning regret metric and iterative algorithm for minimizing it.
result Empirical evidence shows the proposed algorithm outperforms existing methods.
Survey of integrating planning and learning in model-based reinforcement learning.
problem Sequential decision making in AI, formalized as MDP optimization.
method Systematic coverage of dynamics model learning and planning-learning integration.
result Broad conceptual overview of model-based reinforcement learning.
NOT learns optimal transport plans, kernel costs improve performance.
problem NOT algorithm learns non-optimal plans with weak quadratic costs.
method Introduced kernel weak quadratic costs to improve NOT's performance.
result Kernel costs provide improved theoretical and practical guarantees.
Planning has been very successful for control tasks with known environment dynamics. To leverage planning in unknown environments, the agent needs to learn the dynamics from interactions with the world. However, learning dynamics models that are accurate enough for planning has been a long-standing challenge, especiall…
This work clarifies the role of inference types in planning.
problem Lack of consistency in using inference types for planning.
method Variational framework and loopy belief propagation.
result All inference types correspond to different weights in variational problems.
We study an asset allocation stochastic problem with restriction for a defined-contribution pension plan during the accumulation phase. We consider a financial market with stochastic interest rate, composed of a risk-free asset, a real zero coupon bond price, the inflation-linked bond and the risky asset. A plan member…
We introduce a framework for model learning and planning in stochastic domains with continuous state and action spaces and non-Gaussian transition models. It is efficient because (1) local models are estimated only when the planner requires them; (2) the planner focuses on the most relevant states to the current planni…
This paper studies when particle filtering is efficient for planning in partially observed systems.
problem The efficiency of particle filtering for planning in partially observed linear dynamical systems.
method Coupling of ideal and approximate sequences to bound particle complexity.
result Polynomially many particles suffice for stable systems to approximate optimal planning.
Paper uses NMT to predict solutions to stochastic optimization problems quickly.
problem Predicting solutions to stochastic discrete optimization problems under uncertainty.
method Applied a state-of-the-art NMT algorithm with minimal adaptations and hyperparameter tuning.
result NMT can produce accurate solutions in milliseconds with less variability.
Distribution and sample models are two popular model choices in model-based reinforcement learning (MBRL). However, learning these models can be intractable, particularly when the state and action spaces are large. Expectation models, on the other hand, are relatively easier to learn due to their compactness and have a…
Solves risk-aware optimal switching problems in discrete time.
problem Non-Markovian optimal switching problems with risk awareness and general filtration.
method Solves reflected backward stochastic difference equations.
result Existence and uniqueness of solutions for the problems.
Proposes a new model to optimize investment plans with varying terminal times.
problem Improving the classical mean-variance model for continuous time investments.
method Uses stochastic optimal control and varying terminal time to determine optimal strategies.
result Optimal strategies and terminal times can be determined to minimize portfolio variance.
E2C separates planning and execution in LLMs, improving efficiency and performance.
problem Entangled planning and execution in LLMs waste tokens and limit flexibility.
method E2C splits exploration and execution phases, using SFT and RL for training.
result E2C achieves 53.3% accuracy on AIME'2024 with 12.4k tokens, outperforming alternatives.
This paper offers a methodological contribution at the intersection of machine learning and operations research. Namely, we propose a methodology to quickly predict tactical solutions to a given operational problem. In this context, the tactical solution is less detailed than the operational one but it has to be comput…
Study optimal investment under uncertain conditions.
problem Optimal investment in uncertain market conditions.
method Modelled Knightian uncertainty through multiple priors, solved using stochastic backward equations.
result Existence and uniqueness of optimal investment plan derived.
Paper analyzes robust strategies in a pension plan game with ambiguous financial markets.
problem Analyzing robust strategies in a defined benefit pension plan game with ambiguous financial markets.
method Formulated and solved two robust non-zero-sum games using stochastic dynamic programming.
result Explicit forms and optimality of the solutions are shown for the firm and union.
Autonomous robots need to interact with unknown, unstructured and changing environments, constantly facing novel challenges. Therefore, continuous online adaptation for lifelong-learning and the need of sample-efficient mechanisms to adapt to changes in the environment, the constraints, the tasks, or the robot itself a…
In the context of tree-search stochastic planning algorithms where a generative model is available, we consider on-line planning algorithms building trees in order to recommend an action. We investigate the question of avoiding re-planning in subsequent decision steps by directly using sub-trees as action recommender. …
New algorithms for planning with adversarial changes in costs.
problem Planning with adversarial changes in costs over time.
method Developed algorithms for adversarial SSP with high probability regret bounds.
result Obtained sub-linear regret bounds for adversarial SSP.
Optimal transport aims to estimate a transportation plan that minimizes a displacement cost. This is realized by optimizing the scalar product between the sought plan and the given cost, over the space of doubly stochastic matrices. When the entropy regularization is added to the problem, the transportation plan can be…
Study finds optimal retirement timing in uncertain wage scenarios.
problem Optimal retirement timing in presence of uncertain wages.
method Formulated as a free boundary problem in an incomplete market.
result Developed a method to determine optimal retirement timing.
This work tackles maintenance planning with deep reinforcement learning under uncertainty.
problem Optimizing inspection and maintenance policies in deteriorating environments with incomplete information and constraints.
method Joint framework of constrained POMDPs and multi-agent DRL addressing challenges of state/action space, history, uncertainty, and constraints.
result The proposed framework outperforms existing methods in resource and risk-aware decision-making.
In this paper we apply change of numeraire techniques to the optimal transport approach for computing model-free prices of derivatives in a two periods model. In particular, we consider the optimal transport plan constructed in \cite{HobsonKlimmek2013} as well as the one introduced in \cite{BeiglJuil} and further studi…
Optimal Transport (OT) naturally arises in many machine learning applications, yet the heavy computational burden limits its wide-spread uses. To address the scalability issue, we propose an implicit generative learning-based framework called SPOT (Scalable Push-forward of Optimal Transport). Specifically, we approxima…
Investigates risk measures for DC pension decumulation.
problem Develop optimal decumulation strategies for DC plan holders.
method Formulates decumulation as a control problem, studies risk measures (expected shortfall, linear shortfall, probability of shortfall).
result Optimal controls for expected reward and expected shortfall are identical to those for expected reward and linear shortfall.
A moment constraint that limits the number of dividends in the optimal dividend problem is suggested. This leads to a new type of time-inconsistent stochastic impulse control problem. First, the optimal solution in the precommitment sense is derived. Second, the problem is formulated as an intrapersonal sequential dyna…
New algorithm models satiation in recommender systems.
problem Satiation effects in user preferences not modeled by existing algorithms.
method Rebounding bandits, modeling satiation as time-invariant linear dynamical systems.
result Greedy policy optimal for identical deterministic dynamics; EEP algorithm for stochastic dynamics.
This paper solves steady-state planning for multichain MDPs.
problem Specifying constraints on the steady-state behavior of an agent.
method Linear programming solution for multichain MDPs.
result Optimal solutions yield stationary policies with rigorous guarantees.
Study on limits of LLM-based multi-agent planning reliability.
problem Reliability limits of LLM-based multi-agent planning.
method Modeling LLM-based multi-agent architecture as a decision network, showing dominance by centralized Bayes decision maker.
result Optimizing multi-agent directed acyclic graphs under communication budget is equivalent to choosing a constrained experiment.
A new method solves complex hydroelectricity planning problems.
problem Solving multistage stochastic linear programming for hydrothermal dispatch planning.
method Regularized Linear Decision Rules (AdaLASSO) to reduce overfitting and improve out-of-sample performance.
result Significant reductions in non-zero coefficients and improved spot-price profiles.
Optimizes pension fund management under funding risks.
problem Managing DB pension fund under underfunded and overfunded conditions.
method Stochastic model with Ornstein-Uhlenbeck interest rate, geometric Brownian motion for benefits, and cash, bond, stock investments.
result Optimal wealth process, portfolio, and efficient frontier obtained under various tolerance levels for solvency risk.
This paper presents a novel two-step approach for the fundamental problem of learning an optimal map from one distribution to another. First, we learn an optimal transport (OT) plan, which can be thought as a one-to-many map between the two distributions. To that end, we propose a stochastic dual approach of regularize…
PS framework selects best policy from library for CSO problems.
problem Policy selection in CSO with heterogeneous performance across covariate space.
method PS framework constructs library of candidate policies and learns a meta-policy to select the best one.
result PS consistently outperforms best single policy in heterogeneous CSO problems.
Develops a method to plan exploration that learns strong policies with fewer samples.
problem Lack of efficient exploration in reinforcement learning for real-world tasks.
method Plans an action sequence that maximizes information gain about the optimal trajectory.
result 2x fewer samples than exploration baselines and 200x fewer than model-free methods.
Unified approach to path planning using probabilistic inference on factor graphs.
problem Path planning problems using probabilistic inference.
method Unified framework using probabilistic factor graphs and message composition rules.
result Unified approach includes various algorithms like Sum-product, Max-product, Dynamic programming, and mixed criteria.
This paper offers a methodological contribution at the intersection of machine learning and operations research. Namely, we propose a methodology to quickly predict expected tactical descriptions of operational solutions (TDOSs). The problem we address occurs in the context of two-stage stochastic programming where the…
Continuous state spaces and stochastic, switching dynamics characterize a number of rich, realworld domains, such as robot navigation across varying terrain. We describe a reinforcementlearning algorithm for learning in these domains and prove for certain environments the algorithm is probably approximately correct wit…
Optimal timing for borrowing from a 457(b) plan to maximize returns.
problem Deciding the best time to borrow from a tax-advantaged retirement account.
method Formulated and solved the optimal stopping problem for a loan from a 457(b) plan.
result Derived cutoff rules for optimal loan control, showing how to wait until a certain amount of money is accumulated.
Improves RL planning by proposing sub-goals hierarchically.
problem Sequential planning assumption in RL.
method Divide-and-Conquer Monte Carlo Tree Search (DC-MCTS).
result Improves navigation and control tasks.
New method designs fairer transport plans with uncertainty.
problem Designing fair and balanced mass transport plans.
method Hierarchical fully probabilistic design (HFPD) for transport plans.
result Optimal hyperprior for transport plans with uncertain marginals.