New approach learns human planning algorithms for reward inference.
problem Learning reward functions from human demonstrations with biases.
method Data-driven approach to learn planning algorithms directly from demonstrations.
result Mixed results: better reward inference but at cost of differentiability.
End-to-end learnable network for safer self-driving with interpretable intermediate representations.
problem Safe motion planning for self-driving vehicles.
method Differentiable semantic occupancy representation for cost calculation in motion planning.
result Significantly outperforms state-of-the-art planners in imitating human behaviors and producing safer trajectories.
Financial planners helped preserve and increase household net financial assets during the Great Recession.
problem Impact of financial planners on household net financial assets during the Great Recession.
method Utilized 2007-2009 Survey of Consumer Finances (SCF) panel dataset, analyzed 3,862 respondents.
result Starting to use a financial planner during the Great Recession had a positive impact on preserving and increasing household net financial assets.
Automated planning is one of the foundational areas of AI. Since no single planner can work well for all tasks and domains, portfolio-based techniques have become increasingly popular in recent years. In particular, deep learning emerges as a promising methodology for online planner selection. Owing to the recent devel…
Neural A* uses machine learning to improve path planning efficiency.
problem Challenges in applying machine learning to search-based path planning.
method Reformulated A* search as a differentiable network coupled with a convolutional encoder.
result Neural A* outperforms state-of-the-art planners in optimality and efficiency.
We introduce a framework for model learning and planning in stochastic domains with continuous state and action spaces and non-Gaussian transition models. It is efficient because (1) local models are estimated only when the planner requires them; (2) the planner focuses on the most relevant states to the current planni…
Study shows how multiple traders can trade together without excessive price impact.
problem Coordination issues in trading to exploit a common signal.
method Closed-loop Nash competition model for stochastic differential games.
result Excessive trading reduced but not significantly for practical parameters.
Study optimal investment decisions for diverse risk-tolerant agents.
problem Optimizing investment choices for agents with varying risk preferences.
method Characterizes optimal behavior using certainty equivalents and lognormal risks.
result Derives optimal decision menus under known and uncertain preference distributions.
We introduce a variant of Farber's topological complexity, defined for smooth compact orientable Riemannian manifolds, which takes into account only motion planners with the lowest possible "average length" of the output paths. We prove that it never differs from topological complexity by more than 1, thus showing th…
Investment strategies for rank-dependent utility agents are derived in a continuous-time market.
problem Time inconsistency in rank-dependent utility models.
method Study of consistent planners seeking intra-personal equilibrium strategies.
result Explicit final wealth profile replicating equilibrium strategies, with scaling function derived.
New research shows exponential lower bounds for planning in MDPs with linearly-realizable optimal action-value functions.
problem Determining the minimum number of queries needed for sound planners in MDPs with linear function approximation.
method Analyzing fixed-horizon and discounted MDPs with a generative model, showing lower bounds on the number of queries required.
result Sound planners need at least exponential number of queries in both fixed-horizon and discounted settings.
Develops an equilibrium model for securities pricing in a mixed cooperative and non-cooperative market.
problem Equilibrium pricing of securities in a market with cooperative and non-cooperative agents.
method Conditional extended mean-field control for cooperative agents, mean-field model for both cooperative and non-cooperative agents.
result Existence of a unique equilibrium for both finite-agent and mean-field models under certain conditions.
The paper optimizes pension policies with guarantees and sustainability constraints.
problem Designing optimal pension policies with guarantees and sustainability constraints.
method Dynamic utility model, stochastic domain, overlapping generations, time-consistent decision criterion.
result Optimal investment/pension policy computed for a general framework.
A reinforcement learning framework combining value function and tree search planner for strategic and tactical decisions.
problem Strategic and tactical decision-making in discrete environments.
method Combines value function and tree search planner, using uncertainty modeling and risk measurement.
result Improves performance and learning speed on hard exploration environments.
Value Iteration Networks (VINs) are effective differentiable path planning modules that can be used by agents to perform navigation while still maintaining end-to-end differentiability of the entire architecture. Despite their effectiveness, they suffer from several disadvantages including training instability, random …
Turnpike property applies to optimal control of PDEs and ResNets.
problem Optimal control of PDEs and ResNets.
method Mathematical formalization and controllability analysis.
result Optimal controls and states are nearly constant over most of the time.
Agent learns to navigate uncertain 3D maps using a hybrid planner.
problem Planning in 3D environments with uncertain topological maps.
method Hierarchical strategy combining graph planner and local policy, data-driven learning with neural network.
result Machine learning can overcome missing information in probabilistic topological maps.
Study presents a low-cost local motion planner for vineyard navigation.
problem Autonomous navigation in vineyards with limited resources.
method RGB-D camera, dual layer control algorithm, deep learning synergy.
result Robust motion planning for vineyard navigation achieved.
Algorithm learns fair division from noisy feedback in uncertain markets.
problem Learning fair division in uncertain markets with noisy feedback.
method Wrapper algorithms using dual averaging to learn item and agent values from bandit feedback.
result Asymptotically achieves optimal Nash social welfare in linear Fisher markets.
Improved model-based reinforcement learning for interactive dialogue tasks reduces sample needs and improves performance.
problem Limited data and high sample cost in interactive dialogue systems.
method Model-based actor-critic approach with an environment model and planner.
result 70 times fewer samples required compared to baseline model-free algorithm, with 2x better asymptotic performance.
XLVINs improve data efficiency in implicit planning by leveraging latent space.
problem Improving data efficiency in implicit planning algorithms.
method XLVINs use a high-dimensional latent space to perform planning computations, breaking the algorithmic bottleneck.
result XLVINs significantly improve data efficiency across various settings compared to value iteration-based implicit planners and model-free baselines.
Bayesian uncertainty from deep learning improves robot safety in unknown areas.
problem Improving robot safety in unknown or dangerous environments.
method Bayesian approximations of uncertainty from deep learning in a robot planner.
result Incorporating uncertainty leads to 18% less risky paths.
Model analyzes optimal interbank networks during liquidity shocks, revealing core-periphery structures and co-investment requirements.
problem Formation of optimal interbank networks during liquidity shocks.
method Solves system-wide optimal control problem in two settings: decentralized and centralized.
result Decentralized setting leads to less cash reserves and greater vulnerability to shocks; core banks have highest co-investment requirements.
New approach improves black-box planning efficiency by discovering focused macros.
problem Difficulty of deterministic planning increases exponentially with depth.
method Discovering macro-actions with focused effects to improve goal-count heuristics.
result Focused macros dramatically improve black-box planning efficiency.
New algorithm learns FMDP structure while minimizing regret.
problem Regret minimization in FMDPs with unknown structure.
method Optimism in face of uncertainty principle combined with statistical structure learning.
result First algorithm to learn FMDP structure while minimizing regret.
Neural planners for RDDL MDPs produce deep reactive policies in an offline fashion. These scale well with large domains, but are sample inefficient and time-consuming to train from scratch for each new problem. To mitigate this, recent work has studied neural transfer learning, so that a generic planner trained on othe…
In model-based reinforcement learning, the agent interleaves between model learning and planning. These two components are inextricably intertwined. If the model is not able to provide sensible long-term prediction, the executed planner would exploit model flaws, which can yield catastrophic failures. This paper focuse…
DDPD separates generation into planning and denoising for improved efficiency.
problem Efficiently denoise corrupted data during generation.
method Separates generation into a planner and denoiser, selecting denoising positions based on corruption severity.
result DDPD outperforms traditional methods on language and image generation benchmarks.
TIM framework uses LLMs and domain experts to infer DeFi user transaction intents.
problem Challenges in understanding user intent in DeFi transactions due to complex interactions and opaque logs.
method TIM framework leverages a DeFi intent taxonomy, multi-agent LLM system, and a Meta-Level Planner.
result TIM significantly outperforms existing methods in inferring user transaction intents.
In this paper we study continuous-time stochastic control problems with both monotone and classical controls motivated by the so-called public good contribution problem. That is the problem of n economic agents aiming to maximize their expected utility allocating initial wealth over a given time period between private …
With a point of departure in the concept "uncomfortable knowledge," this article presents a case study of how the American Planning Association (APA) deals with such knowledge. APA was found to actively suppress publicity of malpractice concerns and bad planning in order to sustain a boosterish image of planning. In th…
New algorithm for traffic routing in congested conditions.
problem Optimal routing in congested traffic networks.
method Congested Bandits model, UCB algorithm, iterative least squares planner.
result No-regret algorithms for congestion-aware routing.
A method uses RL to learn abstractions for planning, improving robot navigation and manipulation tasks.
problem Planning requires suitable abstractions for states and transitions, which RL struggles with for temporally extended tasks.
method Goal-conditioned policies learned with RL are incorporated into planning, with a latent variable model representing valid states.
result Our method significantly outperforms prior work on image-based robot navigation and manipulation tasks.
Semi-analytical approach for optimal wealth management contributions.
problem Optimizing contributions to achieve a financial goal with uncertain returns.
method Controlled backward Kolmogorov equation and Schrodinger equation solution.
result Semi-analytical solutions for efficient frontiers in control space.
Study optimal stopping for group with diverse discount rates using an attitude function.
problem Optimal stopping for a group with diverse discount rates under an aggregation preference.
method Develop iterative approach using consistent planning for time-consistent equilibria.
result Characterize all time-consistent mild equilibria as fixed points of an operator.
This article presents results from the first statistically significant study of traffic forecasts in transportation infrastructure projects. The sample used is the largest of its kind, covering 210 projects in 14 nations worth US$59 billion. The study shows with very high statistical significance that forecasters gener…
The study shows how probability weighting can lead to betting in a risk-averse economy.
problem Understanding how probability weighting affects economic behavior and risk aversion.
method Examining a von Neumann-Morgenstern economy with an RDU agent to model probability weighting effects.
result Probability weighting can lead to endogenous betting in an economy with common beliefs.
This paper presents a general framework for studying diverse beliefs in dynamic economies. Within this general framework, the characterization of a central-planner general equilbrium turns out to be very easy to derive, and leads to a range of interesting applications. We show how for an economy with log investors hold…
Adaptive financial dataflow system improves model robustness in dynamic markets.
problem Static historical data leads to poor performance in dynamic financial markets.
method Drift-aware dataflow system with adaptive control and optimization.
result Enhanced model robustness and improved risk-adjusted returns.
Recent research in economic theory attempts to study optimal economic growth and spatial location of economic activity in a unified framework. So far, the key result of this literature - asymptotic convergence, even in the absence of decreasing returns to capital - relies on specific assumptions about the objective of …
We present the perceptor gradients algorithm -- a novel approach to learning symbolic representations based on the idea of decomposing an agent's policy into i) a perceptor network extracting symbols from raw observation data and ii) a task encoding program which maps the input symbols to output actions. We show that t…
Optimal scheme minimizes deviation in federated transfer learning for kernel regression.
problem Minimizing cumulative deviation in federated transfer learning across multiple datasets.
method Regret-optimal iterative scheme for continual communication between nodes and server.
result Explicit updates for the regret-optimal algorithm in finite-rank kernel regression.
In this paper, we consider the problem of optimal reinsurance design, when the risk is measured by a distortion risk measure and the premium is given by a distortion risk premium. First, we show how the optimal reinsurance design for the ceding company, the reinsurance company and the social planner can be formulated i…
Italy and the Eurozone are heading in the year 2012 into a financial depression of unprecedented magnitude, with a forthcoming multitude of often contradictory public economic and financial stability emergency interventions whose ultimate endogenous and exogenous effects on public and private health spending and on the…
Naive investors make riskier choices than optimal strategies in continuous-time finance.
problem Continuous-time Markowitz portfolio selection with naive reoptimization.
method Analytical derivation of naive policies from discretely naive policies.
result Naive policies are always riskier and less efficient than equilibrium policies.
We study Nash equilibria for inventory-averse high-frequency traders (HFTs), who trade to exploit information about future price changes. For discrete trading rounds, the HFTs' optimal trading strategies and their equilibrium price impact are described by a system of nonlinear equations; explicit solutions obtain aroun…
Study minimizes market inefficiency in systemic economies.
problem Minimizing deviations of market prices from fundamental values.
method Characterized market inefficiency and developed a matrix of holdings to minimize it.
result Portfolio holdings should deviate more from diversification if banks have similar systemic significance.
Partially observable Markov decision processes (POMDPs) are a powerful abstraction for tasks that require decision making under uncertainty, and capture a wide range of real world tasks. Today, effective planning approaches exist that generate effective strategies given black-box models of a POMDP task. Yet, an open qu…