Deep RL drone trained to compete against classical path planning in drone racing.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
In model-based reinforcement learning, the agent interleaves between model learning and planning. These two components are inextricably intertwined. If the model is not able to provide sensible long-term prediction, the executed planner would exploit model flaws, which can yield catastrophic failures. This paper focuse…
The hierarchical structure of production planning has the advantage of assigning different decision variables to their respective time horizons and therefore ensures their manageability. However, the restrictive structure of this top-down approach implying that upper level decisions are the constraints for lower level …
This work tackles long-term visual planning by goal-conditioned hierarchical predictors.
Study integrates reliability constraints into generation planning models.
Due to the threat of climate change, a transition from a fossil-fuel based system to one based on zero-carbon is required. However, this is not as simple as instantaneously closing down all fossil fuel energy generation and replacing them with renewable sources -- careful decisions need to be taken to ensure rapid but …
For any business, planning is a continuous process, and typically business-owners focus on making both long-term planning aligned with a particular strategy as well as short-term planning that accommodates the dynamic market situations. An ability to perform an accurate financial forecast is crucial for effective plann…
Learning and inference movement is a very challenging problem due to its high dimensionality and dependency to varied environments or tasks. In this paper, we propose an effective probabilistic method for learning and inference of basic movements. The motion planning problem is formulated as learning on a directed grap…
We address the problem of Bayesian reinforcement learning using efficient model-based online planning. We propose an optimism-free Bayes-adaptive algorithm to induce deeper and sparser exploration with a theoretical bound on its performance relative to the Bayes optimal policy, with a lower computational complexity. Th…
New AI model optimizes personalized care for elderly residents.
This work tackles maintenance planning with deep reinforcement learning under uncertainty.
We generalize the classic Shiller cyclically adjusted price-earnings ratio (CAPE) used for prediction of future total returns of the stock market. We treat earnings growth as exogenous. The difference between log wealth and log earnings is modeled as an autoregression of order 1 with linear trend 4.6% and Gaussian inno…
A new approach for deep exploration in sparse reward reinforcement learning.
Reward tweaking optimizes behavior for long-term goals by adjusting the reward function.
Framework learns useful subgoals from demonstrations and instructions.
This paper proposes a framework to predict long-term trends and short-term fluctuations in multivariate time series.
TD-Flow improves long-term predictions in agent learning.
An accurate load forecasting has always been one of the main indispensable parts in the operation and planning of power systems. Among different time horizons of forecasting, while short-term load forecasting (STLF) and long-term load forecasting (LTLF) have respectively got benefits of accurate predictors and probabil…
Framework uses RL and simulation for optimal microgrid energy storage planning.
Trial-and-error based reinforcement learning (RL) has seen rapid advancements in recent times, especially with the advent of deep neural networks. However, the majority of autonomous RL algorithms require a large number of interactions with the environment. A large number of interactions may be impractical in many real…
The study improves Monte Carlo simulations for long-term investments using advanced financial models.
Proposes a new model for predicting future motion of road actors in autonomous vehicles.
Behavioral theories posit that investor sentiment exhibits predictive power for stock returns, whereas there is little study have investigated the relationship between the time horizon of the predictive effect of investor sentiment and the firm characteristics. To this end, by using a Granger causality analysis in the …
We propose a simple technique for encouraging generative RNNs to plan ahead. We train a "backward" recurrent network to generate a given sequence in reverse order, and we encourage states of the forward model to predict cotemporal states of the backward model. The backward network is used only during training, and play…
UCRL-WVTR tackles long-term reinforcement learning with general approximations, achieving horizon-free and instance-dependent regret bounds.
New method for selecting clusters in residential electricity data.
We explore the use of deep learning and deep reinforcement learning for optimization problems in transportation. Many transportation system analysis tasks are formulated as an optimization problem - such as optimal control problems in intelligent transportation systems and long term urban planning. Often transportation…
The paper predicts travel times using tree-based ensembles.
Develops hierarchical reinforcement learning value function approximators.
A new method solves complex hydroelectricity planning problems.
The increasing penetration level of energy generation from renewable sources is demanding for more accurate and reliable forecasting tools to support classic power grid operations (e.g., unit commitment, electricity market clearing or maintenance planning). For this purpose, many physical models have been employed, and…
LemonadeBench evaluates LLMs' economic intuition through a simulated lemonade stand.
We consider the portfolio choice problem for a long-run investor in a general continuous semimartingale model. We suggest to use path-wise growth optimality as the decision criterion and encode preferences through restrictions on the class of admissible wealth processes. Specifically, the investor is only interested in…
The paper analyzes how mutable blockchain protocols affect miner behavior and strategic stability.
New method designs fairer transport plans with uncertainty.
Improves RL planning by proposing sub-goals hierarchically.
We aim to reduce the burden of programming and deploying autonomous systems to work in concert with people in time-critical domains, such as military field operations and disaster response. Deployment plans for these operations are frequently negotiated on-the-fly by teams of human planners. A human operator then trans…
We introduce here for the first time the long-term swap rate, characterised as the fair rate of an overnight indexed swap with infinitely many exchanges. Furthermore we analyse the relationship between the long-term swap rate, the long-term yield, see Biagini et al. [2018], Biagini and Härtel [2014], and El Karoui et a…
A framework combining HSMM and survival analysis for lifecycle-oriented mobility analysis.
Kernel method estimates long-term effects from short-term data.
New approach improves black-box planning efficiency by discovering focused macros.
Study motion planning for points avoiding obstacles in a plane.
Selective planning with imperfect models reduces harmful effects of model inadequacy.
CoMPNetX uses neural networks to efficiently solve constrained motion planning problems.
This article asks how planning scholarship may effectively gain impact in planning practice through media exposure. In liberal democracies the public sphere is dominated by mass media. Therefore, working with such media is a prerequisite for effective public impact of planning research. Using the example of megaproject…
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning. Our architecture learns to dynamically construct plans using a learned state-transition model by selecting and traversing between simulated states and…
New method combines heuristics and search techniques to speed up cooperative planning for autonomous vehicles.
This paper presents a unifying framework for reinforcement learning and planning.