The paper tackles model-based RL's inaccuracy issue by dynamically adjusting planning horizons.
problem Model-based RL's failure due to model inaccuracy over long planning horizons.
method State-dependent planning horizon, learning cumulative model errors with Temporal Difference methods.
result The proposed method successfully adapts planning horizons to state-dependent model accuracy, improving policy learning efficiency.
New algorithm reduces reinforcement learning complexity, approaching contextual bandits.
problem Episodic reinforcement learning's difficulty compared to contextual bandits.
method Proposes MVP algorithm with a new Bernstein-type bonus for episodic reinforcement learning.
result Achieves near-optimal regret bound of $O\left(\left(\sqrt{SAK} + S^2A
ight) \poly\log \left(SAHK
ight)
ight)$, improving state-of-the-art results.
A framework combining HSMM and survival analysis for lifecycle-oriented mobility analysis.
problem Understanding individual metro usage dynamics over multi-year horizons.
method A state-based lifecycle modeling framework integrating HSMM and discrete-time survival analysis.
result Identification of interpretable mobility states, transition dynamics, and state-dependent exit and re-entry processes.
Agents learn state ambiguity from non-linear sensor data using Gaussian approximations.
problem Learning state representation from non-linear sensor data.
method Second-order Taylor approximation of Gaussian distribution for non-linear measurement functions.
result Induces a preference for states based on inferability from observations.
Proposes a new framework for optimizing utility with state-dependent benchmarks.
problem Various interpretations of benchmarks in utility functions.
method General framework of state-dependent utility optimization with stochastic benchmarks.
result Provides optimal solutions and addresses issues of well-definedness and feasibility.
Unique optimal strategy identified for state-dependent risk aversion.
problem Consistency of optimal portfolio choice for varying risk aversion.
method Analysis of state-dependent exponential utilities in arbitrage-free markets.
result Uniqueness of optimal strategy across any time horizon.
A Hawkes process with state-dependent factor models order flows in limit order books.
problem Modeling order flows in limit order books for better market prediction.
method A Hawkes process with a state-dependent factor for conditional intensity estimation.
result State-dependent formulations improve the fit of LOB models to financial data.
Study of SGD with state-dependent noise, improving escape from local minima.
problem Understanding and improving the dynamics of SGD in non-convex optimization.
method Formal study on SGD with state-dependent noise, proposing power-law dynamic with state-dependent diffusion.
result Power-law dynamic can escape from sharp minima exponentially faster than flat minima.
We study statistical aspects of state-dependent Hawkes processes, which are an extension of Hawkes processes where a self- and cross-exciting counting process and a state process are fully coupled, interacting with each other. The excitation kernel of the counting process depends on the state process that, reciprocally…
A softmax operator applied to a set of values acts somewhat like the maximization function and somewhat like an average. In sequential decision making, softmax is often used in settings where it is necessary to maximize utility but also to hedge against problems that arise from putting all of one's weight behind a sing…
Study controlled contagion with state-dependent killing, proving a comparison principle.
problem Analyzing controlled McKean--Vlasov contagion with state-dependent killing.
method Proof of a comparison principle using Wasserstein smooth-gauge comparison and killing-jump absorption estimates.
result Established a comparison principle for the two-population killed-particle HJB.
Theory integrates loss aversion into expected utility for monetary returns.
problem Modeling loss aversion in expected utility theory.
method Develops state-dependent linear utility functions incorporating loss aversion.
result Contracts from monopolists in insurance markets.
This paper improves credit risk analysis by incorporating state-dependent recovery rates into a factor model.
problem Accurate default forecasting in credit risk analysis.
method Extends a one-factor Gaussian copula model to include state-dependent recovery rates and a common factor.
result The proposed model outperforms other models in default prediction, especially during hectic periods.
Study efficient algorithms for nonconvex optimization with state-dependent Markov data.
problem Stochastic optimization with Markovian data and state-dependent transition kernels.
method Projection-based and projection-free algorithms for constrained nonconvex problems.
result The number of oracle calls to achieve an ε-stationary point is O(1/ε2.5). The paper defines and characterizes conditional nonlinear expectations.
problem Defining and characterizing conditional nonlinear expectations.
method Embedding in decision theory, using state-dependent preferences, and continuous utility representation.
result Consistent backward conditional projections are characterized by the Sure-Thing Principle.
Distributed strategic learning has been getting attention in recent years. As systems become distributed finding Nash equilibria in a distributed fashion is becoming more important for various applications. In this paper, we develop a distributed strategic learning framework for seeking Nash equilibria under stochastic…
Most decision theories, including expected utility theory, rank dependent utility theory and cumulative prospect theory, assume that investors are only interested in the distribution of returns and not in the states of the economy in which income is received. Optimal payoffs have their lowest outcomes when the economy …
The paper extends intensity models for limit order books using marked point processes.
problem Modeling intensity ratios in limit order books with state dependency and clustering.
method Developed a new model combining three multiplicative components for marked point processes.
result The new model outperforms other intensity-based methods in predicting market order signs and aggressiveness.
The paper analyzes fill probabilities in limit order books with varying price levels.
problem Determining the likelihood of limit orders being executed in a limit order book.
method Developed a state-dependent stochastic framework to model limit order book dynamics.
result Derived semi-analytical expressions for fill probabilities and mid-price changes.
Self-regulating annealing improves sampling from heavy-tailed datasets.
problem Sampling from heavy-tailed distributions using diffusion models.
method Proposed an SDE-based sampler with a state-dependent diffusion coefficient.
result State dependence induces a self-regulating annealing mechanism.
A new model for forward curves captures behavior through a single equation.
problem Modeling forward curves in a complex function space.
method Developed a stochastic partial differential equation with locally state-dependent coefficients.
result The model retains simplicity while capturing entire forward curve behavior.
We develop a new method to estimate failure probabilities in complex systems.
problem Estimating failure probabilities in safety-critical autonomous systems is challenging due to the rarity of failures and large state spaces.
method We propose an adaptive importance sampling algorithm that minimizes forward Kullback-Leibler divergence and uses Markov score ascent methods.
result Our method provides more accurate failure probability estimates than existing techniques.
In cargo logistics, a key performance measure is transport risk, defined as the deviation of the actual arrival time from the planned arrival time. Neither earliness nor tardiness is desirable for customer and freight forwarders. In this paper, we investigate ways to assess and forecast transport risks using a half-yea…
In a dual risk model, the premiums are considered as the costs and the claims are regarded as the profits. The surplus can be interpreted as the wealth of a venture capital, whose profits depend on research and development. In most of the existing literature of dual risk models, the profits follow the compound Poisson …
In this paper, we consider the asset-liability management under the mean-variance criterion. The financial market consists of a risk-free bond and a stock whose price process is modeled by a geometric Brownian motion. The liability of the investor is uncontrollable and is modeled by another geometric Brownian motion. W…
Innovative extensions to option pricing models using asymmetric Brownian motion and random walk approaches.
problem Capturing empirical phenomena like return skewness, heavy tails, and volatility asymmetry in option pricing models.
method Developing the Geometric Asymmetric Brownian Motion (GABM) within the Bachelier--Black--Scholes--Merton framework.
result Deriving closed-form option pricing formulas and a discrete-time binomial tree algorithm that converges to the GABM limit.
New algorithm speeds up MCMC for complex distributions.
problem Efficient sampling from complex, high-dimensional distributions.
method Numerical Generalized Randomized Hamiltonian Monte Carlo with state-dependent event rates.
result Approximates Hamiltonian trajectories for robust sampling.
The paper solves a consumption-investment problem with state-dependent lower bounds.
problem A life-time consumption-investment problem with a state-dependent lower bound on consumption.
method Transformed the problem into a state-independent control problem to apply standard theory.
result Explicit optimal strategies provided for both homogeneous and non-homogeneous constraints.
New volatility model for option pricing with time-varying risk premium.
problem Volatility risk premium is time-varying and not well captured by existing models.
method Combines Markov switching with Realized GARCH framework to derive a state-dependent pricing kernel.
result The model reduces option pricing errors by 15% or more compared to competing models.
The paper proposes a machine learning approach for state-dependent asset allocation.
problem Market conditions cause performance deviations from long-term averages.
method Analyzes historical market states and asset returns to directly relate state variables to portfolio weights.
result The proposed approach generates a more efficient portfolio compared to traditional methods.
In this article, we consider a Markov process X, starting from x and solving a stochastic differential equation, which is driven by a Brownian motion and an independent pure jump component exhibiting state-dependent jump intensity and infinite jump activity. A second order expansion is derived for the tail probability …
Bitcoin reacts positively to USDT minting but not burning, showing state-dependence.
problem Understanding Bitcoin's response to Tether's supply changes.
method Analyzing Bitcoin's intraday price movements in response to USDT minting and burning events.
result Bitcoin's response to USDT minting events declines after 60 minutes and is influenced by investor sentiment and public announcements.
Measures price impact in order-driven markets without relying on averages.
problem Measuring price impact in order-driven markets without relying on averages.
method Modeling the limit order book using state-dependent Hawkes processes and defining price impact profile as a function of the compensator of a stochastic process.
result The clustering of sell child orders has a bigger impact on price than their sizes.
Probabilistic proof of smooth boundaries in optimal stopping problems.
problem Continuous differentiability of time-dependent optimal boundaries in optimal stopping problems.
method Local probabilistic arguments for a wider range of conditions.
result First probabilistic proof of continuous differentiability under general conditions.
A new model predicts bid-ask spread dynamics in financial markets.
problem Capturing the self-exciting nature of bid-ask spread changes.
method State-dependent Spread Hawkes model (SDSH) incorporating various spread jump sizes and current state impact.
result The SDSH model accurately forecasts spread values at short-term horizons.
New method designs fairer transport plans with uncertainty.
problem Designing fair and balanced mass transport plans.
method Hierarchical fully probabilistic design (HFPD) for transport plans.
result Optimal hyperprior for transport plans with uncertain marginals.
Improves RL planning by proposing sub-goals hierarchically.
problem Sequential planning assumption in RL.
method Divide-and-Conquer Monte Carlo Tree Search (DC-MCTS).
result Improves navigation and control tasks.
We consider reinforcement learning in input-driven environments, where an exogenous, stochastic input process affects the dynamics of the system. Input processes arise in many applications, including queuing systems, robotics control with disturbances, and object tracking. Since the state dynamics and rewards depend on…
We aim to reduce the burden of programming and deploying autonomous systems to work in concert with people in time-critical domains, such as military field operations and disaster response. Deployment plans for these operations are frequently negotiated on-the-fly by teams of human planners. A human operator then trans…
Study integrates reliability constraints into generation planning models.
problem Challenges in integrating reliability constraints with generation planning models.
method Leverages a weighted oblique decision tree (WODT) technique to embed reliability verification constraints.
result Demonstrates effectiveness in achieving reliable and optimal planning solutions.
Study shows how sentiment shocks affect equity markets, revealing asymmetries and state-dependent effects.
problem Understanding how sentiment shocks propagate through equity markets and their impact on different investor groups.
method Used four independent proxies with sign-aligned kappa-rho parameters, calibrated a structural model to link sentiment to returns.
result A one standard deviation sentiment shock has a 1.06 basis point impact, with effects amplified over 11.2 months and concentrated in retail-tilted stocks.
Neural Lévy model improves risk and density forecasting for financial returns.
problem Financial returns exhibit heavy tails, volatility clustering, and jumps.
method Proposes a neural Lévy jump-diffusion framework that learns conditional drift, diffusion, jump intensity, and size distribution.
result Demonstrates improved calibration, sharper tail control, and risk reduction.
New approach improves black-box planning efficiency by discovering focused macros.
problem Difficulty of deterministic planning increases exponentially with depth.
method Discovering macro-actions with focused effects to improve goal-count heuristics.
result Focused macros dramatically improve black-box planning efficiency.
Study motion planning for points avoiding obstacles in a plane.
problem Avoiding collisions for multiple points in a plane with unknown obstacles.
method Algebraic and topological tools for motion planning.
result New topological complexity for planar motion planning.
Selective planning with imperfect models reduces harmful effects of model inadequacy.
problem Harmful effects of using an imperfect model in reinforcement learning.
method Selective planning with heteroscedastic regression to estimate predictive uncertainty from model inadequacy.
result Effective selective planning requires considering both parameter uncertainty and model inadequacy.
CoMPNetX uses neural networks to efficiently solve constrained motion planning problems.
problem Finding collision-free paths on constraint manifolds efficiently.
method Neural generator and discriminator with neural gradients-based projection operator.
result CoMPNetX finds path solutions with high success rates and lower computation times.
This article asks how planning scholarship may effectively gain impact in planning practice through media exposure. In liberal democracies the public sphere is dominated by mass media. Therefore, working with such media is a prerequisite for effective public impact of planning research. Using the example of megaproject…
We introduce Dynamic Planning Networks (DPN), a novel architecture for deep reinforcement learning, that combines model-based and model-free aspects for online planning. Our architecture learns to dynamically construct plans using a learned state-transition model by selecting and traversing between simulated states and…