Study optimal pricing and inventory control in dynamic settings with censored demand.
problem Optimal pricing and inventory control in dynamic settings with censored demand.
method Approximate optimal policy via high-order MDP, propose novel algorithms for solving Bellman equations.
result Established finite-sample regret bounds and demonstrated efficacy through numerical experiments.
A new method for inventory control using in-context learning and generative models.
problem Inventory control with decision-dependent censoring, focusing on the censored newsvendor problem.
method In-context generative posterior sampling (ICGPS) combining modern generative models and in-context autoregressive generation.
result ICGPS achieves sublinear Bayesian regret for the censored newsvendor problem, outperforming existing methods.
We propose a continuous-time stock-flow consistent model for inventory dynamics in an economy with firms, banks, and households. On the supply side, firms decide on production based on adaptive expectations for sales demand and a desired level of inventories. On the demand side, investment is determined as a function o…
MaxCOSD algorithm tackles non-i.i.d. demands and stateful dynamics in online inventory control.
problem Managing inventory with non-i.i.d. demands and stateful dynamics.
method MaxCOSD, an online algorithm with provable guarantees for non-degeneracy assumptions.
result MaxCOSD achieves optimal performance for non-i.i.d. demands and stateful dynamics.
This paper tackles inventory control with general arrival dynamics and post-processing, improving profitability.
problem Inventory control with arbitrary arrival dynamics and post-processing constraints.
method Formulated as an exogenous decision process, incorporating deep generative models for arrivals, and applying supervised learning techniques.
result Improves profitability over production baselines and real-world A/B test data.
A contextual bandit method evaluates and improves inventory control policies.
problem Evaluating and improving periodic review inventory control policies with nonstationary demand.
method Contextual bandit-based algorithm to evaluate and tweak policies.
result The method achieves favorable guarantees in both theory and practice.
Optimistic pricing algorithm handles online dynamic pricing with censored demand.
problem Online dynamic pricing with censoring of potential demand.
method Optimistic estimates of derivatives for pricing algorithm.
result Achieves i l d e O ( T ) ilde{O}(\sqrt{T}) i l d e O ( T ) optimal regret against adversarial inventory series. Model analyzes RFQ markets using stochastic control to optimize dealer performance and inventory.
problem Optimizing market making in aggregator-routed RFQ markets with varying dealer performance scores.
method Two-tier stochastic control model that separates RFQ-level price competition from macro routing.
result Optimal controls can be expressed through derivatives of reduced Hamiltonians, leading to interpretable mappings from optimal win probabilities to optimal offsets.
Paper combines RL with policy regularization for inventory policies.
problem Optimizing inventory policies using RL and dynamic programming.
method Hybrid approach combining RL with policy regularization.
result Generalization guarantees for inventory policies using VC theory.
Many retailers today employ inventory management systems based on Re-Order Point Policies, most of which rely on the assumption that all decreases in product inventory levels result from product sales. Unfortunately, it usually happens that small but random quantities of the product get lost, stolen or broken without r…
Optimal vehicle repositioning policy found for shared mobility services.
problem Matching fixed supply with spatial customer demand under uncertain and correlated demand.
method Base-stock repositioning policy, asymptotic optimality, regret analysis, adaptive repositioning algorithm.
result Surrogate Optimization and Adaptive Repositioning algorithm achieves optimal regret of O ( n 2.5 T ) O(n^{2.5} \sqrt{T}) O ( n 2.5 T ) . Study on inventory control with changing demand, proposing adaptive algorithms.
problem Inventory control with non-stationary demand distributions.
method Adaptive online algorithms optimizing base-stock policies.
result Sharp separation in adaptability across different inventory models.
A new approach integrates inventory prediction and routing optimization for better supply chain management.
problem Optimizing efficient route selection in supply chain management with uncertain inventory demand.
method Decision-focused learning approach using neural networks to directly integrate inventory prediction and routing optimization.
result Direct integration of inventory prediction and routing optimization leads to better supply chain decisions.
We study the cross-correlation matrix C i j C_{ij} C ij of inventory variations of the most active individual and institutional investors in an emerging market to understand the dynamics of inventory variations. We find that the distribution of cross-correlation coefficient C i j C_{ij} C ij has a power-law form in the bulk followed by …
A PID-based feedback-control system improves multiple KPIs in RTB display advertising.
problem Challenges in simultaneously improving multiple KPIs in RTB campaigns.
method Sequential Control using PID-based feedback and importance metrics.
result Effective in simultaneously controlling multiple KPIs in both simulations and live traffic.
Study visualizes actor-critic loss landscapes for inventory optimization.
problem Difficulties in solving multi-store dynamic inventory control problems.
method Low-dimensional visualizations of actor loss function.
result Loss landscapes favor optimal policies in reinforcement learning.
Unified framework connects two market-making models, revealing their underlying equivalence.
problem Independent calibration of two market-making frameworks (Avellaneda-Stoikov and Cartea-Jaimungal).
method Axiomatic approach to market preference functional, showing equivalence under specific conditions.
result Avellaneda-Stoikov and Cartea-Jaimungal frameworks are equivalent under certain conditions.
Optimal hidden-target learning for online inventory optimization on general convex sets.
problem Online inventory optimization (OIO) on arbitrary bounded convex capacity sets.
method Maintaining a hidden target and projecting it onto the feasible order-up-to set.
result The method improves the best known regret guarantee for OIO on general convex sets from inverse to inverse-square-root dependence on the common-demand probability.
Investigates market dynamics with informed traders and high-frequency traders.
problem Trading large orders in a market with multiple high-frequency traders.
method Analyzes a three-period Kyle's model with a normal-speed informed trader and multiple anticipatory high-frequency traders under different inventory pressures.
result Surprising results: improving HFTs' speed or prediction can harm them but benefit the informed trader.
A new algorithm avoids worst-case outcomes in risky contexts.
problem Risk-averse behavior in contextual bandits is challenging.
method Developed a first risk-averse contextual bandit algorithm with online regret guarantees.
result First algorithm with an online regret guarantee for risk-averse contextual bandits.
Develops methods for dynamic pricing in incomplete data settings.
problem Incomplete historical data makes optimal pricing difficult.
method Nonparametric partial identification framework for offline dynamic pricing.
result Pessimistic and opportunistic policies with regret bounds.
Apparently random financial fluctuations often exhibit varying levels of complexity, chaos. Given limited data, predictability of such time series becomes hard to infer. While efficient methods of Lyapunov exponent computation are devised, knowledge about the process driving the dynamics greatly facilitates the complex…
New method learns credit prices offline without interaction.
problem Dynamic pricing of consumer credit.
method Offline deep reinforcement learning with Q-Learning.
result Effective personalized pricing policy learned without online interaction.
Machine learning reveals inventory effects on VSTOXX futures pricing.
problem Understanding how inventory affects VSTOXX futures pricing.
method Combining stochastic processes and machine learning, we formulate and calibrate a Heston model for VSTOXX futures pricing.
result Machine learning models show that inventory significantly impacts VSTOXX futures prices.
New method reduces inventory inaccuracies by 10x, saving retailers 4% annually.
problem Inaccurate inventory records cost retailers 4% annually, and manual detection is impractical.
method Proposes a new anomaly detection method for low-rank Poisson matrices using cross-sectional data.
result Our approach reduces anomaly detection costs by up to 10x compared to existing methods.
Paper analyzes sample complexity for offline RL with deep ReLU networks.
problem Theoretical analysis of sample complexity for offline RL with deep ReLU networks.
method Establishes sample complexity for offline RL with deep ReLU networks, considering Besov dynamic closure and correlated structure.
result First theoretical characterization of sample complexity for offline RL with deep neural network function approximation.
The paper proposes autoregressive models for better offline RL.
problem Offline RL policy evaluation and optimization challenges.
method Autoregressive dynamics models for sequential state and reward prediction.
result Autoregressive models outperform standard methods in log-likelihood and RL tasks.
A dealer manages quotes and rejection rules to control slippage risk in FX markets.
problem Managing inventory risk and latency risk in OTC FX market making.
method Dynamic programming and adiabatic-quadratic approximation to optimize quotes and rejection rules.
result Developed a method to optimize quotes and rejection rules for managing slippage risk.
BOMS enhances offline MBRL by improving model selection with Bayesian optimization.
problem Inaccurate model selection in offline MBRL due to distribution shift.
method Proposes BOMS, an active model selection framework using Bayesian optimization.
result Improves model selection with only a small amount of online interaction.
We consider a repeated newsvendor problem where the inventory manager has no prior information about the demand, and can access only censored/sales data. In analogy to multi-armed bandit problems, the manager needs to simultaneously "explore" and "exploit" with her inventory decisions, in order to minimize the cumulati…
Algorithm learns optimal dynamic mechanisms from data.
problem Designing optimal mechanisms for dynamic settings with unknown reward functions.
method Offline reinforcement learning with pessimism principle.
result Learned mechanisms are efficient, individually rational, and truthful.
Study on inventory management under uncertainty using smooth ambiguity preference.
problem Managing inventory under Knightian uncertainty with smooth ambiguity preference.
method Demonstrates continuous-time smooth ambiguity as the infinitesimal limit of Kalman-Bucy filtering with recursive robust utility. Solves forward-backward stochastic differential equations with quadratic growth to determine cost function. Derives value function and optimal control policy using variational inequalities and viscosity solutions. Transforms problem into two-dimensional singular control.
result Ambiguity drives decision-makers to act earlier, reducing the continuation region.
Modeling option market making with hedging-induced price impact.
problem Tackles the challenge of market making in options markets with price impact.
method Models option order flow using Cox processes and studies the dynamics of inventory and price under hedging-induced impact.
result Establishes the well-posedness of the mixed control problem involving quoting and hedging.
Proposes EDESH-SA for better inventory management under uncertainty.
problem Inventory management under uncertainty.
method Ensemble Differential Evolution with simulation-based hybridization and self-adaptation.
result Improves financial performance and optimizes search spaces.
Bayesian nonparametrics enables dynamic skill discovery from expert data.
problem Fixed K for offline skill discovery in reinforcement learning.
method Variational inference, continuous relaxations, Bayesian nonparametrics.
result Nonparametric model with dynamically-changing number of options.
MOOSE improves offline RL robustness by using dynamics models.
problem Low robustness of model-free offline RL algorithms in industrial settings.
method MOOSE uses dynamics models to assess policy performance, keeping policies within data support.
result MOOSE outperforms state-of-the-art model-free offline RL algorithms in robust performance.
Identifies latent actions and dynamics from offline data with diverse demonstrators.
problem Recovering latent actions and environment dynamics from action-free trajectories.
method Assumes distinct policies for each demonstrator, identifies latent transitions and policies via matrix factorization.
result Identifies latent transitions and demonstrator policies up to permutation.
Bayesian optimization improves forest inventory sampling using remote sensing data.
problem Optimizing forest inventory sampling in large areas with limited data.
method Bayesian optimization applied to RS data for improved sampling design.
result The proposed method outperforms baseline methods in terms of MSE values.
In this paper we extend the market-making models with inventory constraints of Avellaneda and Stoikov ("High-frequency trading in a limit-order book", Quantitative Finance Vol.8 No.3 2008) and Gueant, Lehalle and Fernandez-Tapia ("Dealing with inventory risk", Preprint 2011) to the case of a rather general class of mid…
New method learns policies from offline data using operator models.
problem Limited understanding of approximation errors in offline reinforcement learning.
method Linking reinforcement learning to Hamilton-Jacobi-Bellman equation, proposing operator-theoretic algorithm.
result Global convergence of the value function and finite-sample guarantees derived.
Study finds inventory inaccuracies are linked to store activity and product perishability.
problem Inventory record inaccuracy in grocery retailing environments.
method Analysis of 24,000 SKUs across 11 stores, field quasi-experiment on audits.
result Inventory audits can boost sales by 11%, especially for perishable items.
Optimal market making strategy with price forecasts reduces inventory costs and spreads.
problem Optimal market making strategy with price forecasts reduces inventory costs and spreads.
method Modeling market making strategy with linear price impact, random slope and intercept, and simultaneous order arrivals.
result Simultaneous order arrivals and price forecasts reduce inventory costs and spreads.
Model analyzes competitive pricing strategies in large markets of perishable products.
problem Maximizing profits in a competitive market of perishable products.
method Mean-field competition model, Hamilton-Jacobi-Bellman equation, iterative numerical algorithm.
result Properties of equilibrium pricing strategies and market dynamics.
A new algorithm uses IVs to learn optimal policies from observational data.
problem Learning optimal policies from unobserved variable confounded data.
method IV-aided Value Iteration (IVVI) algorithm based on conditional moment restrictions.
result First provably efficient algorithm for instrument-aided offline RL.
New algorithms tackle robust RL with linear models, revealing unique challenges.
problem Distributionally robust offline RL with uncertainty in dynamics.
method Proposes minimax optimal and computationally efficient algorithms using novel function approximation mechanisms.
result Function approximation in robust offline RL is distinct and harder than in standard offline RL.
Proposes a method to learn policies from offline data with reduced bias.
problem Learning policies from offline data with reduced bias and complexity constraints.
method Cross-fitted debiasing device for policy learning from offline data.
result Achieves N \sqrt N N regret for complex policy classes with a product-of-errors nuisance remainder. The paper tackles reward-relevance in offline RL with sparse decision dynamics.
problem Offline reinforcement learning with sparse decision dynamics and estimation sparsity.
method Reward-filtered least-squares policy evaluation using thresholded lasso.
result The method provides theoretical guarantees with sample complexity dependent on sparse component size.
MOPO optimizes offline RL by penalizing dynamics uncertainty.
problem Learning policies from offline data with distributional shift.
method Modify model-based RL to avoid distributional shift.
result MOPO outperforms model-free and standard model-based RL.