Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,742 papers · 148 categories

Trend · papers per month

6.3%12.5%18.8%25.0% · Jul 199319922001200920172026
48 results for periodic policies

The paper tackles non-stationary MAB with periodic rewards.

problem Non-stationary mean rewards over time in a business context.
method Combines Fourier analysis with confidence-bound learning to estimate periods and minimize regret.
result Proposes a near-optimal policy with a regret bound of O(Tk=1KTk)O(\sqrt{T\sum_{k=1}^K T_k}).

New framework finds periodic policies in reset-free MDPs with sublinear regret.

problem Reset-free reinforcement learning with unknown dynamics and terminal law constraints.
method Periodic framework, periodic policies, periodic regret.
result First non-asymptotic guarantees for reset-free learning in multi-agent settings.

Step-DAD improves BED by periodically updating a design policy during experiments.

problem Improving flexibility and robustness in Bayesian experimental design.
method Semi-amortized, policy-based approach that updates a design policy during data collection.
result Consistently superior decision-making and robustness compared to current BED methods.

News on inflation and monetary policy impacts US household inflation expectations.

problem Understanding how news affects inflation expectations.
method Monthly disaggregated US data from 1978 to 2016, controlling for various factors.
result News on rising inflation and easier monetary policy has a stronger impact on inflation expectations.

Policy gradient methods with aggregated states can achieve better performance than approximate policy iteration.

problem Approximation errors in policy and value function approximations.
method State-aggregated representations and policy gradient methods.
result Policy gradient methods can achieve a per-period regret bounded by ε, while approximate policy iteration and value iteration have a higher regret.

A new pricing strategy minimizes regret by controlling strategic buyer behavior.

problem Designing a pricing policy for strategic buyers with limited seller information.
method Phased-structure policy with randomized isolation periods.
result Regret of TT-period O~(T)\widetilde{\mathcal{O}}(\sqrt{T}) against a benchmark policy.

A contextual bandit method evaluates and improves inventory control policies.

problem Evaluating and improving periodic review inventory control policies with nonstationary demand.
method Contextual bandit-based algorithm to evaluate and tweak policies.
result The method achieves favorable guarantees in both theory and practice.

Optimizes trading strategy considering alpha decay and transaction costs.

problem Maximizing reward in a multi-period portfolio with transaction costs and alpha decay.
method Formulated as an infinite horizon Markov Decision Process, solved using a modified value iteration algorithm with convergence proof and asymptotic analysis.
result Characterized optimal trading policy that maximizes average expected reward.

PROPO tackles non-stationary MDPs with efficient policy optimization.

problem Non-stationary MDPs with varying reward and transition kernels.
method PROPO, a periodic restarted optimistic policy optimization algorithm with sliding-window-based policy evaluation and improvement.
result PROPO achieves near-optimal performance in non-stationary MDPs.

The OGY method is one of control methods for a chaotic system. In the method, we have to calculate a stabilizing periodic orbit embedded in its chaotic attractor. Thus, we cannot use this method in the case where a precise mathematical model of the chaotic system cannot be identified. In this case, the delayed feedback…

2019-07-16abs ↗pdf ↗

Improved algorithms solve multi-period multi-class packing problems with bandit feedback.

problem Optimizing item packing under budget constraints with class-dependent rewards and bandit feedback.
method Developed a new estimator and a closed-form bandit policy for linear contextual multi-class multi-period packing problems.
result The proposed policy achieves sublinear regret in non-degenerate contexts, significantly outperforming benchmarks.

This paper presents empirical evidence using recently developed techniques in econophysics suggesting that the degree of long-range dependence in interest rates depends on the conduct of monetary policy. We study the term structure of interest rates for the US and find evidence that global Hurst exponents change dramat…

2006-07-26abs ↗pdf ↗

This paper pretends to analyze the importance which the natural advantages and local resources are in the manufacturing industry location, in relation with the "spillovers" effects and industrial policies. To this, we estimate the Rybczynski equation matrix for the various manufacturing industries in Portugal, at regio…

2011-10-25abs ↗pdf ↗

A new insurance and reinsurance pricing scheme based on realized loss.

problem Determining fair and risk-adjusted insurance premiums.
method Performance-based variable premium scheme with random initial premium adjusted based on realized loss.
result The variable premium scheme reduces reinsurer's total risk exposure compared to expected-value premium.

The paper tackles auction market design flaws by randomizing closing times and optimizing transaction fees.

problem Strategic traders exploit accumulated information to delay their orders, distorting auction efficiency.
method Randomizing auction closing times and designing optimal transaction fees policies.
result Policies encourage strategic traders to send orders earlier, improving auction market efficiency.

Payments data and machine learning improve nowcasting accuracy for macroeconomic indicators.

problem Lagged indicators in linear models are insufficient during crisis periods.
method Non-traditional payments data, nonlinear machine learning, and tailored cross-validation.
result Improved macroeconomic nowcasting accuracy up to 40% during crises.

Dynamic pricing policy converges to Nash equilibrium with low regret.

problem Sequential price competition among sellers over multiple periods.
method Semi-parametric least-squares estimation of s-concave demand functions.
result Prices converge to Nash equilibrium with rate O(T1/7)O(T^{-1/7}) and sellers incur regret O(T5/7)O(T^{5/7}).

Study shows how macroprudential policies affect credit growth in Israel, especially in housing and business sectors.

problem Impact of macroprudential policies on credit growth in Israel.
method Bank-level panel data analysis for Israel, 2004-2019; interaction of monetary and macroprudential policies.
result Accommodative monetary policy interacts with macroprudential policies to increase total credit growth.

Study evaluates reinforcement learning for trading diverse stocks, finds Q-learning outperforms.

problem Evaluating reinforcement learning for trading diverse stocks.
method Implemented Value Iteration (VI), State-action-reward-state-action (SARSA), and Q-Learning on a diverse stock portfolio dataset.
result Q-learning performs better than VI and SARSA during testing, but performance varies based on market conditions.

The paper examines how macroeconomic control tools lost effectiveness, leading to a 'dark ages' period.

problem Loss of effectiveness of control tools in macroeconomic stabilization policy.
method Historical analysis of macroeconomic stabilization policy from 1948 to 1993.
result The overstatement of the Lucas critique and Kydland and Prescott's time-inconsistency led to a period of ineffective stabilization policy.

Enhances financial time series forecasting with a multi-period learning framework.

problem Accurate financial time series forecasting requires considering both short-term and long-term trends.
method Proposes a Multi-period Learning Framework (MLF) with three modules: Inter-period Redundancy Filtering, Learnable Weighted-average Integration, and Multi-period self-Adaptive Patching.
result Improves financial time series forecasting accuracy and efficiency.

Paper addresses OPE for dependent bandit samples using MDS and batch updates.

problem Evaluating policies from non-i.i.d. historical data in contextual bandits.
method Constructs an MDS-based estimator for dependent samples, solves batch update and deficient support issues.
result Derives an asymptotically normal estimator for evaluation policy value.

New policy optimizes product assortment in the presence of unpredictable customers.

problem Optimizing product assortment in the presence of outlier customers.
method Developed a robust online assortment optimization policy using an active elimination strategy.
result Established upper and lower bounds on regret, showing optimality up to logarithmic factor in TT.

New bounds assess policy evaluation under unobserved confounders, showing model-based methods are more effective.

problem Policy evaluation under unobserved confounders in uncertain causal environments.
method Developed worst-case bounds for sensitivity to unobserved confounders, demonstrating model-based methods are more effective.
result Model-based approaches with robust MDPs provide sharper lower bounds for policy evaluation.

Optimizes multi-period portfolios with tail-risk constraints using neural networks.

problem Maximizing expected return while managing tail-risk constraints over multiple periods.
method Recurrent neural network approach to approximate optimal policy.
result Validated in financial and insurance models, capturing long-term risk dynamics.

Examines how central bank policies affect stock markets and asset prices.

problem Understanding the impact of monetary policy on stock markets and asset prices.
method Used Taylor rule equations to analyze data from 1990 to 2020 for US and UK, testing with various econometric methods.
result Monetary policy can explain asset price volatility and output gap better than just inflation rate.

Recent work on imitation learning has generated policies that reproduce expert behavior from multi-modal data. However, past approaches have focused only on recreating a small number of distinct, expert maneuvers, or have relied on supervised learning techniques that produce unstable policies. This work extends InfoGAI…

2017-10-13abs ↗pdf ↗

PER-ETD improves ETD by reducing variance to polynomial complexity.

problem Large variance in ETD leading to exponential sample complexity.
method Periodically restart and update the follow-on trace for a finite period.
result PER-ETD converges to the same fixed point as ETD but with improved sample complexity.

The paper solves multi-period portfolio selection with constraints using a dynamic factor model.

problem Multi-period mean-variance portfolio selection with constraints.
method Dynamic factor model, dynamic programming, piecewise linear feedback policy.
result Optimal portfolio policies determined by two stochastic processes.

We approximate the distribution of total expenditure of a retail company over warranty claims incurred in a fixed period [0, T], say the following quarter. We consider two kinds of warranty policies, namely, the non-renewing free replacement warranty policy and the non-renewing pro-rata warranty policy. Our approximati…

2010-08-05abs ↗pdf ↗

Study Federated RL with diverse constraints, proposing new optimization methods.

problem Solving reinforcement learning with multiple constraints in federated learning.
method Federated primal-dual policy optimization methods based on policy gradient methods.
result FedNPG achieves global convergence with an ildeO(1/T) ilde{O}(1/\sqrt{T}) rate.

Study shows oil prices but not COVID-19 cases affect US economic policy uncertainty.

problem Effect of COVID-19 and crude oil prices on US economic policy uncertainty.
method Used ARDL model with daily data from January 21-March 13, 2020.
result Crude oil price dynamics increase US economic policy uncertainty, while COVID-19 cases have mixed effects.

Investors target specific regions of payoff distributions for portfolio optimization.

problem Optimizing portfolio performance across different return distribution regions.
method Developed a dynamic portfolio-choice framework targeting downside or upside quantiles.
result Policies focused on downside regions provide stronger left-tail protection and higher Sharpe ratios.

An ensemble method enhances cryptocurrency trading strategies using deep reinforcement learning.

problem Improving generalization performance in stochastic cryptocurrency trading environments.
method Model selection and mixture distribution policy to ensemble deep reinforcement learning models.
result Improved out-of-sample performance compared to benchmarks.