Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,694 papers · 148 categories

Trend · papers per month

57115172229 · Jun 202019922001200920172026
48 results for Index Policy

Study on indexability of restless multi-armed bandits and rollout policy performance.

problem Maximizing discounted rewards in finite state restless multi-armed bandit problems.
method Decouple the problem into single-armed restless bandits, analyze using value iteration, and compare with Whittle index policy.
result Demonstrates conditions for indexability and compares performance of index policy and rollout policy.

This paper measures the intensity of implicit government guarantees using PMC index model.

problem Excessive local government debt due to implicit government guarantees.
method Text mining of policy documents related to municipal investment bonds, PMC index model.
result Recent policies have reduced the intensity of implicit government guarantees.

New method optimizes policies without assuming known link functions between preferences and rewards.

problem Policy alignment with unknown and unrestricted link functions.
method Formulates an ff-divergence-constrained reward maximization problem, learning policies directly.
result Induces a semiparametric single-index binary choice model for policy alignment.

Neural Index Policy for multi-action bandits with heterogeneous budgets.

problem Real-world settings often involve multiple interventions with heterogeneous costs and constraints, breaking classical assumptions.
method Introduces a Neural Index Policy (NIP) that learns to assign budget-aware indices to arm-action pairs using a neural network and differentiable knapsack layer.
result Empirically achieves near-optimal performance while strictly enforcing heterogeneous budgets and scaling to hundreds of arms.

Optimal policy for multi-armed multi-action bandits with unknown parameters.

problem Optimal sequential action selection for multi-armed multi-action bandits with unknown parameters.
method Occupancy-Measured-Reward Index Policy (OMRIP) and R(MA)^2B-UCB algorithm.
result Asymptotically optimal policy with sub-linear regret and low computational complexity.

The paper evaluates index-based allocation policies using data from randomized control trials.

problem Evaluating index-based allocation policies in resource-scarce scenarios.
method Using data from randomized control trials, the paper introduces an efficient estimator and methods for computing asymptotically correct confidence intervals.
result Valid statistical conclusions can be drawn for index-based allocation policies.

The study analyzes how neural reward models learn features for policy optimization in a Gaussian single-index model.

problem Reward modeling in policy optimization and its impact on downstream value.
method Two-stage neural reward model: first learns hidden direction, then fits readout layer.
result For any feature-learning temperature above a dimension-free threshold, a constant fraction of neurons recover the hidden direction.

Study improves policy search in continuous control by using heavy-tailed distributions.

problem Challenges in continuous space policy search due to non-convexity and myopic-farsighted incentives.
method Introduced heavy-tailed policy parameterizations and analyzed convergence rates and stability.
result Convergence rate to stationarity depends on policy's tail index and exploration tolerance.

Paper constructs a CRRIX index to assess cryptocurrency market risks from regulatory changes.

problem Lack of indices quantifying regulatory risks in cryptocurrencies.
method CRRIX index based on news coverage frequency, using Latent Dirichlet Allocation and Hellinger distance.
result CRRIX successfully captures major policy-changing moments and synchronizes with market volatility.

The paper analyzes the sliding regret of stochastic bandit algorithms.

problem Measuring the one-shot behavior of no-regret algorithms in stochastic bandits.
method Introducing sliding regret to measure the worst pseudo-regret over a time-window.
result Randomized methods have optimal sliding regret, while index policies have the worst possible sliding regret.

New algorithm uses Whittle index to improve Q-learning for restless bandits.

problem Optimizing decision-making in multiarmed restless bandits with average reward.
method Introduces a novel reinforcement learning algorithm combining Q-learning and Whittle index policy.
result Demonstrates significant computational gains and excellent empirical performance.

We consider a finite-horizon multi-armed bandit (MAB) problem in a Bayesian setting, for which we propose an information relaxation sampling framework. With this framework, we define an intuitive family of control policies that include Thompson sampling (TS) and the Bayesian optimal policy as endpoints. Analogous to TS…

2019-02-12abs ↗pdf ↗

Paper extends quantile factor analysis with probabilistic methods for better economic policy and financial condition prediction.

problem Improving accuracy in economic and financial condition prediction.
method Probabilistic quantile factor analysis with regularization and variational approximations.
result The probabilistic estimator outperforms a recent loss-based estimator in many cases.

Study creates a global living index to assess quality of life.

problem Long-term impacts of global economic changes on living conditions.
method Machine learning framework combining socio-economic factors.
result Developed a practical tool for policymakers to identify areas needing improvement.

Study Whittle index learning algorithms for restless bandits with constant stepsizes.

problem Optimizing decisions in restless multi-armed bandits with constant stepsizes.
method Developed Q-learning algorithms with constant stepsizes for index learning in restless bandits, extending to DQN and function approximations.
result The algorithms learn the Whittle index effectively.

This paper is about index policies for minimizing (frequentist) regret in a stochastic multi-armed bandit model, inspired by a Bayesian view on the problem. Our main contribution is to prove that the Bayes-UCB algorithm, which relies on quantiles of posterior distributions, is asymptotically optimal when the reward dis…

2016-01-06abs ↗pdf ↗

Study finds market inefficiencies vary by time scale, with news uncertainty key.

problem Evaluating scale-dependent informational efficiency of stock markets.
method Tensor-eigenvalue-based Financial Chaos Index, Granger causality, network analysis.
result Semi-strong form of EMH rejected at daily frequency, but not at monthly.

Policy iteration is a family of algorithms that are used to find an optimal policy for a given Markov Decision Problem (MDP). Simple Policy iteration (SPI) is a type of policy iteration where the strategy is to change the policy at exactly one improvable state at every step. Melekopoglou and Condon [1990] showed an exp…

2019-11-28abs ↗pdf ↗

Configuration spaces for computer systems can be challenging for traditional and automatic tuning strategies. Injecting task-specific knowledge into the tuner for a task may allow for more efficient exploration of candidate configurations. We apply this idea to the task of index set selection to accelerate database wor…

2019-09-16abs ↗pdf ↗

A RL framework for hedging equity index options with realistic costs.

problem Dynamic hedging of equity index option exposures under transaction costs.
method Reinforcement Learning (RL) with a leak-free environment, cost-aware reward function, and stochastic actor-critic agent.
result The RL policy improves risk-adjusted performance compared to no-hedge, momentum, and volatility-targeting baselines.

The study examines how global economic policy uncertainty affects crude oil futures volatility.

problem Predicting crude oil futures volatility using global economic policy uncertainty.
method Established single-factor and two-factor models under the GARCH-MIDAS framework, tested with rolling-window and fixed-span specifications.
result GEPU changes have stronger predictive power than the GEPU index for crude oil futures volatility.

We consider the restless Markov bandit problem, in which the state of each arm evolves according to a Markov process independently of the learner's actions. We suggest an algorithm that after TT steps achieves O~(T)\tilde{O}(\sqrt{T}) regret with respect to the best policy that knows the distributions of all arms. No ass…

2012-09-12abs ↗pdf ↗

We introduce in this paper a new algorithm for Multi-Armed Bandit (MAB) problems. A machine learning paradigm popular within Cognitive Network related topics (e.g., Spectrum Sensing and Allocation). We focus on the case where the rewards are exponentially distributed, which is common when dealing with Rayleigh fading c…

2012-04-07abs ↗pdf ↗

New algorithm optimizes decision-making for complex systems with varying parameters.

problem Optimizing decisions in systems with varying parameters and heterogeneous restlessness.
method Model Predictive Control (MPC) approach with randomized rounding for heterogeneous RMABs.
result Achieves an O(logN1/N)O(\log N\sqrt{1/N}) optimality gap in infinite time average reward problems.

Develops an anytime-valid framework for optimal policy identification from logged contextual bandit data.

problem Selecting the optimal policy from a candidate policy class while monitoring evidence continuously.
method Constructs a time-indexed set that retains the true optimal policy set uniformly over time.
result The procedure allows the analyst to monitor policy values, eliminate clearly suboptimal policies, and stop at data-dependent times without invalidating inference.

NeurWIN learns Whittle indices for restless bandits using deep reinforcement learning.

problem Finding optimal solutions for restless bandits with complex transition kernels.
method NeurWIN uses a neural network to learn Whittle indices, leveraging the Whittle indices' mathematical properties.
result NeurWIN outperforms other RL algorithms in solving three recently studied restless bandit problems.

The paper uncovers the mathematical structure enabling value decomposition in multi-agent systems.

problem Theoretical justification for why value decomposition works effectively in multi-agent systems remains underexplored.
method The paper introduces the concept of Markov entanglement to measure the underlying structure and demonstrates how it can be used to bound the decomposition error.
result The widely-used class of index policies is weakly entangled and enjoys a sublinear O(N)\mathcal O(\sqrt{N}) scale of decomposition error for NN-agent systems.

Bayesian optimization with cost-awareness using Gittins index.

problem Optimizing unknown functions with limited data evaluations and costs.
method Developed a connection between cost-aware Bayesian optimization and the Pandora's Box problem, using the Gittins index as an acquisition function.
result The Gittins index-based acquisition function performs well in cost-aware Bayesian optimization, especially in high dimensions.

Investors target specific regions of payoff distributions for portfolio optimization.

problem Optimizing portfolio performance across different return distribution regions.
method Developed a dynamic portfolio-choice framework targeting downside or upside quantiles.
result Policies focused on downside regions provide stronger left-tail protection and higher Sharpe ratios.

Study finds multifractal cross-correlations between agricultural markets and external uncertainties.

problem Investigating relationships between agricultural spot markets and external uncertainties.
method Multifractal detrending moving-average cross-correlation analysis (MF-X-DMA).
result Maize exhibits intrinsic joint multifractality with all uncertainty proxies.

This study shows how trade policy uncertainty affects stock-T bill correlations.

problem The impact of trade policy uncertainty on stock-T bill relationships.
method Extended Dynamic Conditional Correlation (DCC) framework incorporating exogenous variables.
result Trade policy uncertainty significantly alters stock-T bill correlations, especially under specific political conditions.

Kernelized bandit algorithm tackles adaptive contextual bandits with single-index models.

problem Adaptive contextual bandits with single-index models and unknown link functions.
method Kernelized ε-greedy algorithm combining Stein-based index estimation and kernel ridge regression for reward functions.
result Unified framework for simultaneous learning and inference in single-index contextual bandits.

This study uses deep learning to analyze stock market sentiment from financial forums.

problem Improving stock market prediction accuracy through emotional analysis.
method Crawling financial forum data, training Bert model on financial corpus, using MIC for comparison.
result BERT model's emotional analysis of financial texts correlates with stock market fluctuations.