Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

2.9%5.7%8.6%11.5% · May 202619922001200920182026
48 results for action-dependent baselines

Action-dependent baselines reduce policy gradient variance in deep RL.

problem High variance in policy gradient methods, especially in long-horizon or high-dimensional action spaces.
method Derive a bias-free action-dependent baseline that fully exploits policy structure without additional assumptions.
result Demonstrates and quantifies the benefit of action-dependent baselines through theoretical and numerical results.

Action-dependent baselines do not reduce variance in reinforcement learning.

problem The effectiveness of action-dependent baselines in reducing variance and improving sample efficiency in reinforcement learning.
method Decomposed the variance of the policy gradient estimator and reviewed the implementation details of prior papers.
result Action-dependent baselines do not reduce variance over a state-dependent baseline in commonly tested benchmark domains.

Survey of three geometric frameworks for action-dependent field theories.

problem Understanding action-dependent field theories through geometric structures.
method Introduction and analysis of three geometric frameworks: k-contact, k-cocontact, and multicontact.
result Analysis of relationships among these geometric structures and comparison with other definitions.

Study shows observing order book can significantly improve online market making performance.

problem Online market making with private valuations and limited feedback.
method Introduces action-dependent feedback model and proposes elimination-based and explore-then-perturb algorithms.
result Achieves O(T)O(\sqrt{T}) regret bounds with high probability in various settings.

IDAC improves reinforcement learning efficiency by modeling implicit distributions.

problem Improving sample efficiency in reinforcement learning algorithms.
method IDAC uses two DGNs for a distributional critic and a semi-implicit actor to model implicit policy distributions.
result IDAC outperforms state-of-the-art algorithms on OpenAI Gym environments.

New algorithm tackles multi-agent reinforcement learning with optimal convergence rate.

problem Multi-agent reinforcement learning with large state spaces and linear function approximations.
method Refined AVLPR framework with data-dependent pessimistic estimation and action-dependent bonuses.
result First algorithm with optimal O(T1/2)O(T^{-1/2}) convergence rate and no poly(AmaxA_{\max}) dependency.

In a previous paper we constructed classical spin Chern-Simons for any compact Lie group GG: a gauge theory whose action depends on the spin structure of the 3-manifold. Here we apply geometric quantization to the classical Hamiltonian theory and investigate the formal properties of the partition function in the Lagra…

2006-05-09abs ↗pdf ↗

Linear recurrent networks explain reinforcement learning performance in partially observable settings.

problem Understanding why linear recurrent networks work in reinforcement learning with partial observability.
method Constructed and studied two linear filters for HMMs and action-controlled HMMs.
result Linear filters serve as sufficient statistics and reduce state ambiguity, explaining empirical reinforcement learning success.

New algorithm for online learning with noisy side observations.

problem Online learning with noisy side feedback and graph-structured dependencies.
method Proposes an algorithm using a weighted directed graph to model dependencies and guarantees a regret bound of O(√α* T).
result Guarantees a regret of O(√α* T) after T rounds, where α* is the effective independence number.

New method combines heuristics and search techniques to speed up cooperative planning for autonomous vehicles.

problem Efficient cooperative planning for autonomous vehicles in complex traffic scenarios.
method Combining learned heuristics with Monte Carlo Tree Search (MCTS) to guide search towards promising actions.
result Better solutions at lower computational costs achieved through accelerated planning.

This paper analyzes risk-sensitive reinforcement learning with Conditional Value-at-Risk (CVaR) for robust Markov Decision Processes.

problem Risk-sensitive reinforcement learning for robust Markov Decision Processes (RMDPs) with state-action-dependent ambiguity sets.
method The paper establishes a connection between robustness and risk sensitivity, defining a new risk measure NCVaR and proposing value iteration algorithms.
result The proposed approach using NCVaR optimization and value iteration algorithms can solve problems with state-action-dependent ambiguity sets.

The study predicts trader actions and price movements in forex markets using lead-lag networks.

problem Predicting trader behavior and price movements in foreign exchange markets.
method Infer lead-lag networks from trader-resolved data in the foreign exchange market.
result Trader actions and price movements can be predicted from past prices and trader behavior.

New algorithm reduces linear contextual bandit regret with adversarial corruption.

problem Linear contextual bandit with adversarial reward corruption.
method Optimism in the face of uncertainty principle, weighted ridge regression.
result Achieves nearly optimal regret for both corrupted and uncorrupted cases.

Study shows how repetition affects learning in bandit settings, providing algorithms with sublinear regret.

problem Effect of persistence of engagement on learning in stochastic multi-armed bandit settings.
method Novel algorithms that achieve sublinear regret under temporal constraints.
result Additive effect of priming on regret upper bound, matching popular algorithms in absence of priming.

New framework for fair online allocation in continuous time with deadlines.

problem Fair allocation under deadlines in continuous-time online learning.
method Continuous-time utility maximization, dual ascent optimization for time averages.
result Achieves ildeO(B1/2) ilde{O}(B^{-1/2}) regret bound in the absence of statistical knowledge.

Cheshire optimizes social network activity by incentivizing users to post.

problem Maximize overall activity in social networks through user incentives.
method Modelled user actions with Hawkes processes and SDEs with jumps; used stochastic optimal control.
result Optimal incentivized actions are linearly related to current activity levels.

Optimizes assortment decisions with a new OFU scheme for online choice problems.

problem Online assortment optimization under stochastic choice with revenue performance and inference quality considerations.
method Forced-exploration OFU scheme combining regularized estimators for decision making and inference.
result Explicit regret bound and error bounds for approximate optimistic actions, showing Pareto optimality.

Enhanced visual feature attribution via adaptive baseline weighting.

problem IG's sensitivity to baseline images leads to noisy or unstable explanations.
method Weighted Integrated Gradients (WG) evaluates and weights baselines for improved reliability.
result WG improves over Expected Gradients (EG) by up to 36% across various models.

The study introduces backward baselines to distinguish past prediction from future prediction in machine learning models.

problem Differentiating between past and future prediction in machine learning models.
method Theoretical, empirical, and normative arguments support a family of simple and efficient statistical tests called backward baselines.
result The study provides a meaningful backward baseline for auditing black-box prediction systems.

Algorithm safely learns from sub-optimal baseline policies while satisfying constraints.

problem Safe reinforcement learning with constraints when baseline policy is sub-optimal.
method Iterative policy optimization alternating between return maximization, baseline distance minimization, and constraint projection.
result Consistently outperforms baselines, achieving 10x fewer constraint violations and 40% higher reward.

Bayesian Deep Learning experiments often use weak baselines, leading to misleading conclusions.

problem Misleading conclusions in Bayesian Deep Learning due to weak baselines in experiments.
method Used a fixed number of iterations for baselines and compared them with models trained to convergence.
result Monte Carlo dropout baseline outperforms or performs competitively with superior methods.

We reduce variance in RL with input-dependent baselines.

problem High variance in RL with standard baselines in input-driven environments.
method Derive and use a bias-free, input-dependent baseline; propose a meta-learning approach.
result Input-dependent baselines improve training stability and policy quality.

New approach to avoid bad incentives in reinforcement learning agents.

problem Designing safe reinforcement learning agents that avoid unnecessary disruptions.
method Break down side effects penalties into baseline state and deviation measure; introduce new stepwise inaction baseline and relative reachability deviation measure.
result Combination of new design choices avoids undesirable incentives, while simpler alternatives fail.

A new model predicts discrete events with flexible, nonparametric baseline and excitation.

problem Limited flexibility in discrete Hawkes models for event prediction.
method Gaussian Process Discrete Hawkes Process (GP-DHP) with collapsed latent representation.
result Improves predictive log-likelihood for diverse event patterns.

An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well as a given baseline strategy. In this paper, we develop and analyze a new model-based approach to compute a safe policy when we have access …

2016-07-13abs ↗pdf ↗

Bayesian Scattering offers a simple baseline for image data uncertainty.

problem Lack of interpretable, mathematically grounded uncertainty quantification methods for image data.
method Coupling wavelet scattering transform with a simple probabilistic head.
result Bayesian Scattering provides sensible uncertainty estimates under distribution shifts.

Partial-input models fail to detect dataset artifacts, even when they perform poorly.

problem The effectiveness of partial-input models in detecting dataset artifacts is questionable.
method Design artificial datasets and identify trivial patterns in the SNLI dataset.
result Partial-input models can solve examples previously considered hard, indicating potential dataset artifacts.

Investigates Q value evolution in Stable Baselines for DQL in simple vs complex environments.

problem DQL in Stable Baselines struggles with simple non-game environments.
method Comparison of TrafficLight and FrozenLake environments; Q value decomposition analysis.
result Q values meander far from optimal in complex relationships between states.

Statistical mechanics models node-perturbation learning with noisy baselines.

problem Understanding learning dynamics in node-perturbation algorithms with noisy baselines.
method Developed statistical mechanics to model node-perturbation learning with noisy baselines and derived coupled differential equations.
result Derived coupled differential equations of order parameters to depict learning dynamics and calculated generalization error.

Paper identifies problematic baselines in Shapley value explanations and proposes a reweighting mechanism.

problem Identifying and addressing the suboptimality of baselines in Shapley value feature importance analysis.
method Analyzed suboptimality of baselines, identified problematic baseline, generalized uninformativeness, and designed a reweighting mechanism.
result Proposed uncertainty-based reweighting mechanism effectively accelerates computation and improves explanation quality.