Action-dependent baselines reduce policy gradient variance in deep RL.
problem High variance in policy gradient methods, especially in long-horizon or high-dimensional action spaces.
method Derive a bias-free action-dependent baseline that fully exploits policy structure without additional assumptions.
result Demonstrates and quantifies the benefit of action-dependent baselines through theoretical and numerical results.
Action-dependent baselines do not reduce variance in reinforcement learning.
problem The effectiveness of action-dependent baselines in reducing variance and improving sample efficiency in reinforcement learning.
method Decomposed the variance of the policy gradient estimator and reviewed the implementation details of prior papers.
result Action-dependent baselines do not reduce variance over a state-dependent baseline in commonly tested benchmark domains.
Survey of three geometric frameworks for action-dependent field theories.
problem Understanding action-dependent field theories through geometric structures.
method Introduction and analysis of three geometric frameworks: k-contact, k-cocontact, and multicontact.
result Analysis of relationships among these geometric structures and comparison with other definitions.
Efficient algorithms for second-price auctions with action-dependent censoring.
problem Sequential bidding strategies for repeated auctions with incomplete information.
method Proposed novel UCB-like algorithms for second-price auctions in a stochastic setting.
result Significant improvement in worst-case regret, especially for low item values.
New method reduces variance in policy gradient estimation for reinforcement learning.
problem Large variance issue in policy gradient estimation.
method Action-dependent control variates using Stein's identity.
result Significantly improves sample efficiency of policy gradient methods.
We show that J. Lott's equivariant higher analytic torsion for compact group actions depends only on the equivariant Euler characteristic.
Study shows observing order book can significantly improve online market making performance.
problem Online market making with private valuations and limited feedback.
method Introduces action-dependent feedback model and proposes elimination-based and explore-then-perturb algorithms.
result Achieves O ( T ) O(\sqrt{T}) O ( T ) regret bounds with high probability in various settings. Proposes a neural network for efficient deep hedging strategies.
problem Hard training of optimal hedging strategies due to action dependence.
method Introduces no-transaction band network, a neural architecture.
result Demonstrates faster and more precise hedging strategies.
IDAC improves reinforcement learning efficiency by modeling implicit distributions.
problem Improving sample efficiency in reinforcement learning algorithms.
method IDAC uses two DGNs for a distributional critic and a semi-implicit actor to model implicit policy distributions.
result IDAC outperforms state-of-the-art algorithms on OpenAI Gym environments.
Efficiently handles contextual bandits with diffusion models.
problem Challenges in online decision-making with contextual bandits.
method Leverage pre-trained diffusion models as priors to capture action dependencies.
result Developed an algorithm for efficient posterior approximation.
New algorithm tackles multi-agent reinforcement learning with optimal convergence rate.
problem Multi-agent reinforcement learning with large state spaces and linear function approximations.
method Refined AVLPR framework with data-dependent pessimistic estimation and action-dependent bonuses.
result First algorithm with optimal O ( T − 1 / 2 ) O(T^{-1/2}) O ( T − 1/2 ) convergence rate and no poly( A max A_{\max} A m a x ) dependency. Trajectory-wise CVs reduce variance in policy gradient methods.
problem High variance in estimating policy gradient estimates.
method Proposes trajectory-wise control variates to reduce variance without bias.
result Trajectory-wise CVs are optimal for variance reduction under reasonable assumptions.
In a previous paper we constructed classical spin Chern-Simons for any compact Lie group G G G : a gauge theory whose action depends on the spin structure of the 3-manifold. Here we apply geometric quantization to the classical Hamiltonian theory and investigate the formal properties of the partition function in the Lagra…
BRPO optimizes batch RL policies to better exploit state-action differences.
problem Batch RL's conservatism limits exploitation of state-action differences.
method Proposes residual policies and derives BRPO to maximize policy performance.
result BRPO achieves state-of-the-art performance in various tasks.
Let G be a group and let M be a CAT(0) proper metric space (e.g. a simply connected complete Riemannian manifold of non-positive sectional curvature or a locally finite tree). Isometric actions of G on M are (by definition) points in the space R := Hom(G, Isom(M)) with the compact open topology. Sample theorems: 1. The…
Linear recurrent networks explain reinforcement learning performance in partially observable settings.
problem Understanding why linear recurrent networks work in reinforcement learning with partial observability.
method Constructed and studied two linear filters for HMMs and action-controlled HMMs.
result Linear filters serve as sufficient statistics and reduce state ambiguity, explaining empirical reinforcement learning success.
New algorithm for online learning with noisy side observations.
problem Online learning with noisy side feedback and graph-structured dependencies.
method Proposes an algorithm using a weighted directed graph to model dependencies and guarantees a regret bound of O(√α* T).
result Guarantees a regret of O(√α* T) after T rounds, where α* is the effective independence number.
New method combines heuristics and search techniques to speed up cooperative planning for autonomous vehicles.
problem Efficient cooperative planning for autonomous vehicles in complex traffic scenarios.
method Combining learned heuristics with Monte Carlo Tree Search (MCTS) to guide search towards promising actions.
result Better solutions at lower computational costs achieved through accelerated planning.
This paper analyzes risk-sensitive reinforcement learning with Conditional Value-at-Risk (CVaR) for robust Markov Decision Processes.
problem Risk-sensitive reinforcement learning for robust Markov Decision Processes (RMDPs) with state-action-dependent ambiguity sets.
method The paper establishes a connection between robustness and risk sensitivity, defining a new risk measure NCVaR and proposing value iteration algorithms.
result The proposed approach using NCVaR optimization and value iteration algorithms can solve problems with state-action-dependent ambiguity sets.
The study predicts trader actions and price movements in forex markets using lead-lag networks.
problem Predicting trader behavior and price movements in foreign exchange markets.
method Infer lead-lag networks from trader-resolved data in the foreign exchange market.
result Trader actions and price movements can be predicted from past prices and trader behavior.
New algorithm reduces linear contextual bandit regret with adversarial corruption.
problem Linear contextual bandit with adversarial reward corruption.
method Optimism in the face of uncertainty principle, weighted ridge regression.
result Achieves nearly optimal regret for both corrupted and uncorrupted cases.
Study shows how repetition affects learning in bandit settings, providing algorithms with sublinear regret.
problem Effect of persistence of engagement on learning in stochastic multi-armed bandit settings.
method Novel algorithms that achieve sublinear regret under temporal constraints.
result Additive effect of priming on regret upper bound, matching popular algorithms in absence of priming.
New framework for fair online allocation in continuous time with deadlines.
problem Fair allocation under deadlines in continuous-time online learning.
method Continuous-time utility maximization, dual ascent optimization for time averages.
result Achieves i l d e O ( B − 1 / 2 ) ilde{O}(B^{-1/2}) i l d e O ( B − 1/2 ) regret bound in the absence of statistical knowledge. Cheshire optimizes social network activity by incentivizing users to post.
problem Maximize overall activity in social networks through user incentives.
method Modelled user actions with Hawkes processes and SDEs with jumps; used stochastic optimal control.
result Optimal incentivized actions are linearly related to current activity levels.
Khovanov homology for pro-tangles and spectral sequences
problem Developing a framework for Khovanov homology for pro-tangles and spectral sequences
method Using pro-tangles, simplicial presheaves, and spectral sequences
result Establishing a fully faithful embedding and an algebraic spectral sequence for pro-tangles
New methods correct spectral distortions using known analyte concentrations.
problem Distorted spectral shapes from absorbing and scattering contributions.
method Modified penalized baseline correction methods that incorporate known analyte concentrations.
result Improved prediction performance on near infra-red data sets.
Optimizes assortment decisions with a new OFU scheme for online choice problems.
problem Online assortment optimization under stochastic choice with revenue performance and inference quality considerations.
method Forced-exploration OFU scheme combining regularized estimators for decision making and inference.
result Explicit regret bound and error bounds for approximate optimistic actions, showing Pareto optimality.
Paper finds Dutch Draw optimal baseline for binary classification.
problem Need a proper baseline for binary classification validation.
method Examined all input-independent baseline methods.
result Dutch Draw is optimal baseline under given conditions.
Paper improves safe policy improvement with estimated baseline policy.
problem Unreliable batch Reinforcement Learning algorithms in real-world applications.
method Apply SPIBB algorithms with an estimated baseline policy.
result Safe policy improvement guarantees over true baseline without direct access.
Enhanced visual feature attribution via adaptive baseline weighting.
problem IG's sensitivity to baseline images leads to noisy or unstable explanations.
method Weighted Integrated Gradients (WG) evaluates and weights baselines for improved reliability.
result WG improves over Expected Gradients (EG) by up to 36% across various models.
Paper presents a universal baseline for binary prediction models.
problem Need a robust baseline to evaluate model performance.
method Dutch Draw (DD) baseline method for binary classification models.
result Reduces to almost always predicting zero or one in most situations.
The study introduces backward baselines to distinguish past prediction from future prediction in machine learning models.
problem Differentiating between past and future prediction in machine learning models.
method Theoretical, empirical, and normative arguments support a family of simple and efficient statistical tests called backward baselines.
result The study provides a meaningful backward baseline for auditing black-box prediction systems.
Algorithm safely learns from sub-optimal baseline policies while satisfying constraints.
problem Safe reinforcement learning with constraints when baseline policy is sub-optimal.
method Iterative policy optimization alternating between return maximization, baseline distance minimization, and constraint projection.
result Consistently outperforms baselines, achieving 10x fewer constraint violations and 40% higher reward.
Bayesian Deep Learning experiments often use weak baselines, leading to misleading conclusions.
problem Misleading conclusions in Bayesian Deep Learning due to weak baselines in experiments.
method Used a fixed number of iterations for baselines and compared them with models trained to convergence.
result Monte Carlo dropout baseline outperforms or performs competitively with superior methods.
Paper uses tensor completion to estimate HVAC fan power baselines.
problem Estimating HVAC fan power without demand response.
method Tensor completion for multi-dimensional data analysis.
result Tensor completion outperforms existing baselining methods.
A new baseline for Shapley values in MLPs considers model use.
problem Lack of a robust baseline for Shapley values in neural networks.
method Proposes a neutrality-based baseline for Shapley values.
result Empirically validated the proposed baseline for binary classification tasks.
Bayesian algorithm detects changes in fluctuating baselines.
problem Detecting change points in time series with a shifting baseline.
method Extended Bayesian online change point detection (BOCPD) algorithm.
result The extended algorithm can detect changes in fluctuating baselines.
We reduce variance in RL with input-dependent baselines.
problem High variance in RL with standard baselines in input-driven environments.
method Derive and use a bias-free, input-dependent baseline; propose a meta-learning approach.
result Input-dependent baselines improve training stability and policy quality.
New approach to avoid bad incentives in reinforcement learning agents.
problem Designing safe reinforcement learning agents that avoid unnecessary disruptions.
method Break down side effects penalties into baseline state and deviation measure; introduce new stepwise inaction baseline and relative reachability deviation measure.
result Combination of new design choices avoids undesirable incentives, while simpler alternatives fail.
Simple baselines improve multimodal utterance learning.
problem Learning rich multimodal utterance representations.
method Conditional factorization of utterances into unimodal factors; extending to bimodal and trimodal factors.
result Optimal embeddings can be derived in closed form.
A new model predicts discrete events with flexible, nonparametric baseline and excitation.
problem Limited flexibility in discrete Hawkes models for event prediction.
method Gaussian Process Discrete Hawkes Process (GP-DHP) with collapsed latent representation.
result Improves predictive log-likelihood for diverse event patterns.
An important problem in sequential decision-making under uncertainty is to use limited data to compute a safe policy, i.e., a policy that is guaranteed to perform at least as well as a given baseline strategy. In this paper, we develop and analyze a new model-based approach to compute a safe policy when we have access …
Bayesian Scattering offers a simple baseline for image data uncertainty.
problem Lack of interpretable, mathematically grounded uncertainty quantification methods for image data.
method Coupling wavelet scattering transform with a simple probabilistic head.
result Bayesian Scattering provides sensible uncertainty estimates under distribution shifts.
Partial-input models fail to detect dataset artifacts, even when they perform poorly.
problem The effectiveness of partial-input models in detecting dataset artifacts is questionable.
method Design artificial datasets and identify trivial patterns in the SNLI dataset.
result Partial-input models can solve examples previously considered hard, indicating potential dataset artifacts.
Investigates Q value evolution in Stable Baselines for DQL in simple vs complex environments.
problem DQL in Stable Baselines struggles with simple non-game environments.
method Comparison of TrafficLight and FrozenLake environments; Q value decomposition analysis.
result Q values meander far from optimal in complex relationships between states.
Strong baseline for medical imaging domain adaptation.
problem Improving generalization in medical imaging across different datasets.
method Training on diverse chest X-ray datasets.
result Empirical demonstration of model generalization to out-of-sample domains.
Statistical mechanics models node-perturbation learning with noisy baselines.
problem Understanding learning dynamics in node-perturbation algorithms with noisy baselines.
method Developed statistical mechanics to model node-perturbation learning with noisy baselines and derived coupled differential equations.
result Derived coupled differential equations of order parameters to depict learning dynamics and calculated generalization error.
Paper identifies problematic baselines in Shapley value explanations and proposes a reweighting mechanism.
problem Identifying and addressing the suboptimality of baselines in Shapley value feature importance analysis.
method Analyzed suboptimality of baselines, identified problematic baseline, generalized uninformativeness, and designed a reweighting mechanism.
result Proposed uncertainty-based reweighting mechanism effectively accelerates computation and improves explanation quality.