Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,291 papers · 148 categories

Trend · papers per month

3672108144 · Jun 202019922001200920182026
48 results for expected sarsa

The paper analyzes SARSA with linear function approximation, providing finite-sample error bounds.

problem Finite-sample analysis of SARSA with linear function approximation under non-i.i.d. data.
method Developed a novel technique to characterize stochastic bias of SARSA with time-varying kernels, enabling finite-sample convergence analysis.
result Provided finite-sample analysis on mean square error of SARSA and fitted SARSA algorithms.

EPG unifies SPG and DPG for reinforcement learning.

problem Improving reinforcement learning algorithms for policy optimization.
method EPG integrates across actions for gradient estimation, using practical results for Gaussian policies and extending to broader classes of policies.
result EPG reduces gradient variance without deterministic policies and with minimal overhead.

Implicit Q-learning and SARSA adjust step-sizes automatically, improving stability and performance.

problem Numerical instability and slow progress in Q-learning and SARSA due to step-size calibration.
method Reformulate iterative updates as fixed-point equations, scaling step-sizes inversely with feature norms.
result Implicit methods maintain stability over broader step-size ranges and achieve comparable convergence rates.

FedSARSA converges with heterogeneous agents, achieving linear speed-up.

problem Convergence analysis of Federated SARSA with heterogeneous agents.
method Linear function approximation, local training, multi-step error expansion.
result FedSARSA achieves linear speed-up with respect to the number of agents.

Unified algorithm for reinforcement learning with function approximation.

problem Limited scalability of Q(σ,λ) for large-scale learning.
method Proposes GQ(σ,λ) with linear function approximation to extend tabular Q(σ,λ).
result Empirical results show GQ(σ,λ) outperforms full-sampling and pure-expectation methods.

A new method for marine robot navigation using sparse Gaussian Process TD learning.

problem Data efficiency and online prediction challenges in marine robot navigation.
method Sparse Gaussian Process Temporal Difference (SPGP-TD) learning with a sparse approximation.
result SPGP-SARSA outperforms state-of-the-art sparse methods and replicates exact prediction quality.

This study evaluates multi-step reinforcement learning methods in the Mountain Car environment.

problem Lack of statistical significance in evaluating multi-step reinforcement learning methods.
method Combines n-step action-value algorithms with DQN architecture, tests in Mountain Car environment.
result Performance varies with off-policy correction, backup length, and target network update frequency.

Study evaluates reinforcement learning for trading diverse stocks, finds Q-learning outperforms.

problem Evaluating reinforcement learning for trading diverse stocks.
method Implemented Value Iteration (VI), State-action-reward-state-action (SARSA), and Q-Learning on a diverse stock portfolio dataset.
result Q-learning performs better than VI and SARSA during testing, but performance varies based on market conditions.

A new approach to reinforcement learning by adjusting stepsize based on state criticality.

problem The bias-variance trade-off in n-step reinforcement learning algorithms.
method Adapting the stepsize (n) for each state based on a human trainer's criticality measure.
result Improved reinforcement learning performance by dynamically adjusting the stepsize.

Paper proposes a new RL approach combining IL and RL methods to improve decision-making.

problem Challenges in RL with large state and action spaces, and difficulty in reward determination.
method Combines Imitation Learning and RL methods (SARSA and A3C) to learn sequential decision-making policies.
result Significantly decreases human effort and exploration time in learning decision-making policies.

TDprop uses Jacobi preconditioning to improve adaptive optimizers in Deep RL.

problem Improving performance of adaptive optimizers in Deep RL.
method TDprop computes per-parameter learning rates based on Jacobi preconditioning of the TD update rule.
result TDprop matches or exceeds Adam's performance in Deep RL experiments, suggesting Jacobi preconditioning can improve adaptive methods.

Unified algorithm Q(σ,λ)Q(σ,λ) combines eligibility traces to improve memory and computation efficiency.

problem Memory and computation inefficiency in multi-step temporal-difference learning.
method Introduced Q(σ)Q(σ) algorithm, combined with eligibility traces to propose Q(σ,λ)Q(σ,λ).
result Proposed algorithm Q(σ,λ)Q(σ,λ) converges to optimal value function exponentially.

We propose a reinforcement learning solution to the \emph{soccer dribbling task}, a scenario in which a soccer agent has to go from the beginning to the end of a region keeping possession of the ball, as an adversary attempts to gain possession. While the adversary uses a stationary policy, the dribbler learns the best…

2013-05-28abs ↗pdf ↗

This paper examines methods to estimate uncertainty in neural networks for dialogue policy optimization.

problem Efficient exploration in neural network-based dialogue policy optimization.
method Extensive benchmark of deep Bayesian methods to extract uncertainty estimates from DQN.
result Combining uncertainty estimation methods with DQN improves sample efficiency and user experience.

This paper proposes a benchmarking environment for RL-based dialogue management.

problem Lack of a common benchmarking framework for RL models in dialogue management.
method Developed a set of simulated environments and compared RL models (DQN, A2C, Natural Actor-Critic, GP-SARSA).
result Demonstrated the effectiveness of various RL models in dialogue management.

A two-stream reinforcement learning model improves decision-making across human and neuropsychiatric studies.

problem Improving reinforcement learning models to better simulate human decision-making and neuropsychiatric conditions.
method Proposes a two-stream reinforcement learning model that processes positive and negative rewards and incorporates reward-processing biases.
result The two-stream model outperforms standard Q-learning and SARSA methods on various tasks and datasets.

Transformers can implement reinforcement learning algorithms from data without updates.

problem Training reinforcement learning algorithms from data without parameter updates.
method Design a teacher-mimicking training procedure for transformers to implement policy-improvement methods.
result Gradient flow converges to an optimal parameter manifold corresponding to the desired RL update.

Tabular Q-learning outperforms advanced RL methods in monetary policy.

problem Dynamic setting of short-term interest rates to stabilize inflation and unemployment under uncertain macroeconomic conditions.
method Discrete-action Markov Decision Process with tabular Q-learning, SARSA, Actor-Critic, Deep Q-Networks, Bayesian Q-learning, POMDP formulations.
result Standard tabular Q-learning achieved the best performance (-615.13 +- 309.58 mean return) compared to advanced RL methods and traditional policy rules.

Projective simulation outperforms standard reinforcement learning in navigation tasks.

problem Benchmarking projective simulation in navigation problems.
method Projective simulation model applied to grid world and mountain car problems, compared to Q-learning and SARSA.
result Projective simulation outperforms standard reinforcement learning in terms of parameter simplicity and computational efficiency.

Develops hyperfinite GG-expectation theory for continuous-time processes.

problem Creating a discrete model for continuous-time GG-expectation.
method Introduces hyperfinite GG-expectation and develops its theory, proving existence of liftings.
result Establishes existence theorem for liftings of continuous-time GG-expectation.

Introduces a new conditional expectation under distorted probabilities, addressing time-inconsistency.

problem Time-inconsistency in nonlinear expectations under probability distortion.
method Localizes probability distortion and constructs a time-consistent conditional expectation.
result Constructs a conditional expectation that is time-consistent and corresponds to a parabolic differential equation.

Active inference minimizes expected free energy for optimal behavior.

problem Understanding and optimizing behavior in complex systems.
method Combines Bayesian decision theory, optimal Bayesian design, and the free energy principle.
result Active inference emerges as a unified framework for information-seeking, utility maximization, and goal-directed behavior.

We provide a general construction of time-consistent sublinear expectations on the space of continuous paths. It yields the existence of the conditional G-expectation of a Borel-measurable (rather than quasi-continuous) random variable, a generalization of the random G-expectation, and an optional sampling theorem that…

2012-05-11abs ↗pdf ↗

Paper proposes using expectation models for planning in stochastic environments.

problem Intractability of learning distribution and sample models in large state and action spaces.
method Proposes using approximate expectation models for MBRL, analyzes linear and non-linear parametrizations, and presents a policy evaluation algorithm.
result Planning with an expectation model is equivalent to planning with a distribution model under certain conditions.

Dual representation and properties of expectile-based expected shortfall studied.

problem Studying the expectile-based expected shortfall as a risk measure.
method Provided dual representation in terms of Bochner integral, showed boundedness properties, and computed for selected distributions.
result Explicit dual representation and boundedness properties of expectile-based expected shortfall.

The paper proposes a new risk measure, Expected Downside Risk, to explain risk-preference.

problem Contradictory empirical findings between risk and reward.
method Introducing Expected Downside Risk (EDR) as a new risk measure.
result EDR better explains investors' utility perception and can model both positive and negative risk-reward relationships.

We refine Expected Shortfall by controlling different tail portions, offering tailored risk assessments.

problem Risk assessment in financial positions, especially in tail regions.
method Introducing adjusted Expected Shortfall measures that control different tail portions.
result Adjusted Expected Shortfall measures ensure risk does not exceed specified thresholds for various probability levels.