Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,181 papers · 148 categories

Trend · papers per month

113226338451 · Jun 202019922001200920182026
48 results for direct policy search

Direct policy gradients optimize policies in discrete action spaces using sampling.

problem Optimizing policies in discrete action spaces with direct methods.
method Combining direct optimization and A^\star sampling for policy gradient approximation.
result DirPG algorithms can incorporate domain knowledge and have higher probability of sampling informative gradients.

Direct contextual policy search methods learn to improve policy parameters and simultaneously generalize these parameters to different context or task variables. However, learning from high-dimensional context variables, such as camera images, is still a prominent problem in many real-world tasks. A naive application o…

2016-11-10abs ↗pdf ↗

Bayesian optimization uses priors to speed up robot learning.

problem Identifying the best prior when multiple exist for a new task.
method Introduces MLEI, a new acquisition function combining prior likelihood and expected improvement.
result MLEI effectively identifies and exploits priors in new situations.

Bayesian optimization with directionally constrained search improves efficiency within a budget.

problem Optimizing expensive functions with limited computational resources.
method Directionally constrained search to allocate model capability efficiently.
result Our approach outperforms in finding the optimum within a prescribed evaluation budget.

New algorithm converges to optimal filter for predicting linear dynamical systems.

problem Direct policy search for optimal dynamic filters in partially observable systems.
method Regularizer enforcing informativity over filter states.
result Gradient descent converges to globally optimal solution at rate O(1/T).

The paper proposes an iterative approach to batch reinforcement learning for safer and more informative data collection.

problem Learning policies that are too rigid and do not adapt to new data.
method Safe diversified model-based policy search in an iterative batch reinforcement learning framework.
result Improved learned policies through continuous data collection and adaptation.

Improves neural network search in combinatorial spaces of mathematical symbols.

problem Early commitment and initialization bias limit exploration in neural network search.
method Entropy regularization and distribution initialization methods.
result Improves performance, increases sample efficiency, lowers solution complexity.

A new imitation learning method uses random search for simple policies, outperforming complex models.

problem Complexity and reward dependency issues in imitation learning.
method Derivative-free optimization with simple linear policies and random search.
result The proposed method achieves competitive performance on MuJoCo locomotion tasks without a direct reward signal.

PGS uses neural networks to improve policies online without search trees.

problem Limited scalability of Monte Carlo Tree Search (MCTS) for high branching factor games.
method Adapts a neural network simulation policy via policy gradient updates, avoiding search trees.
result PGS achieves comparable performance to MCTS and defeats strong Hex agents.

Paper presents a new policy gradient theorem using weak derivatives for reinforcement learning.

problem Continuous state-action reinforcement learning problems.
method Introduced an alternative policy gradient theorem using weak derivatives.
result The new approach yields algorithms that converge almost surely to stationary points of the value function.

Trust-region methods and natural gradients are equivalent in certain policy search scenarios.

problem Improving policy search methods in continuous control tasks.
method Introducing compatible policy search (COPOS) that uses natural parameterization and compatible value function approximation to control entropy loss.
result COPOS yields state-of-the-art results in challenging tasks and reduces entropy loss.

A new method optimizes neural sequence models for better task performance.

problem Training neural sequence models with maximum likelihood estimation ignores task losses.
method Maximum likelihood guided parameter search (MGS) in the parameter space.
result MGS optimizes sequence-level losses, reducing repetition and non-termination.

This paper introduces a new method to evaluate multiple policies simultaneously.

problem Estimating the value of many policies for a single set of states.
method Developed a scalable, differentiable fingerprinting mechanism to represent complex policies.
result The method can produce policies that outperform those that generated the training data, in zero-shot manner.

Enhances RL performance with a population-guided parallel learning scheme.

problem Improving off-policy reinforcement learning performance.
method Population-guided parallel learning scheme with shared experience replay buffer and soft policy update.
result Monotone improvement of the expected cumulative return proved theoretically and demonstrated in practice.

CF-GPS learns policies from logged data by considering counterfactual outcomes.

problem Learning policies from limited real experience in complex environments.
method Assumes logged real experience and models counterfactual outcomes. Uses structural causal models for evaluation.
result Improves policy evaluation and search results on a grid-world task.

New approach learns walk and trot gaits from simulated quadruped using strategic exploration.

problem Learning symmetric gaits (walk and trot) from high-dimensional action spaces.
method Introduced symmetry properties into initial covariance of Gaussian search distribution for strategic exploration. Used episode-based likelihood ratio policy gradient and relative entropy policy search.
result Significant performance enhancement in learning walk and trot gaits compared to random gaits.

The search space of Bayesian Network structures is usually defined as Acyclic Directed Graphs (DAGs) and the search is done by local transformations of DAGs. But the space of Bayesian Networks is ordered by DAG Markov model inclusion and it is natural to consider that a good search policy should take this into account.…

2013-01-10abs ↗pdf ↗

EffAcTS uses active learning to select model parameters for robust policy search.

problem Learning robust policies that perform well in unseen environments.
method EffAcTS is an active learning framework that selects model parameters using Linear Bandits to minimize data collection.
result Efficiency gains in sample collection and improved policy performance.

Methods for learning to search for structured prediction typically imitate a reference policy, with existing theoretical guarantees demonstrating low regret compared to that reference. This is unsatisfactory in many applications where the reference policy is suboptimal and the goal of learning is to improve upon it. Ca…

2015-02-08abs ↗pdf ↗

Simple policy search outperforms advanced learnable test-time augmentation techniques.

problem Improving predictive performance through test-time data augmentation.
method Greedy policy search (GPS) for learning test-time augmentation policies.
result Augmentation policies learned with GPS achieve superior predictive performance and robustness.

Paper proposes MCTSPO for better reinforcement learning policy optimization.

problem Local optima and saddle points in gradient-based methods and poor initialization in gradient-free methods.
method Monte-Carlo tree search combined with gradient-free optimization.
result Improved performance on reinforcement learning tasks with deceptive or sparse reward functions.

Introduces GAMPS for better model-based policy learning.

problem Misspecified model classes lead to poor policy estimates.
method Exploits current policy to learn approximate transition model, focusing on relevant parts of the environment.
result Empirically validated GAMPS on benchmark domains, demonstrating improved properties.

Efficient neural architecture search by sampling structure and operations.

problem Efficiently searching for optimal neural architectures.
method Decouples structure and operation search, using reinforcement learning with policy vectors.
result Significantly improved efficiency compared to traditional methods.

New methods use vector search and nearest-neighbor matching for policy learning in causal inference.

problem Learning optimal policies in causal inference with limited data.
method RAG-based policy learning with vector search and nearest-neighbor matching.
result The methods bound the within-candidate choice regret and evaluate the one-step method directly as a policy.

This paper improves self-play learning in games by manipulating experience distributions.

problem Improving self-play learning in games through better experience sampling.
method Three approaches: weighted sampling, Prioritized Experience Replay, and diversifying trajectories.
result Major improvements in early training performance in some games, minor improvements overall.

New algorithm scales model-based policy search for robotics to high-dimensional systems.

problem Inefficient scaling of model-based policy search algorithms in high-dimensional state/action spaces.
method Introduces parameterized black-box priors to scale up model learning and improve robustness to prior inaccuracies.
result Significantly more data-efficient than previous algorithms, learning gaits in 16-30 seconds.

Paper proposes a policy-search algorithm to learn entropy-maximizing exploration policies in reward-free environments.

problem Reward-free learning in high-dimensional, continuous-control domains.
method Maximum Entropy POLicy optimization (MEPOL) algorithm that maximizes a non-parametric state entropy estimate.
result MEPOL learns a maximum-entropy exploration policy that facilitates learning various reward-based tasks.

Contextual policy search allows adapting robotic movement primitives to different situations. For instance, a locomotion primitive might be adapted to different terrain inclinations or desired walking speeds. Such an adaptation is often achievable by modifying a small number of hyperparameters. However, learning, when …

2015-11-13abs ↗pdf ↗

Genie optimizes search marketplaces by estimating policy impacts without risky experiments.

problem Optimizing search marketplaces with frequent policy changes and limited randomized experiments.
method Genie uses an open box simulation engine and click calibration model to estimate KPI impacts.
result Genie outperforms existing approaches in optimizing Bing Ads Marketplace.

Paper develops efficient nonmyopic active search methods for drug and materials discovery.

problem Efficiently identifying many members of a class in high-throughput screening.
method Bayesian decision framework, nonmyopic approximations, efficient computation.
result Proposed policy outperforms sequential and batch settings in drug and materials discovery.