Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

169,051 papers · 148 categories

Trend · papers per month

2.3%4.6%6.9%9.2% · May 202619922001200920182026
48 results for policy evolution

A novel meta-learning method using ES for efficient reinforcement learning.

problem Sample inefficiency in reinforcement learning.
method Evolution strategies (ES) for exploration in parameter space, deterministic policy gradients for adaptation.
result Demonstrates improved performance in high-dimensional control tasks compared to gradient-based methods.

The paper optimizes pension policies with guarantees and sustainability constraints.

problem Designing optimal pension policies with guarantees and sustainability constraints.
method Dynamic utility model, stochastic domain, overlapping generations, time-consistent decision criterion.
result Optimal investment/pension policy computed for a general framework.

We study the optimal trading policies for a wind energy producer who aims to sell the future production in the open forward, spot, intraday and adjustment markets, and who has access to imperfect dynamically updated forecasts of the future production. We construct a stochastic model for the forecast evolution and deter…

2016-09-07abs ↗pdf ↗

New approach uses Wasserstein distances to score and optimize policy behaviors.

problem Comparing reinforcement learning policies and guiding policy optimization.
method Dual formulation of Wasserstein distances in latent behavioral space, learning score functions, smoothed WDs, stochastic gradient descent, on-policy algorithms.
result Demonstrated improved performance over existing methods in various environments.

A new method simulates large, diverse populations of learning agents evolving in games.

problem Limited scalability and efficiency of Multi-Agent Reinforcement Learning.
method Parallelizable implementation of Policy Gradient and Opponent-Learning Awareness for evolutionary simulations.
result Simulated large, diverse populations of learning agents evolve under various strategies.

Paper proposes MCTSPO for better reinforcement learning policy optimization.

problem Local optima and saddle points in gradient-based methods and poor initialization in gradient-free methods.
method Monte-Carlo tree search combined with gradient-free optimization.
result Improved performance on reinforcement learning tasks with deceptive or sparse reward functions.

Current economic theories miss most of economic dynamics.

problem Accuracy of economic theories and policies depend on economic variables and processes.
method Identify and analyze overlooked economic variables and processes.
result Many economic variables and processes not accounted for in current theories.

This work uses a recommendation system to enhance exporting countries' fitness.

problem Predicting and optimizing the evolution of international trade networks.
method A recommendation system was used to identify overlooked products for countries.
result Countries can improve their national competitiveness by diversifying exported products.

Optimizes antenna tilt for better QoS in cellular networks.

problem Hard to learn optimal antenna tilt policies in real networks due to risk and simulation gap.
method Uses off-policy Contextual Multi-Armed-Bandit (CMAB) techniques to learn from existing data.
result Trained policies show consistent improvements over existing logging policies.

Deep Reinforcement Learning (DRL) algorithms have been successfully applied to a range of challenging control tasks. However, these methods typically suffer from three core difficulties: temporal credit assignment with sparse rewards, lack of effective exploration, and brittle convergence properties that are extremely …

2018-05-21abs ↗pdf ↗

This work develops agents to learn generalizable policies for dynamic network environments.

problem Real-world network topologies change due to attackers, defenders, or system failures, leading to failures in adaptive ACD systems.
method Developing agents to learn generalizable policies across dynamic network environments.
result Agents can learn robust policies for dynamic network topologies and diverse attackers.

New algorithms solve robust MDPs efficiently, significantly faster than existing methods.

problem Computing robust MDP solutions with uncertainty in transition probabilities is computationally expensive.
method Partial policy iteration and fast robust Bellman operator computation methods.
result The proposed methods are many orders of magnitude faster than state-of-the-art approaches.

MetaGenRL learns a general objective function from diverse agents.

problem Generalizing to new environments in reinforcement learning.
method MetaGenRL distills experiences from many agents into a low-complexity neural objective function.
result MetaGenRL can generalize to new environments and outperforms human-engineered algorithms.

The success of popular algorithms for deep reinforcement learning, such as policy-gradients and Q-learning, relies heavily on the availability of an informative reward signal at each timestep of the sequential decision-making process. When rewards are only sparsely available during an episode, or a rewarding feedback i…

2018-05-25abs ↗pdf ↗

Quantum-enhanced metrology aims to estimate an unknown parameter such that the precision scales better than the shot-noise bound. Single-shot adaptive quantum-enhanced metrology (AQEM) is a promising approach that uses feedback to tweak the quantum process according to previous measurement outcomes. Techniques and form…

2016-08-22abs ↗pdf ↗

NGE uses neural graphs to efficiently design robots.

problem Designing robots is hard due to combinatorial search space and evaluation costs.
method Formulated as graph search, NGE uses neural networks for policy parameterization and graph mutation with uncertainty.
result NGE significantly outperforms previous methods, discovering kinematically preferred structures.

Study shows activist board representation improves Japanese companies' performance.

problem Lack of innovation and improvement in Japanese companies.
method Examined two Japanese companies with activist board representation, analyzing performance metrics.
result Companies with activist board representation experienced significant improvements in stock returns and operational metrics.

Study optimal retirement time and consumption with habitual persistence.

problem Understanding retirement consumption patterns with habitual persistence.
method Established concise habitual evolution, used martingale and duality methods.
result Optimal consumption declines sharply at retirement but excess consumption increases.

The paper examines how slightly biasing towards under-represented groups in sequential selection processes can lead to long-term fairness.

problem Designing fair sequential decision-making processes for long-term social fairness.
method Proposes Multi-agent Fair-Greedy policy to balance score maximization and fairness.
result Proves convergence to long-term fairness target set by agents when score distributions are identical.

Model predicts time evolution of supply chain networks under varying costs.

problem Regulating downstream relationships for sustainable SMEs.
method Time varying SCN model based on Lagrangian mechanics, incorporating EDES cost kernels.
result Model predicts bankruptcy and break-even states under different cost scenarios.

Robust policies improve ICU transfer outcomes by predicting patient deterioration.

problem Higher mortality rates for unplanned ICU transfers.
method Markov Decision Process model to predict patient severity and optimize transfer policies.
result Robust policies are more aggressive in transferring patients than nominal policies, improving overall patient care.

CPS methods improve sample efficiency in robotics.

problem Improving sample efficiency in reinforcement learning for robotics.
method Empirical evaluation of C-CMA-ES with active covariance matrix adaptation and comparison-based surrogate model.
result Improvements in sample efficiency with C-CMA-ES extensions.

Study best arm identification in restless Markov multi-armed bandits with state-dependent transitions.

problem Identify the best arm in a multi-armed bandit with time-varying states.
method Propose a sequential policy to select arms without knowing their exact TPMs.
result Upper and lower bounds on expected time to find the best arm match in a special case.

Novel RL-based NPG improves multi-objective NAS efficiency and performance.

problem Discovering optimal neural architectures with multiple conflicting objectives.
method Non-stationary policy gradient with adaptive reward functions and shared model.
result Framework efficiently approximates full Pareto front and achieves superior performance.

Optimizes consumption under regime-switching economic states with risk-sensitive preferences.

problem Optimizing consumption in an economy with uncertain states and random shocks.
method Risk-sensitive optimization of consumption-utility with a Markov chain model of economic states and i.i.d. random shocks.
result Existence of unique optimal policy and value function in stationary policies.

Modeling option market making with hedging-induced price impact.

problem Tackles the challenge of market making in options markets with price impact.
method Models option order flow using Cox processes and studies the dynamics of inventory and price under hedging-induced impact.
result Establishes the well-posedness of the mixed control problem involving quoting and hedging.

Analyzes COVID-19 data to predict mortality, forecast spread, and optimize resource allocation.

problem Challenges in patient triage, treatment, and care management during the pandemic.
method Integrated four-step approach combining descriptive, predictive, and prescriptive analytics.
result Optimized resource allocation and informed policy decisions.

Unified coverage analysis for linear off-policy evaluation in reinforcement learning.

problem Lack of a unified understanding of coverage parameters in linear off-policy evaluation.
method Developed a novel finite-sample analysis for LSTDQ algorithm, introducing feature-dynamics coverage.
result Unified understanding of coverage parameters in linear off-policy evaluation.

We have modeled the employment/population ratio in the largest developed countries. Our results show that the evolution of the employment rate since 1970 can be predicted with a high accuracy by a linear dependence on the logarithm of real GDP per capita. All empirical relationships estimated in this study need a struc…

2011-07-24abs ↗pdf ↗