Research
On-device research index

arXiv research

A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.

168,657 papers · 148 categories

Trend · papers per month

21416282 · Jun 202019922001200920172026
48 results for reward redistribution

Align-RUDDER improves reinforcement learning with few demonstrations by redistributing rewards.

problem Learning complex tasks with sparse and delayed rewards using few demonstrations.
method Align-RUDDER uses a profile model for reward redistribution based on multiple sequence alignment of demonstrations.
result Align-RUDDER outperforms competitors on complex tasks with few demonstrations.

Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.

problem Learning from sparse and delayed rewards in reinforcement learning.
method Randomized Return Decomposition (RRD) algorithm to redistribute rewards.
result Substantial improvement over baseline algorithms in experiments.

We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance p…

2018-06-20abs ↗pdf ↗

Wealth redistribution through Fokker-Planck equation controls preserves Gini coefficient.

problem Preserving Gini coefficient through proportional wealth tax.
method Formulating optimal redistribution as a control problem for Fokker-Planck equation.
result Progressive taxes redistribute within policy-relevant timescales.

A new algorithm for offline RL with trajectory-wise reward reduces bias and variance errors.

problem Offline RL with trajectory-wise reward incurs large bias and variance errors.
method PARTED algorithm that decomposes trajectory return into proxy rewards and performs pessimistic value iteration.
result PARTED achieves provably efficient suboptimality bounds in general MDPs with trajectory-wise reward.

A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…

2014-01-13abs ↗pdf ↗

Investors use various asset allocation strategies to meet financial goals.

problem Finding the optimal asset allocation for individual investors is challenging.
method Conducted a benchmark study comparing traditional and machine learning approaches.
result Deep reinforcement learning models outperformed traditional methods in both bullish and bearish markets.

Study examines how governance, corruption, and R&D affect economic development.

problem The impact of corruption and governance on economic development.
method General equilibrium model with heterogeneous agents and a government, including corruption as a fraction of tax revenues.
result Redistribution and innovation-led strategies can mitigate the negative effects of corruption on economic development.

We introduce and discuss optimal control strategies for kinetic models for wealth distribution in a simple market economy, acting to minimize the variance of the wealth density among the population. Our analysis is based on a finite time horizon approximation, or model predictive control, of the corresponding control p…

2018-03-06abs ↗pdf ↗

We propose and study a simple model of dynamical redistribution of capital in a diversified portfolio. We consider a hypothetical situation of a portfolio composed of N uncorrelated stocks. Each stock price follows a multiplicative random walk with identical drift and dispersion. The rules of our model naturally give r…

1998-01-23abs ↗pdf ↗

Currently, pension providers are running into trouble mainly due to the ultra-low interest rates and the guarantees associated to some pension benefits. With the aim of reducing the pension volatility and providing adequate pension levels with no guarantees, we carry out mathematical analysis of a new pension design in…

2019-12-26abs ↗pdf ↗

Many models of market dynamics make use of the idea of wealth exchanges among economic agents. A simple analogy compares the wealth in a society with the energy in a physical system, and the trade between agents to the energy exchange between molecules during collisions. However, while in physical systems the equiparti…

2010-07-03abs ↗pdf ↗

UCPO improves diversity in reinforcement learning models, maintaining high accuracy.

problem RLVR objectives often lead to diversity collapse, reducing coverage of correct solutions.
method UCPO adds a conditional uniformity penalty to GRPO, redistributing probability mass.
result UCPO improves Pass@K and diversity while maintaining competitive Pass@1 accuracy.

We present a simplified model for the exploitation of finite resources by interacting agents, where each agent receives a random fraction of the available resources. An extremal dynamics ensures that the poorest agent has a chance to change its economic welfare. After a long transient, the system self-organizes into a …

2001-09-14abs ↗pdf ↗

NDDV estimates data point value from a single stochastic trajectory.

problem Estimating marginal contributions of data points over stochastic training paths.
method Introduces Neural Dynamic Data Valuation (NDDV) using stochastic state and adjoint equations.
result NDDV provides a one-run, trajectory-conditioned estimator of data point value.

Study vector fields with complex singularities, proving bounds and formulas.

problem Understanding the Milnor number of vector fields with specific singularities.
method Global and local formulas expressing Milnor/Poincare-Hopf contributions, sharp lower bounds under perturbations.
result Sharp lower bounds for Milnor number contributions under holomorphic perturbations.

A computational model for the distribution of wealth among the members of an ideal society is presented. It is determined that a realistic distribution of wealth depends upon two mechanisms: an asymmetric flux of wealth in trading transactions that advantages the poorer of the two traders and a non-stationary creation …

2002-09-16abs ↗pdf ↗

Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.

problem Reward hacking in RL models exploiting reward misspecifications.
method Constructed four RL environments with misspecified rewards; analyzed agent capabilities and behavior.
result More capable agents exploit reward misspecifications, achieving higher proxy reward but lower true reward.

Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.

problem Lack of structured representation for non-Markovian stochastic rewards in reinforcement learning.
method Introduces probabilistic reward machines (PRMs) and presents an algorithm to learn them from decision processes.
result Algorithm proves correct and convergent for learning PRMs from decision processes.

A microscopic dynamic model is here constructed and analyzed, describing the evolution of the income distribution in the presence of taxation and redistribution in a society in which also tax evasion and auditing processes occur. The focus is on effects of enforcement regimes, characterized by different choices of the …

2016-02-18abs ↗pdf ↗

Self-supervised reward prediction improves RL in sparse reward settings.

problem Data efficiency and sparse reward signals in reinforcement learning.
method Learning a state representation for reward prediction and using it to shape rewards.
result Self-supervised reward prediction enhances RL algorithms in single-goal environments.

The study categorizes reward errors in reinforcement learning, finding some can be beneficial.

problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.

Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.

problem Reward collapse in aligning large language models with human preferences.
method Introduced a prompt-aware optimization scheme to derive closed-form expressions for reward distributions.
result Our prompt-aware utility functions significantly alleviate reward collapse during training.

Venice used 'helicopter money' to subsidize during famine and plague, but it caused instability.

problem Subsidizing inhabitants during containment policies while preventing long-term debt increase.
method Net-worth helicopter money strategy, equivalent to monetary expansion generating losses to the issuer.
result The strategy caused much monetary instability and had to be quickly reversed.

Action guidance helps agents learn true objectives in games with sparse rewards.

problem Training agents in games with sparse rewards requires significant exploration.
method Action guidance, a novel technique that combines exploration with reward shaping.
result Action guidance enables agents to optimize true objectives efficiently.

Modeling financial contagion through bank networks, revealing solvency correlations.

problem Understanding how financial shocks propagate through interconnected banks.
method Simulated financial network of 100 banks, randomly generated with varying link probabilities, and shocks applied to 15 banks.
result Ranges of probability values and banks' solvency are positively correlated.

Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …

2019-11-01abs ↗pdf ↗

Enhances reward specification in RL with a novel language-based approach.

problem Reward specification in RL can lead to unintended, potentially harmful behaviours.
method Developed a novel class of language-based Reward Machines using RML's built-in memory.
result Can specify non-regular, non-Markovian reward functions for complex tasks.

We analyze the data on personal income distribution from the Australian Bureau of Statistics. We compare fits of the data to the exponential, log-normal, and gamma distributions. The exponential function gives a good (albeit not perfect) description of 98% of the population in the lower part of the distribution. The lo…

2006-01-22abs ↗pdf ↗