Align-RUDDER improves reinforcement learning with few demonstrations by redistributing rewards.
arXiv research
A locally-built, LLM-digested index of recent arXiv papers in quant finance, geometry/topology, and statistical ML — keyword search served straight from SQLite on this machine.
Trend · papers per month
Paper proposes RRD to learn proxy rewards for sparse delayed rewards in episodic reinforcement learning.
RIVCoin stabilizes cryptocurrency portfolios through a DAO and redistributes income.
We propose RUDDER, a novel reinforcement learning approach for delayed rewards in finite Markov decision processes (MDPs). In MDPs the Q-values are equal to the expected immediate reward plus the expected future rewards. The latter are related to bias problems in temporal difference (TD) learning and to high variance p…
Income redistribution is the transfer of income from some individuals to others directly or indirectly by means of social mechanisms, such as taxation, public services and so on. Employing a spatial public goods game, we study the influence of income redistribution on the evolution of cooperation. Two kinds of evolutio…
Wealth redistribution through Fokker-Planck equation controls preserves Gini coefficient.
Study shows wealth distribution tails near criticality are not universal.
A new algorithm for offline RL with trajectory-wise reward reduces bias and variance errors.
Paper introduces a new principle for fair redistribution of insurance surplus.
We demonstrate by mathematical analysis and systematic computer simulations that redistribution can lead to sustainable growth in a society. The human capital dynamics of each agent is described by a stochastic multiplicative process which, in the long run, leads to the destruction of individual human capital and the e…
A government has to finance a risk for its population. It shares the charges among the population with a fixed scale based on economic criteria. Various organisms have to collect and to redistribute fairly the subsidies. Under these conditions, when the size of the organisms is varied, the distribution's laws of the cr…
In this work we use an inelastic scattering process of particles to propose a model able to reproduce the salient features of the wealth distribution in an economy by including taxes to each trading process and redistributing that collected among the population according to a given criterion. Additionally, we show that…
Investors use various asset allocation strategies to meet financial goals.
Model shows significant income inequality emerges from equal opportunities in a simple economy.
Study examines how governance, corruption, and R&D affect economic development.
We present here a general framework, expressed by a system of nonlinear differential equations, suitable for the modelling of taxation and redistribution in a closed (trading market) society. This framework allows to describe the evolution of the income distribution over the population and to explain the emergence of c…
We introduce and discuss optimal control strategies for kinetic models for wealth distribution in a simple market economy, acting to minimize the variance of the wealth density among the population. Our analysis is based on a finite time horizon approximation, or model predictive control, of the corresponding control p…
We discuss a family of models expressed by nonlinear differential equation systems describing closed market societies in the presence of taxation and redistribution. We focus in particular on three example models obtained in correspondence to different parameter choices. We analyse the influence of the various choices …
We propose and study a simple model of dynamical redistribution of capital in a diversified portfolio. We consider a hypothetical situation of a portfolio composed of N uncorrelated stocks. Each stock price follows a multiplicative random walk with identical drift and dispersion. The rules of our model naturally give r…
Financial portfolio management is the process of constant redistribution of a fund into different financial products. This paper presents a financial-model-free Reinforcement Learning framework to provide a deep machine learning solution to the portfolio management problem. The framework consists of the Ensemble of Ide…
Currently, pension providers are running into trouble mainly due to the ultra-low interest rates and the guarantees associated to some pension benefits. With the aim of reducing the pension volatility and providing adequate pension levels with no guarantees, we carry out mathematical analysis of a new pension design in…
The study revises GDPpc trends and redistributes economic power among countries.
We demonstrate the possibility of what we call sparse learning: accelerated training of deep neural networks that maintain sparse weights throughout training while achieving dense performance levels. We accomplish this by developing sparse momentum, an algorithm which uses exponentially smoothed gradients (momentum) to…
MC-LSTM extends LSTM to conserve mass in neural networks.
Many models of market dynamics make use of the idea of wealth exchanges among economic agents. A simple analogy compares the wealth in a society with the energy in a physical system, and the trade between agents to the energy exchange between molecules during collisions. However, while in physical systems the equiparti…
UCPO improves diversity in reinforcement learning models, maintaining high accuracy.
We present a simplified model for the exploitation of finite resources by interacting agents, where each agent receives a random fraction of the available resources. An extremal dynamics ensures that the poorest agent has a chance to change its economic welfare. After a long transient, the system self-organizes into a …
We present a method for constructing the log-optimal portfolio using the well-calibrated forecasts of market values. Dawid's notion of calibration and the Blackwell approachability theorem are used for computing well-calibrated forecasts. We select a portfolio using this "artificial" probability distribution of market …
In recent work, Boltzmann and Fokker-Planck equations were derived for the "Yard-Sale Model" of asset exchange. For the version of the model without redistribution, it was conjectured, based on numerical evidence, that the time-asymptotic state of the model was oligarchy -- complete concentration of wealth by a single …
NDDV estimates data point value from a single stochastic trajectory.
Study vector fields with complex singularities, proving bounds and formulas.
A computational model for the distribution of wealth among the members of an ideal society is presented. It is determined that a realistic distribution of wealth depends upon two mechanisms: an asymmetric flux of wealth in trading transactions that advantages the poorer of the two traders and a non-stationary creation …
The so-called "Yard-Sale Model" of wealth distribution posits that wealth is transferred between economic agents as a result of transactions whose size is proportional to the wealth of the less wealthy agent. In recent work [B.M. Boghosian, "Kinetics of Wealth and the Pareto Law," {\it Phys. Rev. E} {\bf 89} (2014) 042…
Reward hacking exploits misspecified rewards, affecting agent capabilities and true performance.
Paper introduces PRMs to learn non-Markovian stochastic rewards for reinforcement learning.
This study is a detailed analysis of Speculation Game, a minimal agent-based model of financial markets, in which the round-trip trading and the dynamic wealth evolution with variable trading volumes are implemented. Instead of herding behavior, we find that the emergence of volatility clustering can be induced by the …
A microscopic dynamic model is here constructed and analyzed, describing the evolution of the income distribution in the presence of taxation and redistribution in a society in which also tax evasion and auditing processes occur. The focus is on effects of enforcement regimes, characterized by different choices of the …
This work analyzes the value of future reward information in RL.
Self-supervised reward prediction improves RL in sparse reward settings.
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
Reward collapse occurs when ranking-based reward models yield uniform rewards for different prompts.
Venice used 'helicopter money' to subsidize during famine and plague, but it caused instability.
Reward models need more than just accuracy for effective RLHF.
Action guidance helps agents learn true objectives in games with sparse rewards.
Modeling financial contagion through bank networks, revealing solvency correlations.
Learning reward functions from data is a promising path towards achieving scalable Reinforcement Learning (RL) for robotics. However, a major challenge in training agents from learned reward models is that the agent can learn to exploit errors in the reward model to achieve high reward behaviors that do not correspond …
Enhances reward specification in RL with a novel language-based approach.
We analyze the data on personal income distribution from the Australian Bureau of Statistics. We compare fits of the data to the exponential, log-normal, and gamma distributions. The exponential function gives a good (albeit not perfect) description of 98% of the population in the lower part of the distribution. The lo…