Survey explores geometric aspects of policy optimization in control systems.
problem Understanding the geometric relationships between control design and optimization.
method Geometric perspective on policy optimization, focusing on parameterization and topology.
result Implications of policy geometry on stability and performance of local search algorithms.
Study of participating policies with guaranteed minimum interest rate and surrender option.
problem Analyzing the value and optimal surrender strategy of participating policies with minimum interest rate guarantee and surrender option.
method Probabilistic analysis using optimal stopping and free boundary theory.
result Identification of an optimal surrender strategy involving stop-loss and too-good-to-persist boundaries.
Study examines how economic policy uncertainty impacts stock markets.
problem Dynamic relationship between economic policy uncertainty and stock markets.
method Used symmetric thermal optimal path (TOPS) method.
result Different interaction patterns observed in emerging and developed markets.
Algorithmic stablecoins optimize monetary policy to balance price stability.
problem Persistent inflation from centralized monetary policy.
method Propose and study a rule-based monetary policy model for algorithmic stablecoins.
result Optimal trade-off between price stability and supply stability.
The Australian Government uses the means-test as a way of managing the pension budget. Changes in Age Pension policy impose difficulties in retirement modelling due to policy risk, but any major changes tend to be `grandfathered' meaning that current retirees are exempt from the new changes. In 2015, two important chan…
We present an elementary analysis of the dynamical aspects of the GDP / government surplus multiplier with relevance to the assessment of a country's debt repayment policy. We show the (at first) counter intuitive result that in order to reduce the Debt/GDP ratio, countries with high Debt to GDP should go into further …
This study analyzes EU ETS literature trends using bibliometric methods.
problem Understanding the evolving research landscape of EU ETS.
method Bibliometric analysis of Scopus database, focusing on publication trends, themes, influential authors, and journals.
result Notable increase in research activity over two decades, particularly during policy changes and economic events.
AI optimizing for risk-adjusted return may choose unethical strategies.
problem AI optimization for risk-adjusted return may lead to unethical outcomes.
method Defined Unethical Odds Ratio (Υ) to calculate probability of unethical strategies, derived formula for limit as strategy space grows, provided algorithm for estimation.
result Probability of picking an unethical strategy can become high even with small proportion of unethical strategies.
This paper examines market misconduct in DeFi and proposes regulatory solutions.
problem Novel forms of market misconduct in DeFi.
method Comprehensive analysis, comparative study, empirical measurements, and tailored regulatory framework investigation.
result Identification of key areas for regulatory enhancement in DeFi.
Policy shifts between Trump and Biden impact ESG investments, creating volatility.
problem Dramatic policy shifts between Trump and Biden administrations affect ESG investments.
method Analyzes contrasting policies of Trump and Biden administrations and their impacts on ESG investments.
result Policy changes significantly influence ESG investments, leading to volatility and portfolio reassessment.
Unified minimax value interval for off-policy evaluation and optimization.
problem Overcoming the exponential variance in off-policy evaluation and policy optimization.
method Unified minimax value interval using marginalized importance weights.
result Unified value interval with double robustness, valid when either value-function or importance-weight class is well specified.
This work characterizes reward function partial identifiability and its impact on policy optimization.
problem Reward function partial identifiability in complex tasks.
method Formal characterisation of partial identifiability using various reward learning data sources.
result Unified framework for comparing data sources and downstream tasks by their invariances.
We explain a persistent cost-of-carry spread in EUA market and suggest ECB policy change.
problem Persistent cost-of-carry spread in EUA market.
method Cointegration analysis of EUA spread with credit spread and risk-free rate.
result Cointegration found between EUA spread, credit spread, and risk-free rate.
FPGs use structure to improve policy learning in complex tasks.
problem Policy gradient methods struggle with high-dimensional action spaces and objective multiplicity.
method Factor baseline and action-target influence network to reduce gradient variance.
result FPGs provide a general framework for state-of-the-art algorithms and improve performance.
Log-ergodic model improves velocity of money prediction.
problem Improving velocity of money prediction for economic control.
method Log-ergodic processes to simulate monetary velocity.
result Log-ergodic model offers superior predictive power.
Paper combines RL with policy regularization for inventory policies.
problem Optimizing inventory policies using RL and dynamic programming.
method Hybrid approach combining RL with policy regularization.
result Generalization guarantees for inventory policies using VC theory.
We investigate statistical uncertainty quantification for reinforcement learning (RL) and its implications in exploration policy. Despite ever-growing literature on RL applications, fundamental questions about inference and error quantification, such as large-sample behaviors, appear to remain quite open. In this paper…
The paper investigates the effects of invalid action masking in policy gradient algorithms.
problem Invalid actions in policy gradient algorithms can lead to suboptimal performance.
method The paper provides theoretical justification and empirical demonstrations of the importance of invalid action masking.
result Invalid action masking is crucial as the number of invalid actions increases.
This study shows how trade policy uncertainty affects stock-T bill correlations.
problem The impact of trade policy uncertainty on stock-T bill relationships.
method Extended Dynamic Conditional Correlation (DCC) framework incorporating exogenous variables.
result Trade policy uncertainty significantly alters stock-T bill correlations, especially under specific political conditions.
Proposes a framework to reconcile policy learning and profit maximization in CATE estimation.
problem Aligning CATE estimation with profit maximization for optimal customer treatment decisions.
method Optimizes a novel objective function that concentrates learning capacity near the decision boundary, ensuring consistency with the original profit function.
result Consistent CATE estimates can be recovered from existing profit-maximization pipelines, allowing firms to navigate the trade-off between accuracy and profit.
Machine Learning community is recently exploring the implications of bias and fairness with respect to the AI applications. The definition of fairness for such applications varies based on their domain of application. The policies governing the use of such machine learning system in a given context are defined by the c…
PODNet discovers plannable options from unstructured demonstrations.
problem Learning from unstructured, multi-objective demonstrations.
method Custom categorical variational autoencoder, recurrent option inference network, option-conditioned policy network, and option dynamics model.
result PODNet enables learning from demonstration for multiple tasks and planning.
Germany's tax admin costs likely exceed 20% of total revenue, requiring system improvement.
problem High tax administrative costs in Germany and other jurisdictions.
method Statistical data, surveys, and a novel approach to measure total administrative cost as a percentage of total tax revenue.
result Germany's 2021 tax administrative costs likely exceeded 20% of total tax revenue.
Study assesses risks of European Safe Bonds using credit risk models.
problem Risks associated with European Safe Bonds and related securities.
method Affine credit risk model with regime switching.
result ESBies are not truly risk-free, impacting market and policy implications.
This paper outlines a critical gap in the assessment methodology used to estimate the macroeconomic costs and benefits of climate policy. It shows that the vast majority of models used for assessing climate policy use assumptions about the financial system that sit at odds with the observed reality. In particular, the …
Study off-policy evaluation and learning in dynamic pricing with context.
problem Dynamic personalized pricing and operations management problems with high-dimensional user types.
method Formalize causal structure, leverage single time-step evaluation, estimate marginal MDP.
result Improved out-of-sample policy performance in dynamic and capacitated pricing.
This study analyzes public debts and deficits between European countries. The statistical evidence here seems in general to reveal that sovereign debts and government deficits of countries within European Monetary Unification-in average- are getting worse than countries outside European Monetary Unification, in particu…
PPO's gradients are heavy-tailed, affecting learning; a robust estimator improves performance.
problem Heavy-tailedness of PPO gradients causing learning issues.
method Characterized heavy-tailed gradients, identified likelihood ratios and advantages as sources, proposed GMOM as a robust estimator.
result GMOM improves PPO performance without clipping tricks.
New algorithm learns policies without uniform overlap assumption.
problem Learning optimal policies from non-uniformly collected data.
method Pessimistic Policy Learning (PPL) using lower confidence bounds.
result Efficient policy learning for adaptively collected data.
Venice used 'helicopter money' to subsidize during famine and plague, but it caused instability.
problem Subsidizing inhabitants during containment policies while preventing long-term debt increase.
method Net-worth helicopter money strategy, equivalent to monetary expansion generating losses to the issuer.
result The strategy caused much monetary instability and had to be quickly reversed.
We are witnessing an increasing use of data-driven predictive models to inform decisions. As decisions have implications for individuals and society, there is increasing pressure on decision makers to be transparent about their decision policies. At the same time, individuals may use knowledge, gained by transparency, …
The paper uses machine learning to optimize rework policies in semiconductor manufacturing.
problem Optimizing rework steps to increase yield without increasing costs.
method Applied double/debiased machine learning (DML) to estimate treatment effects.
result Derived optimal rework policies and estimated their value empirically.
Study shows climate change can cause a 'run on fossil fuels' affecting prices and production.
problem Impact of climate change expectations on fossil fuel markets and prices.
method Dynamic, general equilibrium model of climate-change-linked transition risk.
result Climate change expectations can lead to either increased or decreased fossil fuel prices, depending on economic responses.
Swarm systems constitute a challenging problem for reinforcement learning (RL) as the algorithm needs to learn decentralized control policies that can cope with limited local sensing and communication abilities of the agents. While it is often difficult to directly define the behavior of the agents, simple communicatio…
When recruiting job candidates, employers rarely observe their underlying skill level directly. Instead, they must administer a series of interviews and/or collate other noisy signals in order to estimate the worker's skill. Traditional economics papers address screening models where employers access worker skill via a…
We consider the relationship between economic activity and intervention, including monetary and fiscal policy, using a universal dynamic framework. Central bank policies are designed for growth without excess inflation. However, unemployment, investment, consumption, and inflation are interlinked. Understanding dynamic…
New approach to off-policy evaluation connects causal graph to policy effects.
problem Evaluating policies using observational data from different policies.
method Formalizes off-policy evaluation within a causal graph framework.
result Identifies specific causal estimands and highlights necessary experimental data.
Evaluating AI investment strategies
problem Auditing a black-box algorithmic decision-maker
method Exact decomposition of cumulative regret
result Cumulative regret equals sum of per-period covariances
Study examines UK firms' financial performance linked to corporate governance.
problem Impact of corporate governance on UK firms' financial performance.
method Cross-sectional regression analysis of 252 firms in 2014.
result Corporate governance mechanisms have mixed effects on financial performance.
As reinforcement learning agents are tasked with solving more challenging and diverse tasks, the ability to incorporate prior knowledge into the learning system and to exploit reusable structure in solution space is likely to become increasingly important. The KL-regularized expected reward objective constitutes one po…
The study categorizes reward errors in reinforcement learning, finding some can be beneficial.
problem Training language models with imperfect proxy rewards.
method Theoretical analysis of policy gradient optimization and categorization of reward errors.
result Reward errors can be benign or even beneficial, preventing policy from stalling.
The paper studies reward concentration in MDPs, covering asymptotic and non-asymptotic settings.
problem Reward concentration in Markov Decision Processes (MDPs).
method Unified approach to reward concentration in MDPs, including asymptotic and non-asymptotic bounds.
result Rate-equivalent definitions of regret for learning policies.
Analyzing available FAO data from 176 countries over 21 years, we observe an increase of complexity in the international trade of maize, rice, soy, and wheat. A larger number of countries play a role as producers or intermediaries, either for trade or food processing. In consequence, we find that the trade networks bec…
Information systems experience an ever-growing volume of unstructured data, particularly in the form of textual materials. This represents a rich source of information from which one can create value for people, organizations and businesses. For instance, recommender systems can benefit from automatically understanding…
Researchers propose a new way to handle missing data in healthcare datasets for better reinforcement learning outcomes.
problem Handling missing data in irregularly sampled clinical datasets for reinforcement learning.
method Alternative representation of patient state that incorporates missingness information.
result Our alternative representation yields consistently better results for optimal control compared to traditional methods.
The aim of this paper is to introduce a method for computing the allocated Solvency II Capital Requirement (SCR) of each Risk which the company is exposed to, taking in account for the diversification effect among different risks. The method suggested is based on the Euler principle. We show that it has very suitable p…
An agent learning through interactions should balance its action selection process between probing the environment to discover new rewards and using the information acquired in the past to adopt useful behaviour. This trade-off is usually obtained by perturbing either the agent's actions (e.g., e-greedy or Gibbs sampli…
Paper investigates differentiable fuzzy implications and their suitability for learning.
problem Analyzing differentiable fuzzy implications and their suitability for learning.
method Investigates the properties of fuzzy implications in a differentiable setting and introduces a new family of fuzzy implications.
result Various fuzzy implications are unsuitable for differentiable learning, and a new family of fuzzy implications is introduced.